arXiv:2607.21633v1 Announce Type: new
Abstract: Logic Gate Networks (LGNs) implement computation through compositions of Boolean operations, yet unlike classical Boolean circuits, existing LGNs do not reliably benefit from increased depth. We identify two distinct causes: optimization collapse in d...
Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions
arXiv:2607.21635v1 Announce Type: new
Abstract: Personal agents maintain memories, learned skills, tool configurations, and policy state that evolve with each user. Existing agent benchmarks often evaluate these capabilities in isolation: tool benchmarks test invocation under fixed APIs, memory ben...
Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models
arXiv:2607.21636v1 Announce Type: new
Abstract: Synthetic tabular data is valued for preserving not only each column's marginal distribution but the dependencies between columns -- structure that carries much of the discriminative signal for minority classes in imbalanced domains such as fraud and ...
FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills
arXiv:2607.21596v1 Announce Type: new
Abstract: Large language model agents increasingly solve complex tasks by constructing inference-time workflows that combine reasoning, tool use, and code execution. While such workflows enable flexible problem solving, the useful procedures discovered during e...
Risk Is Not the Target: A Monotonic Framework for Evaluating Wildfire Operational Risk Signals
arXiv:2607.21597v1 Announce Type: new
Abstract: Evaluating wildfire risk systems using standard machine-learning metrics such as F1-score or IoU is fundamentally flawed: these metrics assess event prediction accuracy, not the operational coherence of a continuous risk signal. This work proposes a n...
Securing Multimodal AI through Internal Information Decomposition
arXiv:2607.21600v1 Announce Type: new
Abstract: Multimodal large language models introduce attack surfaces absent in unimodal systems: adversaries can distribute malicious intent across modalities to evade unimodal safeguards. This motivates using cross-modal consistency as a detection signal rathe...
Transferable Latency Prediction for Fast LLM Screening on Heterogeneous Edge Devices
arXiv:2607.21602v1 Announce Type: new
Abstract: Accurate latency prediction is critical for deploying large language models (LLMs) on heterogeneous edge devices, where inference latency is affected by model architecture, prompt behavior, runtime backend, hardware utilization, dynamic voltage and fr...
GH-ESD: Grounded Hypothesis-Driven Error Slice Discovery for Instance-Level Vision Tasks
Systematic failures of vision models on semantically coherent subsets, known as error slices, reveal limitations in robustness and evaluation. Existing slice discovery approaches largely model slices as clusters in representation space or combinations of predefined attributes. While effective for im...
Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction
-->
-->
Overview of ABBEL compared to traditional recursive summarization. Beliefs replace the full interaction history as the agent’s working context, and belief grading improves performance by
supervising the contents of each belief state..
As task horizons grow, LLM contexts can’t s...
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators
arXiv:2607.20465v1 Announce Type: new
Abstract: The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how well LLMs, agents, and data-centric workflows actually prepare training data end to end. We view LLM...
PhantomFill: When the Form Demands an Answer, Language Models Invent One
arXiv:2607.20492v1 Announce Type: new
Abstract: Language models in production do not write prose. They fill forms: JSON fields, function arguments, extraction templates. We show that the form itself causes hallucination.
We ask thirteen models the same question about the same input and change onl...
Scaling Closed-Loop Feature Channel Configuration with LLMs
arXiv:2607.20516v1 Announce Type: new
Abstract: Promising initial results in closed-loop large-language-model-based channel-configuration search demonstrated that neural-network widths can be optimized directly through executable code generation and accuracy feedback. However, those results were ob...
AINTMA: Agentic AI Architecture for Autonomous Test Management with Generative Intelligence, Secure Cloud Communication and Adaptive Quality Analytics
arXiv:2607.20452v1 Announce Type: new
Abstract: Modern software quality assurance demands intelligent, autonomous systems capable of adaptive decision-making across distributed cloud environments. This paper presents AINTMA (Agentic Intelligent Test Management Architecture), a multi-agent agentic A...
Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts
arXiv:2607.20462v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly integrated into clinical workflows, stressing the need for reliable traceability of model-generated output with watermarking. Yet, most watermarks are evaluated on general-purpose benchmarks, leaving domai...
ClickGuard: Detecting and Spoiling Clickbait News with Informativeness Measures and Large Language Models
arXiv:2607.20463v1 Announce Type: new
Abstract: This paper presents an AI-driven browser extension that identifies clickbait to help users avoid misleading Internet articles. Moving beyond traditional detection, the application employs a hybrid machine learning architecture that combines transforme...
Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs
arXiv:2607.20464v1 Announce Type: new
Abstract: When a language model gives different answers on repeated runs, does that variation reveal what it does not know? Self-consistency turns the variation into a per-question uncertainty estimate via majority voting. But does the same variation reveal cro...
arXiv:2607.20466v1 Announce Type: new
Abstract: Rigorous benchmarks have driven progress in autonomous GPU kernel performance optimization by establishing a shared target to hillclimb on, but no equivalent exists for TPUs. We present JAXBench, a TPU-native benchmark suite for AI-generated kernel op...
LEAD: Breaking the No-Recovery Bottleneck in Long-Horizon Reasoning
Long-horizon execution in Large Language Models (LLMs) remains unstable even when high-level strategies are provided. Evaluating on controlled algorithmic puzzles, we demonstrate that while decomposition is essential for stability, extreme decomposition creates a “no-recovery bottleneck”. We show th...
Policymakers, academics, healthcare providers, AI developers, and patient advocates convened by Stanford HAI identify critical gaps in how we regulate AI tools used for therapy and emotional support.
Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions
arXiv:2607.19378v1 Announce Type: new
Abstract: Subquadratic alternatives to attention require compromises when applied to multi-dimensional data: standard convolutions lack global receptive fields and input dependency, while recurrent models require rasterizing data such as images, volumes, and pa...
arXiv:2607.19379v1 Announce Type: new
Abstract: Prior work has shown that transformers can perform exact Bayesian filtering within a fixed
hypothesis class. Can they also perform Bayesian model selection -- identifying the correct
hypothesis class from data? We introduce model-selection Bayesia...
CruiseBench: A Real-Flight-Aligned N-CMAPSS Benchmark for Engine RUL Prediction
arXiv:2607.19380v1 Announce Type: new
Abstract: Remaining useful life (RUL) prediction estimates how long an engine can continue safe operation and is central to maintenance planning. N-CMAPSS extends C-MAPSS by simulating run-to-failure aero-engine trajectories using recorded real-flight profiles ...
Air Quality Arena: A Large-Scale Multi-Region Ground Monitoring Dataset and Benchmark for Air Quality Forecasting with Time-Series Foundation Models
arXiv:2607.19381v1 Announce Type: new
Abstract: Air pollution causes an estimated 7.9 million premature deaths annually, making accurate forecasting a critical public health priority. Machine learning is increasingly being applied to forecast air pollution levels, yet existing benchmarks remain nar...