Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions
arXiv:2608.18078v1 Announce Type: new
Abstract: This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to exhibit collusive behavior and should be required to obtain behavioral certification before making decisions that affect economic markets. This i...
Position: Profiling Game Worlds by Transition Complexity
arXiv:2608.18079v1 Announce Type: new
Abstract: Game world modeling (GWM) and reinforcement learning (RL) are often confounded because research papers rarely quantify how difficult the underlying transition prediction problem is at the declared interface (pixels/tokens/latents with finite history)....
Position: Current Model Cards Are Insufficient for Downstream Governance of Open-Weight Foundation Models
arXiv:2608.18086v1 Announce Type: new
Abstract: The growth of open-weight foundation models (OWFMs) has prompted the AI community to re-evaluate strategies for effective downstream governance. Although model cards have been widely adopted as transparency artifacts in model repositories, existing fr...
Towards Reversible Forgetting: Managing Obsolete Knowledge in Continual Enterprise AI Agents
arXiv:2608.18177v1 Announce Type: new
Abstract: Continual learning has traditionally treated forgetting as a failure, emphasizing preservation of previously acquired knowledge as environments evolve. We argue that this objective is incomplete for enterprise AI agents operating in non-stationary env...
arXiv:2608.18147v1 Announce Type: new
Abstract: Adaptive stochastic quantization (ASQ) is a recently introduced quantization approach that optimizes the Mean Squared Error (MSE) for a given input while preserving unbiasedness. It is designed to alleviate the communication and memory bottlenecks of ...
H$^2$EDL: Hyper Evidential Deep Learning for Hierarchical Classification
arXiv:2608.18185v1 Announce Type: new
Abstract: Fine-grained recognition often involves hierarchical label spaces, where a model may be confident about a coarse semantic concept while remaining uncertain among its descendant classes. Such structured ambiguity requires uncertainty representations th...
Runtime Governance for Agentic AI: Action-Boundary Control with Trusted Provenance and Fail-Closed Execution
arXiv:2608.16891v1 Announce Type: new
Abstract: Agentic AI systems request tool actions that can modify files, send messages, launch jobs, or change workflow state. This shifts the safety problem from harmful text generation to harmful operational side effects. Prompt-level governance can shape mod...
The Problem Is the Problem: Towards Scalable Mathematical Discovery
arXiv:2608.16977v1 Announce Type: new
Abstract: AI systems are increasingly capable of contributing to mathematical research. In research practice, frontier-model reasoning is a limited resource, and expert mathematical review is even more sharply constrained. Allocating these scarce resources well...
GxP-Agent: Process-DAG Topology for Reliable Clinical Trial Programming with LLM Agents
arXiv:2608.16890v1 Announce Type: new
Abstract: Clinical trial programming -- transforming study protocols into analysis-ready datasets under CDISC standards -- is a bottleneck in regulatory submissions, yet LLM-based code generation fails catastrophically on this task: across 11 single-shot attemp...
The Price of Thinking: Reasoning Effort as a Model-Specific API Contract
arXiv:2608.16956v1 Announce Type: new
Abstract: API buyers purchase a dated contract, not a model name alone: the contract includes the requested and served model, reasoning-effort term or its omission, output rail, service product, prompt, and price schedule. We study the reasoning-effort term thr...
Benchmarking Classical and Transformer-Based Models for Document Sensitivity Classification
arXiv:2608.16928v1 Announce Type: new
Abstract: Automatic sensitivity classification of organizational documents is a critical yet underserved problem, where the consequences of misclassification range from regulatory violations to security breaches. While AI-based approaches offer a scalable alter...
Data-DPO: Direct Preference Optimization for Target Model Data Selection in LLM Post-Training
arXiv:2608.16926v1 Announce Type: new
Abstract: Data selection in supervised fine-tuning aims to select a small set of effective samples from large-scale candidate data, reducing training cost while preserving model performance. However, existing methods usually treat data value as a relatively sta...
Detecting and Discriminating Operator Misspecification in Hybrid PDE-Parameter Learning: a Reference-Free Instrument, with Discrimination Bounded In Sample
arXiv:2608.16925v1 Announce Type: new
Abstract: We build an instrument that reads, from a single fit and with no oracle, whether the operator a hybrid PDE-parameter estimator postulates is wrong-and separates that from a merely unidentifiable parameter. On one self-adjoint parabolic inverse problem...
Hierarchical Data Selection via Manifold Coverage and Sparse Feature Coverage in LLM Post-training
arXiv:2608.16927v1 Announce Type: new
Abstract: As supervised fine-tuning data continues to scale, selecting high-value subsets from large candidate pools is crucial for reducing training cost and improving model performance. Existing methods often measure diversity directly in the original embeddi...
The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning
arXiv:2608.14558v1 Announce Type: new
Abstract: Current multimodal models have demonstrated remarkable proficiency in recognizing static visual and auditory content. However, their capacity for abstract perceptual reasoning, inferring unseen information from dynamic, generative processes, remains a...
Coarse-to-Fine Multi-Resolution Diffusion Models for Trajectory Generation in Urban Systems
arXiv:2608.14570v1 Announce Type: new
Abstract: Understanding human mobility is critical for a wide range of urban applications, including traffic management, epidemic control, and urban planning. However, due to privacy concerns, the availability of large-scale public trajectory data remains limit...
Learning Discrete Riemannian Metrics for Physical Fields with Cochain-Frame Equivarianc
arXiv:2608.14556v1 Announce Type: new
Abstract: Physical fields on meshes require a separation between topology and geometry: conservation laws are topological and should be exact, while geometry, material response, and anisotropic coupling must be learned from data. Existing neural surrogates ofte...
DumpsterCluster: From Dumpster Diving to Serving LLaMA-70B on $60 GPUs
arXiv:2608.14614v1 Announce Type: new
Abstract: As AI datacenters retire functional GPUs, vast quantities of still capable accelerators enter secondary markets. This paper investigates whether these retired GPUs can find a productive afterlife to form a DumpsterCluster that can serve modern LLM inf...
When to Communicate: Belief Distributions and KL Divergence for Principled Gating in Multi-Agent RL
arXiv:2608.14559v1 Announce Type: new
Abstract: Effective communication in multi-agent reinforcement learning requires agents to decide not only \textit{what} to communicate, but when? Existing approaches either communicate at every timestep or learn a binary gate through REINFORCE policy gradients...
arXiv:2608.14563v1 Announce Type: new
Abstract: Forward-Pass-Only MLP training (FPO) adapts large language models without a backward pass through the model body, achieving 2.7--3.2x the throughput of standard fine-tuning at ~40% less peak training memory, while leaving off-domain benchmarks within ...
Large Language Models Show Metacognitive Sensitivity in Medical Reasoning
arXiv:2608.14552v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly evaluated and used in medicine, but clinical usefulness depends on answer accuracy and whether confidence tracks evidence quality and uncertainty. We developed a controlled, psychophysics-inspired clinical...
Hard Cases, Bad Labels: Testing Error Exposure and Error Location in Uncertainty Sampling Under Bounded Label Noise
arXiv:2608.13601v1 Announce Type: new
Abstract: Active learning can reduce labeling cost by selecting informative examples, but the most uncertain examples may also be the hardest to label correctly. This study tests whether uncertainty sampling fails because it acquires more corrupted labels or be...
Modular Cognitive Architecture Emerges in Large Language Models
arXiv:2608.13567v1 Announce Type: new
Abstract: The human brain exhibits a striking degree of functional specialization, with distinct networks supporting language, formal reasoning, reasoning about other minds, and reasoning about the physical world. Is this modular organization a fundamental prin...
L-FNO: Lorentzian Fourier Neural Operator for Stochastic Event Dynamics
arXiv:2608.13562v1 Announce Type: new
Abstract: Modern operational systems face uncertainty even in routine conditions, where rare, bursty, and self-exciting events emerge from both exogenous covariates and endogenous event dynamics. Standard neural operators are typically trained as regression-sty...