Modular Cognitive Architecture Emerges in Large Language Models
arXiv:2608.13567v1 Announce Type: new
Abstract: The human brain exhibits a striking degree of functional specialization, with distinct networks supporting language, formal reasoning, reasoning about other minds, and reasoning about the physical world. Is this modular organization a fundamental prin...
Depth-Aware Sensitivity Analysis of Mixture-of-Experts Models via Magnitude-Based Expert Masking
arXiv:2608.13565v1 Announce Type: new
Abstract: Mixture-of-Experts (MoE) architectures scale large language models (LLMs) while preserving computational efficiency through sparse activation. Despite their widespread adoption, the relative importance of individual MoE layers remains insufficiently c...
Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation
arXiv:2608.13564v1 Announce Type: new
Abstract: Evaluating language-model agents at scale increasingly relies on a second language model as an automatic judge, because the gold signal, an executable environment reward, is expensive, slow, or unavailable at deployment time. Such a judge is a reward-...
A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing
arXiv:2608.13573v1 Announce Type: new
Abstract: Large Language Model (LLM) serving has become a critical cloud workload, and realistic traces are essential for motivating and benchmarking serving systems. However, existing LLM serving workload studies remain limited in scale and scope. They often o...
Agentao: A Governed Local-First Runtime for Tool-Using LLM Agents
arXiv:2608.13574v1 Announce Type: new
Abstract: LLM agents increasingly operate as execution systems that invoke tools, modify local state, use persistent memory, and interact with external protocols. These capabilities make agents useful, but they also introduce risks related to over-privileged ac...
Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists
arXiv:2608.12345v1 Announce Type: new
Abstract: Language models are increasingly deployed as co-scientists, yet their ability to uphold research integrity under institutional pressure remains unmeasured. We introduce IntegrityBench, a benchmark evaluating misconduct classification, ethical action r...
Position: The Alignment Community is Unintentionally Building a Censor's Toolkit
arXiv:2608.12346v1 Announce Type: new
Abstract: This position paper argues that modern AI alignment methods - originally designed to prevent harmful output - are dual-use technologies that may easily be misused by malicious actors for censorship and manipulation. By mapping current alignment techni...
Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments
arXiv:2608.12368v1 Announce Type: new
Abstract: Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs). Yet agreement in final labels does not show that human annotators and models rely on the same moral grounds. Two agents may reach the same j...
Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing
arXiv:2608.12371v1 Announce Type: new
Abstract: Stream-processing systems increasingly operate across heterogeneous mobile edge--cloud infrastructures, where workload volatility, resource contention, and stringent quality-of-service (QoS) requirements complicate decentralized scheduling. This paper...
Position: Reasoning is a Learnable Rule-Based Process
arXiv:2608.12325v1 Announce Type: new
Abstract: Autonomous reasoning is among the most scientifically and economically motivating topics in AI today. Historically the purview of symbolic AI, recent advances have mainly emerged from deep probabilistic generative models. Despite immense interest and ...
arXiv:2608.12438v1 Announce Type: new
Abstract: We formulate generative modeling as a path integral in which flow-based, diffusion-based, variational, and adversarial models arise as different evaluation principles for a single master action. Its Martin-Siggia-Rose-Janssen-de~Dominicis (MSRJD) form...
LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining
arXiv:2608.12419v1 Announce Type: new
Abstract: Large language models (LLMs) have achieved remarkable breakthroughs across various applications. However, their architectures remain inefficient in pretraining due to two main limitations: (i) self-attention lacks an explicit inductive bias for locali...
Basin: Efficient and Extensible Numerical Optimization in Rust
arXiv:2608.11279v1 Announce Type: new
Abstract: Basin is a numerical optimization library for the Rust programming language. Numerical optimization is the task of finding the inputs that minimize a function, and it is a fundamental element across the sciences: fitting a model to data, calibrating a...
FarSky: Task-Aware Latent-Space Coupling for Generative Intra-Hour Solar Forecasting
arXiv:2608.11254v1 Announce Type: new
Abstract: Accurate solar irradiance forecasting is essential for the reliable integration of photovoltaic power into modern electricity grids. All-sky imagers (ASI) provide high-resolution observations of clouds, making them well suited for intra-hour forecasti...
Federated Learning for Distributed CNC Tool Wear Prediction
arXiv:2608.11281v1 Announce Type: new
Abstract: Tool wear prediction is an important task in CNC machining, where accurate monitoring of tool condition supports product quality and process reliability. Machine learning methods have shown potential for this task, but their use in industrial environm...
A Forced-Structure Reduction and Verifiable Bounds for Conway's 99-Graph
arXiv:2608.11211v1 Announce Type: new
Abstract: Conway's 99-graph problem asks whether a strongly regular graph with parameters $\mathrm{srg}(99,14,1,2)$ exists. We report a systematic, fully reproducible attack by an autonomous AI research agent, scored under the track's partial-credit metric. Our...
Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes
arXiv:2608.11207v1 Announce Type: new
Abstract: When two LLM agents with structurally opposed objectives interact across multiple turns, the absence of a shared goal function produces not competition but collapse: the visitor capitulates, the site agent stops varying its approach, and the conversat...
Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop
arXiv:2608.11215v1 Announce Type: new
Abstract: Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase behaviour, stylised facts, and scaling with the number of agents $N$, not the cognition of any sing...
Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration
arXiv:2608.11210v1 Announce Type: new
Abstract: Bayesian calibration of process-based models requires a prior distribution for each model parameter. Despite decades of methodological work, researchers almost always fall back on uniform priors. The main reason is that building informative priors fro...
CurveFP: Rational-Radix Logarithmic Datatypes with Closed Products for Language Models
arXiv:2608.10010v1 Announce Type: new
Abstract: Low-precision datatypes reduce language-model cost, but most formats optimize scalar fidelity while leaving the arithmetic induced by their products unchanged. We introduce CurveFP, a closed-product codebook family that distributes quantized magnitude...
DOCSCHISEL: Adaptive Tool Documentation Optimization Framework for LLM Agents
arXiv:2608.10037v1 Announce Type: new
Abstract: Large language models (LLMs) increasingly rely on external tools to accomplish complex real-world tasks, making tool documentation a critical grounding resource for LLM agents. Existing studies mainly focus on improving the tool-use capabilities of LL...
arXiv:2608.10016v1 Announce Type: new
Abstract: Heterogeneous federated systems require agents to learn and exchange informative representations despite differences in data distributions, sensing modalities, model architectures, latent dimensionalities, and local learning objectives. To address thi...
SPOTting the Future: Lookahead Explanations for Deep Reinforcement Learning
arXiv:2608.09967v1 Announce Type: new
Abstract: Deep reinforcement learning (DRL) agents achieve strong performance in complex environments, yet their decision-making processes remain difficult to interpret. We introduce SPOT (Sampling Policy Observation Tree), a novel model-agnostic, sampling-base...