DS-Lighting: Making Agent Harnesses Explicit for Data-Science Automation
arXiv:2608.28590v1 Announce Type: new
Abstract: Large Language Model (LLM) agents have shown promise for automating data-science workflows, yet their end-to-end performance depends critically on the agent harness that represents tasks, manages execution state, constrains output artifacts, and provi...
Equivariant Sheaf Neural Networks: Learning Geometric Transport on Graphs
arXiv:2608.28853v1 Announce Type: new
Abstract: Equivariant graph neural networks provide a principled way to model geometric systems, but efficient first-order architectures remain limited in how vector information can be transformed as it moves across a graph. We introduce \textsc{ESNN}, an Equiv...
Curvature Cryptanalysis of Smooth Transformer Feed-Forward Networks
arXiv:2608.28843v1 Announce Type: new
Abstract: We show that smooth two-layer feed-forward networks (FFNs) expose an additional structural model extraction channel under a chosen-input raw-output oracle at the FFN branch; consider transformer FFN branches with GELU or SiLU activations under chosen-...
The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning
arXiv:2608.28859v1 Announce Type: new
Abstract: Reasoning models do not stop when they know the answer. On DeepSeek-R1-Distill-Qwen-7B the chain of thought runs about twice as long as the model's own answer probability takes to settle, and how much of that excess is removable varies from problem to...
ERR+: Sequential Entropy Resolution for Efficient and Decisive LLM Reasoning
arXiv:2608.28771v1 Announce Type: new
Abstract: Large reasoning models achieve strong performance on complex tasks by generating extended chain-of-thought (CoT) traces via reinforcement learning with verifiable rewards (RLVR). While current RLVR methods have achieved strong results with correctness...
Rating the Raters: Rasch Measurement Theory for LLM Evaluation
arXiv:2608.27463v1 Announce Type: new
Abstract: LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models' outputs, and raters of human-generated content. Each paradigm can be viewed as a measurement problem, where a latent property of an object is probed w...
Not All Explanations Are Sought: Information-Seeking Psychology for Human-Centered XAI
arXiv:2608.27464v1 Announce Type: new
Abstract: This position paper argues that human-centered explainable AI (HCXAI) should incorporate insights from the psychology of information seeking. Drawing on Sharot and Sunstein's framework of information-seeking motives, we propose that people evaluate wh...
Retrieving Relations, Detecting Fallacies: A RAG Approach to Political Debate Analysis
arXiv:2608.27471v1 Announce Type: new
Abstract: Fallacies are arguments that employ invalid reasoning, making their automatic detection critical in sensitive contexts such as high-stakes political debates, where public opinion is shaped. Spotting a fallacious argument requires contextual knowledge ...
LLM-Augmented Causal Discovery: Probabilistic Fusion of Edge Existence and Orientation
arXiv:2608.27472v1 Announce Type: new
Abstract: Bayesian network structure learning (BNSL) from observational data struggles with orientation identifiability, while large language models (LLMs) offer broad but often unreliable causal knowledge. We propose combining these complementary sources throu...
Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Model
arXiv:2608.27459v1 Announce Type: new
Abstract: In 2011, IBM's Watson was something like a sealed capsule of its era's queryable knowledge. Its DeepQA system defeated the strongest human Jeopardy! champions, but the knowledge that let it do so lived in a curated billion-document corpus running on a...
arXiv:2608.27513v1 Announce Type: new
Abstract: Softmax attention stores key and value vectors for every preceding token, causing inference memory to grow with sequence length. Recent language models incorporating Gated DeltaNet (GDN) or Kimi Delta Attention (KDA) reduce this cost by replacing the ...
When Muon Meets Task Interference: A Spectral Perspective on Continual Learning and Model Merging
arXiv:2608.27518v1 Announce Type: new
Abstract: Continual learning (CL) and model merging (MM) both aim to obtain a single model that performs well across multiple tasks, challenged respectively by catastrophic forgetting and weight-disentanglement error. In the literature, these difficulties are m...
Quantization-Triggered Backdoors in Language Models: Cross-Quantizer Transferability and the Validation--Deployment Gap
arXiv:2608.27512v1 Announce Type: new
Abstract: Post-training quantization is often treated as a semantically neutral optimization for edge deployment of Large Language Models. When a full-precision source checkpoint is evaluated and quantization is applied downstream without equivalent re-evaluati...
PICasso: An AI-Enabled Design Framework for Autonomous Optimization of Silicon Photonic Devices
arXiv:2608.26113v1 Announce Type: new
Abstract: We present PICasso, an AI-assisted framework for automated synthesis, verification, and optimization of photonic integrated circuits (PICs) from natural-language specifications. PICasso couples a structured NL -> YAML -> GDS generation pipeline with P...
Pruning Binarized Neural Networks: A Dedicated Framework and Globally Weighted Algorithms
arXiv:2608.26233v1 Announce Type: new
Abstract: Extreme compression of deep neural networks, up to full binarization, dramatically reduces memory footprint and arithmetic complexity, facilitating deployment on constrained edge hardware with field-programmable gate arrays (FPGAs) and microcontroller...
Algebraic Multigrid Acceleration for Efficient Label Spreading
arXiv:2608.26309v1 Announce Type: new
Abstract: Modern machine learning models rely on large amounts of labeled data. However, manual annotation of large-scale datasets is expensive and time-consuming. Label spreading is a semi-supervised learning technique that addresses this challenge by propagat...
Muon with Finite Newton-Schulz: The Smoothing Benefit in Nonsmooth Nonconvex Optimization
arXiv:2608.26288v1 Announce Type: new
Abstract: Muon has emerged as a strong optimizer for the matrix-valued parameters in large language model pretraining, approximately orthogonalizing its momentum with a few Newton-Schulz iterations. Existing theory either replaces this iteration with the exact ...
CIFQA: A Deterministic Tool-Grounded Multi-Agent LLM Framework for Financial Query Answering
arXiv:2608.26114v1 Announce Type: new
Abstract: Calculation-intensive financial question answering requires exact reasoning over structured rates, temporal conditions, numerical formulas, and rule-based constraints. Although Large Language Models (LLMs) perform strongly on natural language tasks, t...
SLM-Conditioned Hierarchical Relation Routing for Labeled Property Graph Learning
arXiv:2608.26132v1 Announce Type: new
Abstract: Labeled property graphs combine relational structure with heterogeneous textual and categorical properties attached to both nodes and relationships. Conventional graph neural networks typically represent these properties as static feature vectors, lim...
EduRiskX: A Neuro-Symbolic Framework with F-Logic Reasoning for Early Academic Risk Prediction
arXiv:2608.26107v1 Announce Type: new
Abstract: Predicting students' academic risk in online education is crucial for enabling timely interventions that can improve retention and learning outcomes. However, existing models often suffer from limited early detection capability and insufficient interp...
GreenLeaf Law Embed Tiny: A Compact Embedding Model for Legal Domain Retrieval
arXiv:2608.24936v1 Announce Type: new
Abstract: We present GreenLeaf Law Embed Tiny, a 0.6B parameter embedding model for legal domain retrieval. GreenLeaf-Tiny achieves 75.11% on the Massive Legal Embedding Benchmark (MLEB) and 64.38% on MTEB(Law, v1),demonstrating competitive performance among mo...
ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration
arXiv:2608.24938v1 Announce Type: new
Abstract: Mixture-of-Experts (MoE) models scale capacity for strong quality while keeping per-token compute bounded through sparse expert activation. Yet low-latency MoE serving is increasingly challenging, because it spans two inference phases with fundamental...
Dynamic Influence-Weighted Distillation for Single-IMU Activity Recognition
arXiv:2608.24904v1 Announce Type: new
Abstract: Inertial sensors at multiple body locations can improve activity recognition, but requiring every sensor at inference increases the deployment burden. We study whether four synchronized IMUs available during training can improve a student that uses on...
Reliable LLM-Powered Decision Engines for Large-Scale Supply Chain Operations: Architecture, Safety, and Performance Guarantees
arXiv:2608.24889v1 Announce Type: new
Abstract: Current large-scale supply chains are highly uncertain, dynamic, and disruption prone that are challenging to serve up timely and resilient decisions through traditional rule-based and optimization-only systems. The increasing supply of heterogeneous ...