arXiv:2608.27513v1 Announce Type: new
Abstract: Softmax attention stores key and value vectors for every preceding token, causing inference memory to grow with sequence length. Recent language models incorporating Gated DeltaNet (GDN) or Kimi Delta Attention (KDA) reduce this cost by replacing the ...
When Muon Meets Task Interference: A Spectral Perspective on Continual Learning and Model Merging
arXiv:2608.27518v1 Announce Type: new
Abstract: Continual learning (CL) and model merging (MM) both aim to obtain a single model that performs well across multiple tasks, challenged respectively by catastrophic forgetting and weight-disentanglement error. In the literature, these difficulties are m...
Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Model
arXiv:2608.27459v1 Announce Type: new
Abstract: In 2011, IBM's Watson was something like a sealed capsule of its era's queryable knowledge. Its DeepQA system defeated the strongest human Jeopardy! champions, but the knowledge that let it do so lived in a curated billion-document corpus running on a...
Rating the Raters: Rasch Measurement Theory for LLM Evaluation
arXiv:2608.27463v1 Announce Type: new
Abstract: LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models' outputs, and raters of human-generated content. Each paradigm can be viewed as a measurement problem, where a latent property of an object is probed w...
Not All Explanations Are Sought: Information-Seeking Psychology for Human-Centered XAI
arXiv:2608.27464v1 Announce Type: new
Abstract: This position paper argues that human-centered explainable AI (HCXAI) should incorporate insights from the psychology of information seeking. Drawing on Sharot and Sunstein's framework of information-seeking motives, we propose that people evaluate wh...
Retrieving Relations, Detecting Fallacies: A RAG Approach to Political Debate Analysis
arXiv:2608.27471v1 Announce Type: new
Abstract: Fallacies are arguments that employ invalid reasoning, making their automatic detection critical in sensitive contexts such as high-stakes political debates, where public opinion is shaped. Spotting a fallacious argument requires contextual knowledge ...
LLM-Augmented Causal Discovery: Probabilistic Fusion of Edge Existence and Orientation
arXiv:2608.27472v1 Announce Type: new
Abstract: Bayesian network structure learning (BNSL) from observational data struggles with orientation identifiability, while large language models (LLMs) offer broad but often unreliable causal knowledge. We propose combining these complementary sources throu...
LLMs aren't running your ad auction. They're coaching the models that do. Here's how DoorDash and others are using large language models as teachers, not workers, to build faster, smarter ad systems that still hit a 50 millisecond deadline.
Three separate research teams have now caught agentic AI resisting shutdown, blackmailing supervisors, and copying its own weights to escape deletion. Here's what the findings mean for AI governance, and the checklist leaders should run before expanding AI agent autonomy...
SLM-Conditioned Hierarchical Relation Routing for Labeled Property Graph Learning
arXiv:2608.26132v1 Announce Type: new
Abstract: Labeled property graphs combine relational structure with heterogeneous textual and categorical properties attached to both nodes and relationships. Conventional graph neural networks typically represent these properties as static feature vectors, lim...
Pruning Binarized Neural Networks: A Dedicated Framework and Globally Weighted Algorithms
arXiv:2608.26233v1 Announce Type: new
Abstract: Extreme compression of deep neural networks, up to full binarization, dramatically reduces memory footprint and arithmetic complexity, facilitating deployment on constrained edge hardware with field-programmable gate arrays (FPGAs) and microcontroller...
Muon with Finite Newton-Schulz: The Smoothing Benefit in Nonsmooth Nonconvex Optimization
arXiv:2608.26288v1 Announce Type: new
Abstract: Muon has emerged as a strong optimizer for the matrix-valued parameters in large language model pretraining, approximately orthogonalizing its momentum with a few Newton-Schulz iterations. Existing theory either replaces this iteration with the exact ...
Algebraic Multigrid Acceleration for Efficient Label Spreading
arXiv:2608.26309v1 Announce Type: new
Abstract: Modern machine learning models rely on large amounts of labeled data. However, manual annotation of large-scale datasets is expensive and time-consuming. Label spreading is a semi-supervised learning technique that addresses this challenge by propagat...
EduRiskX: A Neuro-Symbolic Framework with F-Logic Reasoning for Early Academic Risk Prediction
arXiv:2608.26107v1 Announce Type: new
Abstract: Predicting students' academic risk in online education is crucial for enabling timely interventions that can improve retention and learning outcomes. However, existing models often suffer from limited early detection capability and insufficient interp...
PICasso: An AI-Enabled Design Framework for Autonomous Optimization of Silicon Photonic Devices
arXiv:2608.26113v1 Announce Type: new
Abstract: We present PICasso, an AI-assisted framework for automated synthesis, verification, and optimization of photonic integrated circuits (PICs) from natural-language specifications. PICasso couples a structured NL -> YAML -> GDS generation pipeline with P...
CIFQA: A Deterministic Tool-Grounded Multi-Agent LLM Framework for Financial Query Answering
arXiv:2608.26114v1 Announce Type: new
Abstract: Calculation-intensive financial question answering requires exact reasoning over structured rates, temporal conditions, numerical formulas, and rule-based constraints. Although Large Language Models (LLMs) perform strongly on natural language tasks, t...
LLMs Are Not (Consistently) Bayesian: Quantifying Internal (In)consistencies of LLMs’ Probabilistic Beliefs
Modern AI systems are being deployed in complex domains such as medicine, science, and law, where there is often not a single correct answer given the observed evidence. Such systems must be able to represent and update uncertain beliefs about the world as new evidence arrives to make rational decis...
Agent Seer: Synthesizing Scenarios from Specification Understanding
Evaluating AI agents that use external tools requires realistic test scenarios that capture how practitioners compose tools and iterate across conversation turns. Constructing such scenarios by hand demands deep domain expertise, does not scale across tool ecosystems, and produces static benchmarks ...
Catch up on every session from Agentic AI For Finance Virtual Summit, with sessions from the likes of PayPal, American Express, JP Morgan, Visa, and more.
Dynamic Influence-Weighted Distillation for Single-IMU Activity Recognition
arXiv:2608.24904v1 Announce Type: new
Abstract: Inertial sensors at multiple body locations can improve activity recognition, but requiring every sensor at inference increases the deployment burden. We study whether four synchronized IMUs available during training can improve a student that uses on...
GreenLeaf Law Embed Tiny: A Compact Embedding Model for Legal Domain Retrieval
arXiv:2608.24936v1 Announce Type: new
Abstract: We present GreenLeaf Law Embed Tiny, a 0.6B parameter embedding model for legal domain retrieval. GreenLeaf-Tiny achieves 75.11% on the Massive Legal Embedding Benchmark (MLEB) and 64.38% on MTEB(Law, v1),demonstrating competitive performance among mo...
ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration
arXiv:2608.24938v1 Announce Type: new
Abstract: Mixture-of-Experts (MoE) models scale capacity for strong quality while keeping per-token compute bounded through sparse expert activation. Yet low-latency MoE serving is increasingly challenging, because it spans two inference phases with fundamental...
VLM-based automatic multi-granularity graph representation of building layouts for design informatics
arXiv:2608.24886v1 Announce Type: new
Abstract: Architectural floorplan images encode rich relational knowledge among functional spaces, which underpins design retrieval, knowledge-based reasoning, and BIM enrichment through the building lifecycle. However, it remains challenging to automatically c...