Dynamic Influence-Weighted Distillation for Single-IMU Activity Recognition
arXiv:2608.24904v1 Announce Type: new
Abstract: Inertial sensors at multiple body locations can improve activity recognition, but requiring every sensor at inference increases the deployment burden. We study whether four synchronized IMUs available during training can improve a student that uses on...
ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration
arXiv:2608.24938v1 Announce Type: new
Abstract: Mixture-of-Experts (MoE) models scale capacity for strong quality while keeping per-token compute bounded through sparse expert activation. Yet low-latency MoE serving is increasingly challenging, because it spans two inference phases with fundamental...
TRACE: Transition-Aware Residual Control for Multi-Objective Materials Discovery
arXiv:2608.23631v1 Announce Type: new
Abstract: Multi-objective materials discovery with LLM agents is often limited not only by how many candidates can be proposed, but by how effectively each costly property evaluation informs the next search step. Existing agents mainly store evaluated candidate...
A survey detection channel overrides the pixels in an astronomical foundation model, and biases tomographic mean redshifts
arXiv:2608.23626v1 Announce Type: new
Abstract: Foundation models for astronomy are trained on survey pixels together with the catalogue products derived from those pixels. Those catalogues are incomplete at a measurable rate, and a model trained on both inherits that incompleteness as a systematic...
LLM Agents Perform Controlled Experiments Using Simulation Models
arXiv:2608.23622v1 Announce Type: new
Abstract: Large language models (LLMs) have shown strong capabilities in reasoning, planning, and tool use, but many scientific and engineering tasks require more than plausible text and code generation. They require understanding how a system responds to inter...
RENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation
arXiv:2608.23568v1 Announce Type: new
Abstract: Memory and RAG evaluations often treat the answering model's input as an implementation detail, even though systems may render the same history as a memory entry, summary, typed record, or raw excerpt. We introduce RENDER, a benchmark control that fix...
ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergence
arXiv:2608.23569v1 Announce Type: new
Abstract: State-of-the-art Natural Language to SQL (NL2SQL) models report execution accuracy exceeding 89 percent on established benchmarks such as Spider and BIRD. However, these benchmarks rely on simplified academic schemas and open-source SQL dialects that ...
Renormalization Group Flow Matching for Scalable Local Generative Modeling
arXiv:2608.23696v1 Announce Type: new
Abstract: Despite their remarkable success in modeling complex data, generative models face a fundamental tradeoff. Global approaches can capture full structural coherence but suffer from high computational costs, while local models are efficient but often fail...
From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers
arXiv:2608.23660v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used to provide prior causal knowledge for structural causal discovery, yet whether their direct-edge judgments and confidence can be trusted remains unclear. We systematically evaluate 12 instruction-tune...
Data Predictability Shapes Weibull Weight-Scale Growth in Transformer Training
arXiv:2608.23573v1 Announce Type: new
Abstract: A trained transformer's weight magnitudes can be summarized by a two-parameter Weibull distribution whose shape $k \approx 1.2$ is stable across layers and models, so the scale $\lambda$ carries most training-induced movement. What corpus property set...
Response Renormalization for Critical Deep Equilibrium Models
arXiv:2608.23725v1 Announce Type: new
Abstract: Deep Equilibrium Models (DEQs) compute predictions from a hidden representation unchanged by the model update. Training through this equilibrium uses implicit differentiation and requires solving an adjoint system built from the residual Jacobian. If ...
Equivariant Cellular Sheaves for Molecular Electronic Structure: Bridging Sheaf Cohomology and E(3)-Equivariant Hamiltonian Learning
arXiv:2608.23571v1 Announce Type: new
Abstract: Equivariant message-passing networks are the standard model for molecular property and interatomic-potential prediction, and recent work predicts the electronic Hamiltonian itself in an E(3)-equivariant way. Separately, topological deep learning has e...
AI Learning and Conceptual Transfer in the Game of Hidden Rules
arXiv:2608.21372v1 Announce Type: new
Abstract: This report summarizes the work conducted on the Game of Hidden Rules (GOHR), focusing on reinforcement learning agents trained to infer hidden rules from trial-and-error feedback, representation design, rule difficulty analysis, transfer learning, ge...
Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification
arXiv:2608.20378v1 Announce Type: new
Abstract: Safety alignment in Large Language Models (LLMs) is often superficial, relying on refusal mechanisms that trigger only at the final stages of generation without erasing the foundational knowledge of harmful concepts acquired during pretraining. This s...
Bankruptcy Prediction via Hybrid Resampling and Stacking Ensemble Techniques with Explainable Artificial Intelligence (XAI)-Driven Analysis
arXiv:2608.20343v1 Announce Type: new
Abstract: This study develops and evaluates a bankruptcy prediction framework that integrates consensus-based feature selection, hybrid resampling, stacking ensembles, and explainable artificial intelligence to improve minority-class detection in severely imbal...
From Thermal Preference Prediction to Adaptive Thermal Intervention: A Reinforcement Learning Approach Using Physiological and Environmental Sensing
arXiv:2608.20423v1 Announce Type: new
Abstract: Personalised thermal comfort is essential for occupant wellbeing and for the development of more responsive building-control strategies, yet conventional Heating, Ventilation, and Air Conditioning (HVAC) systems rely on static setpoints and population...
A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Applications
arXiv:2608.20379v1 Announce Type: new
Abstract: Advances in large language models (LLMs) have fueled a wave of research into agency: the ability to reason, plan, and act. This effort has produced agentic frameworks that orchestrate perception, memory, and decision-making around powerful LLM backbon...
PrimeAgentOrchestrator: Memory-Primed Agent Spawning for Personal AI Infrastructure
arXiv:2608.20342v1 Announce Type: new
Abstract: Large language model (LLM) coding agents start each session with an empty context window, discarding accumulated knowledge from prior work. We present PrimeAgentOrchestrator (PAO), a system that spawns new instances of Claude Code -- Anthropic's termi...
SDAD: Spec-Driven Agentic Development for the AI-Native SDLC
arXiv:2608.20341v1 Announce Type: new
Abstract: Frontier coding agents backed by large language models with context windows from hundreds of thousands to millions of tokens are restructuring the Software Development Life Cycle (SDLC). Rich context handling and multi-step reasoning now allow substan...
Towards On-Board Implementation of ML-Based Helicopter Weight Estimator
arXiv:2608.19210v1 Announce Type: new
Abstract: This paper focuses on the implementation of a novel supervised Machine Learning model for estimating helicopter weight during takeoff, utilizing extensive datasets from Airbus's global in-service fleet. The study details a learning assurance process a...
Holtercare-Bench: A Multimodal Benchmark for Evaluating Long-Term Dynamic ECG Analysis
arXiv:2608.19297v1 Announce Type: new
Abstract: While multimodal large language models (MLLMs) excel in medical applications, most of them favor static images or short-term signals. In the critical field of dynamic electrocardiograms (ECG), models struggle with complex temporal reasoning and diagno...
Improved Confidence Estimates for Black-Box Large Language Models
arXiv:2608.19323v1 Announce Type: new
Abstract: Uncertainty quantification (UQ) is essential for the safe deployment of large language models (LLMs). Existing methods, from verbalized confidence to ones requiring multiple generations, are often zero-shot and produce scores quantifying uncertainty w...
Position: Profiling Game Worlds by Transition Complexity
arXiv:2608.18079v1 Announce Type: new
Abstract: Game world modeling (GWM) and reinforcement learning (RL) are often confounded because research papers rarely quantify how difficult the underlying transition prediction problem is at the declared interface (pixels/tokens/latents with finite history)....
Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions
arXiv:2608.18078v1 Announce Type: new
Abstract: This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to exhibit collusive behavior and should be required to obtain behavioral certification before making decisions that affect economic markets. This i...