arXiv:2605.04050v1 Announce Type: new
Abstract: We introduce Lossless Context Management (LCM), a deterministic architecture for LLM memory that outperforms Claude Code on long-context tasks. When benchmarked using Opus 4.6, our LCM-augmented coding agent, Volt, achieves higher scores than Claude C...
Actionable Real-Time Modeling of Surgical Team Dynamics via Time-Expanded Interaction Graphs
arXiv:2605.04169v1 Announce Type: new
Abstract: Surgical team performance arises from complex interactions between technical execution and non-technical skills, including communication and coordination dynamics. However, current surgical AI systems predominantly model visual workflow signals, lacki...
Pro$^2$Assist: Continuous Step-Aware Proactive Assistance with Multimodal Egocentric Perception for Long-Horizon Procedural Tasks
arXiv:2605.04227v1 Announce Type: new
Abstract: Procedural tasks with multiple ordered steps are ubiquitous in daily life. Recent advances in multimodal large language models (MLLMs) have enabled personal assistants that support daily activities. However, existing systems primarily provide reactive...
Jensen Huang called it "the ChatGPT moment for robotics." Deloitte says 80% of businesses plan to use physical AI within two years. Here is what you actually need to know, and do, to prepare…
StateSMix: Online Lossless Compression via Mamba State Space Models and Sparse N-gram Context Mixing
arXiv:2605.02904v1 Announce Type: new
Abstract: We present StateSMix, a fully self-contained lossless compressor that couples an online-trained Mamba-style State Space Model (SSM) with sparse n-gram context mixing and arithmetic coding. The model is initialised from scratch and trained token-by-tok...
eOptShrinkQ: Near-Lossless KV Cache Compression Through Optimal Spectral Denoising and Quantization
arXiv:2605.02905v1 Announce Type: new
Abstract: We show that the key-value (KV) cache in transformer attention heads admits a natural decomposition into a low-rank \emph{shared context} component and a full-rank \emph{per-token} residual, well described by the spiked random matrix model. This obser...
An End-to-End Framework for Building Large Language Models for Software Operations
arXiv:2605.02906v1 Announce Type: new
Abstract: In the field of software operations, Large Language Models (LLMs) have attracted increasing attention. However, existing research has not yet achieved efficient and effective end-to-end intelligent operations due to low-quality data, fragmented knowle...
Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR
arXiv:2605.02909v1 Announce Type: new
Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become a powerful approach for improving the reasoning capabilities of large language models (LLMs). While RLVR is designed for tasks with verifiable ground-truth answers, real-world verifiers ...
CreativityBench: Evaluating Agent Creative Reasoning via Affordance-Based Tool Repurposing
arXiv:2605.02910v2 Announce Type: new
Abstract: Recent advances in large language models have led to strong performance on reasoning and environment-interaction tasks, yet their ability for creative problem-solving remains underexplored. We study this capability through the lens of creative tool us...
Stable Agentic Control: Tool-Mediated LLM Architecture for Autonomous Cyber Defense
arXiv:2605.03034v1 Announce Type: new
Abstract: Agentic systems involved in high-stake decision-making under adversarial pressure need formal guarantees not offered by existing approaches. Motivated by the operational needs of security operations centers (SOCs) that must configure endpoint detectio...
Programmatic Context Augmentation for LLM-based Symbolic Regression
arXiv:2605.03101v1 Announce Type: new
Abstract: Symbolic regression (SR), the task of discovering mathematical expressions that best describe a given dataset, remains a fundamental challenge in scientific discovery. Traditional approaches, primarily based on genetic algorithms and related evolution...
AI lets chemists design molecules by simply describing them
Creating complex molecules usually requires years of experience and countless decisions, but a new AI system is changing that. Synthegy lets chemists guide synthesis and reaction planning using simple language, while powerful algorithms generate and evaluate possible solutions. The AI doesn’t just c...
From Where Things Are to What They’re For: Benchmarking Spatial–Functional Intelligence for Multimodal LLMs
True spatial intelligence for multimodal agents transcends low-level geometric perception, evolving from knowing where things are to understanding what they are for. While existing benchmarks, such as VSI-Bench, effectively evaluate this foundational geometric stage, they fall short of probing the h...
Normalizing Flows (NFs) are a classical family of likelihood-based methods that have received revived attention. Recent efforts such as TARFlow have shown that NFs are capable of achieving promising performance on image modeling tasks, making them viable alternatives to other methods such as diffusi...
Microsoft at NSDI 2026: Advances in large-scale networked systems
Microsoft researchers share advances in building and operating large-scale distributed systems, spanning datacenters, networking, and the growing intersection with AI during NSDI ’26.
The post Microsoft at NSDI 2026: Advances in large-scale networked systems appeared first on Microsoft Research.
Agentopic: A Generative AI Agent Workflow for Explainable Topic Modeling
arXiv:2605.00833v1 Announce Type: new
Abstract: Agentopic is a novel agent-based workflow for explainable topic modeling that leverages the reasoning capabilities of Large Language Models (LLMs). Existing topic modeling approaches such as Latent Dirichlet Allocation (LDA) and BERTopic often lack tr...
Polynomial-Time Optimal Group Selection via the Double-Commutator Eigenvalue Problem
arXiv:2605.00834v1 Announce Type: new
Abstract: The algebraic diversity framework replaces temporal averaging over multiple observations with algebraic group action on a single observation for second-order statistical estimation. The central open problem in this framework is $\textit{group selectio...
Fast Log-Domain Sinkhorn Optimal Transport with Warp-Level GPU Reductions
arXiv:2605.00837v1 Announce Type: new
Abstract: Entropic regularized optimal transport (OT) via the Sinkhorn algorithm has become a fundamental tool in machine learning, yet existing implementations either suffer from numerical instability for small regularization parameters or incur significant ov...
2026 Roadmap on Artificial Intelligence and Machine Learning for Smart Manufacturing
arXiv:2605.00839v1 Announce Type: new
Abstract: The evolution of artificial intelligence (AI) and machine learning (ML) is reshaping smart manufacturing by providing new capabilities for efficiency, adaptability, and autonomy across industrial value chains. However, the deployment of AI and ML in i...
AI Agents for Sustainable SMEs: A Green ESG Assessment Framework
arXiv:2605.00841v1 Announce Type: new
Abstract: This study presents a novel, AI-driven framework for assessing Environmental, Social, and Governance (ESG) performance in European small and medium-sized enterprises (SMEs). An initial phase established expert-validated ESG baseline scores from a subs...
Understanding Emergent Misalignment via Feature Superposition Geometry
arXiv:2605.00842v1 Announce Type: new
Abstract: Emergent misalignment, where fine-tuning on narrow, non-harmful tasks induces harmful behaviors, poses a key challenge for AI safety in LLMs. Despite growing empirical evidence, its underlying mechanism remains unclear. To uncover the reason behind th...