Do Models Fake Alignment Without Clear Consequences?
arXiv:2607.24758v1 Announce Type: new
Abstract: Large language models are capable of recognizing evaluation contexts and altering their behavior to reflect evaluator expectations rather than typical deployment behaviors, a phenomenon known as alignment faking. The reasons why models fake alignment ...
Beyond Memory: A Templated Substrate for Heterogeneous Collaborative Knowledge Work with LLM Agents
arXiv:2607.24759v1 Announce Type: new
Abstract: Research projects, educational efforts, and adjacent knowledge work accumulate findings, decisions, and reasoning that future collaborators rarely recover. The parts most useful to that work, including dead ends and walked-back claims, are routinely e...
CORVUS: Context Optimization and Reduction Via Underlying Synchronization for LLM Coding Agents
arXiv:2607.22711v1 Announce Type: new
Abstract: LLM coding agents operate by constructing trajectories that accumulate reasoning, tool calls, and results to enable multi-step decision-making. However, the conventional append-only trajectory architecture found in practice tightly couples file-read a...
CausalGate: Causal Importance Distillation for Transformer Module Pruning
arXiv:2607.22720v1 Announce Type: new
Abstract: Existing adaptive inference methods for Large Language Models rely on observational heuristics, such as hidden-state similarity or activation magnitudes, to drop redundant modules. However, these correlation-based metrics often fail to capture subtle,...
QFedPolyp: A Communication- and Inference-Efficient Federated Learning Framework for Polyp Segmentation
arXiv:2607.22743v1 Announce Type: new
Abstract: Background and Objective: Automatic polyp segmentation supports computer-aided diagnosis and early colorectal cancer detec- tion. Centralized deep learning requires hospitals to share sensitive medical data, while federated learning preserves privacy ...
Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B
arXiv:2607.22545v1 Announce Type: new
Abstract: Deploying large language models in financial-services and agentic settings requires safety classifiers that simultaneously handle prompt injection, regulatory compliance, and general harm, a combination no existing open guardrail addresses in a single...
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks
arXiv:2607.22724v1 Announce Type: new
Abstract: Group-based policy optimization has been increasingly used to train large language model (LLM) agents from sparse outcome rewards by comparing trajectories or steps within a group. However, on difficult long-horizon tasks, this comparison can suffer f...
QFoldAgent: An Autonomous Quantum Optimization Multi-Agent System for Protein Structure Prediction
arXiv:2607.22549v1 Announce Type: new
Abstract: Hybrid quantum-classical protein structure prediction depends strongly on Hamiltonian penalty weights, yet existing lattice-based workflows typically fix these coefficients by hand and evaluate only very short fragments in simulation. We present QFold...
SeT-Diff: Towards Semantic Foundation Models for HPC Telemetry and Time-Series
arXiv:2607.22548v1 Announce Type: new
Abstract: Data centers and their compute nodes require accurate and flexible digital twins capable of modeling the complex interplay of workloads, environmental parameters, and physical metrics. Current machine learning approaches for HPC and its telemetry typi...
Concept-based Visual Counterfactual Explanations with Diffusion Models
arXiv:2607.22544v1 Announce Type: new
Abstract: Visual counterfactual explanations aim to answer "what minimal change to this image would flip the model's prediction?", and are increasingly important as vision models are deployed in safety-critical domains (e.g., medicine). Existing diffusion-based...
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy
arXiv:2607.22554v1 Announce Type: new
Abstract: Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge when the same question is phrased in different but equivalent ways. In this work, we study how model answers change...
DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs
arXiv:2607.22555v1 Announce Type: new
Abstract: Medical diagnosis is a multi-stage process: extract facts, consult knowledge, generate a differential analysis, and select the best diagnosis with explanations. Frontier LLMs are strong generalists, but single-shot prompting often yields brittle diagn...
arXiv:2607.21633v1 Announce Type: new
Abstract: Logic Gate Networks (LGNs) implement computation through compositions of Boolean operations, yet unlike classical Boolean circuits, existing LGNs do not reliably benefit from increased depth. We identify two distinct causes: optimization collapse in d...
Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions
arXiv:2607.21635v1 Announce Type: new
Abstract: Personal agents maintain memories, learned skills, tool configurations, and policy state that evolve with each user. Existing agent benchmarks often evaluate these capabilities in isolation: tool benchmarks test invocation under fixed APIs, memory ben...
Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models
arXiv:2607.21636v1 Announce Type: new
Abstract: Synthetic tabular data is valued for preserving not only each column's marginal distribution but the dependencies between columns -- structure that carries much of the discriminative signal for minority classes in imbalanced domains such as fraud and ...
Transferable Latency Prediction for Fast LLM Screening on Heterogeneous Edge Devices
arXiv:2607.21602v1 Announce Type: new
Abstract: Accurate latency prediction is critical for deploying large language models (LLMs) on heterogeneous edge devices, where inference latency is affected by model architecture, prompt behavior, runtime backend, hardware utilization, dynamic voltage and fr...
Securing Multimodal AI through Internal Information Decomposition
arXiv:2607.21600v1 Announce Type: new
Abstract: Multimodal large language models introduce attack surfaces absent in unimodal systems: adversaries can distribute malicious intent across modalities to evade unimodal safeguards. This motivates using cross-modal consistency as a detection signal rathe...
FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills
arXiv:2607.21596v1 Announce Type: new
Abstract: Large language model agents increasingly solve complex tasks by constructing inference-time workflows that combine reasoning, tool use, and code execution. While such workflows enable flexible problem solving, the useful procedures discovered during e...
Risk Is Not the Target: A Monotonic Framework for Evaluating Wildfire Operational Risk Signals
arXiv:2607.21597v1 Announce Type: new
Abstract: Evaluating wildfire risk systems using standard machine-learning metrics such as F1-score or IoU is fundamentally flawed: these metrics assess event prediction accuracy, not the operational coherence of a continuous risk signal. This work proposes a n...
PhantomFill: When the Form Demands an Answer, Language Models Invent One
arXiv:2607.20492v1 Announce Type: new
Abstract: Language models in production do not write prose. They fill forms: JSON fields, function arguments, extraction templates. We show that the form itself causes hallucination.
We ask thirteen models the same question about the same input and change onl...
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators
arXiv:2607.20465v1 Announce Type: new
Abstract: The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how well LLMs, agents, and data-centric workflows actually prepare training data end to end. We view LLM...
Scaling Closed-Loop Feature Channel Configuration with LLMs
arXiv:2607.20516v1 Announce Type: new
Abstract: Promising initial results in closed-loop large-language-model-based channel-configuration search demonstrated that neural-network widths can be optimized directly through executable code generation and accuracy feedback. However, those results were ob...
AINTMA: Agentic AI Architecture for Autonomous Test Management with Generative Intelligence, Secure Cloud Communication and Adaptive Quality Analytics
arXiv:2607.20452v1 Announce Type: new
Abstract: Modern software quality assurance demands intelligent, autonomous systems capable of adaptive decision-making across distributed cloud environments. This paper presents AINTMA (Agentic Intelligent Test Management Architecture), a multi-agent agentic A...
Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts
arXiv:2607.20462v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly integrated into clinical workflows, stressing the need for reliable traceability of model-generated output with watermarking. Yet, most watermarks are evaluated on general-purpose benchmarks, leaving domai...