Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy
arXiv:2607.22554v1 Announce Type: new
Abstract: Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge when the same question is phrased in different but equivalent ways. In this work, we study how model answers change...
DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs
arXiv:2607.22555v1 Announce Type: new
Abstract: Medical diagnosis is a multi-stage process: extract facts, consult knowledge, generate a differential analysis, and select the best diagnosis with explanations. Frontier LLMs are strong generalists, but single-shot prompting often yields brittle diagn...
Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers
Siri Expressive Voices synthesize rich, configurable speech in real time and entirely on device, powered by AFM 3 Core Advanced, Apple’s most powerful on-device foundation model. This work presents the memory-efficient audio synthesis architecture behind that capability: a detokenizer that converts ...
Kimi AI and kvcache-ai Open Sources ‘AgentENV’: A Distributed System that Powers Agentic Reinforcement Learning (RL) Training for Kimi K3
Moonshot AI's Kimi team and kvcache-ai open-sourced AgentENV (AENV) under MIT, as part of Kimi K3 Open Day. It runs agent sandboxes as Firecracker microVMs with millisecond snapshot, resume, and 16-way fork, behind an E2B-compatible API.
The post Kimi AI and kvcache-ai Open Sources ‘AgentENV’: A Dis...
Designing Skill-Driven Financial Analysis Agents with Claude, Python, MCP Connectors, and Automated Deliverables
In this tutorial, we build an advanced workflow around Anthropic’s financial-services repository and reproduce its skill-driven architecture in pure Python. We begin by installing the required libraries, cloning the repository, and programmatically mapping its agents, vertical plugins, partner integ...
OpenAI called the Hugging Face attack unprecedented. But we’ve been here before.
This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Reading OpenAI’s account last week of how some of its models broke their containment and hacked into the computer systems of Hugging Face, another AI company, was...
OpenAI’s Hugging Face breach has reignited the debate over alignment and control
OpenAI's Hugging Face breach has reignited debate over AI alignment and control, exposing competing views on whether increasingly capable AI should be better aligned, better contained, or both.
Perplexity Releases pplx, a Single-Binary CLI That Puts Its Search API in the Terminal for Coding Agents
Perplexity has released pplx, an official command line client for its Search API. The tool exposes two commands — pplx search web and pplx content fetch — and returns exactly one JSON object on stdout. It ships as a checksum-verified single binary for macOS arm64 and Linux, alongside an Agent Skill ...
Google’s AI search is rapidly becoming the default, new data shows
Google’s AI Overviews now appear in 43% of searches, underscoring how quickly AI-generated answers are becoming the default way people discover information online.
Lightbits Labs Strengthens Enterprise Linux With Ubuntu Certification
Native Ubuntu Support Simplifies Deployment of High-Performance Software-Defined Block Storage for Private Clouds Built in Kubernetes and OpenStack Environments Lightbits Labs®, inventor of the NVMe® over TCP storage protocol and Inferra™, the first KV cache prefetch engine for AI acceleration, toda...
ARC Cuts Documentation Time by 18.5% and Optimizes Coding Accuracy Using Suki
One of Texas’s largest multispecialty groups achieves 97% clinician engagement rate — far exceeding industry benchmarks — as ambient clinical intelligence scales across 40 locations Austin Regional Clinic (ARC), one of the largest multispecialty medical groups in Central Texas, serving more than 700...
Coalesce Capital Announces Growth Investment in Workstreet
Coalesce Capital (“Coalesce”), a private equity firm focused on investing in next-generation technology-enabled services companies, today announced a strategic growth investment in Workstreet, (“the Company”) a leading provider of AI-native compliance and cybersecurity solutions to companies in regu...
92% of Healthcare Leaders Demand Clinical Expertise to Trust AI
New national survey finds adoption stalls for structural reasons, even as organizations see value Carta Healthcare, the leader in enterprise clinical data management, today released findings from a national survey of U.S. healthcare leaders showing that AI is proving its worth but failing to expand,...
Claude Opus 5: Near-Frontier Intelligence, On a Dial
Anthropic has released Claude Opus 5. The fourth model in two months, if you are keeping count. Most people are not. This one matters more than the count suggests. Opus is the workhorse tier, the model that does the actual paid work, and it just got a step change rather than a bump. Anthropic’s own ...
Imagine a healthcare system made up of multiple AI agents: one that manages symptom assessment, another scheduling, a third insurance, and a fourth pharmacy. Each is an expert in its domain. But they all have their own distinct knowledge and objectives. Today they can exchange data, but they are not...