Hugging Face hack could indicate cultural issues at OpenAI
This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. By now you’ve probably heard about last month’s major AI security incident, in which OpenAI agents escaped their sandbox and hacked into the AI platform Hugging F...
GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models
What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration.
The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery wit...
Your LLM Can Return Perfect JSON and Still Be Wrong
What I learned after thinking more carefully about Structured Outputs on messy, incomplete data
The post Your LLM Can Return Perfect JSON and Still Be Wrong appeared first on Towards Data Science.
Nvidia’s $3.5B MediaTek bet reveals its plan for tackling Big Tech’s AI chip buildout
Nvidia invests $3.5 billion into Taiwanese chipmaker MediaTek. The deal shows how Nvidia plans to stay essential to AI infrastructure as Big Tech begins to build its own AI chips.
A “quantum bath” puts quantum entanglement on autopilot
Physicists have demonstrated a new way to entangle distant quantum bits without the constant measurements and active control normally required. The team created a “quantum bath,” a shared environment filled with correlated microwave photons that automatically pushes separated qubits into an entangle...
Combining LLM Embeddings with Tabular Features in a Unified Scikit-learn Pipeline
In this article, you will learn how to build a unified scikit-learn pipeline that combines text embeddings generated by a lightweight open-source language model with...
AgentOps Is Not MLOps: What Breaks in Your Monitoring Stack When Agents Go to Production
The five MLOps monitoring assumptions agents break, and which inherited signals now pass failed runs as healthy.
The post AgentOps Is Not MLOps: What Breaks in Your Monitoring Stack When Agents Go to Production appeared first on Towards Data Science.
Quantization-Triggered Backdoors in Language Models: Cross-Quantizer Transferability and the Validation--Deployment Gap
arXiv:2608.27512v1 Announce Type: new
Abstract: Post-training quantization is often treated as a semantically neutral optimization for edge deployment of Large Language Models. When a full-precision source checkpoint is evaluated and quantization is applied downstream without equivalent re-evaluati...
arXiv:2608.27513v1 Announce Type: new
Abstract: Softmax attention stores key and value vectors for every preceding token, causing inference memory to grow with sequence length. Recent language models incorporating Gated DeltaNet (GDN) or Kimi Delta Attention (KDA) reduce this cost by replacing the ...
When Muon Meets Task Interference: A Spectral Perspective on Continual Learning and Model Merging
arXiv:2608.27518v1 Announce Type: new
Abstract: Continual learning (CL) and model merging (MM) both aim to obtain a single model that performs well across multiple tasks, challenged respectively by catastrophic forgetting and weight-disentanglement error. In the literature, these difficulties are m...
Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Model
arXiv:2608.27459v1 Announce Type: new
Abstract: In 2011, IBM's Watson was something like a sealed capsule of its era's queryable knowledge. Its DeepQA system defeated the strongest human Jeopardy! champions, but the knowledge that let it do so lived in a curated billion-document corpus running on a...
Rating the Raters: Rasch Measurement Theory for LLM Evaluation
arXiv:2608.27463v1 Announce Type: new
Abstract: LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models' outputs, and raters of human-generated content. Each paradigm can be viewed as a measurement problem, where a latent property of an object is probed w...
Not All Explanations Are Sought: Information-Seeking Psychology for Human-Centered XAI
arXiv:2608.27464v1 Announce Type: new
Abstract: This position paper argues that human-centered explainable AI (HCXAI) should incorporate insights from the psychology of information seeking. Drawing on Sharot and Sunstein's framework of information-seeking motives, we propose that people evaluate wh...
Retrieving Relations, Detecting Fallacies: A RAG Approach to Political Debate Analysis
arXiv:2608.27471v1 Announce Type: new
Abstract: Fallacies are arguments that employ invalid reasoning, making their automatic detection critical in sensitive contexts such as high-stakes political debates, where public opinion is shaped. Spotting a fallacious argument requires contextual knowledge ...
LLM-Augmented Causal Discovery: Probabilistic Fusion of Edge Existence and Orientation
arXiv:2608.27472v1 Announce Type: new
Abstract: Bayesian network structure learning (BNSL) from observational data struggles with orientation identifiability, while large language models (LLMs) offer broad but often unreliable causal knowledge. We propose combining these complementary sources throu...
Lowest-Latency Inference APIs for Voice and Realtime Agents: A Time to First Token TTFT-First Benchmark
Voice agents fail on latency long before they fail on intelligence. Time to first token is the metric most teams use to choose an inference API, and it is the right starting point and the wrong stopping point. This benchmark works through every layer of the voice stack — LLM, speech-to-text, text-to...
Google AI Introduces EnvHarness: A Programmable Layer That Turns Static Agent Environments Into Adaptive Training Worlds
Google Cloud AI Research, with Washington University in St. Louis and UNC Chapel Hill, has released EnvHarness, an Apache-2.0 layer that turns a static agent benchmark into one that adapts to the policy training on it. It wraps a frozen environment through the standard reset()/step() interface, so t...
Top 7 Free AI Automation Courses with Certificates
You don’t need to know anything about AI automation to get started. There are plenty of free courses that can take you from the basics to building your own automations, even if you’re starting from zero. Some are completely beginner-friendly, while others are better once you’re comfortable with the ...
Quick and simple tips to help you write better agent instructions
The post 8 Tips for Writing Effective Agent Instructions appeared first on Towards Data Science.
Musk’s faster path to more gas turbines comes with pollution problem
Elon Musk says a secretive new SpaceX foundry will let him cast his own turbine blades and get gas power online 18 months faster than anyone else — but it's a bet on a fuel source that's already triggering lawsuits and health studies everywhere his (and others') turbines have gone in.