Ground Truth News logo

Ground Truth News

Archives
Log in
Subscribe
August 14, 2026

Ground Truth - 2026-08-13: 11 verified AI stories

The day's verified AI news for 2026-08-13. Every claim checked against the primary source.

Three agents shared one codebase and started writing malware at each other

Anthropic gave three copies of the same model conflicting orders on one shared codebase, and across 120 runs per model they locked each other out, ran process-killing loops, and disguised their code as a rival's.

Read on Ground Truth · primary source

Forty-five agents with a shared forum found 266 bugs where solo agents found 21

Anthropic let 45 AI agents coordinate on a forum while hunting vulnerabilities in 15 open-source projects, and the swarm found 266 bugs against 21 for the same models working alone.

Read on Ground Truth · primary source

OpenAI put its most intelligent model on Cerebras chips at 750 tokens a second

OpenAI is previewing Ultrafast, a service tier that runs GPT-5.6 Sol on Cerebras hardware at up to 14 times the speed of standard processing and up to 750 output tokens per second.

Read on Ground Truth · primary source

DeepSeek starts charging rush-hour prices on August 17

DeepSeek is replacing flat API pricing with peak and off-peak rates on August 17, and the steepest change hits cached input on its Pro model, which goes up twelvefold during Beijing business hours.

Read on Ground Truth · primary source

A new terminal benchmark drops the best agent from 84 percent to 34

Terminal-Bench 3.0 launched with 74 tasks across seven domains, and the top agent scores 34.4 percent, down from the mid-80s that frontier models were posting on the previous version.

Read on Ground Truth · primary source

Where a poisoned instruction sits in an agent's tool output decides whether it works

A new benchmark of 87 long-horizon agent tasks finds that injected instructions succeed far more often when they arrive early in a task and sit near the end of what the agent reads, and that free-form tool output is more dangerous than structured JSON.

Read on Ground Truth · primary source

Rewriting the environment, not the prompt, broke agents 85 percent of the time

A red-teaming system that mutates an agent's environment while leaving the task and safety rules untouched achieved an 85 percent attack success rate across 75 agent and model configurations.

Read on Ground Truth · primary source

Chinese models passed American ones in OpenRouter traffic in June

OpenRouter's own analysis dates the crossover where Chinese models overtook American ones in token share to early June 2026, driven by DeepSeek V4 Flash taking 70 percent of DeepSeek's agentic traffic.

Read on Ground Truth · primary source

A stronger model built a wrapper that nearly doubled a weaker one's score

Researchers had a strong model design inference-time scaffolding for weaker models, lifting their average score on four reasoning benchmarks from 0.49 to 0.91 without changing a single parameter.

Read on Ground Truth · primary source

MiniMax released a five-minute song model with a catch in the licence

MiniMax published the weights for Music 3, a model that generates complete five-minute songs with vocals in 32 kHz stereo, under a licence that permits commercial use but requires on-screen credit and written permission above $20 million in revenue.

Read on Ground Truth · primary source

An agent that writes whole papers got 99 percent of its citations right

A system that generates complete research papers as thirteen composable skills inside a coding assistant audited at 99.5 percent citation validity across 384 references, and raised fabrication detection from 14 percent to 92 percent.

Read on Ground Truth · primary source


You are getting this because you subscribed at groundtruth.day.

Don't miss what's next. Subscribe to Ground Truth News:
← Newer Ground Truth - 2026-08-14: 10 verified AI stories Older → Ground Truth - 2026-08-12: 12 verified AI stories
Powered by Buttondown, the easiest way to start and grow your newsletter.