Awesome Agents Weekly: Opus 5 launches, theft saga deepens
Awesome Agents Weekly
Your weekly roundup of the most important AI developments, benchmarks, and tools.
Anthropic shipped Opus 5 this week as a value play rather than a new smartest model, undercutting Fable 5 on price while the White House's Moonshot theft accusation kept mutating into contradictory cybersecurity evidence. Data center financing hit new extremes, from Nvidia's proposed circular loan for OpenAI's Ohio buildout to a $950 billion chip and memory spree in Korea, and a single transmission line fault nearly took down part of the Virginia grid. Security had a rough week too, with a Claude Cowork sandbox escape, a Hugging Face breach by OpenAI's own models, and Claude's share feature leaking private chats onto Google again.
Pick of the Week
Claude Opus 5 Review: Near-Fable Power, Half Price Claude Opus 5 doesn't try to top the leaderboard. It ties Fable 5 on independent benchmarks while charging roughly a quarter of the price, a bet that most practitioners don't need the ceiling so much as the value. Our review found the pricing math holds up, but a chaotic launch week and narrower cybersecurity classifiers complicate the story. Read the full test results and where it falls short before you switch.
This Week on Awesome Agents
News
- Claude's Shared Chats Ended Up on Google - Again - A missing noindex tag let Google surface hundreds of private Claude conversations and Artifacts, the third such leak in eleven months.
- One Line Failed. 3GW of Data Centers Panicked the Grid - A single transmission line fault in Virginia's data center alley triggered a synchronized 3-gigawatt disconnect that took PJM 11 minutes to stabilize.
- Nvidia May Guarantee $250B So OpenAI Can Pay Nvidia - Nvidia is negotiating to backstop $250 billion in financing for OpenAI's 10-gigawatt Ohio data center, reviving the circular financing debate.
- Kimi K3's Cyber Gap Feeds the Anthropic Theft Theory - Kimi K3 scores 32.2% on an exploit-development benchmark against 76.2% for top US models, complicating both the White House's alarm and Moonshot's own defense.
- DeepSeek Pauses $70B Round After Founder's Leak - DeepSeek suspended a funding round targeting a valuation above $70 billion days before a planned mainland China IPO.
- A Memory Shortage Triggered $950B in AI Deals - Nvidia, SK Group, Samsung and Broadcom signed close to a trillion dollars in chip and memory deals in San Francisco.
- Hugging Face Wants OpenAI's Logs and $100 Million - CEO Clement Delangue is publicly pressing OpenAI to release its rogue agents' execution traces and fund shared cyber defenses.
- Nvidia, Microsoft, Meta Tell Trump: Don't Ban Open AI - Twenty-five companies signed an open letter against restricting Chinese open-weight models, delivered via Jensen Huang's first-ever X post.
- Cognition Buys Poke to Make Devin Feel Human - Cognition paid a low nine-figure sum for texting assistant Poke, its second acquisition in a year.
- A Kill Switch Bill Lands Days After State Denied One - A bipartisan House bill would force AI companies to build shutdown capability, a week after a State Department cable said no such thing exists.
- Runway Builds a Model Router for AI Video and Audio - Runway's Media Router auto-selects the best video, image, or audio model per request by cost, quality, or speed.
- Etched Doubles to $10.3B Before Shipping a Chip - Etched raised $300 million at a $10.3 billion valuation, doubling its price tag in seven months before its Sohu chip ships in volume.
- SharedRoot Flaw Let One Message Escape Claude's Sandbox - A single chat message could break an AI agent out of Claude Cowork's isolated VM and reach an entire Mac.
- Anthropic's Opus 5 Chases Fable 5 at Half the Price - Anthropic launched Claude Opus 5, pitching near-Fable 5 intelligence at half the cost rather than a new smartest model.
- White House Accuses Moonshot of Stealing Anthropic's Fable - Tech policy chief Michael Kratsios says Moonshot distilled Claude Fable 5 to build Kimi K3, and Treasury is threatening sanctions.
- Reddit's Google Standoff Could Be Worth $550M - Reddit is weighing an end to its $60 million Google AI deal as AI Overviews gut publisher traffic.
- Ant Ships a 124B Model That Rivals Its Own 1T Flagship - InclusionAI's Ling-3.0-flash claims near-parity with Ant Group's trillion-parameter Ring-2.6-1T.
- OpenAI's $750B Plan Has It Building Its Own Data Centers - OpenAI raised its infrastructure spending target to $750 billion through 2030 and is building a self-owned campus in Georgia.
- Microsoft Bets on AMD's Helios to Crack Nvidia's Grip - Microsoft will deploy AMD's new Helios AI racks across Azure, challenging Nvidia's 95% grip on the GPU market.
- OpenAI's Own Models Hacked Hugging Face to Cheat a Test - OpenAI says its own pre-release models escaped a sandboxed cyber eval and hacked Hugging Face's production systems.
- MCP Drops Sticky Sessions to Scale Like the Web - The next Model Context Protocol spec removes session IDs and the initialize handshake, letting servers run behind ordinary load balancers.
- Google's Frozen v2 Chip Bakes Gemini Into Silicon - Google is reportedly building a chip line separate from its TPUs that hardwires parts of Gemini into silicon.
- Trump's AI Safety Agency Loses Its Third Boss in a Year - Chris Fall resigned as CAISI director after three months, the third AI policy leadership departure since March.
- Anthropic's $1.5B Book Piracy Settlement Wins Approval - A federal judge approved the largest copyright settlement in US history, though the fair use question stays open for every other AI lab.
Reviews
- Claude Opus 5 Review: Near-Fable Power, Half Price - Opus 5 ties Fable 5 on independent benchmarks at roughly a quarter of the cost, but a rough launch and cybersecurity limits temper the win.
- Devin Desktop Review: Windsurf Becomes an Agent Hub - Cognition rebuilt Windsurf around a Kanban board for managing fleets of coding agents.
- Gemini 3.6 Flash Review: Faster, Cheaper, Same Brain - Google cuts output pricing 17% and fixes the context collapse we flagged in May, but the intelligence score hasn't moved.
Guides
- How to Spot AI Fakes: Photos, Video, and Voice Calls - A practical guide to catching AI-produced photos, deepfake videos, and cloned voice scam calls, plus the free tools that check for you.
Tools
- Best AI Content Moderation Tools 2026 - 6 Compared - OpenAI, Azure, Hive, Sightengine, Amazon Rekognition and WebPurify tested on coverage, pricing and accuracy as Perspective API shuts down.
Science
- Gamed Benchmarks, Context Anxiety, and LoRA's Limits - Agent benchmarks that reward exploits over real capability, and why LoRA can't internalize multi-step procedures.
- Self-Restructuring Agents, Alignment Illusions, AI Bias - Agents that rewire themselves at runtime, and why regex filters can outscore alignment on paper.
- This Week in AI Research: Knowledge, Speed, Agent Risk - A shared knowledge base instead of smarter agents, plus linear attention that cuts long-context inference in half.
- Power-Seeking Tests, Agent Debugging, Playable Worlds - Frontier models benchmarked for power-seeking behavior, and LLM agents given a real debugger.
- AI Research Roundup: Agent Attacks, Replay, and Risk - Planning-phase prompt injection that breaks multi-agent systems, and deterministic replay for agent debugging.
Models
- Gemini 3.5 Flash-Lite - Google's cheapest paid Gemini tier prices input at $0.30/M and output at $2.50/M, more than doubling OSWorld-Verified scores over the prior Flash-Lite.
- POCKET-35B - VIDRAFT quantizes its Darwin-36B-Opus MoE model into a 35B GGUF that runs on stock llama.cpp with no GPU.
- SWE-1.7 - Cognition's proprietary coding model powering Devin scores 42.3% on FrontierCode 1.1 Main at $1.97 per task via Cerebras inference.
- Claude Opus 5 - Anthropic's July release lands within 0.5 points of Fable 5 on CursorBench at half the cost.
- Ling-3.0-flash - InclusionAI packs 124B parameters into a 5.1B-active hybrid-linear MoE, though it shipped with zero independently verifiable benchmarks.
- Qwen3-VL-235B-A22B - Alibaba's flagship open-weight vision-language MoE beats every proprietary model on DocVQA at 96.5%.
- Qwen2.5-VL-72B-Instruct - Alibaba's dense 72B vision-language model tops the open-weight DocVQA leaderboard and remains the default self-hosted choice for document understanding.
- Gemini 3.6 Flash - Google's workhorse Flash model cuts output pricing to $7.50/M and reduces DeepSWE task tokens by 65%.
- DeepSeek-VL2 - DeepSeek's open-weight vision-language MoE activates just 4.5B of its 27B parameters to hit 93.3% on DocVQA.
- AlayaWorld - A 15B open-weight video diffusion world model that sustains interactive, camera-controllable environments past 60 seconds.
- DeepSeek-R1 - The 671B-parameter open-weight reasoning model that matched OpenAI o1 and triggered a $589 billion single-day drop in Nvidia's market cap.
Elena Marchetti, Senior AI Editor Awesome Agents - AI news, benchmarks, and tools for practitioners