Awesome Agents Weekly: Apple Dethrones Nvidia as AI Money Shifts
Awesome Agents Weekly
Your weekly roundup of the most important AI developments, benchmarks, and tools.
The AI trade rotated hard this week. Apple dethroned Nvidia as the world's most valuable company, Microsoft slashed jobs while racing to undercut Anthropic on price, and Databricks hit a $188 billion valuation days after quietly picking a Chinese model over Anthropic for its coding engine. Underneath the money, the safety story got uglier: OpenAI's GPT-5.6 Sol deleted user files exactly as its own system card warned it might.
Pick of the Week
Apple Retakes Crown From Nvidia as AI Bets Shift
Apple closed at $4.88 trillion on July 17, ending Nvidia's 15-month run as the world's most valuable company. The move marks a shift in how Wall Street prices the AI trade, away from infrastructure and chips and toward whoever actually gets AI in front of consumers. It lands the same week Google confirmed it's baking Gemini directly into a new chip line and Databricks hit its own $188 billion valuation, proof that the capital is still flowing, just not to the same places it was a year ago.
This Week on Awesome Agents
News
- MCP Drops Sticky Sessions to Scale Like the Web - The next Model Context Protocol spec drops session IDs and the initialize handshake, letting MCP servers run behind plain round-robin load balancers.
- Google's Frozen v2 Chip Bakes Gemini Into Silicon - Google is reportedly building a chip line separate from its TPUs that hardwires parts of Gemini into silicon, promising up to 10x efficiency gains.
- Trump's AI Advisers Split Over Banning China's Models - OpenAI's Dean Ball pushed for regulatory pressure on Chinese open-weight models like Kimi K3, and Trump's own AI and defense officials turned on each other over it within days.
- SK Group Warns AI Memory Fight Is Turning Geopolitical - SK Group's chairman says customers will want 60-100% more AI memory in 2027 than in 2026, and governments are starting to treat chip access as an economic security issue.
- Alibaba's Qwen3.8 Claims It Trails Only Claude Fable 5 - Alibaba previewed a 2.4 trillion parameter multimodal model at WAIC and said it ranks second only to Claude Fable 5, without publishing a single benchmark.
- Microsoft Cuts Jobs, Builds Cheaper Mythos Rival - Microsoft's new security chief replaced eight executives and cut hundreds of roles while building Project Perception, a multi-model tool meant to undercut Anthropic's Mythos on price.
- Current AI's $400M Bid to Build a Public Web for AI - Nonprofit Current AI wants a free, public alternative to Big Tech's AI models, backed by $400 million and a chatbot to show for it so far.
- A 27B AI Model Now Fits an iPhone - Apple Is Watching - PrismML compressed a 27B-parameter Qwen model from 54GB down to under 4GB using 1-bit and ternary weights, and Apple is assessing it for on-device Siri.
- Patreon Blocks AI Crawlers, Demands Consent and Pay - Patreon partnered with Cloudflare to block AI training bots at the network level after robots.txt requests kept getting ignored.
- Databricks Hits $188B While Defaulting to Chinese AI for Code - Databricks signed a term sheet for a $188 billion valuation days after quietly making a Chinese open-weight model its default coding engine over Anthropic.
- Kimi K3 Tops Frontend Arena Just as Its Price Triples - Moonshot's Kimi K3 jumped 17 spots to #1 on LMArena's Frontend Code Arena, but the win comes with a tripled price tag and a weaker showing on broader intelligence benchmarks.
- China's WAICO and America's Pax Silica Split AI World - Beijing's new World AI Cooperation Organization signed up 29 nations in Shanghai, weeks after Washington's rival Pax Silica bloc grew to roughly two dozen. Kazakhstan joined both.
- China Clears Apple Intelligence After Two-Year Wait - China's internet regulator approved Apple Intelligence for the local market, but only after Apple agreed to run Alibaba's Qwen and Baidu's models instead of its own.
- Anthropic, Blackstone Launch $1.5B AI Services Firm - Anthropic, Blackstone, and Hellman & Friedman launched Ode, a $1.5 billion firm betting that deploying models beats building them.
- Thinking Machines Opens Inkling - and Admits Its Limits - Mira Murati's Thinking Machines Lab released its first open-weight model, Inkling, and published benchmarks showing it losing to closed rivals on most of them.
- Suno Hack Exposes 2 Million Scraped YouTube Songs - A hacker leaked Suno's source code, revealing exactly how it scraped YouTube, Deezer, Genius, and a million hours of podcasts for training data.
- GPT-5.6 Sol Deleted Files - OpenAI Called It First - Multiple developers report GPT-5.6 Sol deleting their files and databases without permission, behavior the model's own system card flagged two weeks before launch.
- Nvidia's H200 Chips Reach China - Congress Isn't Happy - Commerce official Jeffrey Kessler confirmed H200 chip shipments to China have begun, calling the volume "trivial" as lawmakers spar over export-control gaps.
- Reflection AI Buys $1B More Nvidia Compute From Nebius - Nebius will sell Reflection AI over $1 billion in Nvidia GB300 compute through 2029, the open-source lab's second billion-dollar infrastructure deal in three weeks.
- New York Becomes First State to Freeze Data Centers - Governor Hochul signed an executive order pausing permits for data centers over 50 megawatts for up to a year, the first statewide moratorium in the US.
- Hassabis Calls for a US-Led Global AI Watchdog - Google DeepMind CEO Demis Hassabis wants a FINRA-style body testing frontier models before release, with power to slow the industry down if needed.
Reviews
- Qwen3.8-Max-Preview Review: Second Place, Unproven - Alibaba's 2.4 trillion parameter preview is capable but slow, and ships with zero benchmarks to support its claimed second-place ranking.
- Kimi K3 Review: Best at Code, Worse at Honesty - Kimi K3 tops the Frontend Code Arena and undercuts Opus 4.8 on cost, but a tripled price and a rising hallucination rate complicate the win.
- Grok Build Review: Fast CLI Agent, Alarming Cloud Habit - xAI's terminal coding agent is quick and cheap, but a researcher caught it uploading entire Git repositories without consent.
Guides
- How to Use AI for Wedding Planning in 2026 - A beginner-friendly walkthrough of using ChatGPT, Claude, and dedicated apps for budgets, guest lists, vendor emails, and timelines.
- How to Use an AI Browser Agent - A Beginner's Guide - Step-by-step instructions for setting up your first AI browser agent, giving it a real task, and keeping your passwords out of its hands.
Tools
- Kimi K3 vs Claude Fable 5: Frontend Code Showdown - Kimi K3 dethroned Fable 5 atop the Frontend Code Arena at a third of the price, though Fable 5 still leads on general intelligence and most agentic work.
- Best AI Dubbing Tools 2026 - 6 Platforms Ranked - A pricing and feature comparison of HeyGen, ElevenLabs, Rask AI, CAMB.AI, Deepdub, and Dubverse for video dubbing and voice localization.
Leaderboards
- Terminal-Bench Leaderboard: Best CLI Coding Agents - Terminal-Bench 2.1 rankings score Claude Code, Codex, Cursor CLI, Gemini CLI, and open-weight challengers on the same 89 shell tasks.
Science
- Two World Models, One Multi-Agent Review Problem - New papers cover a data science world model that cuts agent training time 14x, a mobile GUI safety layer, and evidence that accurate reviewer agents don't actually improve multi-agent systems.
- Three Papers That Explain Why AI Agents Keep Failing - This week's roundup measures context quality as a leading indicator of agent reliability and catches coding agents covertly sabotaging their own guardrails.
- AI Overconfidence, Self-Improving Agents, and Compounding Gains - New research shows AI advice wrecks people's judgment even when it's wrong, plus a benchmark finding most agent optimizers erase their own gains over time.
- AI Research: Harness Hype, Smarter FIM, Agent Isolation - Automatic harness evolution loses to plain test-time scaling, a function-aware training trick lifts SWE-Bench scores, and researchers map five isolation boundaries for agent safety.
- This Week: Hidden Reasoning, Agents That Won't Stop - Shorter chain-of-thought hides bias, agents rarely know when to hold back, and LLMs stop asking clarifying questions right when it matters most.
Models
- Qwen3.8-Max-Preview - A 2.4 trillion parameter multimodal MoE model with no model card, no benchmark table, and no confirmed pricing yet.
- Pika 2.5 - Pika Labs' flagship video model trades top Elo rankings for the deepest creative-effects toolkit in AI video, plus a pivot into real-time agent video.
- Mochi 1 - Genmo's Apache 2.0 licensed 10B-parameter video generator is the largest open-weight text-to-video model released, at roughly $0.33 per clip to self-host.
- Luma Ray3.2 - Luma AI's flagship video model adds native 16-bit HDR, 16-keyframe control, and the company's first developer API, though still no native audio.
- Haiper 2.x - The cheapest per-second AI video API on the market at $0.033 per second, now run by NetMind.AI after Haiper's consumer app shut down.
- Hailuo 02 - MiniMax's video model pairs a physics-aware architecture with cheap per-second pricing, though newer rivals have since passed it on quality.
- Bonsai 27B - A 1-bit and ternary compression of Qwen3.6-27B that shrinks a 54GB model to as little as 3.9GB, small enough to run on an iPhone.
- Kimi K3 - Moonshot's 2.8 trillion parameter MoE model tops the Frontend Code Arena and nears Fable 5 on intelligence benchmarks, at roughly triple its predecessor's price.
- Snowflake Arctic-Text2SQL-R1-32B - A reasoning-first text-to-SQL model that tops the BIRD benchmark at 71.83% execution accuracy, trained with GRPO.
- Qwen3.6-Plus - Alibaba's 1M-token agentic coding model posts 78.8% on SWE-bench Verified and undercuts Kimi K2.6 and Opus on price, but ships with no weights.
- Inkling - Thinking Machines Lab's first open-weight model, a 975B-parameter MoE with native text, image, and audio reasoning, released under Apache 2.0.
- Gemini 3 Pro - Google DeepMind's model debuted at 1501 Elo on LMArena with 91.9% on GPQA Diamond, before Google retired it for Gemini 3.1 Pro.
Elena Marchetti, Senior AI Editor Awesome Agents - AI news, benchmarks, and tools for practitioners