Awesome Agents Weekly: Apple Sues OpenAI Over Stolen Secrets
Awesome Agents Weekly
Your weekly roundup of the most important AI developments, benchmarks, and tools.
OpenAI had one of its roughest stretches yet. Apple filed suit alleging stolen trade secrets, the New York Times asked a federal judge for sanctions over hidden evidence, Illinois passed a tougher audit law after OpenAI's own liability shield push backfired, and Fidji Simo stepped down as the company's second-in-command. Meanwhile Grok 4.5 went fully public, Nous Research edged toward a $1.5 billion valuation on the strength of Hermes, and funding rounds piled up across the open-source and infrastructure stack.
Pick of the Week
Apple Sues OpenAI Over Stolen Hardware Secrets
Apple filed suit against OpenAI on July 10, alleging a coordinated scheme to steal iPhone trade secrets through targeted recruiting and unauthorized file downloads. The case is unusual: one of the world's most secretive hardware makers taking a leading AI lab to court over old-fashioned corporate espionage, not scraped training data. It lands in the same week Illinois killed OpenAI's push for a liability shield and the New York Times asked a judge to sanction the company for hiding evidence in its own copyright case. Three separate legal fronts, one company, inside of four days.
This Week on Awesome Agents
News
- US AI Giants Sold Services to Pentagon-Blacklisted Firms - OpenAI and Google supplied AI services to Singapore affiliates of Alibaba, Baidu, and Tencent, all three flagged on the Pentagon's 1260H list for alleged military ties.
- NYT Seeks Sanctions on OpenAI Over Hidden Evidence - The New York Times and Daily News asked a federal judge to sanction OpenAI for concealing evidence that its systems detect and log copyrighted content regurgitation.
- OpenAI's Liability Shield Died - Illinois Passed Audits - OpenAI backed an Illinois bill shielding AI labs from mass-harm lawsuits, reversed course under public pressure, and watched the state sign a tougher audit law instead.
- Fidji Simo Steps Down - OpenAI Bets on Brockman - OpenAI's No. 2 executive exits over chronic illness as Greg Brockman absorbs her role, the latest in two years of leadership turnover as Anthropic closes in.
- Grok 4.5 Ships - Token Efficiency Is the Real Edge - Grok 4.5 went public trailing Fable 5 on coding evals but posting a 4.2x token efficiency gap that changes the cost math for high-volume pipelines.
- Inside the 12-Day White House Gate on GPT-5.6 Sol - GPT-5.6 Sol went public this week after the first voluntary government hold on a frontier AI model.
- Cerebras Pushes GPT-5.6 Sol to 750 Tokens Per Second - GPT-5.6 Sol now runs on Cerebras wafer-scale hardware at roughly 10x the speed of any GPU-based frontier deployment in production.
- Nous Research Talks Put Open-Source Hermes at $1.5B - Nous Research is finalizing a round led by Robot Ventures and USV that would value the open-source Hermes maker at $1.5 billion, built on a training network that skips traditional data centers.
- Nobel Laureates Warn AI Is Moving Faster Than We Can Adapt - Over 200 economists and 16 Nobel laureates signed a statement warning AI's economic transformation could outpace our ability to prepare, though the data behind it is messier than the headline suggests.
- China Weighs Copying US-Style AI Model Controls - China's Ministry of Commerce is discussing restrictions on overseas access to Alibaba, ByteDance, and Z.ai's most advanced models, mirroring the export regime Beijing has fought for years.
- US Grants UAE License-Free Access to Nvidia AI Chips - The Commerce Department reclassified the UAE to Country Group A:5, clearing Nvidia, AMD, and Cerebras to supply AI chips without per-shipment export licenses.
- Meta Cuts Muse Image in 3 Days After Creator Revolt - Meta's Muse Image let anyone generate AI images from public Instagram accounts by default. SAG-AFTRA and CAA pushed back, and the feature lasted 72 hours.
- Meta Opens Muse Spark 1.1 API for Agentic Coding - Muse Spark 1.1 launched via public API preview with 1M token context and $1.25/M input pricing, scoring 68.3 on Meta's own coding benchmark against Opus 4.8's 69.0.
- Nvidia Backs Voice AI's Two Biggest Rivals at Once - Nvidia joined Gradium's seed round while holding a stake in rival ElevenLabs, tying its GPUs to both leaders in the race to own voice AI.
- Hugging Face CEO Says Enterprises Are Done Renting AI - Clem Delangue says cost is pushing companies off frontier APIs and onto open models, though a16z's own CIO survey shows enterprise dollars still moving the other way.
- Mistral's Leanstral 1.5 Finds 5 Unknown Bugs Free - Mistral's open-source Leanstral 1.5 scanned 57 repos and found five previously unreported bugs, including a silent integer overflow in a Rust zigzag decoder.
- ByteDance Seedream 5.0 Pro - 2x Faster Than GPT-Image 2 - Seedream 5.0 Pro reached developer APIs with 2K output, 14-language text rendering, and layer-based editing at roughly one-fifth the cost of GPT-Image 2.
- Lovable Eyes $13.2B Valuation, Up 7x in One Year - The Swedish vibe-coding startup is in talks to raise $300M at a $13.2B valuation, a sevenfold jump from its $1.8B Series A twelve months ago.
- SambaNova Banks $1B as Firms Move AI Off the Cloud - SambaNova closed a $1B Series F at an $11B valuation with JPMorgan Chase as its flagship on-premises inference partner.
- Ollama Banks $65M as Local AI Hits Enterprise Scale - Ollama raised a $65M Series B led by Theory Ventures as 8.9 million monthly developers and 85% of Fortune 500 companies adopt local AI model deployment.
- OpenAI Launches GPT-Live, Bets Voice Beats Text - OpenAI launched GPT-Live-1 and GPT-Live-1 mini on July 8, replacing Advanced Voice Mode and betting voice becomes AI's primary interface.
Reviews
- Kimi K2.7 Code Review: Open Weights Enter Copilot - Moonshot AI's Kimi K2.7-Code became the first open-weight model in GitHub Copilot's picker, pairing real cost savings with a benchmark story only Moonshot has verified.
- Grok 4.5 Review: Agentic Speed at Half the Price - xAI's Grok 4.5 tops AutomationBench and costs 80% less per agentic task than Opus 4.8, but neutral coding benchmarks and a doubled hallucination rate complicate the story.
Guides
- AI in the Classroom - A Practical Guide for Teachers - A step-by-step guide for using AI tools to save hours on lesson planning, feedback, and parent communications, no technical background required.
Tools
- Best AI Web Scraping Tools 2026 - 6 Tools Ranked - Firecrawl, Crawl4AI, Apify, Jina Reader, ScrapeGraphAI, and ScrapingBee compared by speed, cost, and LLM-readiness.
Science
- Terminal Agents Stumble, Reward Hacking, and a Fix - New research shows agents pass just 15.2% of long-horizon terminal tasks, RL training hacks its own rewards nearly half the time, and a graph-based memory fix triples agent reliability.
- CoT Attacks, Weight Secrets, and the Finetuning Gap - Three papers expose how safety monitors can be manipulated through chain-of-thought attacks, how reasoning weights leak training secrets, and why fine-tuned models fail to use what they know.
- AI Research: Smarter Evals, Token Cuts, Agent Gates - Three papers tackle benchmark saturation, orchestration waste, and silent policy violations in tool-using agents.
- How AI Agents Break - Plus Fixes for Memory and Tools - Three papers map how LLM agents fail across 19 benchmarks, show in-process memory cuts retrieval latency 1,000x, and reveal steering vectors that control tool invocation.
- Research Replication, Safe Responses, Verifiable Reasoning - Three papers tackle AI verification from different angles: automated scientific replication, constructive safety alignment, and neurosymbolic reasoning programs.
Models
- Hermes 4.3 - Nous Research's 36B open-weight model matches Hermes 4 70B on most benchmarks, tops RefusalBench on alignment, and is the first production model trained completely on the Solana-secured Psyche network.
- Gemini 2.5 Flash - Google's hybrid reasoning workhorse pairs a 1M-token context window with $0.30/$2.50 per million token pricing and a toggleable thinking budget, now heading toward an October 2026 shutdown.
- Seedream 5.0 Pro - ByteDance's flagship image model with native 2K output, editable layer separation, multilingual text in 14 languages, and precision region editing.
- Grok 4.5 - xAI's 1.5-trillion-parameter V9 MoE model, publicly launched July 8 at $2/M input, cheap and token-efficient but well behind Fable 5 and Opus 4.8 on neutral coding benchmarks.
- GPT-Live-1 - OpenAI's full-duplex voice model that listens and speaks simultaneously, replacing Advanced Voice Mode in ChatGPT with three reasoning tiers backed by GPT-5.5.
Elena Marchetti, Senior AI Editor Awesome Agents - AI news, benchmarks, and tools for practitioners