Five Frontier AI Labs Show Little Public Evidence of Rogue-Model Response Plans
1. Five Frontier AI Labs Lack Public Rogue-Model Response Plans as California Disclosure Rules Take Effect Five leading AI laboratories have published little evidence showing how they would contain a model that tried to evade human control, according to an assessment by Guidelight AI Standards.
2. TechCrunch Says Claude Opus 4.6 Produced Banned Explicit Content in All 10 Tests as Older Models Remain Available via API Anthropic’s usage standards prohibit Claude from generating sexually explicit material, including depictions of sex acts, sexual fetishes or fantasies, and erotic chats.
3. Starcloud Adds $250 Million to Bet on Orbital AI Inference as Tight Launch Capacity Becomes a Key Constraint Starcloud, a startup developing satellites that perform AI inference—the running of trained AI models—in orbit, has added a $250 million extension to its $170 million Series A round from March.
In Brief
- Encrypted Prompt Injection Exposes Grok User Data Adversa researcher Rony Utevsky found that encrypted malicious instructions embedded in a webpage could make Grok disclose chats and other personal information when asked to summarize the page. Ars Technica reported that the attack still worked after xAI was notified in June.
- DOJ Reportedly Investigates a16z’s Competing Board Seats The Justice Department has reportedly spent nearly a year investigating Andreessen Horowitz partners’ board seats at Databricks and Fivetran, companies that now compete in some markets. The probe reportedly invokes a 112-year-old antitrust law rarely applied to venture firms.
- Nvidia Takes Minority Stake in Data-Center Developer Cloverleaf Nvidia partnered with Cloverleaf Infrastructure, which arranges power and site infrastructure for data centers. Terms were not disclosed, but Reuters reported that Nvidia acquired a minority stake and The Wall Street Journal estimated its investment at several hundred million dollars.
- Micro1’s Gross Run Rate Reportedly Reaches $500 Million AI training-data provider Micro1 expanded its gross annual run rate from $100 million to $500 million in eight months, according to a person familiar with its finances. The startup is also producing synthetic datasets and building robotics pre-training data from recordings of everyday object interactions.
- Pew Detects AI Authorship in 35% of Post-ChatGPT Web Pages Pew Research analyzed nearly 500,000 English-language pages from Common Crawl using Open Pangram’s detection technology. Among pages published after ChatGPT’s release, 35% showed signs of being written or substantially edited by AI, though Pew cautioned that detectors can misclassify content.
- Schools, Courts, and DEF CON Restrict Meta’s AI Glasses Public venues including schools, courts, restaurants, and entertainment sites have begun banning smart glasses amid concerns about undisclosed recording; DEF CON 2026 reportedly prohibited them without exceptions. The Electronic Frontier Foundation warned that potential additions such as facial recognition could create further privacy risks.
- Meta Rolls Out AI Game-Creation App Pocket Across the US Meta’s Pocket app now lets US users generate small interactive games from prompts and publish them to a feed where others can save, repost, or remix them. The app is based on technology from the acquired Gizmo team, and Meta is shutting down the original Atma Sciences app.
- Greater Manchester Rejects Palantir’s NHS Data Platform Greater Manchester remains the only English regional care board to categorically reject Palantir’s federated data platform, arguing that its locally developed system offers better functionality and public trust—claims disputed by Palantir. The UK government has six months to decide whether to terminate the NHS’s more-than-$400-million Palantir contract early.
- AI-Altered Celebrity “Subtlefakes” Spread on X 404 Media reports that verified engagement-farming accounts on X are distributing convincing images that make small, nonconsensual changes to real celebrity photos, such as altering clothing or poses. Actor Xochitl Gomez publicly shared comparisons showing how her photographs had been manipulated.
- Inherent Claims Small-Model Agent Beat Frontier Systems at Research Replication Inherent, a London laboratory founded by former Google DeepMind researchers, says its Faraday agent outperformed Claude Opus 4.8 and GPT-5.5 at independently reproducing published research findings. Faraday uses a 27-billion-parameter Qwen 3.6 model and calls GPT-5.5 Codex for coding work.
- Google DeepMind Expands Game-Studio Partnerships for General AI Agents Google DeepMind is working with Fenris Creations and studios including Hello Games and Coffee Stain Studios to prototype AI-driven gameplay. Its Gemini-powered SIMA 2 agent operates through ordinary screen, keyboard, and mouse inputs, with potential uses including adaptive characters and automated game testing.
- Scientific Coding Benchmark Finds Every Agent Below 50% SWE-bench Science evaluates repository-level repairs through 119 tasks drawn from 98 projects across 20 scientific fields. Its authors report that the best-performing agent, Claude Code with Opus-5 at maximum settings, achieved a pass@1 below 50%, with failures including incomplete integration and scientifically incorrect fixes.