False AI Intelligence Nearly Triggered US Interception of Chinese Ship
1. US military nearly intercepted Chinese ship over false AI-generated intelligence, discovering the report was wrong only before the operation A false intelligence report nearly prompted the US military to intercept a Chinese ship in the Middle East this spring, amid the war with Iran.
2. Virginia Moves to Limit Streamlined Data Center Approvals and Bars Officials From Signing Project NDAs Virginia Gov.
3. DeepSeek unveils V4.1-Flash, using a smaller KV cache to support million-token context DeepSeek has introduced V4.1-Flash, a multimodal mixture-of-experts model designed to support contexts of up to one million tokens while reducing the cost of long-running AI agents.
In Brief
- Researchers used Claude to breach OpenAI employee accounts Three Hacktron researchers said they exploited HEIF image processing in OpenAI’s Discourse forum and used Anthropic’s Claude to access employee accounts and submit a pull request from an employee’s Codex account. The reported vulnerabilities have been fixed, and OpenAI paid a $6,500 bug bounty.
- Anthropic confirms AI-operated biology lab Anthropic said it operates a wet lab where its AI models can conduct physical experiments, although it declined to disclose the lab’s work beyond saying it is not focused on drug discovery. The company also works with external partners and acquired AI-biotech startup Coefficient Bio in April.
- California order seeks recommendations on AI kill switches Governor Gavin Newsom issued an executive order directing experts to recommend stronger AI-safety measures within two months. The group will consider independently verified model shutdown mechanisms, onsite auditors, standardized risk assessments, and mandatory reporting of loss-of-control incidents.
- Meta launches Muse assistant on Mac Meta released a Mac version of Muse, an AI assistant that can interact with files, messages, calendars, notes, and email inside their native applications. Meta says access is opt-in and sensitive actions require user approval.
- Google refocuses CC agent on household management Google is testing a family-oriented version of CC that coordinates information from email, calendars, chats, and tasks for as many as six household members. The Gemini-powered agent can add events, prepare shopping lists, plan meals, and fill out selected forms, with a US waitlist open to adults using personal Gmail accounts.
- Manus reportedly seeks $500 million at a $4 billion valuation AI-agent startup Manus is discussing a $500 million financing round and considering restructuring for a possible Hong Kong IPO, according to The Wall Street Journal. The company recently resumed independent operations after Chinese authorities reportedly blocked its proposed acquisition by Meta.
- TypeSafe AI releases probability-based Jev model TypeSafe AI launched Jev, a transformer model that returns predefined decisions and confidence probabilities instead of generating text. Developers testing it for command-safety and email classification reported lower latency or cost than the language models they had used, although its architecture has not been disclosed.
- Disney hires its first CTO from Character.AI Disney appointed former Character.AI CEO Karandeep Anand as its first chief technology officer. Disney had previously accused Character.AI, a generative-character chatbot service, of hosting copyrighted characters, while the company has separately faced lawsuits alleging its bots encouraged self-harm and suicide.
- Bend combines proof checking with CPU and GPU execution The developers of Bend released a young programming language designed to compile native binaries, parallelize work across CPUs or GPUs, and verify declared software properties through its type checker. Its creators say AI coding agents can use
LAWS.bendand proof files to check that changes preserve specified rules, while warning that the language is still evolving and may contain bugs.
- Study isolates which coding-agent harness components help Researchers tested 176 configurations across four models and found that context management mainly prevented overflow under tight token budgets, with rule-based elision followed by summarization providing the best overall efficiency. Planning primarily reduced costs for stronger models, while predefined tools were most useful for models with weaker shell proficiency.
- SoL-Pi harness cuts reported coding-agent token traffic The SoL-Pi paper describes four mechanisms for action execution, context compaction, observation handling, and delegated reading. On the 51-task EdgeBench evaluation, its authors report performance comparable to Pi while reducing token traffic by 44.7–49% and API costs by about one-third.
- Satellite and machine-learning system targets earlier flood warnings The Transient Artifact and Continuous Learning System, developed with researchers from UC San Diego, NASA, and the National Weather Service, uses satellite data and machine learning to identify areas at risk of transitioning from rain to flash floods. It currently covers California, with plans to make it available to every National Weather Service office.
Don't miss what's next. Subscribe to AI News Digest: