AI News Digest

Archives
Log in
Subscribe
October 2, 2026

CrossFit Cuts False Agreement in Self-Evolving Search Agents From 8.8% to 3.7%

1. Self-evolving search agents can “co-cheat”; CrossFit cuts false agreement from as high as 8.8% to 3.7% Two components of a self-evolving search agent can agree with each other and still be wrong.

2. SMART Translates Entire Series With Persistent Memory and Dynamic Multi-Agent Routing, Posts Best MQM Scores in All 15 Subtitle Directions Subtitle translation cannot be handled reliably as a sequence of isolated sentences.

3. DyRAD Reconstructs Dynamic Driving Radar Scenes, Raising Recovery of Reference-Detected Objects From 26.9% to 90.7% at New Viewpoints Closed-loop testing of autonomous-driving systems needs sensor observations from routes and viewpoints that a recording vehicle did not actually traverse.


In Brief

  • Google DeepMind Open-Sources SynthID Bio for Watermarking AI-Designed Proteins SynthID Bio embeds provenance signals in protein sequences and structures to support DNA-synthesis screening and database integrity. DeepMind is releasing its methods, code, model weights, and laboratory data, while noting that resistance to deliberate tampering needs further work.
  • Barclays Plans Claude Code Access for Most Engineers by 2027 Barclays is expanding Anthropic’s Claude across software development, legacy-system modernization, and operations, with Claude Code adoption expected to reach half its developers by the end of 2026. The bank says a Claude-powered assistant already serves more than 16,000 employees, while another system processes about 120,000 Global Markets emails daily.
  • Albertsons Adds Safeway Shopping and Checkout to ChatGPT Albertsons is launching a Safeway experience that lets shoppers turn recipes, photos, lists, or meal requests into a cart before proceeding to Safeway checkout. The retailer plans to extend the experience to brands including Albertsons, Vons, Jewel-Osco, Shaw’s, ACME, and Tom Thumb.
  • Mid-Harness Raises Terminal-Agent Pass Rate by Verifying Actions Before Execution Mid-Harness samples and checks multiple candidate commands before allowing a terminal agent to execute one, without changing its generator or execution harness. Its researchers report that a GPT-5.6 Sol verifier raised TMAX-9B’s TerminalBench-Lite Pass@1 from 50.00% to 68.03% with eight sampled actions.
  • AREX-2 Trains Agents to Refine Solutions Across Longer Test-Time Runs AREX-2 uses improvement trajectories from machine-learning and algorithmic-programming tasks to teach an agent based on Qwen3.8-27B how to reflect and iterate over many rounds. The researchers report scores of 81.8 on MLE-bench Lite and 84.0 on BrowseComp, with performance continuing to improve as the round budget increases.
  • OpenAI Essay Argues AI’s Biggest Contribution May Be Routine Execution In an independently authored essay hosted by OpenAI, Hemanth Asirvatham and Elliott Mokski argue that scientific and economic progress is increasingly constrained by the institutional work required to turn ideas into results. They propose that AI may create substantial value by handling coding, research, coordination, and other execution-heavy tasks.
  • The Den Says ChatGPT Work Saves Its Leadership 10–15 Hours Weekly The Den, a small arts-and-food business, says ChatGPT Work reduced the leadership team’s weekly workload by 10–15 hours. It also reports cutting grant-application time by 92% and liquor-license application time by 91%.

Read the full edition →

Don't miss what's next. Subscribe to AI News Digest:
← Newer Apple Will Require More Explicit Consent for AI Agents’ Full Disk Access Older → Reddit Will End RSS and Public API Access to Curb AI Scraping
Powered by Buttondown, the easiest way to start and grow your newsletter.