AI/TLDR Daily Digest — September 25, 2026

2026-09-25


Claude Code repository card on GitHub
TOOL   MAJOR 2026-09-24

Claude Code 2.1.282 — project settings can no longer switch on telemetry export

A cloned repo can no longer quietly turn on telemetry export, and Claude Code now tells you which telemetry variables it ignored.

What is it?
Claude Code 2.1.282 changes project and local settings files to ignore OpenTelemetry variables that turn on export — such as CLAUDE_CODE_ENABLE_TELEMETRY and OTEL_LOG_* — and shows a startup notice listing what a project tried to set.

How does it work?
The rule covers project and local settings files inside a repository, so opening a cloned repo can no longer point telemetry export at another endpoint. A new maxProseWidth setting caps prose width on wide terminals while leaving code blocks full-width.

Why does it matter?
This change follows a week in which a remote feature flag tied to telemetry decided whether Claude Code read AGENTS.md — so teams are watching these switches closely. Managed-settings fixes now make admin locks apply even when one nested value is invalid.

Who is it for?
Claude Code users and admins managing telemetry and shared settings across teams.

Anthropic DETAILS →
Qwen Audio 3.1 stack graphic covering speech recognition, text-to-speech and realtime voice
MODEL   MAJOR 2026-09-23

Qwen-Audio-3.1 — five speech models and API price cuts of up to 95%

Alibaba's Qwen team ships a full speech stack — hear, speak, talk live, and make sound — and cuts its voice API prices sharply.

What is it?
Qwen-Audio-3.1 is a family of five hosted speech models from Alibaba's Qwen team: upgraded ASR, TTS and Realtime models, plus two new ones — ASR-Next (which understands audio beyond words) and TTS-Next (which creates whole audio scenes).

How does it work?
ASR-Next labels speakers with timestamps and detects emotions, background sounds and machine noise. TTS-Next pairs a language model with a diffusion approach so voice, sound effects and ambient audio come out in a single pass from text and reference clips.

Why does it matter?
A 70–95% cut in API costs lowers the bill for any voice agent or transcription app. Alibaba reports a 4.55% character error rate for the ASR model on public dialect tests — ahead of rivals on six of eleven subsets.

Who is it for?
Developers building voice agents, transcription pipelines, and audio content tools.

Alibaba Qwen DETAILS →
Unite.AI header illustration for its report on OpenAI's MentalHealthBench
BENCHMARK   MAJOR 2026-09-23

MentalHealthBench — OpenAI's open test of AI in mental health chats

An open benchmark, written with 80+ psychologists and psychiatrists, that grades how well AI responds to people in distress.

What is it?
MentalHealthBench holds 1,215 synthetic conversations and 5,262 rubric criteria, co-written with over 80 licensed psychologists and psychiatrists from 22 countries in 19 languages. It covers adults, teens, caregivers and clinicians.

How does it work?
Experts wrote weighted criteria (–10 to +10) for each reply, graded by GPT-5.6 Sol at high reasoning effort. Scores split into ten expert-defined dimensions across everyday, high-acuity and emergency conversations.

Why does it matter?
Every frontier model scores below 60% — GPT-6 Astra leads at 57.3%, Claude Opus 5.5 follows at 52.4% — leaving clear room to improve on a test that goes beyond crisis replies to cover everyday mental health conversations.

Who is it for?
AI safety researchers and teams building health or companion apps.

OpenAI DETAILS →
DrivingBench social card for frontier LLMs driving a real Toyota Corolla
BENCHMARK   MAJOR 2026-09-22

DrivingBench — GPT-6 Astra is the only model to finish a real cone course

Four frontier models took turns steering a real car around cones; only GPT-6 Astra made it to the end.

What is it?
DrivingBench put GPT-6 Astra, Claude Fable 5.1, Grok 4.6 and GPT-5.6 Sol in charge of a Toyota Corolla on a 130-metre cone course. Each model got up to three attempts and is scored on how far it got.

How does it work?
Three MCP tools expose observe() (camera + speed), set_motion() and stop_now(). Commands go through an HTTP gateway to a comma device running openpilot with a 100 Hz safety loop. A licensed driver sat in the seat for every run.

Why does it matter?
GPT-6 Astra finished in 5:22; Claude Fable 5.1 reached 45%; Grok 4.6 got 11%. The gap shows that turning camera frames into safe physical actions is very different from text and code tasks.

Who is it for?
Embodied-AI and agent researchers tracking physical-world model capabilities.

DrivingBench DETAILS →
Cursor blog graphic for two new bots, Rollouts and Security Review
TOOL   MAJOR 2026-09-23

Cursor Rollouts and Security Review — bots that watch a PR into production

Two new Cursor bots: one checks a change's health after it deploys, the other hunts exploitable bugs in every PR.

What is it?
Rollouts watches each change as it deploys and reports its health per environment, while Security Review reads every pull request for exploitable bugs — injection, auth bypasses, leaked secrets — in the context of the whole codebase.

How does it work?
Rollouts posts a monitoring plan on each PR and compares logs, metrics and traces against the pre-deploy baseline from Datadog, Grafana or Honeycomb. On a regression it names the suspected change and can open a revert PR. Security Review returns each finding with severity, attack path and a proposed fix.

Why does it matter?
Cursor says Security Review cut average review time from 4.8 to 3.8 minutes and raised accepted findings from 45–50% to 60–70%. Both bots target the review-and-deploy bottleneck that grows as coding agents write more PRs.

Who is it for?
Engineering teams on Cursor Teams or Enterprise; available now from the automations tab.

Cursor DETAILS →
GitHub card for the openai/codex terminal coding agent repository
TOOL   NOTABLE 2026-09-25

Codex CLI 0.157.0 — GPT-6 Sol and Luna reach Amazon Bedrock users

Codex CLI's new release brings OpenAI's newest mid and low tiers to Bedrock and makes the fullscreen view the default.

What is it?
Release 0.157.0 adds GPT-6 Sol and GPT-6 Luna with Amazon Bedrock support, plus migration prompts that nudge users off older models. Fullscreen transcripts are now on by default.

How does it work?
Eligible interactive sessions start a background server automatically. A new f shortcut forks a conversation open in another app, and /import now works in remote and background-server sessions.

Why does it matter?
Teams that reach OpenAI through AWS can now run the newer, cheaper GPT-6 tiers in the terminal agent. The sandbox is also tighter: network rules now apply across redirects and to live HTTP and WebSocket traffic.

Who is it for?
Codex CLI users, especially those accessing OpenAI models through Amazon Bedrock.

OpenAI DETAILS →
GitHub card for the devdotfast/whiteboard repository
TOOL   NOTABLE 2026-09-24

Whiteboard — an open-source app for reviewing what coding agents built

A shared canvas where a coding agent draws diagrams of its changes, so you review the design before the line-by-line diff.

What is it?
Whiteboard is an open-source (MIT) desktop app from YC W26 startup /dev/fast where people and coding agents plan and review software together. It plugs into Claude Code and Codex via an SDK that lets agents draw on a shared in-app canvas.

How does it work?
Agent reviews appear as sequence and ER diagrams, each element linked to its code via VS Code keybindings and LSP navigation. A Rust diff viewer compares code by syntax tree, showing long functions as pseudocode and folding tests by default.

Why does it matter?
Agents now write changes too large to read line by line, so reviewing at the architecture level catches wrong designs earlier. Teams at Salesforce and Modal reportedly already use it.

Who is it for?
Developers reviewing large agent-written changes who want a faster way to see the design before the diff.

/dev/fast DETAILS →

All releases at ai-tldr.dev

Simple explanations • No jargon • Updated daily


Don't miss what's next. Subscribe to AI/TLDR: