AI/TLDR Daily Digest — September 05, 2026

2026-09-05


Claude Code v2.1.261 release page on GitHub
TOOL   MAJOR 2026-09-04

Claude Code 2.1.261 — /skill-doctor shows which skills waste your context

Claude Code 2.1.261 adds a skill audit, much bigger inline output limits, and a stricter rm -rf safety check.

What is it?
A new /skill-doctor command reports which loaded skills a session never used and what each one costs in context, so you can drop the ones that are only taking up room. The release also adds --append-subagent-system-prompt-file for subagent prompts too large to pass on the command line.

How does it work?
Two new settings, bashOutputMaxChars and taskOutputMaxChars, raise the ceiling for inline command and background-task output to 128K characters. /status and claude doctor also gain an "Organization policy" line that explains why a managed policy could not be loaded.

Why does it matter?
The dangerous-rm prompt now also catches rm -rf on positional parameters and inside double-quoted sh -c scripts, and auto mode stops auto-approving links that pack content into a public diagram renderer's URL. Prompt word-editing keys switch to Bash behaviour, superseding the keybindingFlavor setting.

Who is it for?
Claude Code users and the admins who manage them. Install with: npm i -g @anthropic-ai/[email protected]

Anthropic DETAILS →
OpenEvidence medical AI model family launch with Darwin research preview
MODEL   MAJOR 2026-09-03

OpenEvidence model family — Osler, Sackett and Snow ship free to clinicians

Four medical AI models that all aim at the same accuracy — what changes is how long each one thinks.

What is it?
OpenEvidence now ships a named model family instead of one search box. Osler answers in about five seconds and is the new default; Sackett takes about thirty seconds for evidence-strength questions; Snow takes about five minutes and reads through the medical literature before writing anything.

How does it work?
Every model is held to the same standard of clinical accuracy — the difference is time and search depth. Snow replaces the older Deep Consult feature and runs a full literature investigation. A separate oncology sub-agent draws on a precision oncology knowledge base alongside NCCN treatment algorithms and ASCO guidelines.

Why does it matter?
Clinicians can now match the wait to the question. Osler, Sackett and Snow are free for the platform's 1.12 million license-verified U.S. clinicians. The fourth model, Darwin, achieves a perfect 660/660 on MedQA but stays in research preview because of dual-use risk in virology and genetics work.

Who is it for?
Practising clinicians and clinical AI researchers. Available at openevidence.com and on iOS and Android.

OpenEvidence DETAILS →
OpenAI wordmark on a textured background
SECURITY   MAJOR 2026-09-03

Daybreak for Frontline Defenders — $1B of OpenAI cyber credits for utilities

OpenAI will subsidize $1 billion of Daybreak cyber-model use for under-resourced defenders of essential services.

What is it?
Daybreak for Frontline Defenders opens OpenAI's cyber models to organizations that keep essential services running but cannot pay enterprise rates. Water utilities, electric grids, local governments, community banks, nonprofits and open-source maintainers can apply for subsidized credits, training and hands-on technical support.

How does it work?
Defenders are routed to two existing Daybreak tiers: Daybreak Blue for common defensive work on mainline models, and Daybreak Red for specialized cyber tasks. Teams use them to review legacy code, analyze suspicious activity, find and rank vulnerabilities, and write fixes. A pilot with the Multi-State Information Sharing and Analysis Center adds training for state, local, tribal and territorial defenders.

Why does it matter?
Small utilities and city IT teams defend complex, aging systems with very little staff and budget, while frontier cyber models have been priced for large enterprises. The MS-ISAC pilot alone covers utilities in 40 states plus DC that serve more than half the US population. OpenAI expects the $1B to be used over about six months.

Who is it for?
Critical-infrastructure and public-sector security teams. US organizations first; partner countries in coming weeks.

OpenAI DETAILS →
EEBench post asking whether AI can design circuit boards yet
BENCHMARK   MAJOR 2026-09-04

EEBench — atopile's benchmark scores frontier models on circuit design

A simulation-graded benchmark that asks whether AI models can design circuit boards that actually work.

What is it?
EEBench V1 puts frontier models through 13 circuit-design tasks spanning analog and digital design, and grades answers with physics — not a human vote. Claude Opus 5 leads the first leaderboard at 61.6%, ahead of Grok 4.6 at 57.1% and GPT-5.6 Sol at 39.4%.

How does it work?
Models write atopile design code — so the grader can read components, connections and electrical constraints directly — then each submission is run through SPICE simulation and design checks at worst-case component corners. A final score weights technical performance at 0.65 and cost efficiency against a reference BOM at 0.35.

Why does it matter?
Hardware has had no SWE-bench of its own, so claims that a model could design a working circuit were hard to verify. EEBench turns it into a number, and the spread is wide enough to be meaningful — 22 percentage points separate the leader from GPT-5.6 Sol.

Who is it for?
Hardware engineers and model evaluators. Full leaderboard, per-task scores and cost breakdowns at eebench.org.

atopile DETAILS →
HUMAIN announces humain-m3, its frontier Arabic language model built by MiniMax
MODEL   MAJOR 2026-09-03

humain-m3 — a 428B Arabic model HUMAIN commissioned from MiniMax

Saudi Arabia's HUMAIN commissioned a 428B Arabic frontier model from MiniMax and opened it as a limited preview.

What is it?
humain-m3 is a 428-billion-parameter mixture-of-experts model (23B active per token) announced by HUMAIN at LEAP in Riyadh. Built on the MiniMax-M3 lineage, it was further pre-trained on more than a trillion tokens of Arabic-native content.

How does it work?
HUMAIN paid MiniMax to continue pre-training M3 on Arabic data rather than training from scratch. The preview period is used to check capability, safety and alignment across Arabic dialects. On seven public Arabic benchmarks the model averages 89.37%, leading five of the seven — its widest margin is on AraTrust (97.53% vs 93.42% for GPT-5.6 SOL).

Why does it matter?
Arabic is spoken by hundreds of millions yet stays thinly covered by frontier models. The other signal is geopolitical: a Gulf state's national AI programme licensed a Chinese lab's model rather than building its own. HUMAIN plans to release weights under the MiniMax Community License once safety work is done.

Who is it for?
Teams building Arabic-language products. Request limited preview access at node.humain.com.

HUMAIN DETAILS →
Simon Willison's SVG pelican-riding-a-bicycle drawings compared across GPT-6 Astra and GPT-5.6 models
ARTICLE   NOTABLE 2026-09-04

Simon Willison — GPT-6 Astra draws far better pelicans than GPT-5.6

A side-by-side grid puts GPT-6 Astra against three GPT-5.6 models at every reasoning level.

What is it?
Simon Willison published a comparison grid of SVG pelicans riding bicycles, drawn by GPT-6 Astra at low, medium, high, xhigh and max reasoning effort, next to GPT-5.6 Sol, Terra and Luna. The pelican prompt is his informal test for every new model.

How does it work?
Every cell records input tokens, output tokens and cash cost, priced from published rates. At max effort GPT-6 Astra cost 63.21 cents for one drawing; its cheapest run (low effort) cost 9.55 cents — and still beat every GPT-5.6 Sol result.

Why does it matter?
Reasoning effort moves the bill as much as the model choice — one pelican ranges from 1.57 to 63.21 cents across the grid. Willison also spotted that Astra and Luna both used 16 input tokens while Sol and Terra used 26, and speculates the two may be more closely related than OpenAI has said.

Who is it for?
Developers choosing a model and reasoning level — useful quick signal on how much quality Astra's low-effort tier buys at 9.55 cents vs paying 6× more for max.

Simon Willison DETAILS →
ProgramAsWeights — define functions in English and run them locally
PAPER   NOTABLE 2026-09-03

Compile by Training — turn an English spec into a local neural function

Describe a fuzzy text function in English, wait about a minute, and get a .paw file that runs locally with no API calls.

What is it?
Compile by Training is a new compiler for ProgramAsWeights (PAW) that turns a natural-language description into a reusable neural program you can run offline. It reaches 83.6% semantic accuracy on FuzzyBench-Hard, where the older fast compiler scored 22.4%.

How does it work?
Teacher models generate example input-output pairs from the spec, which finetune a compact adapter over a small fixed interpreter. The result is packaged as a .paw file that needs no further model calls at runtime. Compilation takes about 51 seconds on a B300 GPU.

Why does it matter?
Small fuzzy text jobs — sentiment labels, PII typing, format cleanup — usually mean one API call per invocation. ProgramAsWeights turns each into a local file with a name and a version, storable and reusable like any other module, with no ongoing API cost.

Who is it for?
Developers shipping small text-processing features. Try it: uv run compile.py "Classify sentiment. Return only positive, negative, or neutral." -o sentiment.paw

ProgramAsWeights DETAILS →

All releases at ai-tldr.dev

Simple explanations • No jargon • Updated daily


Don't miss what's next. Subscribe to AI/TLDR: