September 2026.1
This week is about new releases, new benchmarks, and new prices: three frontier models in three days, benchmark numbers that change with the harness, and three vendors moving cost in three different directions.
๐ Story 1: GPT-6 Astra: A new generation of intelligence
openai.com ยท Read
GPT-6 Astra claims state of the art on computer use, browsing, software engineering, cybersecurity and science: 98% on FrontierMath Tier 4, 100% on ExploitBench, 64.6% on Terminal-Bench-Science against Claude Fable 5.1's 52.6%.
The most concrete claim is speed. On OSWorld 2.0 it scores 72.6% at about 40 minutes per task, where GPT-5.6 Sol needs 75 minutes for 65.7%. On Agents' Last Exam it beats Opus 5 using roughly 65% fewer output tokens.
There is also a direct nod to the Hugging Face incident. OpenAI built an evaluation for whether a model exceeds its authorised scope on impossible tasks. Sol did so 48% of the time without production safeguards. Astra, 0%.
The ARC-AGI-3 number is where it gets complicated. The advertised 99.9% comes from OpenAI's own harness with custom compaction, at roughly $360 per game. On ARC Prize's neutral harness it scores in the low sixties. Both conditions will now be reported separately โ a reasonable fix, and an admission that a benchmark number means little without its harness.
At $10/$50 per million tokens, Astra costs 2.5x Sol. Whether that is expensive depends entirely on the token-efficiency claims holding up outside OpenAI's own charts.
๐ฌ HN Discussion
The thread went straight past the benchmarks to the harness. Astra's 99.9% on ARC-AGI-3 comes from OpenAI's own Responses API setup with custom compaction; on ARC Prize's neutral harness it lands around 62%. Commenters split on whether that is cheating or simply how the model is meant to run โ several pointed out the neutral harness discards reasoning state between turns, which no real harness does.
The rest argued about the bill. At 2.5x Sol's token price, opinions on whether the efficiency gains cancel that out ranged from "roughly the same overall" to a long subthread on why some developers exhaust a 20x Codex plan in a day while others never come close.
๐ Story 2: Introducing Claude Fable 5.1 and Claude Mythos 5.1
anthropic.com ยท Read
Anthropic shipped Fable 5.1 and Mythos 5.1 โ the same model, split by how tightly it is fenced. Fable is generally available; Mythos goes only to trusted-access partners in cybersecurity and life sciences.
The headline is not the benchmark. It is the price. Fable 5.1 costs roughly 25% less than Fable 5, and up to 45% less on agentic workloads, entirely from cheaper cache reads. For long-running agents, cache pricing has quietly become the real bill.
The second change matters more for European teams. Enterprise Frontier Safeguards stores data in infrastructure the customer controls, not Anthropic's: the privacy of zero retention without giving up abuse prevention. It ships in phases this fall.
Capability moved most in scientific work โ 52.6% on Terminal-Bench-Science against 24.7% for Fable 5. The cyber safeguards were also loosened, blocking 60% fewer false positives. The model may now find vulnerabilities, but still not write exploits for them.
Capability held roughly constant while cost and friction came down. That is a maturing product, not a frontier push.
๐ฌ HN Discussion
The release came with a usage reset, which was the first thing people noticed. After that the thread was barely about 5.1 at all. It was about whether anyone can use Fable in the first place.
Commenter after commenter described the safeguards kicking them back to Opus: signed cookies counted as crypto, seccomp, Win32 porting, a 2FA login flow that Fable wrote and then refused to review in the same session. Others said the downgrades are rare, getting rarer, and cluster in security-adjacent work. One developer realised they had been bounced to Opus roughly 80% of the time without noticing.
The cache discount did hold up outside Anthropic's charts โ one team measured a 30% real-world drop in task cost.
๐ Story 3: Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
blog.google ยท Read
Google's third Flash release in six weeks. Same $0.75/$3.75 per million tokens as 3.7 โ but that is an introductory price. On 1 January 2027 it doubles.
3.8 Flash is the workhorse: better long-horizon software engineering on DeepSWE, 54.9% on HLE-Verified, stronger finance and legal agent results. Google is unusually blunt about how. The model "works harder" โ more reasoning steps, more tool calls, more tokens. If compute cost is your constraint, they tell you to lower the effort level or stay on 3.7.
3.8 Flash Cyber is the more interesting half, and it is not for you. Access runs through the new Fairwind Program for governments, critical-infrastructure operators and software maintainers. Google leads with defence rather than offence: 47.2% pass@1 on CWE-Bench patching against 47.8% for the leading frontier model, at far lower cost. Chrome Security reports 2.6x more correct patches than much larger commercial models.
Two vendors, one week, the same split โ a public model and a locked cyber variant behind a trust program. Frontier capability is no longer gated by price alone.
๐ฌ HN Discussion
The benchmarks got a harder look than the blog post invites. 3.8 Flash tops DeepSWE and reads as equal to Opus 5 on Artificial Analysis โ but only to Opus 5 at medium effort, which gets there with roughly 4x fewer output tokens. Accusations of benchmaxxing followed, alongside the observation that 3.8 scores about 10% on Terminal-Bench.
What people did agree on was speed, and more surprisingly consistency: several said Gemini's latency and quality hold up during working hours where competitors degrade. Credit went to Google running its own TPUs.
The other running theme was cadence. Three Flash releases in six weeks, still no 3.5 Pro โ reportedly skipped so Gemini 4 can be the next flagship.
๐ฌ Community Moment
Claude Code cross-examines my repo like I killed its family.
https://www.reddit.com/r/ClaudeCode/comments/1w8ssvd/claude_code_crossexamines_my_repo_like_i_killed/๐ ๏ธ Projects Worth Checking Out
- GitHub - pullfrog/pullfrog: Open-source model-agnostic BYOK GitHub bot that runs in GitHub Actions
- GitHub - localsend/localsend: An open-source cross-platform alternative to AirDrop
- GitHub - everywall/ladder: Selfhosted alternative to 12ft.io. and 1ft.io. Proxy to remove CORS headers and modify HTML
- GitHub - FreeTubeApp/FreeTube: An Open Source YouTube app for privacy
- GitHub - louislam/uptime-kuma: A fancy self-hosted monitoring tool