The Signal - daily AI evolution  logo

The Signal - daily AI evolution

Archives
Log in
Subscribe
September 23, 2026

The Signal — September 23, 2026

The Read

Anthropic shipped Claude Opus 5.5 and roughly an hour later OpenAI shipped GPT-6 Sol and Luna, and every one of them came with a price cut: Opus 5.5 at $4/$20 against Opus 5's $5/$25, Sol at $2/$10 — exactly half GPT-5.6 Sol — and Luna at $0.10/$0.50, half on input and 58% cheaper on output. On the same day Epoch AI published the rate underneath all of it — the cost of a fixed level of AI performance has been falling about 47% per quarter, roughly 13x a year, since 2023, faster than electricity, batteries, compute or DNA sequencing ever fell. That is the unusual part. The market event and the measurement of the market event landed on the same date, which is what it looks like when a trend stops being a narrative and becomes an engineering constant. Meanwhile Alibaba put a 20GW datacenter target and its own training chip on the table, METR published a rare quantified number for how much AI is accelerating AI research, and researchers at Cisco Talos documented a Windows malware family that asks four different commercial model APIs what to do next. The intelligence is getting cheaper on a schedule. The question that pays is which old problem you point it at when the price drops another 13x.

🌊 Tide

Confirmed and strengthened — the cost-collapse tide now has a published rate. Every prior confirmation of this tide was an event: a vendor cut, a challenger undercut, an in-house silicon program, a demand curve. Tuesday produced two more events and, for the first time, a peer-style measurement of the slope itself published the same day. Epoch AI's finding that the cost of a given capability falls about 47% per quarter across math, science and games-of-skill benchmarks turns a pattern the Signal has logged item by item since July into a number an operator can put in a model. No tide shift — the direction was already established. What changed is that the tide is now measurable, and a measurable tide is a plannable one.

Two frontier price cuts in one hour, and a published decline rate to explain them

Anthropic released Claude Opus 5.5 at $4 per million input tokens and $20 per million output — a 20% cut from the $5/$25 that Opus 5 held — with cache reads down 60% to $0.20. Anthropic says the model costs about 40% less to run than Opus 5 on typical workloads once you count both cheaper tokens and fewer tokens per task, and generates output more than 30% faster. About an hour later OpenAI released GPT-6 Sol at $2/$10 — exactly half GPT-5.6 Sol's $4/$20 — and GPT-6 Luna at $0.10/$0.50 against $0.20/$1.20, half on input and 58% cheaper on output, and said the reduction is permanent rather than promotional. Simon Willison, who tracked both releases the same evening, notes Luna is among the cheapest models OpenAI has ever shipped, beaten only by the far weaker Nano tiers. Underneath it, Epoch AI published "The plunging price of thought," by Luke Emberson and David Roodman, measuring the cost of a fixed performance level as falling roughly 47% per quarter — about 13x a year — since 2023 across five benchmarks, a decline steeper than any the authors found in electricity, compute, batteries or DNA sequencing.

So what: Stop treating model pricing as a negotiation and start treating it as a schedule. If capability-adjusted cost falls 13x a year, then the correct question about any workload you rejected on unit economics in the last twelve months is not whether it is affordable now — it is what quarter it becomes affordable, and whether you want to be the one already standing there with the distribution and the data when it does. Build the backlog of problems that are obviously worth solving at one-tenth of today's price and start the ones that clear at today's. The compounding advantage is not in routing tokens more cheaply; it is in having pointed intelligence at an old problem early enough that your version is the one with a year of production feedback when everyone else's becomes viable.

Sources: Anthropic — Claude Opus 5.5 · OpenAI — Introducing GPT-6 Sol and Luna · Epoch AI — The plunging price of thought · Simon Willison — Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war

🌊 Waves

External evaluation stops being a courtesy and starts being architecture

OpenAI published "Priorities and principles for effective third party assessments," proposing that outside groups run technical safety assessments during training and evaluation rather than only in the days before launch. It names four priority areas — independent assessment of safety cases, assessment of critical safeguards, review of the capability evaluations behind Preparedness categories, and independent investigation of critical misalignment incidents, citing the OpenAI–Hugging Face incident by name — and seven operating principles covering pre-registered scope, proportionate access, conflict-of-interest disclosure, remediation time before publication and editorial independence for the evaluator. No partner is named and no access terms are set. The same day, METR published its predeployment evaluation of Claude Opus 5.5: ten business days of API access, five tasks, an unpaid agreement, and text Anthropic was allowed to review and edit. Its conclusion was that Opus 5.5 would accelerate AI R&D slightly more than Fable 5.1 but is unlikely to automate it, with persistent weakness in research judgement. A separate METR team with elevated access estimated roughly 1.5x overall acceleration in capabilities due to AI, with perhaps a 30% chance of 2x.

Roadmap implication: The audit market the Signal has been tracking since July just got its reference implementation, and the roadmap implication is a procurement one. Within a year, "which independent group evaluated this model, with what access, under what publication terms" will be a line in enterprise diligence the way SOC 2 is now — and the labs are writing that standard in public right now, while it is still cheap to influence. If you buy frontier capability at any scale, read the METR report and the OpenAI principles side by side and decide what you would require in your own contracts. The 1.5x number is the other half: it is the first credible quantification of AI-on-AI-R&D acceleration from a party that does not benefit from it being large, and it is the figure to plan capability timelines against rather than the vendor narratives on either extreme.

Sources: OpenAI — Priorities and principles for effective third party assessments · METR — Summary of METR's predeployment evaluation of Claude Opus 5.5

Alibaba puts the whole stack on one roadmap: 20GW, its own chip, and Qwen 4 in training

At the Apsara Conference in Hangzhou, Alibaba Group CEO Eddie Wu laid out a full-stack plan: Alibaba Cloud global datacenter capacity to exceed 20GW by 2032; a new T-Head training and inference chip, the Zhenwu V900, with 216GB of memory, 1,200 GB/s inter-chip bandwidth, native FP32-through-FP4 support and clusters of up to 500,000 chips, which Wu called the most powerful AI chip in China today; two Yitian server CPUs for agentic workloads in 2027; and a three-layer agent cloud. TechNode reports mass production in Q1 2027, though The Register notes Wu gave no production date on stage. On models, Qwen 4 is in training, with Qwen 4.5 and Qwen 5 sketched at 5 trillion and 10 trillion parameters. Alibaba also claimed Qwen3.8-Max completed 33 automated self-improvement cycles in a month, raising its Artificial Analysis score from 40 to 45, and in a separate chip-design run made over 10,000 EDA tool calls across 60-plus hours to cut chip area 42% with no performance loss. All of those are Alibaba-reported figures with no independent benchmark behind them. For scale on the headline number, The Register points out that Cushman & Wakefield counts 37.7GW of datacenter capacity under construction in the United States alone.

Roadmap implication: Treat this as a vendor-selection input with a two-year fuse rather than a headline. A Chinese hyperscaler that credibly owns models, silicon, cloud and an agent platform end to end is the first non-US organisation that can offer a full stack at a price set by its own cost structure, and Qwen already anchors the open-weight ecosystem most teams quietly build on. The planning move is to know, now, which of your workloads could run on open Chinese weights if the economics moved another 10x, which of them could never leave your jurisdiction, and what the switching cost is between those two buckets. Do that mapping while it is a strategy exercise. And hold the self-improvement claims loosely — 33 cycles for five points on one index is a real, modest, self-reported result, not the knee of anything, and the chip-area number is the more interesting datum precisely because it is a concrete engineering task with a measurable outcome.

Sources: The Register — Alibaba Cloud plans six-year stroll to 20GW of datacenters, reveals chip to power them · TechNode Global — Alibaba targets 20GW cloud capacity by 2032 in full-stack AI push

The distillation fight becomes a regulatory matter — in Beijing, on Beijing's own grounds

The Information reported that China's Cyberspace Administration has summoned all seven labs named in Anthropic's September 10 threat intelligence report and has narrowed its inquiry to DeepSeek and Moonshot AI, questioning staff at both. Anthropic's report alleged that seven China-based labs relayed roughly 190 million customer exchanges through Claude between December 2025 and August 2026 to harvest outputs — Alibaba the largest at over 151 million, Moonshot about 23 million, DeepSeek over 12.1 million inside fourteen days in July. The detail that reportedly moved the regulator was not the totals: it was a specific case in which a user Anthropic assessed as likely PLA-affiliated asked Moonshot's Kimi to analyse surveillance footage tracking a person across hundreds of police cameras in Chengdu, and Moonshot allegedly forwarded the request and the footage to Claude without telling the user. The probe is open, no penalty has been decided, and neither company has commented. Note what Beijing's complaint actually is: not output theft, but Chinese user data reaching a US company's servers. On September 9 China had publicly rejected a US intelligence advisory on distillation as unfounded.

Roadmap implication: The useful read here is not the geopolitics — it is that silent subprocessing is now a named, regulated risk on both sides of the Pacific, which makes it a solvable one. Any team routing user requests through a model API inherits the provenance question: where does the prompt physically go, which providers see it, and can you prove it. That is an answerable engineering question with a logging and contract answer, and the teams that can answer it in writing will win enterprise deals that the teams who cannot will not even be invited to bid on. Put provider-routing disclosure in your own product terms before a regulator writes it for you, and ask your vendors for theirs this quarter.

Sources: The Next Web — China is investigating DeepSeek and Moonshot, The Information reports · Anthropic — Countering misuse of AI: September 2026

Amazon blocks Meta's Muse agent, and agent access to commerce becomes a negotiation

Ben Thompson's Tuesday Stratechery update covered Amazon blocking Meta's Muse agent from its storefront, arguing that the block was predictable, that Amazon's physical-world investment is itself the AI moat, and that there is still room for a deal. The Information's briefing the same week framed the same fight as an aggregator-versus-aggregator brawl and reported that OpenAI is developing features to counter Grok Bot while weighing its own response to Muse. The pattern across all three: the surfaces agents want to act on — carts, catalogues, checkout, logistics — belong to incumbents whose moat is not software, and those incumbents have now started saying no explicitly rather than tolerating agent traffic by default.

Roadmap implication: If your product assumes an agent can reach a third party's commerce surface, that assumption has a counterparty and the counterparty has just been observed exercising a veto. Roadmap implication: treat agent access to any platform you do not own as a partnership to be negotiated on a timeline, not a technical integration to be scheduled, and price the possibility that the answer is no. The corollary is the opportunity — the categories where nobody owns the surface, or where the owner has an incentive to invite agents in, are wide open right now and far less contested than the ones everyone is fighting over.

Sources: Stratechery — Amazon Blocks Muse, Amazon's Moat, Aggregator v Aggregator

🌊 Ripples

Opus 5.5 leads on agentic benchmarks — and its top effort setting can fail to return anything

Anthropic's published numbers put Opus 5.5 at 66.4% on Terminal-Bench 4.0 against 55.8% for Fable 5.1, 52.3% for Opus 5 and 57.9% for GPT-6 Astra, and at 1846 Elo on GDPval-AA v2.1. It claims a 680,000-line migration completed in under a day by a tester, and a 200,000-line codebase audit-and-fix in under three hours against more than twenty for Opus 5. Anthropic also added its own caveat that "benchmark margins have become a less reliable guide to real-world differences," and flagged that the model often appears to suspect it is being evaluated. Simon Willison's hands-on test found the practical edge: at "max" thinking the model failed to return a response at all on his standard test, twice, burning its 128,000-token output cap while still reasoning — each failure costing $2.56 and taking nearly twenty minutes. His conclusion was that max is "effectively useless." Thinking mode can no longer be disabled; Sonnet 5.5 and Haiku 5.5 are promised in the coming weeks.

Do this now: Do this now: if you run Claude in production, move your default to Opus 5.5 for the 20% price cut and 60% cheaper cache reads, but cap effort at medium and add a timeout-and-fallback path before you touch max. Anthropic's own default-effort numbers are the ones to plan against — medium beats GPT-6 Astra's best on FrontierCode at roughly a fifth of the cost per task — and the max tier is where your cost and latency tail lives.

Sources: Anthropic — Claude Opus 5.5 · Simon Willison — Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war

GPT-6 Sol and Luna ship at half price, with caching changes that move the real bill

OpenAI put GPT-6 Sol at $2/$10 and GPT-6 Luna at $0.10/$0.50 per million tokens — Sol exactly half its GPT-5.6 counterpart, Luna half on input and 58% cheaper on output — with cached input reads discounted 90%. The caching changes are the underreported part: higher default hit rates, explicit cache breakpoints, and — significantly — reasoning-effort and tool-toggle changes no longer invalidate the cache. OpenAI says GitHub saw fresh-processed prompt tokens fall by more than half. On capability, Sol at extra-high effort posts 33.2% on AutomationBench 1.0.6 at $0.27 per task against Astra at low effort scoring 30.3% for 3.9x the cost, and 68.8% on DeepSWE v1.1 within 1.1 points of Fable 5's 69.9% at roughly 80% lower cost per task. Artificial Analysis's read was more measured: the cost-efficiency frontier halved, while the Intelligence and Coding Agent indices came in roughly level with GPT-5.6, with gains in some evals and regressions in others. Astra remains OpenAI's strongest model. Willison also notes GPT-5.6 has a scheduled 25% price increase in November.

Do this now: Do this now: re-run your cost-per-completed-task baseline rather than your cost-per-token one — Artificial Analysis's split verdict is the tell that per-token price cuts and per-task value have decoupled again. If you are still on GPT-5.6, the November increase makes migration a dated decision rather than an open one. And restructure prompts around the new cache breakpoints before you benchmark anything; on agentic workloads that change is likely worth more than the headline cut.

Sources: OpenAI — Introducing GPT-6 Sol and Luna · VentureBeat — OpenAI releases GPT-6 Sol and Luna models, slashing API costs 50% or more

Malware that asks four different model APIs what to do next

The Register reported that Cisco Talos had documented a Windows malware family, CLOSEDQUORUM, that queries up to four commercial LLM providers — Google Gemini, DeepSeek, Qwen and Mistral — to autonomously select its post-compromise actions, including credential theft and crypto-wallet theft. The multi-provider design is the notable engineering detail: it is resilient to any single provider cutting off abusive accounts, and it means the malicious decision-making happens outside the binary, where signature-based tooling cannot see it.

Do this now: Do this now: add outbound calls to commercial model APIs from endpoints and servers that have no business making them to your egress monitoring and alerting rules. This is a cheap detection with an unusually clean signal — a finance workstation talking to four different inference providers is not a false positive — and it is the control that generalises, because the next family will pick different providers.

Sources: The Register — Windows CLOSEDQUORUM malware uses AI models to autonomously select post-compromise actions

Tencent prices image generation at two and a half cents

Tencent released Hy Image 3.5 Preview, timed to open on the same day as Alibaba's Apsara conference, at $0.024 per image on its international API — against roughly $0.067 for Google's Nano Banana 2 and $0.134 for Nano Banana Pro — with reference images and failed generations billed at nothing. Reference-image input rises from Hy Image 3.0's limit of 3 — to 5 per Tencent's launch announcement, or up to 20 per its international API guide, which disagree — output goes direct to 4096x4096, and multi-round continuous editing is supported. The quality claim is an internal blind test by several hundred in-house designers placing it on par with ByteDance's Seedream 5.0 Pro and slightly ahead of Nano Banana Pro; no independent benchmark has scored it and no technical report, parameter count or weights have been published, so this is a product and an API, not an open-weight release. Tencent closed up 5.02% in Hong Kong on the day, having traded as much as 6.6% higher intraday.

Do this now: Do this now: if image generation is a line item rather than a demo in your product, re-baseline it — a 3-5x unit cost difference against the Google tiers changes which features are worth shipping, and free failed generations changes how aggressively you can retry. Run your own eval set before switching; the only quality evidence on offer is the vendor's own designers, which is exactly the kind of benchmark to treat as a hypothesis.

Sources: Yahoo Finance / Bloomberg — Tencent Releases AI Image Model to Catch ByteDance, Alibaba

AI cracked the last sporadic group — and lost the 25,000-case race to two humans

Scientific American reported that a six-person team assembled at an American Institute of Mathematics meeting at Caltech in May — Rachel Pries, Bjorn Poonen, Xiaoyu Huang, Blake Jackson, Kyu-Hwan Lee and Shaowu Zhang — solved the inverse Galois problem for the Mathieu group M23 in under three months, producing an explicit degree-23 polynomial family. M23 was the last of the 26 sporadic groups without a known polynomial; the other 25 were found in the 1980s. AI was decisive: it combed M23 for symmetry combinations, and when numerical approximation stalled at 90 digits against a memory limit, a raft of AI agents was set on new coordinate choices, most of which failed and one of which worked — the team called it miraculous. Kyu-Hwan Lee: "We could do it very efficiently. That wasn't really possible five years ago." The aggregators ran a different story. A separate SAIR Foundation open competition to find polynomials for all 25,000 groups acting on 24 roots closed its first phase in late August with every case realised — and the winners were two German mathematicians whose only use of AI was writing an upload script. Organiser Jen Paulhus: "It was open to AI, and it was still these folks who did the best."

Do this now: Do this now: use this pair as your internal calibration story, because it is the cleanest available illustration of where the frontier actually is. AI broke a forty-year holdout that humans could not, and simultaneously lost an open competition to two people with a shell script. Both are true, and any org policy built on only one of them will be wrong. The operator lesson is that the wins come from pointing AI at the specific bottleneck a human has already localised — not from handing it the whole problem and standing back.

Sources: Scientific American — Mathematicians use AI to find mysterious symmetries, solving decades-old problem · SAIR Foundation — Inverse Galois Problem competition discoveries


Read this and every past edition at excelsiorgroup.ai/insights/signal.

Don't miss what's next. Subscribe to The Signal - daily AI evolution :
← Newer The Signal: Human Advancement — Edition #6 — September 24, 2026 Older → The Signal: Bio/Health — Edition #6 — September 22, 2026
LinkedIn
excelsiorgroup.ai
Powered by Buttondown, the easiest way to start and grow your newsletter.