The Signal - daily AI evolution  logo

The Signal - daily AI evolution

Archives
Log in
Subscribe
September 12, 2026

The Signal — September 12, 2026

The Read

Twenty-five Fields Medallists — from Pierre Deligne in 1978 to Yu Deng this year — signed a declaration saying the AI labs' use of famous open problems as benchmarks is "severely misaligned" with what mathematics is actually for. It is not a Luddite document. It says plainly that AI can accelerate genuine mathematical understanding, and that whether it does "will in large part be determined by the decisions of the humans in control of this new technology." That is the day-zero argument, stated by the profession being re-founded first and worth reading by whoever is next: the output of the work and the point of the work are not the same object, and a system optimised for the first can quietly dismantle the second. Washington spent the same day arguing a cruder version of the same question, with Senate negotiators drafting a duty of care that could gate model releases and a President who says his only AI concern is China. And Cognition published the most immediately useful number of the day — a two-model harness that cuts frontier-grade coding cost by up to 39%, with the counterintuitive finding that using a more expensive model made the whole system cheaper.


🌊 Tide

No shift. All four tides hold, with one confirmation on governance-as-market-structure — and it is the first confirmation where the mechanism under discussion is a gate in front of the release rather than a penalty after it. Every prior confirmation this year has been liability applied backwards: courts, attorneys general, disclosure regimes, procurement conditions. A pre-release block turns safety evidence into a market-access asset, which is a different kind of object entirely.

Washington started drafting the gate, not the fine

Reuters reported that Senate negotiators are debating legislation creating a "duty of care" for AI developers — an obligation to design products with the goal of preventing "catastrophic risks" — and aiming to give the federal government power to block the release of models deemed unsafe, with companies able to challenge that decision in federal court. Part of the measure would also preempt state enforcement on those same risks. The talks run through Majority Leader John Thune, Commerce chair Ted Cruz and Amy Klobuchar as the lead Democrat, with Maria Cantwell weighing in; Cruz has framed the scope on X as "catastrophic risks involving biological or nuclear threats," and Klobuchar's statement to Reuters adds "requiring developers to work with government experts to verify and test models." It would apply only to the most capable models; Reuters names Google, Anthropic and OpenAI among the US companies that fit. The calendar is the binding constraint: the House sits one week before the November 3 midterms, the Senate three. Pointing the other way on the same day, the President was reported as dismissing the premise — asked on Thursday whether he shared frontier researchers' extinction concerns, Trump said "No, I don't have any," and put the US lead over China at "a year, which is, you know, considered a lot."

So what: Here is the opening, and it is a procurement opening rather than a policy one. If release timing becomes a regulated variable, then the ability to swap a model out becomes an operational asset with a price — which is the cheapest argument yet for the multi-model harness in today's wave section. Do this now: for each frontier model in your stack, write down what breaks if it is unavailable for a quarter, and how long the substitution takes. Teams that can answer in hours rather than weeks will be the ones that keep shipping through whatever this becomes. Note also what is not being proposed: nothing here constrains capability research, only release, and preemption of state law is a real deregulatory trade sitting inside a safety bill.

  • The Spokesman-Review — U.S. Senate negotiators consider requiring AI firms to mitigate known major risks
  • Gizmodo — Trump Shrugs Off Warnings That AI Could Cause Human Extinction

🌊 Waves

The expensive model made the system cheaper

Cognition brought Fusion — its two-model harness — to Devin Desktop and CLI, and published the economics with it. On the Artificial Analysis Coding Agent Index v1.5, Claude Code running Fable 5.1 scores 62.2 at $12.36; Fusion pairing Fable 5.1 with Cognition's own SWE-2 as sidekick scores 61.7 at $7.90, 36% cheaper. Codex on Astra scores 61.6 at $7.47; Fusion pairing Astra with SWE-2 scores 58.9 at $4.54, 39% cheaper. The architecture is the argument. Rather than routing a task to a cheap model up front and hoping the guess was right, a frontier "lead" agent owns the plan, the ambiguity and the review, and delegates implementation to a "sidekick" — each keeping its own persistent context, so neither breaks the other's prompt cache, and the lead can take the work back when the sidekick is out of its depth. The finding worth reading twice is the inversion. Swapping a cheap sidekick (GPT-5.6 Luna at $0.20 per million tokens) for one costing 275% more (SWE-2 at $0.75) scored higher and cost 2% less, because a stronger sidekick needs fewer correction rounds from the lead. Cognition reports the same effect at the lead position: replacing Opus 4.8 with Fable 5, nominally twice the per-token price, produced sessions averaging 9% cheaper.

So what: Roadmap implication: retire price-per-token from vendor comparisons this quarter and replace it with price-per-completed-task measured on your own work. Cognition's own framing — that models and model-harness combinations "should be evaluated on price per task rather than price per token" — is now backed by a table in which the cheaper input loses on both axes. This is a harness vendor publishing benchmarks for its own harness, so discount the absolute scores accordingly; the direction is the part you can test on your own repository in an afternoon, and it is the part that changes the buying decision.

  • Cognition — Introducing Fusion in Devin Desktop & CLI

The mathematicians wrote down the terms

Twenty-five Fields Medallists published a joint declaration arguing that the push by AI companies to solve mathematical problems as a benchmark "is detrimental to the science of mathematics, and to the mathematical community," and that "the goals of the AI companies and the goals of the mathematical community are severely misaligned." Terence Tao posted it on his blog and at mathandai.org, inviting further signatures; the signatories run from Pierre Deligne (1978) to Yu Deng (2026) and include Scholze, Viazovska, Kontsevich, Villani, Maynard, Bhargava, Huh and Hairer. The argument is structural, not defensive. Famous problems have served as "landmarks and lighthouses" whose real yield was the new methods a community then spent years studying, simplifying and teaching down to textbook form; solutions "announced in a rush" leave no time for a proper writeup, the isolation of new ideas, or citation of prior work, raising "severe attribution and plagiarism questions" — and "the mass production at faster and faster pace of 'true/false' statements could destroy fertile ground instead of breathing life into new ideas." Tao conceded the process was compressed, with no consultative round of the kind behind the Leiden declaration, because "the urgency of the situation was such that we needed to release a statement sooner rather than later." He published a companion guest post the same day on the current state of the Hodge conjecture, framed around what the field would actually gain or lose if a lab pointed vast resources at it.

So what: Roadmap implication, and the declaration makes the generalisation itself: "the issues the mathematical community faces now are similar to issues that other scientific and creative professions are facing." The operator's version is a design question rather than a policy one. When you automate the output of a skilled function, name explicitly which parts of that function were producing something other than output — training the next cohort, building judgement, holding institutional memory — and decide who owns them now. The optimistic reading is the one the medallists themselves offer: they say AI can enhance and accelerate genuine mathematical understanding, and that the outcome turns on decisions made by humans. That is an invitation to design, not a veto, and the professions that write their own terms early will get better terms than the ones that wait.

  • What's new — A Severe Misalignment of AI in Mathematics
  • What's new — On the Hodge conjecture

Nvidia's off-balance-sheet book nearly tripled in one quarter

SemiAnalysis walked Nvidia's 2Q F1/27 10-Q and totalled $530B of gross off-balance-sheet guarantees across six line items, up from $184B disclosed the prior quarter. Three moves did most of it: supply and capacity commitments went $119B to $279B, primarily memory, with 96% due by F1/29; guarantees and land-power-shell guarantees went $3.5B to $108.5B, almost entirely credit support on roughly 4.25GW at SB Energy's PORTS-Pike campus in Ohio, capped at $105B against twenty-year leases to an OpenAI affiliate; and two line items appeared for the first time — $36B of AI cloud agreements and $20B of datacenter leases Nvidia signed as tenant and expects to reassign. That sits against $91B of on-balance-sheet liabilities. SemiAnalysis's read cuts against the circular-financing headlines: the roughly 6.5GW Nvidia currently backstops is just under 3% of the ~240GW of net global AI capacity growth its datacenter model forecasts between end-2025 and end-2030, against Gigascaler leases of more than 35GW by 2028. Its conclusion is that Nvidia should be doing more of this, because the marginal cost of a guarantee in the base case is zero and each one converts a B-rated operator into a financeable one. The constraint it flags is concentration rather than leverage: OpenAI is 97% of the guarantee cap.

So what: Roadmap implication: this is the machinery behind the compute price you are being quoted. Under the AI Cloud Partner programme, Nvidia floors GB300 rental at roughly $2.35 per GPU-hour on average against a $4.50–$4.60 range for five-year contracts — which means new capacity is arriving underwritten by somebody else's credit rating, and the spread is available to operators who can commit to a term they will actually fill. If you are signing compute in the next two quarters, ask who the ultimate credit behind your provider is, and price the answer. And watch the concentration line: SemiAnalysis puts OpenAI at 97% of the guarantee cap, which is the single number that would turn a demand wobble into a correlated one.

  • SemiAnalysis — Nvidia's Backstop Universe – Heads I Win, Tails Who Loses?

The open-weight business model finally has a revenue line

Bloomberg reported that Moonshot AI is targeting $2B in annualised revenue by the end of 2026, roughly double its August run rate. TechCrunch, citing that report, set it against OpenAI at roughly $40B and Anthropic at roughly $65B, and noted OpenRouter data showing as much as 300 billion tokens a day generated by K3 models on that platform — while flagging that Moonshot's margins are structurally far lower, because the weights are free. The point is not the absolute size. It is that a lab giving its weights away has a P&L that compounds, which is the question hanging over open weights since the beginning. Nathan Lambert published an updated open-models reading list the same day putting the open-closed capability gap at "roughly 4-6 months" and stating that "the leading open models have all come from Chinese labs since ~2024," while walking back his own earlier confidence that DeepSeek had not distilled o1 traces. The counterweight came from John Schulman on Dwarkesh Patel's episode that day, warning that post-RL output diversity is falling and that "so many people are distilling, mostly from Claude, that all the open-weight models write the same way as Claude and have the same tics."

So what: Roadmap implication: a four-to-six-month capability lag with a real business behind it is a procurement fact, not a hobbyist one — it means the open tier will keep shipping, keep being supported, and keep being a credible second source rather than a science project you inherit. Pair it with today's Fusion result and the shape of the 2027 stack is visible: frontier model in the lead position where judgement matters, open weights in the sidekick position where volume does. The monoculture warning is the thing to actually monitor — if every open model has been trained toward the same house style, then a diverse portfolio on paper may be a single point of failure in behaviour. Test your fallbacks on disagreement, not just on benchmarks.

  • TechCrunch — Kimi-maker Moonshot AI targets $2 billion in annual revenue
  • Interconnects — Open-Source AI & Open Models Reading List
  • Dwarkesh Podcast — AI researchers debate how close we are to recursive self-improvement

🌊 Ripples

Two engineers and Codex rewrote OpenAI's storage layer in Rust

OpenAI published the first part of an engineering account of Habitat, the online storage platform behind ChatGPT: more than 70 million requests per second, over 500 petabytes, across almost 40 geographic regions, serving more than a billion weekly users after growing more than 10x year over year for three consecutive years. The detail worth the read is the migration. In Q2 2026 two engineers, working with Codex and GPT-5.5, rewrote the entire service from Python into Rust; it now serves 95% of production and is 6x more CPU efficient and 15x more memory efficient. The Python version had peaked above 20 million requests per second. The post is also honest about the unglamorous failure modes that cost them most — a feature-flag config being JSON-parsed every 60 seconds by eight processes per pod, and a metastable failure traced to LIFO connection reuse in aiohttp, fixed by patching it to FIFO.

So what: Do this now: pull out the one service in your estate that everybody agrees should be rewritten and nobody has staffed, and scope it as a two-person-plus-agent job rather than a team-quarter. The reason that rewrite has never happened is that the labour cost exceeded the efficiency gain — and that arithmetic just changed by roughly an order of magnitude. This is the day-zero thesis in its least glamorous and most bankable form: not a new product, an old decision that was correct under the old prices and is wrong under the new ones.

  • OpenAI — Rapidly scaling online storage to serve over 1 billion ChatGPT users

Frontier models became a standing part of a maintainer's security process

Simon Willison and Alex Garcia shipped security releases of Datasette (1.0a39 and 0.65.4) after what Willison describes as "an extensive audit of Datasette using Claude Fable 5.1, GPT-5.6, and GPT-6 Astra," followed by nearly a week of collaborative fixing. The models, he writes, "helped find some very subtle bugs." The human protocol around them is the transferable part: for each issue, one human wrote the failing test and the other implemented the fix, so two people had eyes on every issue in addition to coding agents running different models. His conclusion is the headline for anyone maintaining software: "We'll be incorporating security audits by frontier models into all of our development work going forward." The fixes matter most for public Datasette instances that mix public and private tables.

So what: Do this now: run two different frontier models over your highest-exposure codebase this week and treat the output as a triage queue, not a verdict. The cost of a second and third independent reader of your code has collapsed, and Willison — who is about as resistant to AI hype as anyone still using the tools daily — has moved this from experiment to standing practice. Keep his protocol: a human writes the failing test, a different human writes the fix. The models widen the funnel; they do not close the loop.

  • Simon Willison's Weblog — Datasette 1.0a39 and 0.65.4 security releases

Enflame tripled on debut — and still fabs its chips abroad

Shanghai Enflame Technology surged on its first trading day after raising about 6.12 billion yuan (roughly $911 million) by selling around 43 million shares. Shares opened at 410 yuan against an IPO price of 142.18; reports of where the day closed differ across outlets, ranging from roughly +179% to +206%, and retail demand was oversubscribed thousands of times over — though first-day doubles and triples are common on mainland exchanges for structural reasons. Tencent owns about 20% and is the dominant customer at 84% of revenue last year. The caveat is in Enflame's own prospectus: it still produces its advanced accelerators at overseas foundries, because domestic manufacturers cannot yet deliver them consistently at the volume and quality it needs. The company remains unprofitable and expects another loss for the first nine months of 2026; revenue was 990 million yuan in 2025, up from 722 million the year before.

So what: Do this now, if you are modelling China compute: separate design self-sufficiency from manufacturing self-sufficiency in your assumptions, because this listing demonstrates the first and undercuts the second in the same document. The demand signal is real and the capital is enthusiastic; the fab constraint is the variable that actually governs timelines. The oversubscription tells you what Chinese capital believes about domestic silicon — the prospectus tells you where the bottleneck still is.

  • The Information — Tencent-backed Enflame Triples on Debut as China Pushes for Chip Self-sufficiency
  • South China Morning Post — Enflame shares soar 188% on Shanghai debut as Nvidia challenger taps investor fever for AI

One line of the prompt moved an agent's harm rate by nearly 84 points

CaML and the University of Warwick released HarvestBench, a tractor-farm simulation that measures whether an LLM agent will spend fuel to avoid running over animals. Kill rates across nine models spread enormously — GPT-5.6 Terra at 0.4%, Sol at 0.9%, GPT-5-mini at 5.4%, DeepSeek V3.1 at 2.4%, Claude Haiku 4.5 at 4.5%, Sonnet 5 at 17.8%, Gemini 2.5 Flash at 38.7%, Mistral Small 3.2 at 88.8%, GPT-4o mini at 98.8%. The number that generalises beyond the subject matter: removing the morality line from the prompt took Sol from 0.9% to 84.6%. The researchers also found models kill wild animals more readily than farmed ones — reasoning, in effect, about each animal's worth to the farmer rather than about the animal.

So what: Do this now: treat every constraint sentence in your agent system prompts as load-bearing infrastructure and put them under version control with tests, because this is an ablation study showing an 83.7-point behavioural swing from deleting a single line. The subject here is livestock; the mechanism is any agent making cost-versus-harm tradeoffs on your behalf — collections, moderation, scheduling, pricing. The optimistic read is that the fix is cheap and available today: the guardrail worked, it was just undocumented and unprotected.

  • The Register — AI more likely to kill animals if it saves fuel or money

Google closed the Mechanize deal without filing anything

Google completed its talent-and-licence deal for Mechanize, the startup building reinforcement-learning environments and evaluations that train agents to code and do complex computer work. Co-founder and former CEO Tamay Besiroglu — who previously co-founded Epoch AI — is now a research scientist at Google DeepMind, and Business Insider reports more than a dozen former Mechanize employees have joined Google, most working on midtraining. Business Insider reported the deal's value in August at over $1.5 billion; Mechanize had raised $9.1 million at a $500 million valuation. The structure is the familiar one: a non-exclusive licence plus hiring, no Hart-Scott-Rodino filing, with former chief of staff Guive Assadi now listing himself as CEO of what remains. Google paid $2.4 billion for the equivalent Windsurf arrangement.

So what: Do this now: if you own proprietary process data — real workflows, real failure traces, real environments where work either succeeds or does not — inventory it this month and price it. What Google bought here is not a product and barely a team; it is the capacity to manufacture environments in which agents can be graded. That is the scarce input in the current training regime, and the buyers have now shown their pricing twice. The licence-and-hire structure is also worth noting as the standing template: it is how frontier capability is being consolidated without a merger filing, and it will keep working until someone rules that it does not.

  • Benzinga — Google Boosts AI Coding Armory by Completing Mechanize AI Talent Deal, Former CEO Joins DeepMind

Read this edition and the full archive at excelsiorgroup.ai/insights/signal.

The Signal — The Excelsior Group

Don't miss what's next. Subscribe to The Signal - daily AI evolution :
← Newer The Signal — September 13, 2026 Older → The Signal — September 11, 2026
LinkedIn
excelsiorgroup.ai
Powered by Buttondown, the easiest way to start and grow your newsletter.