The Signal - daily AI evolution  logo

The Signal - daily AI evolution

Archives
Log in
Subscribe
October 7, 2026

The Signal — October 7, 2026

The Read

The cheap end of the AI market went to the capital markets, and the capital markets said yes. DeepSeek is closing on at least $12 billion in a round backed by Tencent and CATL ahead of an early-2027 IPO; Moonshot sealed its final private round at roughly $50 billion on the way to a Hong Kong listing in the first quarter of 2027 that could raise up to $5 billion; Kuaishou's Kling unit picked banks for a Hong Kong float of at least $1 billion; and Mistral shipped a trillion-parameter open-weight model trained on about 3,800 Grace Blackwell GPUs in European datacenters, weights due on Hugging Face inside the month. Epoch AI published the arithmetic underneath all of it the same day: China's six leading AI firms together earn roughly 10% of OpenAI and Anthropic's combined AI revenue, and one of them watched its share of OpenRouter tokens fall from 88% to 22% in twenty days after it released one of its models' weights. That is not a contradiction, it is the mechanism. The floor under the price of intelligence is being held up by equity rather than by margin, which is exactly why it keeps falling — and the operator's job is not to route tokens more cheaply but to go find the old problems that were never worth solving at the old price and are obviously worth solving at this one.


🌊 Tide

No shift. One confirmation of cost-collapse, from a direction this brief has not logged before: the capital markets. Every prior confirmation came from price cuts, inference engineering, in-house silicon, or demand curves — supply-side economics or buyer behaviour. This one is financing. The labs pushing the price of frontier-adjacent capability down are being funded by equity and, within two quarters, by public markets, on revenue roughly a tenth the size of the US leaders'. A price floor supported by permanent capital rather than by margin does not need to rise. governance-as-market-structure is also confirmed — a municipal legislature issued a subpoena to a frontier lab and is pursuing legal action over the no-show — but that evidence is carried in the waves rather than promoted here, because the mechanism is not new, only the jurisdiction. distribution-rewrite held without new evidence strong enough to log.

Cost-collapse, financing side: the suppliers of cheap intelligence just got permanent capital, and the arithmetic of what they earn got published the same day

Four financing events and one measurement landed inside a single day. Bloomberg reported DeepSeek is close to raising at least 80 billion yuan (about $12 billion), potentially approaching $15 billion, in a round backed by Tencent and CATL, targeting an IPO in early 2027; that Moonshot AI, which makes the Kimi models, closed its final private round at a valuation of roughly $50 billion and is heading for a Hong Kong IPO in the first quarter of 2027 that could raise up to $5 billion; and that Kling AI, spun out of Kuaishou, has hired CICC, Goldman Sachs and UBS for a Hong Kong listing of at least $1 billion, after raising nearly $3 billion in July at a $15 billion pre-money valuation. Mistral released Mistral Large 4, a natively multimodal mixture-of-experts model with 1 trillion total and 49 billion active parameters, trained on roughly 3,800 NVIDIA Grace Blackwell GPUs — about 52 NVL72 racks — in Mistral's own European datacenters, priced at $1.36 per million input tokens and $4.18 per million output, covering more than 160 languages, with weights promised on Hugging Face before the end of October. Against all of that, Epoch AI published annualized AI revenue estimates drawn from months ranging from June to September: Anthropic $65 billion and OpenAI $40 billion, against ByteDance $4.0 billion, Alibaba $2.4 billion, Z.ai $1.8 billion, Moonshot $1.0 billion, DeepSeek $1.0 billion and MiniMax $0.8 billion — China's six leading firms at about 10% of the two US leaders combined. These are estimates, not audited accounts, and Epoch's Anthropic and OpenAI figures are its own annualizations.

So what: Read the two halves together and the conclusion is bullish, not bearish. The companies holding the price floor down are not funded by the margin on cheap tokens — they are funded by investors buying the adoption curve, and in two quarters several of them will be funded by public shareholders. That makes the cheap tier structurally durable rather than a promotional phase that ends when someone needs gross profit. For a buyer, it means you can plan multi-year on sub-$2 input pricing from more than one credible supplier, and that a European-domiciled open-weight option at frontier-adjacent quality is now part of the vendor set, not a hypothetical. For a builder, it means the input cost line in your model is a falling line with permanent capital behind it, so the question to answer this quarter is which problem in your industry was priced out at the old number. That is the whole thesis: the opportunity is not in the routing, it is in the re-founding.

Epoch AI — How do Chinese AI companies make money? · Quartz — DeepSeek is raising at least $12 billion in a Tencent-backed funding round · Bloomberg — China's AI Champions Raise Billions Ahead of IPOs · Mistral AI — Introducing Mistral Large 4


🌊 Waves

The open-weight business model finally got measured, and it has exactly one profitable shape

Epoch AI's report is the most detailed public attempt yet to take apart how Chinese AI companies actually earn, and the margin data is more interesting than the revenue data. MiniMax's AI-native consumer apps ran a 4.7% gross margin over the first nine months of 2025 while its Open Platform and enterprise services ran 69.4% over the same period. Z.ai's API and developer platform margin was 18.9% in 2025, improving to 24.6% in the first half of 2026, while its on-premises deployment business — 73.7% of 2025 revenue at a 48.8% margin — fell to 13.5% of revenue by the first half of this year. DeepSeek's V4 API is estimated at a 70-80% gross margin. The sharpest single datapoint is self-inflicted: after Z.ai released the GLM 5.3 Flash weights on 26 August, its share of OpenRouter tokens collapsed from 88%, measured about 33 hours after launch, to 22% within twenty days, as third-party hosts served the same model more cheaply. ByteDance's Volcano Engine, meanwhile, took 49.5% of China's public-cloud model tokens in 2025 on roughly $221 million of model-as-a-service revenue and is targeting $4 billion in 2026.

Roadmap implication: Three roadmap implications, and the first one is the one most Western teams have backwards. Giving away weights is a distribution strategy that destroys your own inference margin on purpose — Z.ai's 88-to-22 collapse is the price of the adoption, not a failure of execution — so when you evaluate an open-weight vendor, price the relationship on the assumption that someone else will host that model cheaper than they will, and plan to be that someone or to buy from them. Second, the profitable line in every one of these businesses is the enterprise and platform side, not the consumer app: a 4.7% consumer gross margin next to a 69.4% platform margin tells you where the vendors will put their roadmap attention, and it will not be the chat subscription. Third, on-prem is shrinking as a share even where it is high-margin, which means the sovereign-deployment pitch is converging on hosted-with-controls rather than shipped-to-your-rack — worth knowing before you architect a data-residency programme around appliances.

Epoch AI — How do Chinese AI companies make money?

Offensive cyber capability became a credentialled product tier, on the same day the monitoring layer was shown to be attack surface

Anthropic expanded its Cyber Verification Program into three named tiers with different blocking behaviour. Defense Access covers security operations, incident response, malware reverse-engineering and vulnerability analysis, and is open to corporate security teams, universities, government bodies, critical-infrastructure operators, open-source maintainers and individual researchers with documented vulnerability reports, with review in a few days. Red Team Access adds authorized penetration testing, is restricted to organizations rather than individuals, takes several weeks to review, and keeps real-time blocks on actions risking physical harm or mass disruption. Specialized Access carries minimal cyber blocks and is reserved for verified organizations testing safety systems for flight operating systems, power grids, telecom networks, interbank transfer infrastructure and government administrative networks. All tiers reach Claude Opus 5.5, Claude Sonnet 5.5 and Claude Mythos 5.1. The published CyScenarioBench numbers make the tiering legible: across 10 challenges at five attempts each, general availability blocked every task on the first prompt, Defense Access blocked 46 of 50 trials, and Red Team Access blocked none and completed 34 of 50. Anthropic also reported that Project Glasswing partners identified 129,000 verified software vulnerabilities between April and July, its own open-source scanning found a further 5,500 between April and October, over 33,000 were rated critical or high-severity, and partners including Booz Allen and Comcast said Claude Mythos pulled discovery timelines in by months or years. Running in the other direction, METR published a finding that the observability tooling used to supervise AI systems is itself attack surface: a researcher found a vulnerability in the Inspect transcript viewer in about ten minutes that would let a model render one thing to a human reviewer while doing another, patched within a day, with an 'untrusted mode' merged on 1 October. Nathan Lambert's argument the same day was that the whole open-versus-closed cyber-risk debate is mis-framed, because restricting open weights while closed API capability keeps scaling does not obviously reduce net risk.

Roadmap implication: This wave has been about security being re-founded around AI; the new information is that access to offensive capability is now a procurement object with an application form, a review queue and a published blocking rate. Practical roadmap consequence: if you run a red team or a security function, getting credentialled is now a vendor-management task with a multi-week lead time, and the 34-of-50 completion figure is the honest capability estimate to plan against, not the marketing number. The METR finding is the more structural one, and it deserves a line in your controls documentation — if your AI oversight story depends on a human reading a transcript, the transcript renderer is in your trust boundary and should be threat-modelled like a production service. The upside reading is real: a published tier structure with measured blocking rates is a better regime than an unwritten one, and the 129,000-vulnerability number is a defender's number first.

Anthropic — Expanding the Cyber Verification Program · METR — AI systems could cover up misbehavior · Interconnects — The Cyber Risk Discourse is Broken

The agent-protocol war acquired a second front, and this one is about the agent's identity rather than the tool's

Sierra announced the Personal Agent Protocol, an open standard for letting a consumer's personal AI agent authenticate itself to a business and be granted scoped permissions, The eight founding partners are Sierra itself plus Meta, Genesys, Instinct, Rocket, Shopify, Stripe and Walmart. The design is built on OAuth: a session can start with the agent as a guest — checking stock or a returns policy — and then escalate to read-only or write access when the customer signs in, with session continuity across channels so actions before and after login belong to the same conversation. Businesses choose how agents reach them: the public website, APIs via existing standards including MCP and OpenAPI, or the company's own agent. A v0.1 specification is due later this month; payments, push notifications and finer-grained permissions are named as future extensions rather than launch features. The partner list is the tell. Shopify and Stripe had already joined Visa's competing Trusted Agent Protocol, announced in June, which covers roughly 175 million merchants — so two of the eight founding partners are now in both camps.

Roadmap implication: MCP settled how an agent reaches a tool. This is the other half of the problem, which commerce cannot run without: how a business knows which agent is standing at its counter, on whose behalf, and with what authority. For anyone operating a storefront or a service desk, the roadmap item is not picking a winner this quarter — it is making sure your authentication and session layer can express 'authenticated human, acting through a delegated agent, with write scope X', because every one of these protocols will ask for that and most commerce stacks cannot say it today. The fact that Shopify and Stripe joined both standards is the useful signal: the payment and checkout layer intends to be protocol-agnostic, which means the differentiation will land on permissioning and audit, not on the wire format. Build toward the capability, stay loose on the spec.

The Next Web — Sierra announces Personal Agent Protocol, an open standard for personal AI agents

A city council subpoenaed a frontier lab, and the labs could not price their own catastrophic risk

The New York City Council's hearing on advanced AI risk, held on Monday, produced two things worth logging. The first is procedural: Speaker Julie Menin said the council had asked the AI unit of Elon Musk's SpaceX to attend, got no response, issued a subpoena compelling attendance, received a letter saying the company wanted to work with the city, and is now pursuing legal action over the failure to appear. A municipal legislature using subpoena power against a frontier lab is a new rung on this ladder. The second is substantive. Representatives from OpenAI, Anthropic, Google and Meta declined to guarantee their systems were safe — OpenAI's Morgan Dwyer said it was not possible for her to commit or guarantee that any technology is without risk, and Google's Alice Friend said promising perfection would not be possible with any product — and council members were visibly more interested in the questions the companies could not answer: the probability of a worst-case outcome, whether the company would bear legal responsibility if its AI did something illegal, and whether they carried insurance for catastrophic-risk scenarios. Asked to estimate the probability of a catastrophic event, Dwyer said she did not know and did not think it mattered whether it was a 1%, 10% or 20% chance because none of those levels is remotely acceptable; Menin called that flippant at best. Three former lab researchers testified first — Jacob Coxon, formerly of Anthropic and OpenAI, Alex Turner, formerly of Google DeepMind, and Daniel Kokotajlo, formerly of OpenAI — with Coxon testifying that it is more likely than not that humanity loses control to these AIs and that it could end in human extinction, and Turner putting AI takeover at roughly one in three. Coxon and Kokotajlo also testified that lab researchers are increasingly letting AI drive code and research and checking its work less frequently; Anthropic has said that as of August, AI was leading more than a quarter of model research and development at the company. Separately, OpenAI had written to the council the previous week recommending safeguards and offering to work with the city's cyber defenders through its Daybreak programme.

Roadmap implication: The roadmap implication is about insurance and indemnity, not about extinction probabilities. The three questions the labs could not answer — legal responsibility, catastrophic-risk cover, probability estimate — are precisely the three questions your general counsel and your broker will ask when you deploy an agent with write access to a customer system, and the vendors have now said in public testimony that they will not answer them. Treat that as the current state of the market and price the residual yourself: scope agent permissions so that the worst single action is survivable, keep the human approval gate on anything irreversible, and get your own cyber and tech E&O policies read against agent-initiated actions this quarter rather than next year. The subpoena is the forward-looking part — governance of AI is moving to jurisdictions with enforcement machinery and local political incentives, which means compliance surface is about to multiply by the number of cities you operate in.

Fox News — OpenAI, Anthropic, Meta, Google stop short of AI safety guarantee · The Information — OpenAI Sent Letter to New York City Council Recommending Safeguards Against AI Risk


🌊 Ripples

Mistral shipped a trillion-parameter open-weight model from European datacenters and called it Le Chonk

Mistral Large 4 is a natively multimodal mixture-of-experts model with 1 trillion total parameters and about 49 billion active, trained on roughly 3,800 NVIDIA Grace Blackwell GPUs in Mistral's European datacenters, spanning more than 160 languages including every official EU language. API pricing is $1.36 per million input tokens and $4.18 per million output, with public preview immediately and weights on Hugging Face by the end of the month. Mistral's published figures include 82% on the Artificial Analysis Cyber Index vulnerability reproduction task, 93% on Cybench, 61.7% on DeepSWE v1.1, 28.3% on Terminal-Bench 4, 59.9% on AutomationBench, 93.3% attack resistance on the B3 AI Security Benchmark, and 42% on Dense 200 visual grounding; the reinforcement-learning stage generated roughly 33 billion tokens a day, of which around 16 billion were trainable completion tokens. Independent placement is less flattering than the vendor deck: Artificial Analysis puts it between DeepSeek V4.1 Flash and OpenAI's entry-level GPT-6 Luna, and Artificial Analysis scored it at 38, as Simon Willison noted, just behind DeepSeek 4.1 Flash — a 552-billion-parameter model — and called it roughly six months behind the frontier and certainly not Fable-class. The same index scored Mistral Large 3, from December 2025, at 9. It does appear to beat Thinking Machines Lab's Inkling, the strongest US open-weight model. The release follows Mistral's €3 billion Series D last month.

Do this now: Do this now if European data residency or EU-language coverage is a real constraint in your stack: this is the first credible frontier-adjacent open-weight model trained end to end inside the EU, and at $1.36 input it is cheap enough to pilot against your existing router this week rather than next quarter. Two caveats to carry lightly and a note on sequencing. The caveats: the benchmark set is the vendor's own, and the independent index says six months behind the frontier, so do not swap it in for a Fable- or Opus-class task without your own evals. The sequencing: wait for the weights before committing to a self-hosted deployment, since a preview API and an Apache-style weight drop are different procurement decisions. The number that actually matters is 9 to 38 on the same index in ten months — that is the rate at which the second tier is closing, and it is the reason a two-supplier routing strategy is now the conservative choice rather than the clever one.

Mistral AI — Introducing Mistral Large 4 · The Register — European AI flag bearer Mistral's new open weights model is 'Le Chonk' · Simon Willison — Introducing Mistral Large 4: Le chonk

Google bought 890 MW of new nuclear capacity by paying Constellation to uprate reactors it already owns

Google and Constellation announced a long-term power deal built on uprates rather than new construction: 890 MW of new nuclear capacity across 11 Constellation-owned units at six sites in Illinois, Pennsylvania and New Jersey, under a 20-year power purchase agreement, with a separate 15-year agreement for 2,700 MW of additional supply. Constellation is putting over $4.3 billion of new investment behind it, sustaining 4,400 existing jobs and creating 7,200 construction jobs, with the first uprate delivery targeted for 2028. All of it lands on the PJM grid, which serves 67 million people; Constellation operates about 55 GW of capacity in total. The agreement also extends a five-year technology alliance covering speed to power — site selection, modelling and permitting — generation optimization including asset monitoring and outage management, and operational-technology network security, with Google Cloud and Gemini Enterprise named as the delivery vehicles.

Do this now: The move worth copying is the engineering one: an uprate adds megawatts to a licensed, already-connected, already-staffed reactor, which is the fastest clean capacity available anywhere in the US and avoids the two things that kill new nuclear, which are greenfield permitting and first-of-a-kind construction risk. If you are siting compute or any other heavy load in the next three years, the question to ask your developer is not which new plant is coming but which existing licensed units in your interconnect have uprate headroom. Note also what Google traded: this is a power deal with a software deal strapped to it, and the operational-technology security line is the interesting one — the grid operators are buying AI tooling from the same counterparty buying their electrons. First delivery in 2028 means this relieves nothing before then.

Google Cloud Press Corner — Google and Constellation Announce Landmark Agreement to Bring 890 MW of New Nuclear Capacity to PJM Grid as Part of Long-Term Power Deal

OpenAI published a batch of new mathematical results from an internal model, with Lean proofs and the compute bill attached

OpenAI released what it describes as a broad range of new mathematical results produced by an internal frontier model, published to a GitHub repository with formalizations in Lean and 10 summaries of the model's reasoning. The disclosure that makes it useful rather than promotional is the cost: the average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking. OpenAI says it consulted the Advisory Group on Mathematics and AI at the Institute for Advanced Study and is committing to fund workshops and conferences on how to understand and verify AI-generated mathematics. The post does not name individual open problems or collaborating mathematicians, which is the main limit on reading it as an independent capability claim.

Do this now: Three hours of Pro-tier thinking per novel mathematical result is the number to write down, because it converts a capability story into a unit cost. Formalization in Lean is the part that makes it checkable — a machine-verified proof does not require you to trust either the model or the vendor's description of it, which is a standard the rest of the AI-for-science field should be held to. If you run any research function with a formalizable core — optimization, verification, cryptography, scheduling, structural engineering — the actionable step is to find the sub-problems in your own backlog that can be stated formally and cost three hours of thinking to attack, because that is now the going rate. The missing collaborator names mean this is evidence, not proof; the Lean files are what will settle it.

OpenAI — Sharing AI progress in mathematics

Google's new embedding model does text, images, audio and video in 191 MB of phone memory

EmbeddingGemma 2 is an open, Apache 2.0-licensed multimodal embedding model: 740 million total parameters, of which 270 million serve text-only workloads, with an optional 170-million-parameter vision encoder and 300-million-parameter audio encoder, all projecting into one 768-dimensional space. Google reports 191 MB of active RAM for the quantized text-only weights on a Pixel 11 Pro and 567 MB for the full multimodal configuration on the same device, an 8,000-token context window — four times the original — and inputs up to 5.5 minutes of audio, 29 images or 58 video frames. Matryoshka representation learning lets the vectors be truncated from 768 to 128 dimensions for roughly a sixfold storage reduction. The headline quality gain is on code retrieval: 78.68 on MTEB Code against 68.76 for the first EmbeddingGemma, with multilingual MTEB essentially flat at 61.36 against 61.15. The original had passed 20 million downloads.

Do this now: If you run retrieval over anything that is not plain text — support tickets with screenshots, call recordings, inspection photos, video evidence — this collapses a multi-model pipeline into one on-device model with one vector space, which removes the cross-encoder plumbing and the embedding-drift problem that comes with maintaining separate indexes. Do this now: re-embed a sample of your messiest multimodal corpus at 256 dimensions and compare recall against your current text-only index plus whatever captioning hack sits beside it. The flat multilingual score is the honest caveat — this is a code and multimodality upgrade, not a general-text upgrade, so there is no reason to re-index pure prose. The on-device numbers are the strategic part: embedding privately-held content without it leaving the handset is a different product than embedding it in someone's cloud.

Google — EmbeddingGemma 2: an open, lightweight multimodal embedding model · Hugging Face — google/embeddinggemma-2

Nvidia is rethinking the revenue-sharing deal it offered AI clouds, because the clouds that wanted it were the ones that needed it most

Nvidia is fundamentally reconsidering the structure of the AI Compute Partnership it announced earlier this summer, under which it would provide credit support to AI cloud providers renting out its chips in exchange for a cut of the rental revenue, according to The Information, which reported it on Monday. Two reasons surfaced from people involved. More established providers such as Nebius declined to take part because they did not want Nvidia eating into their margins and could raise debt by other means. And Nvidia grew concerned that several of the providers that did want in would become too financially dependent on it, and is adjusting term sheets and contracts to avoid that outcome. The credit backdrop is visible in the same window: JPMorgan is marketing a roughly $5 billion leveraged loan at a yield near 11% for Volta Infrastructure, an AI cloud building a Norwegian data centre that Anthropic has already signed a six-year compute agreement to use.

Do this now: This confirms rather than bends the compute-financialization wave, and the useful read is a credit read. An 11% yield on a loan underwritten against a signed six-year offtake from Anthropic tells you what the debt market currently thinks of AI datacentre risk, and that number is your discount rate when a provider offers you a long-dated capacity commitment at an attractive price. Adverse selection is the thing to watch: when vendor financing is available, the operators who take it are disproportionately the ones who cannot raise elsewhere, which is exactly the conclusion Nvidia appears to have reached about its own programme. If you are negotiating multi-year compute, ask how the counterparty funded the building, and price the possibility that you are a creditor of it in all but name.

The Information — Nvidia Is Rethinking Its Revenue-Sharing Deals with Cloud Firms


Read this edition and the full archive at excelsiorgroup.ai/insights/signal.

The Signal · The Excelsior Group

Don't miss what's next. Subscribe to The Signal - daily AI evolution :
Older → The Signal: Bio/Health — Edition #8 — October 6, 2026
LinkedIn
excelsiorgroup.ai
Powered by Buttondown, the easiest way to start and grow your newsletter.