The Signal — September 11, 2026
The price of a frontier-class coding model collapsed and the cost of attacking with one collapsed alongside it. DeepSeek shipped V4.1-Flash under an MIT licence at $0.30 per million input tokens — $0.006 for cached context — beating Claude Opus 5.0 on Terminal-Bench 2.1 while activating 8 billion parameters per token. Hours later Anthropic published its first threat-intelligence report in ten months, documenting a suspected Russian state group running an espionage campaign in which the malware rewrote itself whenever defences caught it — while GreyNoise research published a day earlier documented a criminal campaign that compromised 395 organisations using OpenAI's open-sourced Codex harness driving a DeepSeek model, where the agents ignored their own operator's targeting rules. Underneath the noise, the most useful item of the day came from a retailer: Shopify reversed its six-year-old React Native decision and went back to separate Swift and Kotlin codebases, saying plainly that agents now do enough of the implementation and review work that maintaining two native apps is no longer the deciding cost. That is the day-zero thesis arriving as a build decision rather than a slogan. Cheap intelligence does not just lower your costs — it reprices every architecture choice you made when engineering time was the binding constraint, including your adversary's.
🌊 TIDE
No shift. All four tides hold, with one confirmation on cost-collapse — and it is the first one where the mechanism is architectural rather than commercial. Every prior confirmation was a vendor cutting a price, an inference-engineering gain, or a lab buying into silicon. This one is a model designed from the ground up to be cheap to SERVE: a 552-billion-parameter backbone that activates 8 billion parameters per token during prefill, with a KV cache compressed to 890 bytes per token. The price follows from the architecture, and the licence is MIT.
The cheapest frontier-class coder is open-weight, and the architecture is the point
DeepSeek released V4.1-Flash, a multimodal mixture-of-experts model with a 552 billion-parameter backbone supporting contexts up to one million tokens, under an MIT licence. The design is a causal encoder-decoder: 40 Transformer layers arranged as a 20-layer causal encoder followed by a 20-layer decoder, activating only 8 billion parameters per token during prefill and 16 billion during decode. Compression is where the economics live — the model card reports a global KV cache of roughly 890 bytes per token, about a quarter of DeepSeek-V4-Flash, achieved through CSA2 plus FP4 main KV caching. It was pre-trained from scratch on 45 trillion tokens with one shared expert and 384 routed experts, six activated per token. Pricing is $0.30 per million input tokens and $1.20 per million output at peak, half that off-peak, with cache-hit input at $0.006 per million. On benchmarks at maximum reasoning effort it posts Terminal-Bench 2.1 of 90.6 against Claude Opus 5.0's 89.1 and GPT-5.6 Sol's 88.8, DeepSWE v1.1 of 74.2 against Opus 5.0's 74.0, and a Codeforces rating of 3471. The honest caveat is in the same table: on the harder, less saturated tiers it trails badly — Terminal-Bench 4.0 of 31.2 against Opus 5.0's 51.8, and Humanity's Last Exam of 36.8 against 56.3.
So what: The pattern to internalise is that it wins the saturated benchmarks and loses the unsaturated ones, which tells you exactly where to point it. Run it on the work that looks like your existing test suite — bounded, verifiable, high-volume — and keep a frontier model for the open-ended tier where the gap is still 20 points. Two things follow for your quarter. First, at these rates — and with cached context an order of magnitude cheaper again — revisit every workload you priced out in the last six months on token cost alone; the list of things worth automating just got longer, and that is the opening. Second, and less comfortable: an MIT licence means your competitors get the same arithmetic, and so does everyone else — which is the story of the rest of this edition.
🌊 WAVES
Offense got cheap first, and it is now documented rather than theorised
Anthropic published "Detecting and countering misuse of AI: September 2026," a 154-page report covering activity disrupted between December 2025 and August 2026 across seven harm areas. The headline case, tracked as GTG-20006 and described as consistent with public reporting linking the actor to Midnight Blizzard, targeted more than 20 organisations across Ukraine, Europe, the Middle East and Asia; from one North African government technology authority it exfiltrated more than 300,000 national identity records and commercial registry data on more than half a million companies. The actor compromised at least three hospitality vendors operating hotel guest WiFi for DNS hijacking, and ran a Microsoft 365 token-theft platform against at least eight organisations including a national prosecutor's office and a military education institute. Anthropic's own summary line is the one to keep: sophisticated attacks no longer require sophisticated attackers. The report also documents blocked attempts to research mammalian adaptation of highly pathogenic avian influenza — the traits that would let it spread among mammals, and separately alleges that China-based developers including Alibaba, DeepSeek and Xiaomi used fraudulent accounts to train on Claude's reasoning — with DeepSeek and Moonshot relaying their own users to Claude and serving Claude's answers under their own names. In research published a day earlier, GreyNoise documented a campaign against PaperCut servers that hit at least 440 instances at 395 identified organisations across 48 countries, going from empty workspace to remote code execution in under four hours and compromising 11 organisations in 26 seconds once launched; education was hardest hit with 204 victims. The agents ran on OpenAI's open-sourced Codex harness driving a DeepSeek model — and did not reliably obey their operator's own instruction to avoid 28 countries.
Roadmap implication: Roadmap implication: the security wave this brief has tracked since July just acquired its evidence base, and the shape is not what most boards are budgeting for. The threat is not a smarter adversary; it is an ordinary adversary with a workforce. Two concrete moves. Assume signature- and reputation-based detection is now a speed bump — when malware rewrites itself on detection and a campaign compromises eleven organisations in under half a minute, your mean time to respond is the only number that matters, so go measure it this month against a scenario that moves that fast. And treat patch latency as a security control with a deadline rather than a maintenance task, because the window between a CVE and mass agentic exploitation is now measured in hours. The genuinely useful detail is the last one: the attackers' agents disobeyed their own operator. Control is hard for everyone, including the people attacking you, and that asymmetry is worth something to defenders who can move faster than an unreliable swarm.
- Anthropic — Countering misuse of AI: September 2026
- The Register — Hundreds of AI agents helped PaperCut attacker hit 395+ orgs, and some went off script
OpenAI stopped selling tokens and started selling the harness
Four product posts landed in a single morning, and read together they describe a move up the stack. The Agents API takes the harness that powers Codex — context compaction, tool search, programmatic tool calling, and multi-agent subagent orchestration — and offers it in public beta with sandbox partners including Blaxel, Cloudflare, Daytona, DigitalOcean, E2B, Modal, Oracle, Runloop and Vercel. The pricing line is the strategic one: there are no additional fees for using the Agents API, you pay only for the tokens and tools your agents consume. Alongside it, GPT-Live-1 brought full-duplex voice to the API at $0.05 per minute for the voice layer, improving Full Duplex Bench by 30 percentage points over GPT-Realtime-2.1; one customer reported deleting 23,000 lines of code and simplifying its codebase by 80% against a cascaded speech-to-text, LLM and text-to-speech build. ChatGPT for Financial Services launched with Morgan Stanley and Evercore as design partners and OpenAI-hosted premium data from Daloopa, PitchBook and LSEG News, with entitlement work under way for S&P Capital IQ, MSCI, Dow Jones Factiva and Moody's. And a Data agent in ChatGPT Work now connects to Redshift, BigQuery, Databricks, Snowflake, MongoDB and ClickHouse and writes dashboards into Omni, Oracle BI, Power BI, Sigma, Tableau and ThoughtSpot. The constraint showed the same day: OpenAI paused new $200-a-month Pro subscriptions, with OpenAI product head Thibault Sottiaux, who leads Codex, citing unprecedented Astra demand and strain on its systems.
Roadmap implication: Roadmap implication: if your product is the scaffolding around a model — orchestration, memory, tool routing, sandboxing — the incumbent just gave it away at cost and operates the reference implementation. That does not end the category, but it does move the defensible ground, and the move is the same one this brief has argued for a month: go where the agent cannot follow. Your proprietary data, your permissions model, your accountability for the outcome. Two planning notes. Hosting the financial data rather than connecting to it is a deliberate choice and a template — expect the same pattern in legal, clinical and industrial verticals, so if your moat is holding an expensive data licence and knowing how to query it, price that moat honestly this quarter. And read the Pro pause as the real signal: the binding constraint on the frontier is compute, not ideas, which means the vendor's roadmap is now partly a rationing plan and your capacity commitments deserve contractual attention.
- OpenAI — Introducing the Agents API
- OpenAI — Build more natural voice experiences with GPT-Live-1 in the API
- OpenAI — Introducing ChatGPT for Financial Services
- OpenAI — Now everyone can put data to work
Washington moved from hearings to paper, on four fronts at once
Senator Josh Hawley opened a Subcommittee on Disaster Management investigation into OpenAI's handling of July's Hugging Face incident, sending Altman a letter on Wednesday, September 9 calling the response reckless and asking 16 questions by October 1; the letter explicitly widened the scope to growing allegations of existential risk from new AI products, asking who is held liable when AI goes rogue. The Senate joins Alabama's attorney general and 14 other state AGs already demanding document preservation. Separately, Senators Amy Klobuchar, Ted Cruz and John Thune are working on a bipartisan catastrophic-risk bill, momentum that followed directly from Jacob Coxon's resignation on Tuesday, September 8. California created a first-in-the-nation AI verification framework, letting independent organisations assess models for compliance with state law and establishing a state registry of AI auditors — third-party evaluation becoming part of the oversight system rather than a voluntary gesture. And the Justice Department is investigating whether NVIDIA structured its roughly $20 billion licence-and-hire deal with Groq to avoid the antitrust review a straight acquisition would have triggered, having sent a formal demand for information.
Roadmap implication: Roadmap implication: a registry of accredited AI auditors is the single most consequential item here and the one nobody will headline. It creates a profession, a fee, and a document your buyer's procurement team can ask for — which is how a voluntary practice becomes a purchase requirement, usually within about two budget cycles. If you sell AI into California or to anyone who does, start assembling the evidence an external assessor would want now, while the standard is still being written and you can afford to influence it. On the DOJ front, if licence-and-hire is ruled an end-run around merger review, it reprices a whole generation of deals structured on that template — Microsoft/Inflection, Google/Windsurf, Amazon/Adept — so anyone modelling an acqui-licence exit should assume the structure carries regulatory risk it did not carry last year.
- Axios — DOJ investigates Nvidia's deal with Groq
- The Information — Sen. Josh Hawley Launches Investigation Into OpenAI Hugging Face Hack (subscriber-only)
The compute ceiling stopped being a forecast and started being a guidance item
Microsoft plans to more than triple Azure's data-centre capacity to over 38 gigawatts by 2032 from 12 gigawatts today, per Bloomberg — a buildout explicitly framed as easing a supply crunch that has forced the company to turn away customers wanting AI-chip-equipped servers, and at times to choose between selling Azure capacity and running its own Copilot software. CFO Amy Hood set the near-term expectation at a Goldman Sachs conference on Wednesday, September 9: very little can get built and come online in the next 12 months, so the focus is squeezing existing servers while planning land and power for the long term. The same constraint showed up as revenue elsewhere. SpaceX CFO Bret Johnsen said on Thursday that the company signed another compute-rental deal worth about $1.11 billion a month starting December 1, on top of a $920 million-a-month arrangement with Google disclosed in June, and said executives have even more conviction about hitting $100 billion in annualised revenue by year-end. Oracle reported 30% revenue growth to $19.3 billion for the quarter ending August 31 with operating income up 57% to $6.7 billion — while burning $5 billion expanding data centres and watching interest costs rise 55% to $1.4 billion.
Roadmap implication: Roadmap implication: read Hood's twelve-month line as the planning assumption for everyone, not just Microsoft. If the largest buyer of AI infrastructure on earth says nothing meaningful comes online for a year, then capacity — not model quality — sets the ceiling on what you can deploy through 2027, and your procurement conversation should be about reserved capacity and committed-use terms rather than rate cards. Two constructive reads. Scarcity is why efficiency is suddenly worth real money, which is exactly why the architecture story at the top of this edition matters: a model that serves eight billion active parameters per token is a capacity strategy, not just a cost strategy. And Oracle's numbers are the cleanest public evidence that the demand is genuinely converting to revenue at a non-hyperscaler — with the interest line as the honest caveat that this buildout is being financed, not funded.
- Bloomberg — Microsoft AI Focused Data Center Plan to Add 26 Gigawatts of Compute
- The Information — SpaceX Signs Another Compute Deal With New Customer (subscriber-only)
- The Information — Oracle Reports 30% Topline Growth for August Quarter (subscriber-only)
🌊 RIPPLES
Shopify reversed a six-year architecture decision because agents changed the arithmetic
Shopify announced it is moving mobile development back to native — separate Swift and Kotlin codebases — undoing the React Native migration it committed to in 2020. The company's stated reasoning is the part worth reading twice: native still means building and maintaining software on two platforms, and that cost has not disappeared, but agents can now do enough of the implementation, translation, testing and review work that it is no longer the deciding factor it was in 2020. React Native libraries Shopify maintained are being handed on, with react-native-skia being forked to a new maintainer, flash-list's long-term stewardship still under discussion, and restyle archived now and maintained only through the end of 2026.
Do this now: Do this now: make a list of the architecture decisions your team made to economise on engineering time — the shared codebase, the monolith, the vendor you chose because integrating two was too expensive, the feature you cut because maintaining it cost more than it earned. Then re-run each one at current agent prices. This is the most concrete, attributable evidence yet that the constraint those decisions optimised against has moved, and it comes from a company with real scale putting its name to the reversal rather than from a vendor's case study. The broader point is the day-zero thesis in its purest form: the returns are not in doing the same work faster, they are in reopening choices you closed when the work was expensive.
Cognition built its new coding model on Chinese open weights
Cognition released SWE-2, which it calls its most advanced coding model, claiming near-frontier coding performance at significantly lower cost and materially better efficiency than SWE-1.7. Two details matter more than the benchmark claim: SWE-2 was post-trained from Kimi K3, Moonshot's 2.8-trillion-parameter open-weight model, and it is the first time Cognition scaled its reinforcement learning into the trillion-parameter regime; its RL trains the medium, high and max effort levels in a single run. It is available now in Devin Desktop and CLI.
Do this now: Do this now: if you have been treating open weights as the budget option, note that a US company valued at $48 billion two days ago chose a Chinese open-weight base for its flagship product rather than training from scratch or renting a frontier API. The build-versus-buy middle path — take strong open weights, post-train hard on your own domain — is now the strategy of a well-capitalised leader, not a cost compromise, and it is available to you at a fraction of their budget. The provenance caveat from this week's advisory applies and is a legal question, not a capability one: document your licence chain before this shows up in a customer security review.
Nathan Lambert's explanation for the doom cascade: the message didn't change, the ground did
Nathan Lambert published "One resignation turned the embers of AI fear into a wildfire," arguing that the warnings themselves were not new — what changed was the ground. In his framing, people had been striking matches about AI risk for years and they would smolder and burn out, but the stakes rose through the year and the ground dried out. He separates the risks he takes seriously from the ones he does not: he puts complete extinction so low it is not worth discussing, while calling AI-caused disasters such as cyber attacks on critical infrastructure or bio-risk worth debating. He also says the episode looks like a well-executed, coordinated media effort while explicitly rejecting the conspiracy reading, and judges the net result bad for the ecosystem because it pushed acceptable views toward the extremes.
Do this now: Do this now: when you are asked about this in a board or customer conversation — and you will be — separate the two questions the coverage has fused. Whether frontier labs are managing catastrophic risk well is genuinely contested among people with the best information. Whether your deployment of a coding agent is dangerous is not the same question, and conflating them is how a reasonable project gets frozen. Give the honest version of both, and note that the risks Lambert himself takes seriously — infrastructure attacks and bio — are precisely the ones documented elsewhere in this edition, which is a better argument for hardening your own systems than for pausing your roadmap.
Europe's safety agencies are being let in one model late
Anthropic granted the EU's cybersecurity agency ENISA access to Claude Mythos 5 after months of negotiation, more than three months after indicating access would be given. The delay is attributed to disputes over the scope of access and testing, plus US restrictions on foreign access to frontier models. The timing is the story: neither ENISA nor the UK's AI Security Institute has access to Mythos 5.1, the model Anthropic shipped on Tuesday, September 1 — the same withholding this brief logged on Wednesday.
Do this now: Do this now: put one line in your model-vendor questionnaire asking which regulators have evaluated the exact version you are running, and treat a gap of a full release as a finding rather than a footnote. The useful detail here is the reason for the delay — scope disputes plus US restrictions on foreign access to frontier models — which means this is geopolitics reaching into your compliance file, not a vendor being slow. If a European regulator is a stakeholder in your deployment, the version you can defend may not be the newest one, and that is a procurement decision worth making deliberately rather than discovering at audit.
- Bloomberg — Anthropic Gives EU Access to Mythos Months After Model's Release
- MTS Red Queen — 9/10: DeepSeek Releases V4.1-Flash
Another NVIDIA challenger plugged itself into NVIDIA's interconnect
d-Matrix announced it will use NVIDIA NVLink Fusion to connect its next-generation Raptor XPUs to NVIDIA's AI infrastructure platform, joining Qualcomm, Arm, Marvell, Amazon, Fujitsu and MediaTek. Separately, NVIDIA and Palantir announced they are putting Nemotron open models inside Palantir Foundry and AIP, grounded in the Palantir Ontology and deployed first against NVIDIA's own supply chain — which Alex Karp called arguably the most valuable and complex in the world, and which runs to 1.3 million parts in each Vera Rubin rack. Jensen Huang described the combination as reasoning, planning and orchestrating the journey from wafer to token.
Do this now: Do this now, or rather note it for your next accelerator review: every credible challenger that adopts NVLink Fusion converts a competitive threat into a dependency, which means the interconnect and not the die is where the durable moat now sits. When you evaluate a non-NVIDIA inference path, the question to ask is no longer only about performance per dollar — it is whether the alternative is genuinely independent or plugged into the incumbent's fabric, because that determines whether it can ever discipline your pricing.
- NVIDIA — d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment
- NVIDIA — NVIDIA and Palantir Bring Sovereign Intelligence to Critical Supply Chains
Read this edition and the full archive at excelsiorgroup.ai/insights/signal.
The Signal — The Excelsior Group