The Signal — September 16, 2026
The Read
The safety coordination everybody read as a weekend reaction had already been running for two months. OpenAI global policy chief Chris Lehane confirmed that representatives of Anthropic, OpenAI and Google have been meeting about forming an AI safety standards body — since July, per The Information's reporting — and that OpenAI backs a FRONTIER Act provision compelling top labs to admit independent verification organizations, while saying the companies do not need the antitrust waiver Dario Amodei asked for. Beijing's mirror image reached English-language readers the same day: its Minister of State Security's own AI-risk essay, published over the weekend, and the national governance framework that followed it. So the governance tide keeps doing what this brief has argued it would — arriving as institutions and procurement language rather than as statute. The part worth acting on is the price tag that showed up alongside it: The Information found small-business buyers going cold on an AI app in light of catastrophic-risk coverage, and investors already rotating into the model-security and infrastructure layer a standards regime makes mandatory. Meanwhile, at the AI Infra Summit in Santa Clara, the industry quietly settled on a new thing to argue about: not dollars per GPU, but tokens per megawatt.
🌊 TIDE
Confirmed — governance-as-market-structure. No shift. Yesterday the mechanism was buyers writing procurement language; today it is a standing multi-lab body that has been meeting since July, a legislative vehicle in the FRONTIER Act, a Chinese counterpart framework, and the first measurable demand-side cost. The tide holds and gains a start date.
The coordination has been running since July — and it now has a bill attached
OpenAI global policy chief Chris Lehane told reporters that representatives of Anthropic, OpenAI and Google have been working together on AI safety for weeks, confirming The Information's reporting over the prior weekend that the three have been meeting regularly since July about potentially forming an AI safety standards body. Lehane said OpenAI supports a FRONTIER Act provision that would require top frontier labs to admit independent verification organizations, and that the companies do not need the narrow government antitrust waiver Dario Amodei proposed in his September 12 essay. The mirror move surfaced in English the same day, via The Register: Chen Yixin, Party secretary and Minister of State Security, had published an essay in China Cyberspace Magazine on September 13 calling AI "the main battleground for global technological competition and a new arena for strategic rivalry among major powers," singling out OpenClaw-type products for "structural problems such as remote control of device management permissions and leakage of sensitive user information," and calling for special AI laws covering algorithm security, data protection, ethical norms and privacy. The Cyberspace Administration of China published version 3.0 of its AI Safety Governance Framework on September 14, endorsing regulatory sandboxes "to make room for error and correction." The Register's read is the uncharitable one: that the pacing push is an attempt to entrench US incumbents against improving Chinese open-weight models.
So what: A standards body that has been meeting since July is not a reaction; it is an institution, and institutions write specifications. The FRONTIER Act provision is the part to watch, because independent verification organizations with employee-like access are a new category of vendor that does not exist yet at scale — someone has to staff, tool and insure them. Treat the whole apparatus as a market being created rather than a ceiling being imposed: the compliance surface is a product line, and the labs have just told you what it has to do.
Sources: OpenAI, Anthropic, Google have been in talks on AI safety for weeks · Anthropic, OpenAI, Google Quietly Discussed an AI Safety Standards Body · The latest AI doomsayer is China's intelligence boss · Anthropic and OpenAI look to Uncle Sam to make them too big to fail
The first bill for the pacing posture arrived alongside it
The Information's Dealmaker reported that the slowdown pledges are already having commercial effects well ahead of any regulation. One venture capitalist described a portfolio company selling an AI app to small and medium businesses finding customers newly wary in light of coverage about potentially catastrophic AI. Investors in the piece expect self-monitoring and regulation to favour the largest startups, which can afford the tooling and headcount to prove their models are safe, making it harder for so-called neolabs researching new model approaches to compete. Two counterweights ran in the same piece: an early-stage software investor expects capital to flock to model-security and infrastructure companies — citing Gimlet Labs, which routes AI work across the right chips for each task, and security startup Neo, which raised $100 million from Andreessen Horowitz, Bessemer Venture Partners and others in July — and an applied-AI investor argued that a genuine lull at the frontier buys app-layer companies time to compound on a fixed target.
So what: This is the first evidence of the pacing debate reaching an end customer's purchase decision, and it is small-business software, not a regulated industry — the part of the market with the least patience for nuance. If you sell AI to SMBs, your objection handling now has to answer a question the buyer did not have three weeks ago. The flip side is the opening: a frontier that moves more slowly and more legibly is a better substrate to build a durable product on than one that reprices your differentiation every six weeks. App-layer founders should treat a real lull as the gift it is.
Sources: Anthropic, OpenAI 'Pacing' Spooks Some Startup Customers · Anthropic, OpenAI, Google Quietly Discussed an AI Safety Standards Body
WAVES
Tokens per megawatt becomes the unit of account
NVIDIA used the AI Infra Summit in Santa Clara — 8,000-plus attendees this year against 3,500 last year — to reframe AI-factory economics around energy rather than silicon count. Ian Buck, VP of hyperscale and HPC, presented alongside the first third-party validation of NVIDIA DSX MaxLPS: Lambda reports running 19 Blackwell nodes inside the power budget normally allocated to 16 full-power nodes, lifting cluster-wide token throughput roughly 24% — about 4 million to 5 million tokens per second — and improving performance per watt by 23%. NVIDIA's own forward claims, which are vendor numbers and should be read as such, are that MaxLPS on Vera Rubin NVL72 enables up to 40% more GPU capacity within the same megawatt budget and up to 35% higher token throughput, and that on SemiAnalysis's AgentX benchmark Vera Rubin NVL72 delivers up to 30x higher throughput per megawatt than GB300 NVL72 on DeepSeek V4 Pro. SemiAnalysis's own independent framing, published a day earlier from access to unreleased hardware, measures Rubin NVL72 at 59.4 million tokens per second per megawatt at 100 tokens per second per user against 28.5 million for GB300 — a 2.1x gain, rising with latency target. Do not conflate the two numbers; they measure different things.
So what: Roadmap implication: if your infrastructure plan is still denominated in GPUs, it is denominated in the wrong unit. Power is the binding constraint on every buildout this brief has tracked, which makes tokens per megawatt the number that decides what your inference actually costs and how much of it you can have. Two practical moves — ask your cloud or colo for delivered tokens per megawatt at your latency target rather than GPU-hours, and note that the Lambda result is a software-and-power-management gain on existing Blackwell silicon, so some of this is available to you before any Rubin lands.
Sources: AI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for AI Factories · From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production
Nvidia's scale-up fabric draws funded attackers
Two startups used the same summit to go after the interconnect layer, and both arrived with money. Cornelis Networks, spun out of Intel in 2020, introduced Active Compute Fabric, an open scale-up and scale-out architecture that puts programmable compute inside the network — targeting KV-cache offloading, mixture-of-experts expert dispatch, MPI offload and in-fabric checkpointing, aimed squarely at NVSwitch and SHARP. Cornelis says its modelling of public data shows roughly half of all GPU hours in a 100,000-GPU system are spent waiting for data, which it values at about $1.68 billion a year in wasted capacity and 500 GWh of power; that is a vendor's own model, not an independent measurement. It announced approximately $205 million in funding alongside the launch, with 800 Gbps CN6000-series switches and NICs imminent. Delos Data, founded by former Barefoot Networks and Intel executives, unveiled its Nonstop AI portfolio: a protocol-agnostic I/O die rated above 30 Tbps aggregate — roughly 8 TB/s each direction, against the 3.6 TB/s ceiling on Nvidia's and AMD's latest accelerators — plus near-packaged optics above 10 Tbps and a 400-plus Gbps NIC. Delos said it has now raised more than $100 million from Matrix, Playground and Socratic Partners.
So what: Roadmap implication: the interconnect has been the least contestable part of the Nvidia stack, and it is now the part attracting the most venture money — because if tokens per megawatt is the metric, idle GPU-seconds waiting on the fabric are the largest single line of waste in the building. Nothing here is procurable this year. What is actionable now is measurement: instrument what fraction of your own GPU time is network-stalled before the next capacity commit, because that number is both the size of the prize for these entrants and your strongest lever in the negotiation you already have.
Sources: AI networking startups race to replace Nvidia's NVLink
The datacenter backlash gets audited — loud politics, quiet megawatts
SemiAnalysis published a project-level audit of every US datacenter moratorium and reached a conclusion at odds with the prevailing narrative. Against more than 300 enacted local moratoriums and bans across counties, municipalities and townships — Michigan and Ohio lead with 45 and 40, with North Carolina and Georgia completing a top four that accounts for roughly half of all local instruments nationally — it finds roughly 20 GW of planned capacity sitting inside a restricted boundary but only 1,525 MW genuinely delayed, about 7.6%, concentrated in three projects: an AWS campus in Ohio, a powered-land developer campus in Pennsylvania and a site in Colorado. New York's Executive Order 62, the first statewide permitting pause, touches roughly 1.4 GW but binds on about 0.8 GW, all of it 2028-or-later delivery. Texas's ERCOT audit adds perhaps three to four months of administrative delay for base-load projects and is, on SemiAnalysis's read, a net positive for behind-the-meter developers. Total nationwide: roughly 2.3 GW genuinely delayed, against a forecast of +38 GW of US datacenter IT capacity delivered in 2027, more than double 2026, of which 22 GW is already under vertical construction. The politics are nonetheless real: SemiAnalysis's own August polling found 46% of voters view datacenters unfavourably against 29% favourable, and 46% would oppose one in their own town — while 46% view AI itself favourably. Thirteen statewide bills died in 2026, almost none on a floor vote.
So what: Roadmap implication: this brief has tracked the datacenter backlash as a building wave and should now hold it more precisely — it is real as electoral politics and largely unpriced as megawatts, because moratoriums freeze new applications rather than revoke granted approvals, and roughly 80% of live restrictions cover no modelled planned capacity at all. The buildout is not being legislated to a halt; it is being routed. The winners are already visible and investable: behind-the-meter power, on-site generation equipment, and land carrying existing entitlements or by-right zoning. If your compute plan assumed a capacity crunch from permitting, re-underwrite it.
Sources: Everyone Says Datacenter Moratoriums Are Killing the US Buildout. We disagree
Anthropic's data-retention gap is now moving enterprise deals to OpenAI
The enterprise data-boundary story this brief ran yesterday acquired its commercial consequence. Major Anthropic customers are still waiting on zero data retention for the newest Fable models; in the meantime Palantir and Booz Allen are expanding use of OpenAI's Astra, which they have secured with ZDR, and both have restricted Fable use since Anthropic said in June it would retain customer data for 30 days for security reasons. An executive at a company selling cybersecurity software to the US government and large financial-services firms said it switched its customer-facing apps from Anthropic's models to primarily Astra, specifically because it could get ZDR guarantees from OpenAI that it could not get from Anthropic. Anthropic's answer, Enterprise Frontier Safeguards — activity data held in the customer's own cloud under the customer's own keys — rolls out to eligible customers later this fall, with Salesforce, Comcast and Visa named among the customers who helped design it; the qualification criteria are not public. Palantir began making Astra available to its customers earlier this month, so many enterprises will have tested OpenAI's frontier model weeks before they can test Anthropic's.
So what: Roadmap implication: data-boundary terms are now a first-order determinant of enterprise model share, not a legal afterthought — a thirty-day retention window cost Anthropic default position inside two of the most demanding buyers in the market, and the gap between announcing a safeguards programme and shipping it is being measured in lost weeks. For buyers, the lesson is cheap: put ZDR and log-custody terms in the first conversation, not the redlines, and make model-portability a design requirement so a retention-policy change is a config change rather than a rebuild.
Sources: Why Anthropic's Data Policy Drama Is Good for OpenAI · Palantir, Nvidia, Booz Allen restrict Anthropic and OpenAI models
RIPPLES
Google ships two live-dialogue models, and someone built on them the same night
Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, a pair of speech-to-speech dialogue models. Gemini 3.8 Live is positioned for scale and cost efficiency with visual grounding; Extended Thinking is built for high-complexity tasks needing multi-step reasoning. Both are available to developers in the Gemini API and Google AI Studio, and in private preview for enterprises in Gemini Enterprise; Gemini 3.8 Live is available to everyone in Search Live, while Extended Thinking reaches Gemini Live, Docs for Google AI Pro and Ultra subscribers, and Gmail and Keep for Google AI subscribers. Simon Willison logged the release the same evening, described the pair as a similar shape to OpenAI's GPT-Live family, and had GPT-6 Astra Extra High build him a no-library browser voice interface — WebSocket plus the Web Audio API, with interruption support — from the docs.
So what: Do this now: if you have a workflow that is really a phone call — intake, triage, scheduling, field support, anything where the human is holding a handset and a clipboard — prototype it against Gemini 3.8 Live this week and against GPT-Live for comparison. Willison's evening build is the actual signal: the integration cost of real-time voice has fallen to roughly one session with a coding agent, which means the constraint on voice products is no longer engineering, it is choosing which call to re-found.
Sources: Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking · Build real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe · Archive for Tuesday, 15th September 2026
Google opens its AI-usage atlas, and puts a number on what scientists get back
Google launched an interactive, open-access front end for its AI & Economy ATLAS dataset and published companion research with Google DeepMind and MIT FutureTech. The occupational mix varies more by country than most adoption decks assume: in the US, computer and mathematical occupations account for 30% of work-related AI usage, double the share in the rest of the world; in India, arts, design and media account for 19%, about 1.6x the global average; in Brazil and Germany, 7% of work AI usage goes to manual tasks such as real-time equipment diagnostics, against 4% in Japan; Brazil and the UAE over-index on adoption relative to GDP per capita. The science study analysed 2,600 specialized AI models and surveyed more than 600 US and UK scientists: nearly half use some form of AI daily and report saving just under 7 hours a week — but the paper identifies significant time spent validating AI outputs, a growing backlog of untested hypotheses, and bottlenecks in physical experimentation and clinical validation.
So what: Do this now: if you operate across borders, stop extrapolating your headquarters' adoption curve — ATLAS is free and shows the occupational mix inverting between markets, which changes where enablement spend actually returns. The sharper finding is the scientists' bottleneck: seven hours back per week is real, and the constraint immediately moved downstream to validation and physical experiment throughput. That is the day-zero opening in science — the money is no longer in generating hypotheses, it is in clearing the queue of ones nobody has time to test.
Sources: New insights from Google's AI & Economy ATLAS
Biren lines up another $1 billion for domestic Chinese GPUs
Shanghai Biren Technology, which designs GPUs for training and running AI models, is considering raising around $1 billion through a stock offering, Bloomberg reported, citing people familiar with the situation. Biren listed on the Hong Kong Stock Exchange in January amid a run of AI-sector listings and raised about $900 million in a July share sale. Banks are at an early stage of gauging investor interest; size and structure are not finalised. The raise sits inside Beijing's push to accelerate domestic chip development and reduce reliance on Nvidia.
So what: Do this now: watch whether the raise clears, and at what discount. Roughly $1.6 billion of public equity already raised into a single non-Nvidia designer inside eight months, with another $1 billion now sought, is the clearest available read on how much capital Chinese domestic silicon can absorb — which is the variable that decides whether export controls bind on capability or merely on timeline. If you model China exposure, this number belongs in it.
Sources: Chinese AI Chip Firm Biren Plans $1 Billion Fundraising
ByteDance's AI bill shows up on the P&L
ByteDance's net profit in the first half of 2026 declined by a single-digit percentage to $20 billion as the company ramped up AI investment, The Information reported, citing three people with knowledge of the figures; revenue over the same period rose roughly 30% to about $120 billion. ByteDance is privately held and does not publish results, so the numbers reach the market through reporting rather than filings, and should be read accordingly. Not to be confused with the 70% full-year profit decline reported in April for 2025.
So what: Do this now: note which direction the pressure runs, because it is the opposite of the bubble framing. Revenue up about 30% and profit down slightly is not demand weakness — it is one of the most profitable private companies on earth deliberately trading near-term margin for AI position, from a $20 billion half-year base, with better demand data than almost anyone. If your board is asking whether AI spend is defensible against margin, this is your comparable: the operators with the clearest view of the demand curve are the ones spending into it.
Sources: ByteDance's First-Half Profit Drops to $20 Billion, Weighed Down by AI Spending
Every edition, archived: https://excelsiorgroup.ai/insights/signal/
The Signal — The Excelsior Group