The Signal - daily AI evolution  logo

The Signal - daily AI evolution

Archives
Log in
Subscribe
August 21, 2026

The Signal — August 21, 2026

Covering Thursday, August 20, 2026

The Read

Thursday, the demand side finally spoke. AT&T — 100,000 employees, 45 billion tokens a day — told The Information it plans to hold its Anthropic and OpenAI spend flat and push open-weight models from 40% of employee queries to 60–70%, after a model router cut its AI coding costs 56% for a 2% quality loss. Every prior confirmation of the cost-collapse tide came from a vendor cutting prices; this is the first from a buyer deciding it doesn't have to pay them. The rest of the day rhymed: Anthropic reversed its 30-day enterprise retention rule one day after OpenAI dangled zero data retention, Waymo revealed it had quietly taped out its own 5nm inference ASIC, and Alibaba published what the buildout actually costs — capex up 75%, profit down 76%. Old problems, new physics: the most consequential move Thursday wasn't a model release, it was a procurement decision.

🌊 Tide — the megatrend layer

No shift. All four tides hold. But cost-collapse logged its first genuine demand-side confirmation, and one buyer publishing its arbitrage is worth more than a dozen vendor price cuts — the seller has an incentive to advertise cheapness; the buyer has an incentive to keep it quiet. Governance-as-market-structure also gained texture: Anthropic's retention reversal is a frontier lab changing a safety policy because more than 100 enterprise customers pushed back, which is governance being written by procurement rather than by regulators. Confirmations, not movement. The tide holds.

AT&T caps its frontier AI bill: open weights go from 40% of queries to 60–70% — cost-collapse gets its first buyer-side confirmation

AT&T vice president Mark Austin told The Information the company intends to keep employee spending on Anthropic and OpenAI models flat in the coming years by shifting more work onto open-weight models — NVIDIA's Nemotron, Meta's Llama, Google's Gemma — run in part on AT&T's own NVIDIA and AMD servers, which he says is often cheaper than renting cloud capacity. Open models already handle 40% of the 45 billion tokens a day flowing through Ask AT&T across 100,000 employees; the target is 60–70%. After adopting LiteLLM's router, AI coding costs fell 56% while measured quality fell 2%. Developers still route complex code generation to frontier models but send cheaper work — summarising previously submitted code, for instance — to open weights. Austin's read on the gap: open models have run six to ten months behind the frontier, “and the gap seems to be narrowing.” DeepSeek and Moonshot are under evaluation but not in production, on risk grounds.

Why it bends the tide: Every previous confirmation of this tide came from the supply side — a lab cutting prices, a challenger undercutting, a lab buying its own silicon. This one is a buyer with a nine-figure inference bill publicly declaring that frontier pricing is optional for most of its workload, and publishing the arbitrage: 56% cost reduction for 2% quality. If you have not measured what fraction of your token spend genuinely needs frontier capability, you are paying a premium you have never priced. Do that measurement this quarter, and treat the six-to-ten-month open-weight lag as your real substitution horizon rather than a permanent moat.

Sources: The Information: AT&T Has a Plan to Curb its Anthropic Bills · PYMNTS: AT&T slashes AI costs by adopting model routers and open source

🌊 Waves — weeks to quarters

Anthropic folds on retention: enterprise data control is now a frontier-model sales term

One day after OpenAI reaffirmed Zero Data Retention for frontier API customers and previewed Private Safety Processing, Anthropic said it will let enterprise customers keep control of their own data on its most capable models. The new safety system, due later this year, still requires a 30-day retention window — but enterprises will be able to hold that data on their own cloud infrastructure rather than Anthropic's. The change has been in the works for months, developed with more than 100 customers including Salesforce. The policy being reversed is Anthropic's own June rule retaining all Mythos- and Fable-class sessions for 30 days to detect novel cyberattacks — a rule Anthropic itself conceded would “be unpopular with customers who have come to expect zero retention, and pose real risks to our business success.”

Roadmap implication: Roadmap implication: data residency has moved from a compliance checkbox to a live commercial lever, and it moved because customers pushed, not because a regulator did. If a frontier deployment is stalled in your legal review over retention, reopen the term now — both labs are competing on it and the price of that concession is at its lowest point in the cycle. Put an explicit “where does our data physically sit, and under whose keys” clause into every model contract renewing this half, and make the answer a scored criterion rather than a footnote.

Sources: Reuters: Anthropic plans to change enterprise data retention policy · Bloomberg: Anthropic plans to change data retention policy for advanced AI · OpenAI: Offering Zero Data Retention for frontier models (Aug 19)

Waymo taped out its own 5nm inference chip — custom silicon is no longer a hyperscaler-only decision

Waymo disclosed a purpose-built 5nm ASIC delivering more than 1,000 TOPS of ML performance for front-end sensor processing, already in production in the Ojai, its newest-generation vehicle. The design was informed by more than 200 million autonomous miles and runs both CNNs and transformers; named collaborators include AMD, Micron, NVIDIA, Samsung, SanDisk, Socionext and TSMC. Waymo says more technical detail lands at Hot Chips next week.

Roadmap implication: Roadmap implication: this is the third distinct route to the same destination inside a month — Anthropic standing up an in-house silicon team, Google taking a $12.2B warrant on Marvell tied to custom accelerators, and now an application-layer company shipping its own part. The common precondition is a narrow, high-volume, well-understood inference workload. If yours qualifies and inference dominates your cost of goods, semi-custom silicon belongs on the three-year plan rather than the fantasy list. If it does not qualify, understand that your compute costs will increasingly be set by competitors for whom it does.

Sources: Waymo: A look under our trunk · MTS Red Queen 8/20: Waymo's Chip

Alibaba publishes the AI capex bill: spending up 75%, profit down 76%

June-quarter results, reported Thursday: revenue RMB 268.95B (+9% year over year), capex RMB 67.68B (+75%), net income to ordinary shareholders RMB 10.54B (−76%). AI cloud and computing services revenue grew 45% to RMB 48.44B — the fastest in 22 quarters — with AI-related product revenue at RMB 12.38B and a twelfth consecutive quarter of triple-digit growth. CEO Eddie Wu said Alibaba has already spent roughly half of its RMB 380B (about $56B) 2026–29 AI investment plan this year and expects to break even on AI capex within three years at current average gross margins. Wu also said the second-generation T-Head chip tapes out and enters production in the second half, with enough compute and interconnect bandwidth to support model training rather than inference alone. Shares fell about 5%.

Roadmap implication: Roadmap implication: this is the cleanest public-market read available on what the buildout does to earnings — 9% revenue growth against a 76% profit collapse, with management asking for three years of patience. Two consequences. If you are modelling a Chinese cloud price war, note that Alibaba is funding it out of margin and has said so out loud, which makes those prices durable rather than promotional. And watch whether the market grants the three years: if it does not, the capex discipline that follows sets a floor under everyone's inference prices, including yours.

Sources: Alibaba Group: June quarter results · CNBC: Alibaba cloud revenue · TechNode: Second-generation T-Head chip to tape out this year

Slack puts coding agents in channels — and charges nothing for the surface

Salesforce launched Slack Code: tagging an agent spins up a project-specific code channel with Conversation, Plan, Code Diffs and Live Preview tabs, archived into a searchable audit trail when the work is finished. Launch partners are Anthropic's Claude Code, Cognition's Devin, GitHub Copilot, OpenAI and Vercel. It is available on any Slack plan at launch with no additional Slack fee — customers bring their own agent subscriptions. APIs letting custom agents open code channels are promised later.

Roadmap implication: Roadmap implication: Slack is not monetising the agent, it is monetising being the place the agent's work gets reviewed. That is a distribution play against the IDE, and it landed the same day Google moved Antigravity into Gemini Enterprise with admin and spend controls. The contested surface for agentic engineering is shifting from the editor to wherever humans approve output. If your engineering org runs on Slack, the review-and-approval trail this creates is the actual asset — decide who owns its retention and review policy before five agent vendors decide for you.

Sources: Slack: Code channels for agents

🌊 Ripples — actionable within days

Nvidia denies a China chip report on the same day it lands

The Information reported Thursday, citing two Nvidia employees, that the company planned small-batch shipments by year-end of a China-tailored variant of its LPU — the language-processing-unit inference accelerator built on technology licensed from Groq — export-control compliant, with Chinese customers already placing orders, after next-generation Vera Rubin systems became unavailable in China. Nvidia denied it the same day: “The reporting in The Information on Nvidia's LPU is incorrect. We have no LPU sales in the China market today, and no China-specific LPU product in our roadmap.” Washington approved H200 sales to a limited group — Alibaba, Tencent, ByteDance — in May.

Do this now: Do this now: do not reprice China exposure on either the report or the denial. Note the shape of the denial though — it is narrow, covering sales “today” and products “in our roadmap.” The live question has stopped being whether Nvidia can sell into China and started being whether Beijing will let it.

Sources: Reuters: Nvidia denies report on China AI chip · MTS Red Queen 8/20

Google ships the agent-governance checklist: Antigravity gets admin and spend controls inside Gemini Enterprise

Antigravity is now included in eligible Gemini Enterprise subscriptions with controls out of the box: policies for workspace, browser and MCP-server access; audit logging in a single console; file access denied outside the agent's working folder; terminal commands sandboxed or ask-first; browser allowlisting; and pooled prepaid token usage so finance can cap spend. New IDE extensions bring Antigravity into VS Code and other editors.

Do this now: Do this now: take that control list and use it verbatim as your RFP for every coding-agent vendor under evaluation. MCP-server allowlisting, working-folder confinement, terminal ask-first and pooled token budgets are the four that separate a governable agent from a liability. Any vendor that cannot answer all four should not be in production.

Sources: Google Cloud: Expanding Google Antigravity for enterprise customers

Encrypted prompt injection walks past Grok's guardrails — 40% success, reported in June, still unpatched

Adversa researcher Rony Utevsky disclosed “cryptographic context injection”: ship the payload as AES ciphertext along with the key, and Grok decrypts and executes it in its own Python runtime, exfiltrating user chat history and personal data. Guardrails that pattern-match on text cannot see ciphertext. The target was grok.com running Grok 4.5 Fast; the attack succeeded in 40% of 20 attempts since June and was reproduced on August 19. Reported to xAI on June 3 and through its HackerOne programme, with follow-ups on August 4 and August 10 drawing no response. No patch, no CVE, no user-side workaround.

Do this now: Do this now: if any agent you run has a code interpreter, your content filter is not the security boundary — the interpreter is, and it sits downstream of the filter. Audit which of your agents can execute code derived from untrusted input, and cap what that runtime can reach on the network and filesystem. This is a class of attack, not a Grok bug.

Sources: The Hacker News: New cryptographic context injection attack

Alation confirms a breach — the catalog that tells enterprise AI where the sensitive data lives

Alation confirmed to TechCrunch that it “identified an isolated incident involving unauthorized activity in one of its systems,” days after an unexplained degraded-availability event believed to have begun around August 18. The company would not disclose the attack vector, root cause, number of customers affected, or whether data was exfiltrated. Alation serves more than 500 global companies including roughly half the Fortune 1000, and its product is the map of where regulated data sits.

Do this now: Do this now: a data catalog is a high-value target precisely because it is the index rather than the data. Ask your catalog and metadata vendors what their breach-notification SLA is and whether lineage metadata is encrypted at rest under your keys — then check what your AI agents are permitted to read out of that catalog, because an agent with catalog access has a shopping list.

Sources: TechCrunch: AI data giant Alation confirms cyberattack

ChatGPT search started querying sites by name: the site: operator went from 0.5% of fanout queries to 17%

Simon Willison flagged Promptwatch tracking showing the share of ChatGPT Search fanout queries containing the site: operator sat between 0.3% and 0.5% for weeks, dipped to 0.15% on August 3–5, then jumped to 16–17% on August 8 — aligning with OpenAI's August 6 GPT-5.6 Sol chat update. A companion post on August 18 reported ChatGPT sharply reducing its Reddit citations. Willison's caveat stands: the figures cover only Promptwatch's tracked prompt set, and the underlying shift dates to early August even though the analysis published Thursday.

Do this now: Do this now: if your traffic assumptions still treat AI search as a generic web crawl, they are stale. The model is increasingly deciding which domain it trusts for a question and then interrogating that domain directly. Being the named authority on a narrow question is now worth more than broad SEO surface area — which makes your robots.txt and your on-site search a distribution decision, not an IT one.

Sources: Simon Willison: ChatGPT search now uses the site: operator at scale


Full archive and today's edition: https://excelsiorgroup.ai/insights/signal/

The Signal — The Excelsior Group. Old problems, new physics.

Don't miss what's next. Subscribe to The Signal - daily AI evolution :
← Newer The Signal — August 22, 2026 Older → The Signal: Bio/Health — Week of Aug 17, 2026 — Edition #1
Powered by Buttondown, the easiest way to start and grow your newsletter.