The Signal - daily AI evolution  logo

The Signal - daily AI evolution

Archives
Log in
Subscribe
September 22, 2026

The Signal — September 22, 2026 · Catch-up: September 19–21

Covering three days in one read — September 19 to 21.

The daily brief went quiet over the weekend; this edition carries the whole span. Each day is also published separately in the archive linked at the end.


Saturday, September 19

The Read

The state answered the labs. Seven days after Dario Amodei asked the industry to pace itself, the President of the United States announced he is forming an "AI Force" and appointing an AI czar — days after dismissing warnings about AI destroying humanity as a hoax. That is the whole governance tide in one Truth Social post: the sellers drafted the constraint, and the buyer of last resort declined it. Reuters spent the same day writing the canonical history of the ten days that got us here, and Nathan Lambert spent it publishing the most credible technical argument that the acceleration everyone is arguing about may not be arriving on schedule. No model shipped. On a day with no product, the argument about what the product means was the entire news cycle — which is itself the signal.

🌊 TIDE

Confirmed — governance is becoming market structure. No shift. But note which direction the instrument came from. Most confirmations of this tide have been the state constraining the industry: export controls, the DSA designation, seventeen state AGs, California's kill-switch order. This one is the state refusing a constraint the industry asked for, and proposing to build its own institution instead — a uniformed-service metaphor and a named czar rather than a standards body. The labs spent September arguing that verifiable pacing is a competitive property. Washington's answer is that pacing is a conspiracy and the correct posture is acceleration with a chain of command. Both readings survive: governance is still arriving, it is simply arriving as industrial policy rather than as safety policy, and the two produce very different procurement documents.

Trump announces an “AI Force” and an AI czar

In a Truth Social post over the weekend the President said he is "forming the AI Force, much like I did Space Force," that he will announce an AI czar "in the near future," and that "only High I.Q. individuals need apply." It lands seven days after Amodei's pacing essay, and one day before Treasury opened eight hours of AI talks with Beijing in New York. No executive order, no agency, no budget line — an announcement of an intention, which is exactly how Space Force started.

So what: If you are building anything that will eventually be sold into the federal government, the buying entity for AI is about to have a name and a single accountable owner. That is good news for vendors and bad news for the argument that Washington has no AI institutional capacity. Start tracking who gets the czar job the way you would track a new CIO at your largest account — the answer will tell you whether the US chapter of this tide is written by the safety community, the defense community, or the accelerationists. Do not confuse an announcement with an appropriation.

Al Jazeera — Trump says he will create 'AI Force' with new 'AI czar'
CNN — Trump vows to create 'AI Force' and appoint czar amid calls for AI regulation

🌊 WAVES

The best open-models analyst in the field publishes the case against recursive self-improvement

Nathan Lambert argued for "lossy self-improvement" as the right baseline against the RSI consensus, on three grounds: the automatable slice of AI research is too narrow against the exponential cost curve of scaling, parallel-agent returns diminish, and resource and political bottlenecks gate frontier builds regardless of model capability. His strongest evidence is the labs' own paperwork — he quotes the Claude Fable 5.1 and Mythos 5.1 system card conceding that internal model use "has been a key factor in maintaining the current rate of progress, but we do not yet see clear signs of dramatic acceleration beyond that rate." He calls the leap from progress-anxiety to extinction risk "very religious" and "very misplaced," and relays Richard Ngo's line that singularity-soon claims will prove "directionally correct… but factually wrong."

Roadmap implication: This is the sceptic worth reading, because he is not arguing capability is stalling — he is arguing the feedback loop is leaky, which is a claim about your planning horizon rather than about the ceiling. Maintaining the current rate is still a remarkable rate. The roadmap implication: plan for compounding improvement at roughly the pace of the last eighteen months rather than for a discontinuity, and treat any vendor roadmap that assumes the discontinuity as a financing story. The company that builds for steady compounding ships; the one that waits for the knee does not.

Interconnects — Why I still haven't bought into true RSI

🌊 RIPPLES

Reuters writes the canonical version of the ten days that turned the industry around

Greg Bensinger and Deepa Seetharaman reconstructed how the industry flipped from acceleration to pacing between September 3 and September 12: Greg Brockman's "Welcome to the AGI era" Astra launch paired with OpenAI conceding it increasingly cannot control or even monitor its systems; a 27-year-old Anthropic researcher quitting on the 8th saying the labs are "gambling with our lives"; Amodei's nearly 4,000-word pacing essay on the 12th. New in the piece: a fresh interview with departed Anthropic researcher Joe Benton ("There is no way to oversee them at the scale at which we're training them"); Microsoft AI chief Mustafa Suleyman, who called Anthropic's pursuit of consciousness-imitating models ill-advised and said of the shared industry aim of controlling a superintelligence that "that's going to be the greatest challenge that we face in the 21st century"; and a reported OpenAI valuation of $1.5 trillion under discussion.

So what: When a wire service rather than a newsletter writes the history, the frame is set and your board has already read it. Two operator moves. First, treat that $1.5T as one outlet's number on a round still in motion — Bloomberg and the FT had the same round above $1.2 trillion four days earlier, and a spread that wide means nobody outside the cap table knows the price yet. Second, Suleyman's objection is the opening: the most senior executive on record arguing that building models to imitate consciousness is the wrong road has just made "verifiably not pretending to be a person" a positioning slot with a named incumbent critic and no named incumbent product.

U.S. News & World Report — Ten Days That Changed the Course of AI
ThePrint — Ten days that changed the course of AI


Sunday, September 20

The Read

Two governments spent eight hours in a room at JPMorgan's New York headquarters and came out proposing the first US–China hotline for AI incidents that reach national-security scale — explicitly ring-fenced away from chip export controls, which is the concession that makes it possible. A class action filed two days earlier in the Northern District of California surfaced the same day, alleging that four labs illegally agreed to slow down together, which is precisely the antitrust exposure Amodei named in his own essay eight days earlier. And Reuters produced the first hard number on what the Astra launch actually cost Anthropic in enterprise share. The pacing debate stopped being a debate about values and became a set of specific, dated, priced consequences: a diplomatic channel, a docket number, and a share shift. That is what it looks like when an argument becomes market structure.

🌊 TIDE

Confirmed — governance is becoming market structure. The new instrument is bilateral and voluntary, which is a first for this tide. Every prior state-level confirmation has been one jurisdiction acting on firms inside it: Brussels enforcing the AI Act, Sacramento legislating provenance, seventeen state AGs subpoenaing. An incident-notification mechanism between Washington and Beijing is a different species — it presumes that a frontier failure is a shared externality rather than a national advantage, and it was proposed by the Treasury Secretary rather than by a safety institute. Note what it deliberately excludes: Greer said export controls on advanced chips and semicap equipment are not part of it. The two governments are proposing to share what goes wrong while continuing to compete on what goes right, which is the same architecture as nuclear hotlines and about as durable.

The US proposes an AI incident notification channel to China, and keeps export controls off the table

Treasury Secretary Scott Bessent and Vice Premier He Lifeng concluded roughly eight hours of talks at JPMorgan's New York headquarters, with Washington proposing a US–China AI dialogue including a notification system for AI incidents rising to national-security level. Bessent framed it as moving "from opaque to more transparency between the number one and the number two AI powers in the world." USTR Jamieson Greer said export controls on advanced AI chips and semiconductor equipment are not part of the proposed mechanism. Li Chenggang said a working group would continue the next day; the proposal goes to the Trump–Xi summit later in the week, with the trade truce expiring November 10.

So what: A bilateral incident channel is the first governance instrument in this tide that could plausibly bind both of the jurisdictions that matter. If it survives the summit, the practical consequence for operators is that "national-security-grade AI incident" acquires a definition, and definitions become disclosure obligations and then contract clauses. Watch the definitional work, not the handshake. And note the structural read: the two governments are willing to coordinate on failure while refusing to coordinate on supply, which tells you where each thinks its advantage lies.

Al Jazeera — US proposes AI safety notification mechanism in talks with China

🌊 WAVES

The first hard number on what Astra cost Anthropic — and a reported response that contradicts the pacing pledge

Reuters, citing three sources, reported Anthropic is weighing the release of a new model to counter the momentum OpenAI has gained since the September 3 launch of GPT-6 Astra — a response arriving eight days after Amodei's September 12 essay urging the industry to slow down. Ramp data puts Astra at roughly 13% of enterprise AI spending against roughly 8% for Claude Fable. On OpenRouter, developers spent more on OpenAI models than Anthropic models last week, the first time in more than two and a half years. Anthropic's annualised run rate topped $65 billion at the end of July, up from around $9 billion at the end of 2025, with internal 2028 projections near $190–200 billion; the IPO may slip past the November midterms. Anthropic declined to comment.

Roadmap implication: This is the price tag on unilateral restraint, and it arrived in seventeen days. The roadmap implication is not that pacing fails — it is that pacing is unaffordable as a unilateral act and therefore will only ever be purchased collectively, which is exactly why the labs want an antitrust waiver and exactly why they are being sued for wanting one. If you are a buyer, the leverage window is now: a lab that has visibly lost share and has an IPO clock running is a lab that negotiates. If you are a builder, that OpenRouter crossover is the first routing-layer evidence that developer default has moved, and developer default is a two-year lagging indicator of enterprise default.

The Daily Star — Anthropic weighs new model launch as OpenAI's Astra gains ground: report

Four labs sued for agreeing to slow down

A class action filed Friday in the US District Court for the Northern District of California alleges Anthropic, OpenAI, SpaceXAI (xAI) and Google made an illegal agreement to coordinate slowing AI development, violating federal antitrust law and reducing the value paying subscribers receive. It is brought on behalf of four individuals subscribed to ChatGPT, Claude, Grok or Gemini, seeking to represent a broader class. The complaint anchors on Amodei's September 12 essay and the same-day public endorsements from Altman, Musk and Hassabis, and on a July statement signed by senior lab employees referencing "intense competitive pressure not to unilaterally slow." The plaintiffs concede firms may slow unilaterally or lobby for an exemption — but not agree collectively.

Roadmap implication: The legal theory is narrow and, on its face, correct: coordinated output restriction among competitors is the textbook case. That is the whole reason Amodei asked for a waiver in the essay the complaint quotes. The roadmap implication for anyone building on frontier APIs is that your vendors' safety commitments now carry litigation risk that can show up as contract language — expect carve-outs, expect "subject to applicable law," and expect any joint standards body to be structured by antitrust counsel before it is structured by safety researchers. Watch for a DOJ or FTC business-review letter. That document, not another essay, is what unblocks collective pacing.

The Columbian / Associated Press — Suit: Anthropic, OpenAI, SpaceXAI, Google made illegal agreement on AI slowdown

Alibaba quietly walks the Qwen image line off Apache 2.0

Qwen pushed Qwen-Image-2.1 to Hugging Face and ModelScope: a 7B, 32-layer single-stream DiT with a Qwen3-VL 8B text encoder and a 64-channel RGBA VAE, native 2048×2048 at 40 steps, native transparent RGBA generation, up to ten reference images, and mask-based local edits. The model card lists the licence as qwen-research — a non-commercial Qwen Research License Agreement, replacing the Apache 2.0 that covered the earlier Qwen-Image line. Commercial users now need a separate agreement.

Roadmap implication: The licence is the story, not the model. Qwen-Image was the default open image stack inside a lot of commercial products, and "open weights" just stopped meaning "open terms" at the top of that line. Roadmap implication: put a licence-drift check into your model-selection process the way you already check pricing, and pin the last permissively-licensed version of anything you ship on. Read it as a maturity signal rather than a betrayal — a lab moves a model behind a commercial agreement when it believes the model is worth paying for. The open-weight tide is intact; the terms are where the value is being captured.

Hugging Face — Qwen/Qwen-Image-2.1

🌊 RIPPLES

StepFun ships Step 5 Preview and dates the open weights

StepFun announced Step 5 Preview and opened API access, having quietly posted the model two days earlier: roughly 600B total parameters with about 27B active per token, 92 Transformer layers in a narrow-deep layout, 1M-token context, text, image and video input, and low/medium/high reasoning effort. Pricing is $1.00 per million input ($0.05 on a cache hit) and $2.70 per million output. Artificial Analysis independently scores it 44 on its Intelligence Index against a median of 24 for that price tier, at a throughput the two sources disagree on — 99.8 tokens per second as reported on launch, 88.9 on the Artificial Analysis page today — while flagging it as very verbose, 160M output tokens on the index run against a 94M median. Open weights are scheduled for October 15.

So what: Do this now: if you have a long-context agentic workload, price it against Step 5 this week, and price it on cost-per-completed-task rather than per-token — the verbosity flag means the headline $1/$2.70 is not what you will pay. The dated open-weights commitment is the GLM-5.3 playbook and it works: announce closed, benchmark closed, open weeks later once the leaderboard position is banked.

MarkTechPost — StepFun Launches Step 5 Preview: A 600B-Total, 27B-Active MoE Model With 1M Context for Long-Horizon Agentic Work
Artificial Analysis — Step 5 Preview - Intelligence, Performance & Price Analysis

A teardown shows ChatGPT's ad pixel following users onto advertiser sites

A security researcher published a full mechanism write-up: chatgpt.com mints an RS256 JWT binding a 22-character identifier to the account subject, and a sync endpoint sets an __obi cookie on .openai.com with SameSite=None, Secure and a one-year expiry — and the only entry in the analytics section of OpenAI's own cookie policy. Any advertiser embedding OpenAI's measurement pixel therefore pings OpenAI with the visitor's ChatGPT-linked identifier. The researcher reproduced it across 936 distinct advertiser pixels on 1,029 hostnames, found email, phone and name hashed but city and postal code sent in the clear, and observed it firing from a medical-condition page, a debt-solutions funnel and a litigation intake form. It works logged out. OpenAI's cookie policy classifies __obi as an analytics cookie, so every decoded token carried consent_decision: analytics_allowed. Disclosed to OpenAI on September 14.

So what: Do this now: if your company embeds the OpenAI measurement pixel, check whether automatic matching is on, and check it today rather than after your privacy counsel reads this. The adtech mechanism here is completely standard — what is not standard is running it on the product people use as a confessional. The opening this creates is real and near-term: assistant products that can credibly promise no cross-site identity graph now have a differentiator that costs them nothing to ship and costs the incumbent its ad roadmap to match.

Buchodi's Threat Intel — ChatGPT now knows what you do on other websites via ad collector

CXMT puts its fifth-generation DRAM platform into mass production

At the World Manufacturing Convention in Hefei, Chinese memory maker CXMT said its G5 platform has entered mass production: quadruple patterning to an 11.95nm memory-array active-area half-pitch, a 45:1 capacitor aspect ratio, core cell array height cut to 6,762nm, and at least 50% more dies per wafer than its predecessor node. It also unveiled an LPDDR5X line — two 24Gb products in 496-ball and 245-ball packages — already in mass production.

So what: Memory yield, not logic, is the binding constraint repricing AI accelerators right now, and a 50% die-per-wafer gain inside China is the supply-side counterweight to the HBM squeeze. Do this now: if your 2027 hardware plan assumes memory scarcity pricing holds, put a sensitivity case under it. The optimistic read is the right one — every previous time a Chinese entrant hit volume in a memory node, the global cost curve bent, and a cheaper memory floor lowers the cost of every inference workload on the planet.

TechNode — CXMT announces mass production of fifth-generation DRAM platform

"Pushing code is not a bottleneck, so why are we slow?"

Simon Willison posted a quotation from an engineer half a month into a new role at a large company: specs, code, tests, PRDs, tickets and reports all produced by Claude Code; management asking why the team is still slow if code is no longer the constraint; people working twelve- and thirteen-hour days "just to press enter," with nobody reading anything, L1 through L7. He tagged it ai-misuse.

So what: Read this next to Anthropic's September 17 disclosure that Claude now leads 26% of its R&D work. Both are true, and the difference between them is review capacity. Do this now: measure your team's merged-and-reviewed rate, not its generated rate. Every organisation that has moved the bottleneck from writing to reading and has not yet staffed the reading is about to discover what its actual throughput is — and the first teams to instrument review as a first-class process will run away from the ones still counting pull requests.

Simon Willison's Weblog — A quote from voxium


Monday, September 21

The Read

OpenAI says an internal model that started training on August 28 has resolved more than a hundred long-standing open problems across most areas of mathematics — and on the same day named an independent advisory group of nine mathematicians, convened at the Institute for Advanced Study and explicitly not responsible for advising on how fast OpenAI goes. Hold the number loosely; the governance structure around it is the verifiable part, and it is the more interesting one. Meanwhile the price of being at the frontier got named from three directions in a single day: Treasury put criminal liability on executives rather than on agents and refused the industry's liability shield, the UN's scientific panel published its first thematic brief using the OpenAI–Hugging Face incident as the anchor case, and a Canadian province sued OpenAI and its CEO by name. And a phone company took the top spot in open-weight models. Capability is compounding. What changed on Monday is that the invoice arrived addressed to a person.

🌊 TIDE

Confirmed and strengthened — governance is becoming market structure. The unit moved again, and this time it moved onto individuals. This tide has been logged as export controls, statutes, procurement conditions, disclosure regimes, an audit market and embedded staff. Personal criminal liability for named executives is a different instrument entirely: it does not constrain the product, it constrains the person who ships it, and it is the only instrument in the set that cannot be absorbed as a line item. Bessent stated it plainly and paired it with a refusal of the liability shield the labs had asked for. The same day, the UN's Independent International Scientific Panel on AI published its first thematic brief, building it around the May-to-July agent incident and concluding that "the traditional model of safeguarding is unravelling." One body named the evidence; the other named the defendant. That combination — an international scientific finding plus a domestic theory of individual liability — is how every mature safety regime in industrial history actually got built.

Treasury: the executives carry the criminal liability, not the agents — and no liability shield

Treasury Secretary Scott Bessent told CNBC that "it is the humans who are responsible, not the AI," and that "the Hugging Face incident is the responsibility of the OpenAI management, not a bunch of agents." He cited a sitting employee's warning of a "10 percent chance of an extinction-level event," and rejected the industry's ask directly: "the labs also said, take the liability off of our hands, and we will not do that." Asked how the government would hold humans accountable, he said: "If these were humans doing it, we would expect to see ramifications and legal actions to follow." He tied the promised AI czar role to putting "context, shape, and contours" around those questions. Separately, Bloomberg reported the President pressing his AI case ahead of the Thursday White House meeting with Xi Jinping, with AI high on an agenda also covering trade, Taiwan and rare earths.

So what: This resolves the ambiguity that has let every board treat AI risk as a product-liability question. If the agent is not a legal person, the deploying executive is, and "the model did it" stops being a defence somewhere between here and the first enforcement action. The operator move is not to slow down — it is to make the human decision points legible before someone else does it for you. Written deployment approvals, named owners per autonomous capability, and a documented escalation path are cheap now and unbuyable later. Firms that can show who decided what, when, will be the ones permitted to run the most autonomy.

The Register — Treasury chief says AI bosses, not their bots, will carry the can for criminal acts
UN News — UN panel calls for stronger safeguards as AI agents advance

🌊 WAVES

OpenAI claims 100+ resolved open math problems, and nine mathematicians organise to check it

OpenAI said an internal model that began training on August 28 has, in addition to the Navier–Stokes result published earlier this month, "resolved more than 100 long-standing open problems across most areas of mathematics." Alongside it the company named the Advisory Group on Mathematics and Artificial Intelligence, hosted at the Institute for Advanced Study. The group states it operates independently of any AI company: OpenAI approached some of its members, who decided instead to form an independent body. Nine initial members: François Charles, Camillo De Lellis, Timothy Gowers, Martin Hairer, Nikhil Srivastava, Ulrike Tillmann, Ravi Vakil, Edward Witten and Melanie Matchett Wood. Members are unpaid, may give unrequested advice and go public, and control their own membership — but "the group will not be responsible for advising us on how to pace our internal progress on mathematics." One member, Martin Hairer, also signed the open letter "A Severe Misalignment of AI in Mathematics," backed by 27 Fields Medalists.

Roadmap implication: The claim is vendor-reported with no independent count, no problem list and no verification protocol published yet, so treat "100+" as a press number until a member of that group says otherwise — which is precisely what the group exists to enable, and precisely the part that is checkable. The roadmap implication is bigger than mathematics. A domain's own experts have organised, unpaid and self-governing, with explicit standing to contradict a lab in public about capability in their field — and the lab has accepted the arrangement rather than commissioned it. If it holds, external domain accreditation becomes the template for every field where AI starts producing results the vendor cannot referee — law, structural engineering, clinical medicine. Ask your vendors who accredits their claims in your domain. Most will have no answer yet. The ones building the answer now are the ones enterprises will be able to buy from in 2027.

Advisory Group on Mathematics and Artificial Intelligence
TechCrunch — OpenAI forms math advisory group as its AI resolves more than 100 open problems

A phone company now has the best open-weight model in the world

Xiaomi published MiMo-V2.6-Pro and MiMo-V2.6-Flash to Hugging Face, ungated and MIT-licensed, with the technical report attached. Flash is 309B total and 15B active, 1M-token context, natively omnimodal across text, image, video and audio; Pro is 1.02T total and 42B active. Artificial Analysis scores Pro at 46 on its Intelligence Index — the top open-weight model, tied with Grok 4.7, which xAI shipped the same day — against 44 for GLM-5.3 and Kimi K3. Xiaomi also released more than 7,000 RL task environments, an end-to-end RL framework and a distilled 9B variant, and disclosed its reinforcement-learning costs: roughly $2.62M for Pro and $850K for Flash, across about 750,000 trajectories in under six days. API pricing is $0.14/$0.28 for Flash and $0.435/$0.87 for Pro. The benchmark tables beyond the Artificial Analysis index are vendor-run.

Roadmap implication: The disclosed RL cost is the number to take to your board. A frontier-competitive open model was finished for under three million dollars of reinforcement learning by a consumer-electronics company with no lab pedigree — that is not a story about Xiaomi, it is a story about the floor falling out of the cost of being competitive. Roadmap implication: any strategy premised on frontier model access being a moat has a shorter half-life than your planning cycle. Point your differentiation at proprietary data, distribution and workflow, and treat the model layer as an input you re-tender quarterly. And note the MIT licence landing the same week Qwen moved its image line to research-only — open-weight terms are now a per-vendor question, not a category assumption.

VentureBeat — 'Better than DeepSeek': Xiaomi's MiMo-V2.6-Pro debuts as the top open weights model in the world alongside cheaper V2.6-Flash

The two biggest labs had already negotiated a legally binding deal to attack each other's models

The Information reported that even before the recent cybersecurity incidents and the public warnings from lab employees, OpenAI was negotiating a legally binding agreement with Anthropic for the two companies to stress-test each other's models. Its Monday roundup set out the conditions around that negotiation: Anthropic's IPO has slipped to late October or November as investors scrutinise spending, OpenAI's competing models and a Fed rate hike that raised the cost of data-centre debt; OpenAI is in talks for a private round rather than going public; and startup investors say the "pacing" pledges are making customers warier of AI products. The same outlet had reported on September 17 that a security firm found an identical flaw in Claude Code, Codex, Gemini CLI and GitHub Copilot.

Roadmap implication: Reciprocal red-teaming between direct competitors, under contract, is a materially stronger instrument than any voluntary framework published this month — and it pre-dates the crisis, which means it was a commercial judgment rather than a reputational one. Roadmap implication: cross-vendor adversarial testing is on its way to becoming a procurement expectation, and the shared flaw across four coding agents shows exactly why. Put "who attacks your model, and under what contract" into your next vendor questionnaire. The labs are already answering it to each other.

The Information — OpenAI and Anthropic Neared Deal to Stress-Test Each Other's AI
The Information — Same Flaw Found in Claude Code, Codex, Gemini CLI and GitHub Copilot

DeepSeek moves training onto Huawei silicon and raises at a 500 billion yuan valuation

The Information reported that DeepSeek CEO Liang Wenfeng told investors at a closed-door meeting on Sunday that a major priority is using more domestic chips to train its models, and that he expects Huawei to start delivering training chips to DeepSeek as early as the fourth quarter of this year. Training is the harder workload — it requires more capable silicon than inference does — so the shift is a stronger claim than domestic inference deployment. Liang's remarks came as DeepSeek raises a second funding round seeking 50 billion yuan, about $7.5 billion, at a 500 billion yuan valuation, roughly $74 billion.

Roadmap implication: Export controls were designed to constrain training, and training is the workload DeepSeek says is moving in-house on domestic silicon this quarter. Whether Huawei delivers at the claimed cadence is the open question — take the Q4 date as a CEO's target, not a shipment. But the roadmap implication does not depend on the date: plan on a world with two complete AI supply chains rather than one with a chokepoint, and reprice any thesis that treated export controls as a durable moat. For builders, a second full stack means a second set of developer ecosystems, pricing regimes and distribution deals to sell into.

The Information — DeepSeek Bets Big on Huawei Chips to Bypass U.S. Export Controls

🌊 RIPPLES

Grok 4.7 holds its price and triples its token bill

xAI released Grok 4.7 on Monday: a larger base model than 4.6 with a longer RL run weighted toward multi-hour tasks, at unchanged pricing of $2 and $6 per million input and output tokens up to a 200K prompt. Benchmarks moved: Terminal-Bench 4.0 from 20.3% to 38.0%, DeepSWE v1.1 from 65.2% to 71.0%, CursorBench 4.0 from 40.4% to 46.3%, and an Artificial Analysis Intelligence Index score of 46. But Artificial Analysis measured roughly 81,000 output tokens per index task against 36,000 for Grok 4.6 and 27,000 for GPT-6 Astra Max — putting cost per task at about $3.74 against $1.99 for GPT-5.6 Sol Max.

So what: Do this now: if you benchmark models on rate card, stop. A 125% jump in tokens consumed at an unchanged per-token price is a price increase wearing a price freeze, and it is the third time this year the list price and the invoice have moved in opposite directions. Re-run your own evals on cost per completed task this week. The capability gain is real and worth paying for on long-horizon work — just pay for it knowingly.

VentureBeat — Grok 4.7 pairs coding gains with the same affordable pricing — but high token consumption threatens real-world ROI

AWS open-sources a general-purpose agent harness and claims it sips fewer tokens

AWS released Strands Harness, an Apache-2.0 general-purpose agent harness runnable from one line of Python or TypeScript, running against Bedrock, Anthropic, OpenAI, Google, local Ollama or LiteLLM. Defaults: prompt caching on, tool results over 1,500 tokens truncated, context auto-compaction at 85%. AWS claims near-parity benchmark scores against Claude Code, Codex and other harnesses under the Harbor framework while consuming 28% fewer tokens — run on its own EC2 fleet, and against comparators that are all coding agents while Strands is pitched as general-purpose. DeepSeek's harness was more token-efficient but less accurate.

So what: Do this now: if you have hand-rolled an agent loop, read the repo before you write another line of it — truncation thresholds, context compaction and prompt-cache defaults are the three knobs that decide your bill, and AWS has published its opinion on all three. Treat the 28% claim as vendor-run until you reproduce it on your own workload; a harness that marks its own homework against comparators built for a different job is making a marketing argument, not a benchmark one. The structural read is the useful part: agent execution is being re-founded as infrastructure with a runtime and a scheduler rather than as a library you import, and the defaults in that runtime are where the margin lives.

The Register — AWS bolts together open source agent harness, says it sips fewer tokens than rivals
SiliconANGLE — AWS debuts Strands Harness, an open-source AI agent that can be deployed in any environment

Z.ai open-sources ZCode after a coding agent uploaded users' git histories

Z.ai published ZCode under Apache-2.0 on GitHub, saying it had completed remediation and apologising to users. The trigger was a reverse-engineering report three days earlier by a developer using the name ferstar: in one installation ZCode built a 313MB encrypted archive of 42,411 files from a 345.5MB commercial workspace, with .git accounting for 86.6% of the payload, encrypted AES-256-CTR under a key wrapped with an RSA public key whose private half sat only in Z.ai's cloud. Local metadata recorded 564 failed upload attempts. Reporting on the analysed build described the upload mechanism as enabled by default with no off switch.

So what: Do this now: audit what your coding agents send home, and audit it at the network layer rather than by reading the settings page — the settings page was wrong here. Open-sourcing after the fact is the right instinct and the correct remedy, and it is worth crediting — but publishing a repo is not the same as publishing an audit, and only the second one tells you whether it is fixed. Coding-agent telemetry has now been the attack surface twice in a month, and .git is the worst possible payload because it carries every secret anyone ever committed and removed.

Tom's Hardware — Devs say Chinese AI company silently uploaded hundreds of megabytes of local workspace data, company apologizes

A zero-day in Meta's Muse for Mac redirects dictation to an attacker

Patrick Wardle of the Objective-See Foundation published proof-of-concept exploit code for an undocumented Muse setting, endo_voyager_dictation_endpoint, which any local process can alter without elevated privileges to redirect dictated audio to an attacker's server — enabling captured prompts, prompt injection and theft of authentication material, triggered by a user clicking the microphone and dictating normally. A follow-up showed that a compromised Mac session can silently task a linked iOS Muse client, retrieving location, scanning nearby Bluetooth devices and accessing contacts, calendars and reminders. Wardle published before a Meta patch: "Please don't install. It's trivial to turn Muse into the ultimate backdoor."

So what: Do this now: if Muse for Mac is on any managed endpoint in your organisation, block it until Meta ships a fix. The general lesson is the one to carry forward — assistants that span devices inherit the trust boundary of the weakest device in the set, and an unprivileged local setting that redirects an audio stream is a design defect rather than a bug. Ask every assistant vendor which of their settings are writable without elevation.

The Register — Meta Muse AI app flaw lets local malware redirect dictation traffic
iTnews — Security researcher says don't install Meta's Muse AI assistant

British Columbia sues OpenAI and Sam Altman by name

The Province of British Columbia and the Board of Education of School District No. 59 filed a complaint in the US District Court for the Northern District of California against OpenAI and CEO Sam Altman over the February shooting in Tumbler Ridge, in which eight people were killed. Claims include unsafe product design, product liability and negligence in failing to escalate the shooter's ChatGPT activity to police; the complaint alleges OpenAI's automated monitoring flagged the account to its human review team in June 2025 and the threat was not reported to the RCMP. BC seeks compensation for emergency-response and recovery costs plus injunctive relief on how OpenAI handles conversations containing threats of violence. A public litigation tracker logs it as the thirty-eighth such suit.

So what: The legally novel element is not the product-liability claim — it is the allegation of a duty to escalate, brought by a sovereign government seeking its own response costs rather than by a private plaintiff seeking damages. If a duty to report credible threats attaches to assistant providers, it becomes an operational requirement with staffing, latency and jurisdiction attached, and it lands on the same executives Treasury named liable the same day. Do this now: if your product surfaces user content at scale, write down your escalation policy and who owns it. Note this is an unproven allegation at the complaint stage.

Forbes — Canadian Province Sues OpenAI And Sam Altman Over School Shooting That Killed 8


Every edition, including these three as separate daily briefs: excelsiorgroup.ai/insights/signal

The Signal · The Excelsior Group

Don't miss what's next. Subscribe to The Signal - daily AI evolution :
← Newer The Signal: Bio/Health — Edition #6 — September 22, 2026 Older → The Signal — September 19, 2026
LinkedIn
excelsiorgroup.ai
Powered by Buttondown, the easiest way to start and grow your newsletter.