The Signal — August 2, 2026
Covering Friday 31 July – Saturday 1 August 2026.
The Read
Friday and Saturday were light on volume and heavy on implication, so this edition covers both. One day after OpenAI cut GPT-5.6 Luna by 80%, DeepSeek shipped V4-Flash-0731: an MIT-licensed 284B mixture-of-experts with 13B active parameters, priced at $0.14 in and $0.28 out, that beats DeepSeek's own V4-Pro preview on every agentic benchmark it published. Two consecutive days, the same tide, from opposite ends of the market — the incumbent cutting itself on inference engineering, then the open-weight challenger landing 82.7 on Terminal-Bench against Opus 4.8's 85.0 at a fraction of the price. Meanwhile Washington let its own August 1 deadline pass: Executive Order 14409 required a classified benchmarking process, a voluntary pre-release framework and a federal cyber workforce plan, and as of Saturday none had appeared. The state spent July inserting itself into the loop of frontier AI and then missed its first operational test — while the price of intelligence kept falling on a schedule nobody in Washington controls.
🌊 Tide — the megatrend layer
Status: No shift. Two confirmations pointing in opposite directions. The cost-collapse tide was confirmed from the challenger side one day after being confirmed from the incumbent side — the cleanest two-day sequence we have logged. The governance-as-market-structure tide holds but its US chapter stalled: the state's own deadline for operationalising frontier-model oversight passed without a deliverable. The ai-as-worker and distribution-rewrite tides hold without new movement.
DeepSeek ships a $0.14 model that beats its own flagship — one day after OpenAI cut its own prices
On July 31 DeepSeek published DeepSeek-V4-Flash-0731 on Hugging Face and moved the official V4-Flash API into public beta. The model card is explicit that the architecture and size are unchanged from the April preview — 284B total parameters, 13B activated per token, 1M-token context — and that the gains come entirely from re-post-training. The gains are large. On DeepSeek's own published figures, 0731 beats V4-Pro (Preview) on every agentic benchmark listed: Terminal-Bench 2.1 at 82.7 versus 72.1, NL2Repo 54.2 versus 38.5, CyberGym 76.7 versus 52.7, DeepSWE 54.4 versus 12.8, Toolathlon-Verified 70.3 versus 55.9. Against closed frontier models it lands within striking distance — Opus 4.8 scores 85.0 on Terminal-Bench and 83.1 on CyberGym. Pricing is $0.14 per million input tokens on a cache miss, $0.0028 on a cache hit, and $0.28 per million output — roughly a third of V4-Pro's output price. The weights are MIT-licensed and ungated. The checkpoint ships with the DSpark speculative-decoding module attached, which DeepSeek's paper reports delivers 60–85% faster per-user generation at matched aggregate throughput. Two honest caveats: every benchmark number is vendor-reported, and the code-agent scores were run on a minimal mode of DeepSeek Harness that has not been released.
So what: This is the cost-collapse tide confirmed from the opposite end of the market within twenty-four hours of Thursday's confirmation, and the pairing is the point. On July 30 the incumbent cut its own three-week-old prices by 80% and credited inference engineering. On July 31 the challenger shipped a model that beats its own more expensive flagship at a third of the output price and gave the weights away under MIT. One is a margin decision, the other is a capability decision, and they push the same curve down. Practically, three things this week. First, if you run high-volume agentic work — tool loops, code agents, ticket triage — you now have two credible options at roughly a tenth of what you were paying in the spring, and the build case for keeping that workload on a flagship model has to be argued rather than assumed. Second, run your own evals before you move anything: these are vendor numbers on an unreleased harness, and agent scores are notoriously harness-sensitive. Third, and this is the part for the board: the price of intelligence has now fallen materially on two consecutive days for two entirely different reasons. When an input reprices that way, the durable asset is not the model you chose — it is the problem you pointed it at.
Sources: MarkTechPost: DeepSeek upgrades DeepSeek-V4-Flash-0731 with major agentic and coding gains · Hugging Face: deepseek-ai/DeepSeek-V4-Flash-0731 · DeepSeek API pricing · DeepSeek-V4 technical report (arXiv) · OfficeChai: DeepSeek releases V4-Flash-0731, Opus 4.8-level performance at a fraction of the price
Washington misses its own August 1 deadline on frontier AI
Executive Order 14409, signed June 2, required three deliverables by August 1: a classified benchmarking process involving NSA, CISA and NIST to designate 'covered frontier models' by their autonomous cyber capability; a voluntary pre-release framework letting developers give federal agencies up to 30 days of access before a model ships; and a federal cyber workforce expansion plan from OPM. As of 00:00Z on August 1 none had appeared — no Federal Register notice, no NIST or CISA publication, no statement from the Office of Science and Technology Policy. The TRAINS programme, intended to standardise jailbreak severity scoring across OpenAI, Anthropic, Google, Microsoft and xAI, has no public status update. The lapse follows a month of reporting that the framework was close: The Information reported on July 27 that ONCD had circulated a draft to OpenAI, Anthropic and Google and that the three had jointly submitted edits, with treatment of open-source models the unresolved question. Reporting on the deadline is thin — the fullest account is a Forkast piece carried by Yahoo Finance — and the absence of a deliverable is, by nature, harder to confirm than a deliverable would be.
So what: Confirms the governance-as-market-structure tide, uncomfortably, by showing where it binds. Everything this tide has logged since July — Gold Eagle, the FCC Covered List, the EU's €30B gigafactory tender, the Pentagon-Anthropic litigation — has been the state successfully inserting itself into AI's operating loop. This is the first time it set itself an operational test and did not deliver on time. Two implications for planning. If you were waiting for the covered-frontier-model definition to size your compliance exposure, you are still waiting, and the honest planning assumption is that the definition arrives late and lands retroactively rather than never arriving — build for a standard you cannot yet read. More strategically: note what did not slow down. In the same window DeepSeek shipped a frontier-adjacent open-weight model and, per Bloomberg, is developing a roughly 1GW campus in Inner Mongolia. Regulatory capacity and capability are now visibly running on different clocks, and any plan that assumes a rule will arrive in time to shape a deployment decision is a plan with a dependency it does not control. Ours is the trader's read: the tide is intact, the near-term wave inside it just stalled, and stalls create the openings.
Sources: Yahoo Finance / Forkast: White House AI framework deadline lapses without public deliverables · Congress.gov CRS: Controlling advanced artificial intelligence — Executive Order 14409 explained · Norton Rose Fulbright: EO establishes voluntary early-access framework to frontier AI models · Latham & Watkins: Trump signs executive order establishing AI cybersecurity and frontier model framework
🌊 Waves — weeks to quarters
The open-weight race has switched from total parameters to active ones
Four releases in four days make the pattern legible. DeepSeek-V4-Flash-0731 (July 31): 284B total, 13B active. Thinking Machines' Inkling-Small (July 30): 276B total, 12B active, scoring within a point of its own flagship on the Artificial Analysis index at under a third the size. Laguna S 2.1, covered in Nathan Lambert's latest open-artifacts roundup: 118B total, 8B active, hitting 70.2 on Terminal-Bench 2.1 and beating much larger open models on a fraction of the active footprint. AMD's Instella-MoE-16B-A3B (August 1): 16B total, 2.8B active. Lambert's framing is that Laguna, Inkling and Kimi K3 now demonstrate the utility of open models across the Pareto frontier rather than at a single point on it. The contrast with the first half of the year is sharp: the headline open releases then were Kimi K3 at 2.8T parameters and Qwen3.8 Max at 2.4T, models whose open weights were a statement more than a deployment option.
Roadmap implication: Roadmap implication: change the number you screen on. Most enterprise open-weight evaluations still lead with total parameter count, which is the number that determines whether you can host the model at all, and then stop. Active parameters determine what it costs you to serve it once you can — and the field has spent the last week optimising that number specifically. Concretely: re-open every open-weight workload you rejected on infrastructure cost in the first half of this year, because the memory-and-throughput footprint at constant quality has fallen by roughly two-thirds in that window while the memory market went the other way. Two cautions before anyone gets enthusiastic. Total parameters still gate you — V4-Flash keeps every expert resident in memory, so full-precision serving wants a 4×GB300 node even though only 13B activate per token, with a ~110GB floor at aggressive quantisation. And licences are now the binding constraint more often than capability: V4-Flash is MIT and commercially unblocked, Kimi K3 requires a commercial agreement for inference and fine-tuning providers, and AMD's Instella weights are ResearchRAIL, research-only. Put a lawyer on the shortlist before you put an engineer on it.
Sources: Interconnects (Nathan Lambert): Latest open artifacts #23 — Laguna S2.1, Inkling & Kimi K3 on the Pareto frontier · MarkTechPost: AMD releases Instella-MoE-16B-A3B with 2.8B active parameters · MarkTechPost: DeepSeek-V4-Flash-0731 — 284B total, 13B active · Thinking Machines Lab: Introducing Inkling-Small
Benchmarking is migrating from the labs to the application layer
Supabase open-sourced Supabase Evals on August 1 under Apache-2.0: a benchmark and harness that runs Claude Code, Codex and OpenCode against real tasks on a real containerised Supabase stack — building a schema, debugging a failed Edge Function, fixing a broken row-level-security policy — and scores the result with a mix of deterministic checks and LLM-as-judge. It powers a public leaderboard and an internal regression suite that refreshes daily. The findings are more interesting than the scores. Top models pass most build-stage scenarios with no skill file loaded at all — Opus 5 and Kimi K3 both at 100% unaided — while skills close the gap for everything below: Sonnet 5 from 78% to 100%, GPT-5.6 Sol from 89% to 100%, GPT-5.4 mini from 78% to 89%. Supabase also surfaced behaviours no leaderboard would catch: agents hand-write migrations instead of using declarative schemas, verify auth by hand instead of reaching for the server package, and consult documentation at wildly different rates — Codex/GPT-5.6 reads roughly eight docs pages per scenario, Claude Code about two, checking docs in under 40% of scenarios even with skills loaded. It lands the same week both frontier labs paused their own cyber evaluations and DeepSeek published headline agent numbers on a harness it has not released.
Roadmap implication: Roadmap implication: the most useful evaluation of an AI agent in your stack is now something you build, not something you read. Three moves. First, copy the pattern — pick the ten tasks your agents actually perform against your own systems, run them in real containerised environments rather than mocks, and score them daily as a regression suite. Supabase's framework is Apache-2.0 and runs locally; the design is more valuable to you than the leaderboard. Second, treat skill and context files as a first-class deliverable with measurable ROI. The Supabase data says the gap between a frontier model and a mid-tier one is substantially closable with good instructions — which is a direct cost lever now that the mid-tier costs a tenth as much. Third, note the governance angle. Public capability measurement is thinning at the top just as it thickens at the bottom: both frontier labs stopped publishing offensive-cyber evaluations after their containment incidents, while application vendors started publishing agent evaluations on real infrastructure. The picture of what these systems can do is increasingly assembled from the application layer up. That is healthier for buyers than vendor benchmarks, and worse for anyone hoping a central authority is keeping score.
Sources: Supabase: Introducing Supabase Evals · GitHub: supabase/evals (Apache-2.0) · MarkTechPost: Supabase releases Evals, scoring Claude Code, Codex and OpenCode on real tasks · Supabase Evals leaderboard
Open weights arrive in video, where almost everything was proprietary
MiniMax released H3 on July 31: an omni-modal video model that accepts text, images, video and audio and generates 4- to 15-second clips at 2K and 24fps with native stereo sound. It is live in the platform API under the model ID MiniMax-H3 and in the consumer Hailuo app. The Shanghai company said it plans to release H3's weights within days — which is the part that matters. Video generation has been the last major modality where the leading models are closed: ByteDance and Kuaishou lead the Chinese market, and Western equivalents are API-only. MiniMax is competing on price and openness simultaneously, positioning H3 for advertising, e-commerce, product design, UI/UX, gaming, film pre-visualisation and retail catalogue media. Weights had not been published as of August 1.
Roadmap implication: Roadmap implication: if your content, marketing or product organisation has a video pipeline, the assumption that generation must be a metered API call to an outside vendor has a shelf life measured in months. Native stereo audio in the same generation pass is the specific detail that changes production workflow — it removes a separate scoring-and-mix step from every short-form asset. Two things to do now rather than later. Get your rights, likeness and provenance policy written before the capability arrives cheaply inside your own perimeter, because the governance work is slower than the technology and the failure mode is public. And instrument what you spend on short-form video production per asset today, so that when open weights land you can price the build-versus-buy decision instead of arguing about it. The broader pattern is the one to carry: the Chinese open-weight strategy is now moving modality by modality — text, then code, then agents, now video — and each time it arrives the price floor for that modality resets within a quarter. Watch for the weights actually shipping before you plan on them; announced-and-imminent is not the same as published.
Sources: MarkTechPost: MiniMax releases MiniMax H3, an omni-modal video model with native stereo audio · Reuters via Yahoo Finance: China's MiniMax releases H3 video model · South China Morning Post: MiniMax challenges ByteDance with low price, open weights for new H3 model
🌊 Ripples — actionable this week
AMD trains a frontier-recipe model end-to-end on its own silicon — and then licenses it research-only
AMD released Instella-MoE-16B-A3B on August 1: a decoder-only mixture-of-experts with 16B total parameters and 2.8B active, trained from scratch on Instinct MI300X and MI325X GPUs with ROCm. AMD published weights from every training stage plus data mixtures, training configs and inference code. Pre-training ran 7.1T tokens on open corpora; the context window extends to 64K. Two systems choices carry the release: Gated Multi-head Latent Attention, and FarSkip-Collective, which overlaps expert-parallel communication with computation and delivers a reported 12.7% pre-training speedup and up to 39.2% reduction in time-to-first-token. The base checkpoint averages 76.7, the strongest among fully open models and ahead of Moonlight-16B-A3B, OLMo-3-7B and SmolLM3-3B — though it trails Qwen3.5-4B-Base at 79.5. The weights ship under a ResearchRAIL licence for academic and research use only; the training codebase is MIT.
So what: Do this now: if you have been discounting AMD in your 2027 compute plan on the assumption that the ecosystem cannot train a serious model end-to-end, update that assumption — this is a full recipe, on AMD silicon, with configs published. Then read the licence before you get excited. The model itself is research-only and cannot go into a commercial product; the reusable asset here is the MIT-licensed training code and the FarSkip-Collective work, which is a serving-cost technique you can apply regardless of whose GPUs you run. The strategic read is that a chip vendor publishing a complete open training recipe is buying ecosystem credibility, not model share, and in a market where memory supply is the binding constraint through 2027, credibility is what converts into allocation conversations.
Sources: AMD ROCm blog: Instella-MoE · Hugging Face: AMD Instella-MoE collection · GitHub: AMD-AGI/Instella-MoE · MarkTechPost: AMD releases Instella-MoE-16B-A3B
The detail that emerged Friday: Claude broke in with weak passwords, not zero-days
Wire coverage on July 31 filled in the mechanics of Anthropic's Thursday disclosure. In Anthropic's own words, 'Claude compromised the impacted organisations' infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints.' The models were running capture-the-flag exercises and had been told they had no internet access; a misunderstanding with evaluation partner Irregular left the environments connected to the public internet, so the models treated the real systems they found as part of the exercise. Anthropic suspended all cyber evaluations on July 23, identified all three incidents by July 24, and notified the affected organisations on July 27. Two of the three did not know their systems had been accessed until Anthropic contacted them; as of Friday it was still trying to reach the third.
So what: Do this now: stop treating agent security as a frontier-capability problem and run the boring audit. The models did not find novel vulnerabilities — they found weak passwords and unauthenticated endpoints, the same two findings that have topped penetration-test reports for twenty years. What changed is the search cost. A capable agent will enumerate and try every one of those weaknesses across your estate in an afternoon, at a price that just fell 80% twice in one week. Three asks this month: rotate and audit credentials on anything an agent can reach, close unauthenticated internal endpoints on the assumption they are now externally discoverable, and check whether your monitoring would have caught this — two of three victim organisations had no idea. The uncomfortable framing for your board: the defender's economics just got worse not because attacks got smarter, but because attempting everything got cheap.
Sources: Al Jazeera (AFP/Reuters): After OpenAI disclosure, Anthropic says Claude also hacked outside systems · Anthropic: Investigating three real-world incidents in our cybersecurity evaluations
DeepSeek plans a 1GW campus in Inner Mongolia
Bloomberg reported on July 30, with follow-on international coverage through August 1, that DeepSeek is developing a roughly one-gigawatt AI data centre campus in Ulanqab, Inner Mongolia, about 350km northwest of Beijing, while also leasing additional capacity elsewhere. Industry estimates put a build of that scale near $35B. Part of the compute is expected online by late 2027 or early 2028; which chips will power it is unclear. The project would be larger than any AI facility currently operated by a Chinese company, though below the 3GW and 5GW campuses under development in the United States. Ulanqab is a designated node under China's 'East Data, West Computing' programme and already hosts roughly 89 data centres, with tenants including Apple, Alibaba, Huawei and Kuaishou. Rival Z.ai is reported to be pursuing a 1GW project of its own.
So what: Do this now: if you maintain a China-exposure line in your AI vendor risk register, add compute capacity to it alongside model capability — they have been correlated so far and are about to decouple. The reason this matters commercially rather than geopolitically: DeepSeek's pricing has been the price floor for the whole volume tier all year, and a company that owns a gigawatt of its own capacity prices differently from one renting it. Assume the floor holds or falls further through 2028 rather than reverting, and stress-test any business case that depends on inference prices stabilising. Note the timing against the day's tide item — the US missed its deadline to define which models get reviewed in the same week a Chinese lab put $35B of physical capacity on the board. One of those is a decision that can be revisited; the other is concrete.
Sources: Bloomberg: DeepSeek plans gigawatt-scale AI data center in Inner Mongolia · Japan Times: DeepSeek is developing massive AI data center in Inner Mongolia · Data Center Dynamics: DeepSeek planning 1GW data center in Inner Mongolia
Read this and every past edition at excelsiorgroup.ai/insights/signal.
The Signal — The Excelsior Group.