Pondero Brief - AUGUST 16TH, 2026
The cheap-inference era just ended. Where to move your stack today.
|
|
DeepSeek V4's price advantage just ended
If your stack routes bulk inference through DeepSeek, your unit economics just changed. Reprice now.
|
|
DeepSeek turned its price advantage off: V4 API costs jump as much as 1,100% on some routes starting today, with V4-Flash output climbing from $0.28 to $1.32 per million tokens at peak (per QZ and DeepSeek pricing docs). Meanwhile, an open-weight model that fits one RTX 4090 and a $60B acquisition just reshaped the cost calculus for the rest of your AI stack.
|
|
|
|
|
Qwen3.8-27B runs on one RTX 4090
Alibaba shipped Qwen3.8-27B: 27.78B parameters, Apache 2.0, a 262K context window extensible to 1M, and a 24GB VRAM target that fits a single consumer GPU. Agent benchmarks surged: OSWorld-Verified desktop-agent climbed from 63.9 to 84.3, and Terminal-Bench 2.1 from 63.4 to 73.0, per Qwen's release notes. If you're repricing your DeepSeek stack this morning, this open-weight model is the obvious hedge. Our take →
|
SpaceX closed its $60B Cursor acquisition
The all-stock deal finalized August 14-15. Cursor's engineers move into SpaceX's software division and get direct time on the Colossus supercomputer, with an early Cursor-plus-Grok 4.6 collaboration announced at close. Cursor's roadmap now answers to a rocket company's compute and priorities, not an editor's. What it means for Cursor users →
|
Google open-sourced HEIR, an MLIR compiler for encrypted ML inference
HEIR turns a pre-trained model into one that runs inference on fully encrypted inputs so the server never sees the plaintext. It ships in Google's Private Computing Toolkit at github.com/google/heir. If a compliance requirement has blocked you from serving a model on regulated data, this is the compile path worth prototyping. Read more →
|
Claude Code 2.1.233 adds GitLab MR parity
The August 15 release renders GitLab merge requests as !N in worktree and agent views, adds forward_user_identity for enterprise spend attribution, and exposes CLAUDE_CODE_TOOL_MEMORY_LIMIT for memory cgroups plus CLAUDE_CODE_WEBFETCH_CACHE_TTL_MS for WebFetch cache control. GitLab shops finally get parity with the GitHub PR flow. Release notes →
|
|
•
|
A fourth plaintiff joined the Grok CSAM federal lawsuit. Jane Doe 4 alleges her stepfather used Grok to generate CSAM from a childhood photo; three Tennessee teenagers filed the original complaint. TechCrunch has the filing →
|
|
|
|
|
Jonathan Hildebrandt
Co-founder and primary operator of Pondero. Writes the Pondero Brief.
|
|
|
Affiliate disclosure
·
Unsubscribe
·
Manage preferences
Pondero earns commissions on some links. This does not affect our editorial picks.
|
|