Pondero AI logo

Pondero AI

Archives
Log in
Subscribe
August 15, 2026

Pondero Brief: GPT-5.6 Sol hits 750 tokens a second with no IQ drop

Pondero Brief - AUGUST 15TH, 2026

Plus a Gemini Flash price cut, Microsoft's Copilot cull, and ChatGPT running ads on one in four searches.
pondero. BRIEF · AUG 15

GPT-5.6 Sol Ultrafast: 750 tokens per second, no intelligence tradeoff

OpenAI's Cerebras-powered preview breaks the speed-vs-smarts limit. Plus: Databricks at $190B, Gemini Flash at half the price, and Microsoft's Copilot cull.

OpenAI opened a limited API preview of Ultrafast mode for GPT-5.6 Sol on August 13 via Cerebras wafer-scale chips: 750 output tokens per second, 14x faster than standard according to OpenAI, no intelligence drop. On Humanity's Last Exam it finished in 11 hours against three-plus days for Claude Fable 5.

Also in today's brief

  • GPT-5.6 Sol Ultrafast hits 750 tokens per second
  • Gemini 3.7 Flash launches at half the price of 3.6
  • Microsoft kills consumer Deep Research on August 18
  • ChatGPT ads now on one in four commercial queries
  • Anthropic watermarks every Claude output globally
  • Databricks raises $5B, crosses $7B ARR at $190B
 
Models & Releases

Gemini 3.7 Flash launched at half the price of 3.6 Flash

Google shipped it August 13 at $0.75 per million input tokens and $3.75 per million output, half of 3.6 Flash's launch rate, while agentic code repair on DeepSWE v1.1 jumped to 65.3% from 49.0% (per 9to5Google). The introductory rate expires December 31, 2026, so budget your agent costs on post-holiday pricing you do not have yet. Google changelog.

 
DeepSeek V4-Pro GA illustration

DeepSeek V4-Pro hit general availability and raised its API prices

V4-Pro went GA August 13 with an 87.9 TerminalBench score, and DeepSeek open-sourced its Harness agent framework under the MIT license (per The Decoder). The catch: API prices went up alongside GA, so the cheap-DeepSeek math for your agent stack just changed. Read our take.

 
Money & Moves

Databricks raised $5B and crossed a $7B revenue run rate

The new round lands on $7 billion ARR in Q2, up more than 80% year over year, with over 1,000 customers spending $1M or more each per year (per Databricks). Positive free cash flow at that growth makes an IPO timeline credible, which matters if you are betting a data platform on their roadmap. SiliconAngle report.

 
Policy & Legal
Anthropic watermarks Claude outputs illustration

Anthropic is now watermarking every Claude output

Machine-readable watermarks in generated text and C2PA provenance metadata in generated images went global August 11, spanning Claude.ai, the API, Claude Code, and the AWS, Google Cloud, and Microsoft Foundry deployments (per Anthropic). If you ship Claude-generated text or images, assume they are detectable as AI-authored downstream. Read our take.

 
Tools & How-To

Microsoft is cutting consumer Copilot Deep Research

Four features go dark August 18, including Deep Research, AI-generated podcasts, and Group Chats; Deep Research becomes Researcher, gated to M365 Premium, so Personal and Family users can read old reports but cannot generate new ones (per TechCrunch). Saved podcasts and Group Chat shared images get deleted in the move, so export anything you need before the 18th. Data-loss details at The Register.

 
ChatGPT ads carousels illustration

ChatGPT ads now reach roughly one in four commercial queries

OpenAI rolled out product carousels and conversion-optimized CPC campaigns August 13 across the UK, Mexico, Brazil, Japan, and South Korea; one third-party study logged ads on 28.69% of healthcare queries (per Search Engine Land). If ChatGPT is a discovery surface for your product, it is now a paid channel with an auction behind it. Read our take.

 
Quick Hits
• Turn any URL into LLM-ready markdown for your RAG pipeline or agent scraper. Try Firecrawl →
• Managed cloud hosting for self-hosted AI tools - Open WebUI, n8n, Ollama - without babysitting a VPS. Try Cloudways →
• Build and deploy AI agents without code to automate repetitive work your team keeps redoing by hand. Try Relevance AI →
From the Pondero Stack
Enterprise Agent Reference Architecture illustration

The enterprise agent reference architecture, annotated

Our blueprint maps ingestion, orchestration, the tool layer, the eval harness, and the governance plane to the specific decisions your platform team hits in production. Start here if you are designing an agent stack from scratch. Read the guide.

 
Enterprise Agent Platform RFP Scorecard illustration

The agent-platform RFP scorecard

A scored sheet for grading agent platforms on harness architecture, eval maturity, governance controls, and lock-in risk. Use it to turn a vendor bake-off into a defensible number. Read the guide.

 
Enterprise Agent Deployment Patterns illustration

Six deployment patterns from real enterprise rollouts

Six patterns pulled from every named public enterprise agent deployment in 2026, drawn from case studies, earnings calls, and vendor disclosures on the record. Match your use case to a pattern before you build. Read the guide.

 

How was today's brief?

★★★★★ Nailed it  |  ★★★ Solid  |  ★ Missed

Jonathan Hildebrandt Jonathan Hildebrandt
Co-founder and primary operator of Pondero. Writes the Pondero Brief.

Affiliate disclosure  ·  Unsubscribe  ·  Manage preferences

Pondero earns commissions on some links. This does not affect our editorial picks.

Don't miss what's next. Subscribe to Pondero AI:
← Newer Pondero Brief: DeepSeek V4 hiked API prices up to 1,100%, ending the price war Older → Pondero Brief: Anthropic buys 20 years of compute from a bitcoin miner
pondero.ai
Bluesky
LinkedIn
Twitter
LinkedIn
Powered by Buttondown, the easiest way to start and grow your newsletter.