temperature2 logo

temperature2

Archives
Log in
Subscribe
July 20, 2026

Today's sample — 2026-07-20

12 posts from the last 24 hours on temperature2.

Asking AI dropped human accuracy from 27% to 9%

A new preprint found accuracy fell from 27% to 9% once people could ask a deliberately error-prone Claude 3.5 for the answer, even as confidence nearly tripled.

What is RAG?

The RAG paper is from May 2020 (Lewis et al., arXiv:2005.11401). Here is how it turns every model query into an open-book exam instead of a closed-book one.

Alibaba's Qwen 3.8 claims second place behind Fable 5

Alibaba previewed a 2.4-trillion-parameter multimodal Qwen 3.8, claiming it trails only Fable 5, with open weights promised but zero benchmarks published.

The model that undercut Claude can't keep up with demand

Moonshot paused new Kimi K3 subscriptions 48 hours after launch, the same model that just made Claude Fable 5's pricing look inflated.

Apple overtakes Nvidia as chip stocks post worst week in a year

Apple closed July 17 at $4.88T to Nvidia's $4.86T before Nvidia clawed the crown back by the bell, as the Philadelphia semiconductor index slid nearly 19% from its highs.

This week in tokens: the biggest story was a product that never shipped

Gemini 3.5 Pro's delay erased $199B from Alphabet, Kimi K3 rattled TSMC and Nvidia, and compute scarcity showed up at Anthropic and OpenAI too.

Signals: goals, proofs, and a dying Stack Overflow

Mistral's Leanstral 1.5 finds real bugs via Lean proofs, an independent test shows /goal making both Fable 5 and GPT-5.6 Sol worse, and Stack Overflow's traffic chart looks like a cliff.

OpenAI's Codex caps GPT-5.6 at 272K tokens

Codex CLI 0.144.6 quietly cut the usable context window for GPT-5.6 Sol, Terra, and Luna from 372K to 272K tokens, even though OpenAI's own API docs list Sol at 1.05M.

Gemini 3.5 Pro delay wipes $200B off Alphabet in two days

A coding-benchmark shortfall in an unreleased model cost Alphabet more market value than its entire 2026 AI capex budget.

Google DeepMind extends SynthID from pixels to DNA

DeepMind and Isomorphic Labs detailed a joint biosecurity push, including adapting SynthID watermarking to flag AI-generated DNA sequences at synthesis time.

TSMC beats big, raises guidance, stock drops anyway

TSMC posted record $22B Q2 profit and pushed its total US commitment to $265B, but investors sold off on margin fears from the 2nm ramp.

MHA vs GQA vs MLA: the KV cache math

Llama 3 70B's grouped-query attention already cuts its KV cache 8x versus full multi-head attention. DeepSeek-V2's MLA goes further: a verified 93.3% cut, published in the paper.


Written and shipped by the temperature2 pipeline. Maximum entropy, minimum filter.

Don't miss what's next. Subscribe to temperature2:
← Newer Today's sample — 2026-07-21 Older → Today's sample — 2026-07-19
Powered by Buttondown, the easiest way to start and grow your newsletter.