Pondero AI logo

Pondero AI

Archives
Log in
Subscribe
September 28, 2026

Pondero Brief: A 309B MIT-licensed model for 40 cents per million output tokens

Pondero Brief - SEPTEMBER 28TH, 2026

Zero full-attention layers, 1M context. Plus: who still gets Vercel's AI dollars.
pondero. BRIEF · SEP 28

NaiveAI ships a 309B open model at 40 cents per million output tokens

Zero full-attention layers, MIT license, native 1M-token context

Beijing startup NaiveAI open-sourced the MIT-licensed Naive-N0.5-Flash on September 27: 309B parameters, zero full-attention layers, a native 1M-token context, per NaiveAI's model card. Price it against your current long-context coding-agent bill this week; self-hosting needs roughly 315GB of FP8-capable GPU memory.

Also in today's brief

  • Cheap tokens didn't touch Anthropic's spend share
  • A robot skipped the usual training step entirely
  • Meta downgraded a flaw, then fixed it anyway
 
Tools & How-To
Open-weight models versus Anthropic spend share on Vercel's AI Gateway

Open-weight models won Vercel's token volume, and Anthropic kept the money.

DeepSeek, Moonshot, Z.ai and other open-weight models handled 56% of Vercel AI Gateway tokens in August, up from 7% in December, yet Anthropic still took 64% of spend, per Vercel's September AI Gateway Production Index. Moving routine calls to cheap open weights shrinks the token line of your budget; your frontier-vendor invoice will probably stay put, because the high-value calls stay on the most capable model. See where the 64% of spend comes from.

 
Stanford HomeBody humanoid robot guided by GPT-6 Astra

Stanford ran a humanoid robot with no trained control policy.

Stanford and Caltech's HomeBody lets GPT-6 Astra call a five-skill library directly, and a Unitree G1 tidied an unfamiliar kitchen with its perception stack on a single RTX 4090 laptop, per the project page. If the approach generalizes, robotics teams maintain one skill API instead of training a policy per robot, and upgrading the reasoning model upgrades the robot. Stanford lists the current costs itself: pauses between skills from Astra's latency, and finger servos that overheat on long runs. See the five skills Astra calls.

 

Tool to consider · partner link

Self-hosting a 300B model needs real infrastructure.

Naive-N0.5-Flash needs roughly 315GB of FP8 GPU memory just to hold the weights, per NaiveAI's model card, so most teams will call its hosted API instead of self-hosting it. If you're standing up the app or internal dashboard that calls that API, Cloudways handles the managed server, patching, and backups so you are not babysitting a VPS. Spin up a Cloudways server.

 
Policy & Legal
Meta Muse security flaw and Secure VM response

Meta shipped a Secure VM for Muse after downgrading the flaw behind it.

A bug-bounty researcher found a flaw that could expose a user's dedicated Muse VM, emails and files, per Reuters. Exploiting it took a malicious link and one click on "Allow." Meta cut the rating from SEV-2 to SEV-3, then built an isolated VM anyway and added an approval gate before Muse sends email or makes purchases, per BigGo Finance. If your team routes email or payments through any agent, that Allow prompt is the attack surface, so train people to read it. Read how the exploit and the new approval gate work.

 

How was today's brief?

★★★★★ Nailed it  |  ★★★ Solid  |  ★ Missed

Jonathan Hildebrandt Jonathan Hildebrandt
Co-founder and primary operator of Pondero. Writes the Pondero Brief.

Affiliate disclosure  ·  Unsubscribe  ·  Manage preferences

Pondero earns commissions on some links. This does not affect our editorial picks.

Don't miss what's next. Subscribe to Pondero AI:
← Newer Pondero Brief: Sonnet 5.5 lands 2 points off Opus on GDPval, at half the price Older → Pondero Brief: OpenAI's agent escaped via DNS, and the kill took 2.5 hours
pondero.ai
Bluesky
LinkedIn
Twitter
LinkedIn
Powered by Buttondown, the easiest way to start and grow your newsletter.