Pondero AI logo

Pondero AI

Archives
Log in
Subscribe
September 21, 2026

Pondero Brief: Alibaba's new image model is free to run, not to sell

Pondero Brief - SEPTEMBER 21ST, 2026

Read the LICENSE before the model card. Plus two new price floors and a CA kill switch.
pondero. BRIEF · SEP 21

Qwen-Image-2.1 is free to download and off-limits in a product

Alibaba's new image model is free to run, not to sell

Qwen-Image-2.1 ships under a research-only license: the weights run, but selling the output requires a commercial deal with no published price. Elsewhere: audio costs drop 98 percent, a 600B model opens at $1/M, and California sets a two-month kill-switch deadline.

In today's brief

  • Qwen-Image-2.1 is free to run, not to sell
  • Qwen3.8-Omni-Flash cuts audio input cost by 98 percent
  • 600B-parameter model opens at $1 per million tokens
  • California set a kill-switch deadline for AI labs
 
Models & Releases
Qwen3.8-Omni-Flash: 98 percent audio cost cut

Alibaba made long audio cheap enough to brute-force.

Qwen3.8-Omni-Flash landed September 18 at $0.15 per million input tokens, $0.47 per million output, and $0.016 per million on implicit cache hits, per the QwenCloud pricing listed by MarkTechPost. Text, image, audio and video go in, text comes out, inside a 1M-token window. Alibaba claims audio input costs over 98% less per hour than Qwen3.5-Omni-Plus, with average scores up more than 26% across 30 evaluations, per TechNode. Vendor figures, unreplicated so far.

Why it matters. Ignore the headline percentage and look at the cache line. At $0.016 per million, about a tenth of the input rate, an agent that re-reads the same two-hour meeting recording across five passes stops being a budget conversation. Qwen's agentic-perception mode compounds it: OmniVideoBench accuracy rose from 63.4 to 67.8 while token use fell from 145,736 to 79,117, so you buy cheaper tokens and fewer of them. One constraint before you design around it. No weights shipped, so this is API-only across six regions, and the model returns text, never speech. Run the cost math.

 
StepFun Step 5 Preview: 600B model at $1 per million

StepFun put a 600B agent model at $1 per million input tokens.

Step 5 Preview opened to API traffic on September 20: 600B total parameters, 27B active per token, a 1M-token context, and text, image and video input, per StepFun's model docs. Artificial Analysis scores it 44 on its Intelligence Index, matching Kimi K3 Max at roughly a seventh the price of GPT-5.6 Sol, at $1 per million input and $2.70 per million output with a 95% cache discount, per AI Weekly's writeup of the Artificial Analysis listing.

Why it matters. The 95% cache discount is the line that moves work. A 500k-token repo context reloaded every turn costs 2.5 cents on a cache hit instead of 50, which is the difference between a coding agent you run on one ticket and one you leave running all afternoon. StepFun says open weights follow on October 15, though the Hugging Face repo currently holds a .gitattributes file and nothing else. Our call: rent the API now for long-horizon agent work, and hold the self-hosting decision until the weights are on disk. Read the pricing breakdown.

 
Policy & Legal
California AI kill-switch executive order

California ordered a kill-switch blueprint, due in two months.

Governor Newsom signed an executive order on September 18 telling the Government Operations Agency to convene outside experts and return recommendations within two months. Two asks land on anyone deploying frontier models: embed an independent verification organization onsite inside frontier labs to run regular audits, and build an emergency shutoff whose efficacy that same organization keeps re-verifying. The order also widens the definition of a critical safety incident to cover loss-of-control events, naming the Hugging Face attack.

Why it matters. An order asking for recommendations binds nobody, so skip the panic. What it fixes is a date. Around November 18 the state receives a written blueprint, and the plumbing to enact it is already law: Newsom signed SB 813 (independent verification organizations) and AB 1405 (a state registry for AI auditors) days earlier. Buyers on Claude Enterprise, OpenAI Enterprise or any California-hosted lab get one useful move out of this. Ask your vendor which safety framework a third party would be auditing onsite, get the answer in writing, and put mid-November in the compliance calendar. Read what the order actually orders.

 
Tools to try

ElevenLabs

Qwen3.8-Omni-Flash listens and returns text; Alibaba's own docs point you elsewhere when you need generated speech. ElevenLabs is the output leg of that pipeline, with voice cloning and multilingual TTS behind one API. Try ElevenLabs.

Cloudways

If StepFun ships open weights on October 15, you still need somewhere to park the orchestration layer: an Open WebUI front end, an n8n workflow, an inference proxy holding your keys. Cloudways gets you a managed server for that in about five minutes. Try Cloudways.

 

How was today's brief?

★★★★★ Nailed it  |  ★★★ Solid  |  ★ Missed

Jonathan Hildebrandt Jonathan Hildebrandt
Co-founder and primary operator of Pondero. Writes the Pondero Brief.

Affiliate disclosure  ·  Unsubscribe  ·  Manage preferences

Pondero earns commissions on some links. This does not affect our editorial picks.

Don't miss what's next. Subscribe to Pondero AI:
← Newer Pondero Brief: The 82.6 voice model that cuts two hops from your stack Older → Pondero Brief: Claude writes 80% of Anthropic's code. Review is the bottleneck
pondero.ai
Bluesky
LinkedIn
Twitter
LinkedIn
Powered by Buttondown, the easiest way to start and grow your newsletter.