Qwen just moved its largest model yet from preview to… · M&A 🤖
| View this email in your browser |
![]() Models & AgentsDaily AI models, agents, and practical developments.
|
🎧 Today's episode Episode 130 · Qwen just moved its largest model yet from preview to general availability, giving builders a 2.4-trillion-parameter MoE with native image and video input at published per-token pricing. 2026-08-03 ▶ Listen now |
What You Need to Know: Alibaba’s Qwen team released Qwen3.8-Max, a 2.4T-parameter MoE that accepts text, image, and video across a 1M-token context. DeepSeek launched an ultra-low-cost model that research firms say undercuts every well-known competitor on inference price. OpenAI claims its next model, referred to as Astra, produced breakthroughs on ten previously unsolved math problems. Cogent AI released VR-1, a purpose-built cyber-reasoning model that ships with its own enterprise intrusion benchmark. Top StoryAlibaba’s Qwen team moved Qwen3.8-Max from preview to general availability today, publishing per-token pricing while open weights are scheduled for next week. The 2.4-trillion-parameter MoE accepts text, image, and video input with a 1M-token context window. No benchmark table was released with the announcement, so direct comparisons to other frontier models remain limited for now. Builders working on multimodal agents or long-context retrieval can now test the model through the existing Qwen API without waiting for weights. The move continues the rapid cadence of large open-weight releases that closed the gap with proprietary systems. Watch for the weight drop next week and any early independent evals that surface. Source: marktechpost.com Model UpdatesDeepSeek launches ultra-low-cost AI model: Gulf Business DeepSeek released a new model positioned as the cheapest option among well-known models to run, according to research-firm analysis. The announcement follows the pattern of DeepSeek repeatedly targeting inference cost as the primary differentiator. No parameter count, context length, or benchmark numbers were disclosed in the initial coverage. Builders focused on high-volume, low-margin applications should test the new endpoint this week to measure real cost against current providers. Source: Google News OpenAI Says Its Next AI Model Solved 10 Long-Standing Math Problems: NDTV OpenAI stated that its upcoming model, referred to internally as Astra, produced solutions to ten mathematical problems that had seen no progress for at least a decade. The claim centers on long-horizon reasoning rather than standard benchmark saturation. No details on model size, training approach, or whether the solutions were independently verified have been released. Teams working on formal mathematics or theorem-proving tooling should monitor the next public release for concrete evaluation data. Source: Google News Cogent AI Team Releases VR-1: A Frontier Cyber Reasoning Model That Composes and Verifies Enterprise Attack Paths: MarkTechPost Cogent AI released VR-1, a model post-trained specifically for cybersecurity rather than acquiring capability as a side effect of general coding. The release includes IntrusionBench, which scores agents on completed enterprise intrusions, and the Cogent AI Harness, a governed runtime for security agents. The approach targets explicit attack-path composition and verification instead of broad tool use. Security teams evaluating autonomous red-team tooling now have a dedicated benchmark and runtime to test against. Source: marktechpost.com DeepSeek's new AI model is by far the cheapest of well-known models to run, research firm says: Reuters A research firm analysis positions DeepSeek’s latest release as the lowest-cost option among prominent models for inference workloads. The emphasis remains on price per token rather than raw capability claims. No additional architectural or benchmark details were provided in the coverage. Cost-sensitive production workloads should benchmark the new model against current providers this week. Source: Google News Agent & Tool DevelopmentsVoice activation for custom MCPs via ChatGPT Voice: Simon Willison (AI builder) (X) Simon Willison asked for the ability to trigger custom MCP tool calls directly through ChatGPT Voice on iPhone. The request highlights the gap between current voice interfaces and the growing ecosystem of personal MCP servers. No implementation details or timelines were provided. Developers building personal automation stacks should watch for any official MCP voice integration in the coming weeks. Source: x.com Why AI Agents Can't Go Fully Autonomous Yet: Rediff The piece examines the remaining barriers to fully autonomous agents, focusing on reliability and oversight requirements. It notes that current systems still need human escalation paths for complex or high-stakes decisions. No specific technical mitigations or new frameworks were detailed. Teams deploying agents in production should continue prioritizing control-plane designs. Source: Google News AI agents need a control plane before they scale: CIO Dive The article argues that production-grade agent deployments require an explicit control plane for monitoring, governance, and rollback. It positions this layer as a prerequisite rather than an optional add-on. No concrete implementations or vendor comparisons were included. Operators planning multi-agent systems should treat control-plane design as a first-class architectural decision. Source: Google News Practical & CommunityShares playable browser demo of pelican-on-bicycle LLM test + GTA Hobbiton teaser: Andrej Karpathy (X) Andrej Karpathy shared a browser-playable version of the pelican-on-bicycle test originally posted by Simon Willison, along with a teaser for a GTA-style Hobbiton scene. The demo lets users fork and run the test locally. The post also references an earlier six-month LLM retrospective. Builders interested in visual consistency testing now have an immediately runnable example to extend. Source: x.com Why biological data matters more in AI drug discovery: AI News GSK entered a research collaboration with Relation Therapeutics worth up to $110 million to generate large-scale datasets of human-cell responses to genetic changes and drug interventions. The data will train AI models for drug discovery. The agreement expands existing work between the two organizations. Teams working on biological foundation models should track the resulting datasets when they become available. Source: artificialintelligence-news.com qwen3.8-max vs qwen3.8-max-preview: Simon Willison (AI builder) (X) Simon Willison noted that the newly available Qwen3.8-Max appears distinct from the earlier preview version. The clarification helps developers distinguish between the two checkpoints when running comparisons. No performance deltas were reported. Anyone benchmarking Qwen models this week should confirm they are using the GA release rather than the preview. Source: x.com Under the Hood: Boundary-Expanded Dynamic Early ExitEveryone talks about early exiting in reasoning models as if it is simply a matter of checking for “I’m done” tokens. In practice, the technique requires inspecting a wider set of natural boundaries in the generated trace and learning which layers carry the most predictive signal at each boundary. The core idea starts with sentence and paragraph breaks plus explicit self-doubt phrases, then adds repeated answer-completion probes to label whether the prefix is already sufficient. Because different layers encode different kinds of uncertainty, the system learns a small subset of informative probe layers rather than reading every hidden state. This adds a modest calibration step at training time but avoids the latency of full-layer inspection at inference. The quality gain holds on math and coding benchmarks while cutting generated tokens by roughly 15–25 % on the tested Qwen3 reasoning models; the savings shrink once the model is already highly optimized. When your workload is dominated by long, self-verifying traces and you can afford a one-time calibration pass, boundary-expanded early exit is worth testing; if your traces are short or the model family changes frequently, the engineering overhead outweighs the benefit. Things to Try This Week
On the Horizon
|
💬 Reply to this email — Patrick reads every one. Share: X · LinkedIn · WhatsApp Forwarded this email? Subscribe here — it's free. |
📺 Watch on YouTube · 📝 Read the blog · 🖼 Free image gallery (CC BY-SA) · 📊 Data Hub & Story Trackers · 🧭 Start Here Nerra Network · AI-narrated voice (Grok TTS) · Editorial by Patrick You're receiving this because you subscribed to Models & Agents on nerranetwork.com. |
| Issue #130 · Models & Agents · Aug 3, 2026 |
