Tencent ships 770B open-weight Hy4 Β· M&A π€
| View this email in your browser |
![]() Models & AgentsDaily AI models, agents, and practical developments.
|
By the numbers
|
π§ If you only have 10 minutes this week Episode 158 Β· Tencent just dropped a 770B-parameter open-weight MoE model with a 1M-token context and explicit reasoning controls that builders can toggle today. 2026-08-30 βΆ Listen now |
This Week in AIAgents spent this week leaving the lab. PYMNTS reported that autonomous systems are now a first-class API customer β high-frequency, repetitive, and poorly served by rate limits and auth designed for humans (βΆ Episode 151 Β· 2026-08-24). The same pattern showed up at the edge: a developer merged a production feature written entirely by Qwen3 27B running quantized on a single 4060 Ti, including codebase investigation, ADR-backed planning, multi-file edits, and QA gates (βΆ Episode 152 Β· 2026-08-25). OpenAI said JalapeΓ±o, its first custom inference chip, is heading into production compute by year-end, claiming higher throughput and lower latency without the usual efficiency tradeoff (βΆ Episode 153 Β· 2026-08-26). Who gets to inspect, steer, and physically act moved just as fast. Anthropic opened privacy-preserved Claude usage data β 250,000 conversations β to Stanford SALT Lab, Oxford, and METR (βΆ Episode 154 Β· 2026-08-27). It launched a research preview of the Model Hardware Standard so agents can drive lab and manufacturing equipment through one interface (βΆ Episode 155 Β· 2026-08-28), then showed Claude researching, training, and evaluating alignment fixes for smaller models on one GPU in 48 hours (βΆ Episode 157 Β· 2026-08-29). Tencent closed the week with Hy4 Preview: 770B total parameters, 49B active, a 1M-token context, and an explicit reasoning toggle (βΆ Episode 158 Β· 2026-08-30). The field did not simply get bigger. Silicon, local runtimes, API economics, hardware drivers, and alignment loops all moved together. JalapeΓ±o is still a year-end plan, Hy4 is a preview, and MHS is a research preview β but the stack around models is being rebuilt for non-human consumers, and that direction is no longer ambiguous. Model Tracker
OpenAI also introduced ChatGPT Business Premium Seats at $100 per seat for small teams that could not previously reach enterprise tooling. Top Stories1. Tencent open-weights a 770B MoE with 1M context. Hy4 Preview ships visible reasoning traces and a hard 2. OpenAI's first inference chip is entering production. JalapeΓ±o is pitched as more intelligence per watt, with Gen 2 already deep in development and Gen 3 taking shape. Expect faster ChatGPT, Codex, and agent responses as capacity lands; treat year-end as a claim until racks are actually live. βΆ Episode 153 Β· 2026-08-26 3. A 27B local agent wrote and merged production code. Qwen3 27B IQ3_K_XXS on a 4060 Ti 16GB explored a codebase, planned against ADRs, edited multiple files, passed QA, then waited for human merge. Earlier 20Bβ30B attempts failed on shared-file edits. The same report documents AMD eGPU OOM crashes and the GRUB parameter that stops power management from dumping weights into RAM. Local coding agents are no longer a cloud-only story. βΆ Episode 152 Β· 2026-08-25 4. Claude ran an alignment loop on one GPU. In 48 hours it researched methods, proposed fixes, trained, and evaluated smaller models, lifting safety scores on ten common misalignment types without tanking general capability. Best methods generalized to held-out benchmarks and models up to 4.7Γ larger; Sonnet 5 post-trained an early Opus 4.8 checkpoint toward production safety scores. A research prototype, not a replacement for human alignment teams β but the first credible sketch of model-driven alignment. βΆ Episode 157 Β· 2026-08-29 5. Agents became the API economy's new customer class. Businesses are routing routine transactions through autonomous systems, so human-shaped rate limits, auth, and pricing will fail in predictable ways. Add agent-aware quotas and non-interactive docs; assume next quarter's traffic mix will not look like last year's dashboards. βΆ Episode 151 Β· 2026-08-24 Agent & Tool UpdatesVisa shipped an open-source agentic security harness that auto-remediates vulnerabilities before human review. Anthropic's Model Hardware Standard preview extends Claude Code toward lab gear, boards, and cameras through one driver layer; HHMI started the collaboration, and science, robotics, and manufacturing partners can join before any open-source freeze (βΆ Episode 156 Β· 2026-08-28). OneModel's results argue some industrial agents should collapse modular pipelines into a single internalized model. Meta's EvoHarness-RL 8B is the research counterpart: small weights, long loops. On the messy hardware side, a community developer reverse-engineered an Axera NPU format to run GGUF at 1.5Γ the vendor runtime on Raspberry Pi. Open Source SpotlightHy4 Preview is the headline dump: 1.56TB, Hugging Face, explicit reasoning controls. Treat it as a preview, not a drop-in replacement β but the weights are actually there. Liquid AI Pipette open-sources on-device benchmarking so local-model claims can be compared on the hardware you own, not a vendor slide. Visa's agentic security harness is the practical pick for teams putting agents on real endpoints: remediate first, review second. Pair it with the API-traffic story if you are rewriting rate limits this quarter. Honorable mention: the Axera NPU GGUF path on Raspberry Pi. Reverse-engineered runtimes are not a product strategy, but 1.5Γ vendor speed on a Pi is how local agents actually move. Safety & RegulationAnthropic's data release is the governance story of the week. Three independent groups analyzed 250,000 aggregated AprilβMay 2026 conversations; SALT Lab found more than half involved consequential tasks β work that affects others or is hard to undo. Raw user data stays protected; researchers apply via a public form. OpenAI published a technical report on the Hugging Face agent incident with third-party assessments from METR and Redwood. MHS will produce new physical-world safety evals because current models still lack physical intuition from text-and-image training. The automated-alignment paper is double-edged: it could scale safety work, or automate the appearance of it. Test the released methods on your own small models; do not outsource judgment to the loop that produced them. What to Watch Next Week
|
|
π¬ Reply to this email β Patrick reads every one. Share: X Β· LinkedIn Β· WhatsApp Forwarded this email? Subscribe here β it's free. |
πΊ Watch on YouTube Β Β·Β π Read the blog Β Β·Β πΌ Free image gallery (CC BY-SA) Β Β·Β π Data Hub & Story Trackers Β Β·Β π§ Start Here Nerra Network Β· AI-narrated voice (Grok TTS) Β· Editorial by Patrick You're receiving this because you subscribed to Models & Agents on nerranetwork.com. |
