Microsoft just gained the freedom to build its own… · M&A 🤖
![]() Models & AgentsDaily AI models, agents, and practical developments.
|
🎧 Today's episode Episode 72 · Microsoft just gained the freedom to build its own frontier models after a contract change with OpenAI, and the first MAI family is already shipping. 2026-06-06 ▶ Listen now |
What You Need to Know: Microsoft announced seven in-house MAI models spanning reasoning, code, image, transcription, and voice, trained from scratch on licensed data without distillation. A major open-weight release wave also landed this week across LLMs, VLMs, TTS, and world models. Builders should watch how quickly the MAI models reach competitive performance on agentic and coding tasks now that Microsoft can iterate independently. Top StoryMicrosoft revealed that a contract change roughly six months ago removed prior restrictions, allowing its AI Superintelligence Team to pursue superintelligence with its own researchers, data, and custom silicon. The company shipped its first substantial in-house model family under the MAI brand, including the 35B-active-parameter MAI-Thinking-1 reasoning model and specialized models for code, image generation, transcription across 43 languages, and multilingual voice. All models were trained from scratch on commercially licensed data without relying on outputs from other labs. They are available through Microsoft Foundry with support for third-party weight tuning on platforms such as OpenRouter, Fireworks, and Baseten. Enterprise customers can already use Frontier Tuning to customize the models on their own workflows inside governed environments. The move signals Microsoft is building a vertically integrated stack alongside its continued OpenAI partnership rather than replacing it. Source: venturebeat.com Model UpdatesAnthropic Science Blog — @AnthropicAI Anthropic released new research showing Claude Opus 4.7 matches or exceeds dedicated NMR spectroscopy software on molecular structure interpretation tasks. The work focuses on helping chemists manipulate molecules by first understanding their structure through NMR data. Builders working on scientific tooling can now test whether Opus 4.7 can replace or augment existing NMR analysis pipelines without custom fine-tuning. Source: x.com DeepSeek V4 Flash — r/LocalLLaMA DeepSeek V4 Flash is now runnable via an early llama.cpp PR, with users reporting strong intelligence for its size, native FP4-FP8 hybrid quantization that holds up well under quantization, and efficient context scaling that uses less KV cache. The model is being positioned as a contender for the 80-140GB local inference space. Early testers created custom 3-bit quants to match the original tensor layout and noted reliable correctness despite slow speeds and missing GPU/FA support. Source: reddit.com Gemma 4 family updates — r/LocalLLaMA Google shipped Gemma 4 12B as a fully open dense any-to-any model supporting text, image, audio, and video with 256k context and coverage of 140+ languages. Quantization-aware training (QAT) variants of the Gemma 4 family are showing speed and VRAM gains on AMD hardware with no measurable quality loss versus standard quants on tested prompts. Source: reddit.com NVIDIA Nemotron 3 Ultra — r/LocalLLaMA NVIDIA released Nemotron 3 Ultra, a 550B hybrid Mamba-MoE model with 55B active parameters, 1M context, and reported MMLU of 89.1. An NVFP4 variant claims roughly 5x throughput on Blackwell hardware and is described as the first openly weighted 550B hybrid Mamba-Transformer. Source: reddit.com Agent & Tool DevelopmentsQwen3.7-Plus — the-decoder.com Alibaba positioned Qwen3.7-Plus as a multimodal model explicitly built to function as a full autonomous agent rather than a chat interface. The release emphasizes agentic capabilities across modalities. Developers exploring agent frameworks should test whether the model’s native agent behaviors reduce the need for heavy scaffolding compared with prior Qwen releases. Source: Google News dots.tts 2B — r/LocalLLaMA RedNote open-sourced dots.tts, a 2B-parameter continuous (no codec) TTS model under Apache 2.0 that performs direct text-to-speech at 48 kHz with zero-shot voice cloning. The architecture skips the traditional phoneme pipeline entirely. Teams building local voice agents now have a fully open, continuous alternative to codec-based TTS stacks. Source: reddit.com OpenLumara agent framework — r/LocalLLaMA A new modular agent called OpenLumara was released with a ~4k token default system prompt, full modularity down to core features, and built-in security controls including sandboxed shell access and HTTP black/whitelists. It is designed specifically for local models and runs efficiently on modest hardware without the token bloat of skill.md-style systems. The project is GPL2 licensed and available on GitHub. Source: reddit.com Practical & Communitymicropython-wasm sandbox — Simon Willison's Weblog Simon Willison released micropython-wasm, an alpha PyPI package that runs MicroPython inside a WebAssembly sandbox with memory/CPU limits, controlled host function access, and no filesystem or network access by default. A CLI mode lets users try it immediately with Gemma 4 QAT benchmarks — r/LocalLLaMA Users running Gemma 4 models on an AMD 7900 XTX reported that QAT versions deliver 1.3–1.5x speedups and meaningful VRAM savings versus standard quants while preserving output quality on long-context and creative tasks. The 12B QAT variant cut generation time by 45% with identical constraint-following behavior. Source: reddit.com KV cache offload to RAM — r/LocalLLaMA A user running Qwen3.6 27B on an RTX 5060 Ti found that enabling Under the Hood: KV Cache Quantization TradeoffsEveryone talks about KV cache quantization as a simple “turn it on and save VRAM” toggle. In practice it is a precision-versus-throughput negotiation that depends on model architecture, context length, and downstream task. The core insight is that attention keys and values are not equally sensitive to quantization; value vectors often tolerate lower precision better than keys because they are summed rather than used for similarity lookups. Lower-bit KV formats therefore introduce small per-token errors that accumulate across long contexts, which is why some teams see reasoning degradation only after 32k–64k tokens even when perplexity looks acceptable. New methods such as KVarN attempt to close this gap by learning per-channel scales that preserve the distribution of important value dimensions, delivering q5-level KLD at roughly 4-bit storage in early llama.cpp tests. The practical tradeoff is that aggressive KV quantization can still hurt multi-step agent trajectories more than single-turn chat, because small embedding drift compounds when the model must maintain consistent state across tool calls. When you are running agents with 50k+ context or long-horizon planning, keep at least 5–6 bits on the value cache unless you have measured task-specific tolerance; for shorter chat workloads the memory win is usually worth it. Things to Try This Week
On the Horizon
|
💬 Reply to this email — Patrick reads every one. |
📺 Watch on YouTube · 📝 Read the blog Nerra Network · AI-narrated voice (Grok TTS) · Editorial by Patrick You're receiving this because you subscribed to Models & Agents on nerranetwork.com. |
| Issue #72 · Models & Agents · Jun 6, 2026 |
