Next Tool

Archives
Log in
Subscribe
July 16, 2026

When a 13-year-old Xeon beats your cloud bill (and 4 other AI shifts)

When a 13-year-old Xeon beats your cloud bill (and 4 other AI shifts)

A/B alternates: - "Inkling goes open-weights, and your old server can finally run a real LLM" - "The week's AI: a 26B model on bare metal, an open Grok Build, and Qwen inside your iPhone"

TL;DR: A real open-weights MoE from Mira Murati's lab dropped this week (975B total / 41B active, 1M context), xAI open-sourced Grok Build under Apache 2.0 after a privacy backlash, and someone got Gemma 4 26B running at 5 tok/s on a $300 box from 2013. If you've been waiting for permission to skip the GPU rental, this is it.


1. Inkling by Thinking Machines — open-weights MoE for fine-tuning

Thinking Machines Lab (Mira Murati's startup) released its first model on July 15: a mixture-of-experts transformer with 975B total parameters, 41B active, a 1M-token context window, and native reasoning over text, images, and audio. Pretrained on 45T tokens. Full weights are on Hugging Face, with an NVFP4 checkpoint for NVIDIA Blackwell. Alongside it, a lighter Inkling-Small (12B active) is previewed. The catch: the team itself says Inkling "is not the strongest overall model available today." Treat it as a base for customization, not a Claude-killer.

Source: https://thinkingmachines.ai/news/introducing-inkling/


2. Grok Build is now Apache 2.0 — xAI's coding CLI goes open-source

xAI's grok CLI got caught uploading entire user directories — including SSH keys and password vaults — to its cloud by default. After the backlash, xAI disabled data retention, deleted previously retained data, and on July 15 released the full Grok Build codebase under Apache 2.0. You can now run Grok Build fully local-first with your own inference. Trust is still up to you, but the code is auditable.

Source: https://github.com/xai-org/grok-build


3. Gemma 4 26B on a 2013 Xeon — 5 tok/s for $300 of hardware

Neomind Labs got Google's Gemma 4 26B-A4B (MoE, Q8_0) running on a repurposed HP StoreVirtual with dual Ivy Bridge Xeons — no GPU, AVX1 only — at ~5.2 tok/s decode and ~16 tok/s prompt eval. The box cost under $300. Built on ik_llama.cpp with 25 carefully chosen flags. If you've been pricing "real" local inference in GPU-hours, this is the proof you can run a modern MoE on commodity iron.

Source: https://www.neomindlabs.com/2026/06/08/running-gemma-4-26b-at-5-tokens-sec-on-a-13-year-old-xeon-with-no-gpu/


4. Apple Intelligence ships in China on Alibaba's Qwen — the first real Apple AI model swap

Apple got approval to launch Apple Intelligence in China powered by Alibaba's Qwen models — the first time Apple has plugged a third-party LLM into its consumer AI stack. The deal signals how fragmented "global AI" really is now: the model in your iPhone depends on the country you bought it in. Useful signal for anyone building cross-region AI products: plan for model swaps at the OS layer.

Source: https://techcrunch.com/2026/07/15/apple-intelligence-approved-for-launch-in-china-with-alibabas-qwen-ai/


5. Most "AI agents" in the enterprise are still chatbot wrappers

VentureBeat surveyed 101 enterprises on agent orchestration. Headline findings: Anthropic's Claude leads platform choice by a wide margin, but most deployed "agents" are still chatbot skins, real-time token-cost controls remain rare, and the control plane is deliberately hybrid to avoid lock-in. If you're selling "agentic AI" into enterprises, the gap between ambition and what's actually running in prod is still your biggest opportunity — and your biggest credibility risk.

Source: https://venturebeat.com/ai/agentic-orchestration-enterprise-ai-organizations-have-a-deployment-problem-not-a-platform-problem-and-most-are-calling-chatbots-agents


Subscribe to Next Tool for a weekly AI-tools roundup that respects your time → https://buttondown.com/nexttool

Know someone who'd want this? Forward it — referrals are how Next Tool grows.

Don't miss what's next. Subscribe to Next Tool:
← Newer Inkling drops, Grok goes open, and Reelful edits your camera roll Older → A 27B model on your phone, Cursor's silent 0-day, and a Sol that deletes your files
sauteri.com
Powered by Buttondown, the easiest way to start and grow your newsletter.