Local inference on Apple Silicon just got faster with… · M&A 🤖
| View this email in your browser |
![]() Models & AgentsDaily AI models, agents, and practical developments.
|
🎧 Today's episode Episode 186 · Local inference on Apple Silicon just got faster with a tuned fork of the Splash engine delivering up to 1.5 times the speed on M5 Max chips. 2026-09-27 ▶ Listen now |
What You Need to Know: A community developer released Splish, an optimized fork of the Splash inference engine for 40-core M5 Max hardware that improves single-request speed by roughly 25 percent and multi-request throughput by up to 50 percent while keeping output quality identical. KT's Auto Model Router placed second on the Router Arena benchmark that evaluates accuracy, cost, and robustness across 8400 queries. A new Towards Data Science guide walks through six advanced GraphRAG architectural patterns. llama.cpp gained a 42 times speedup in prompt lookup drafting. Hippocratic AI highlighted clinical LLM safety practices in a recent webinar. Simon Willison used Claude Opus 5.5 to generate a pixel-art animation of kākāpō parrots that became the closing slide in his keynote. Top StoryA developer released Splish, a fork of the Splash inference engine tuned specifically for 40-core M5 Max chips. The changes include kernel selections measured with Splash's own tuner, new verify kernels that use lighter barriers and compute row sums once per projection, and an attention tweak for the 27B shape. On the same hardware and models, Splish runs 1.25 times faster on single requests with gains between 11 and 35 percent and up to 1.5 times faster at two to four concurrent requests. Real-world short-story generation moved from 45-51 tokens per second to 56-64 tokens per second on 4-bit models. Quality remained unchanged across all measured outputs. The tuned settings target the 40-core M5 Max while other M5 chips fall back to Splash defaults. Kernel choices measured specifically for the 40-core M5 Max produced the biggest single-request win when Swift-1.5 moved from 74.7 to 89.8 tokens per second. Loading those choices from a file via SPLASH_KERNEL_CHOICES allows anyone to retune without rebuilding. New verify kernels for the M5 tensor units delivered an additional 5 percent at one request and 10 to 19 percent at two to four requests. Extending the same kernels to more projections added a further 1 to 3 percent at three to four requests. Tuned Qwen3.6-35B-A3B with the new kernels showed speedups at one to four requests of 5 percent, 18 percent, 22 percent, and 18 percent. GGUF files in Q8_0, Q4_K_M, and Q6_K formats received faster input loads in the decode kernel. A copy rule borrowed from TensorFold for coding agents improves whole-file edits by 24 to 42 percent while keeping output exact. An auto-tuner and reporting script are included for other M5 machines. Source: reddit.com Model UpdatesKT AI model routing technology ranks second in global benchmark : digitaltoday.co.kr KT's in-house Auto Model Router placed second overall on the Router Arena benchmark developed by Rice University researchers. The platform evaluates routers on roughly 8400 queries for response accuracy, cost efficiency, and robustness to input changes. Auto Model Router analyzes task type, difficulty, and knowledge domain before linking each request to the best model from a pool while balancing quality and usage cost. For example it links cost-efficient models for translation or simple information checks and high-performance models for specialised analysis or tasks requiring advanced reasoning. KT already uses the router inside its Token Factory service that integrates multiple AI models and token-usage environments. The company plans to keep adding new models to the routing environment. Jun-seok Kim, executive director and head of KT's Agentic AI Lab, stated that intelligently selecting the optimal model for each situation is important in the AI era and that Auto Model Router will become core technology supporting the competitiveness of KT's agentic AI services. Source: digitaltoday.co.kr Agent & Tool DevelopmentsClaude creates celebratory Kākāpō breeding season video : Simon Willison (AI builder) (X) Simon Willison prompted Claude Opus 5.5 to generate an HTML5 canvas pixel-art animation of at least twenty kākāpō parrots jumping with confetti. He supplied three reference photos from Google image search and asked for an interactive page that could later be recorded as a 15-second video. Claude Code then used Playwright to drive a browser, perform timed clicks across the canvas at nine specific moments spread over sixteen seconds, and produce the final video file. The resulting animation served as the closing slide in Willison's WeAreDevelopers keynote that referenced this year's record-breaking kākāpō breeding season. The prompt transcript and the interactive HTML page are available on his site along with the full Playwright script used to generate the video. Source: x.com Practical & CommunityGraphRAG: A Practitioner's Guide to 6 Advanced Architectural Patterns : Towards Data Science The guide details six architectural patterns for building production GraphRAG systems. It focuses on retrieval and generation workflows that combine graph databases with large language models. The patterns address practical challenges in scaling graph-enhanced retrieval while maintaining generation quality. Source: towardsdatascience.com 42x Faster Prompt Lookup Drafting in llama.cpp : r/LocalLLaMA A contributor added optimizations that deliver a 42 times speedup for prompt lookup drafting inside llama.cpp. The change targets the drafting step used in speculative decoding pipelines. The update is available in the latest builds and improves throughput for workloads that rely on prompt-based token prediction. Source: reddit.com Hippocratic AI Highlights Clinical LLM Safety and Differentiation in Becker’s Webinar : TipRanks Hippocratic AI presented its approach to clinical large language model safety during a Becker's webinar. The session covered differentiation strategies for healthcare deployments. The company emphasized safety mechanisms and model behavior tailored to clinical environments. Source: tipranks.com Under the Hood: Speculative Decoding TradeoffsSpeculative decoding runs a smaller draft model to guess several tokens ahead, then verifies them in one forward pass of the larger target model. The draft model adds almost no extra memory because it stays in the same batch as the target. When the draft matches the target, the system accepts multiple tokens per verification step and cuts total latency by roughly half on typical workloads. When the draft diverges, the system discards the wrong tokens and falls back to standard autoregressive generation, so the worst-case slowdown stays under 10 percent. The quality of the draft model matters most below 30 billion parameters; above that size the acceptance rate plateaus and extra draft compute stops paying off. Teams therefore choose a draft model one quarter the size of the target and accept rates between 60 and 80 percent as the practical operating range. Use speculative decoding when generation length exceeds 200 tokens and the draft model fits comfortably in the same accelerator memory as the target; skip it for short prompts or when memory headroom is already tight. Things to Try This Week
On the Horizon
|
💬 Reply to this email — Patrick reads every one. Share: X · LinkedIn · WhatsApp Forwarded this email? Subscribe here — it's free. |
📺 Watch on YouTube · 📝 Read the blog · 🖼 Free image gallery (CC BY-SA) · 📊 Data Hub & Story Trackers · 🧭 Start Here Nerra Network · AI-narrated voice (Grok TTS) · Editorial by Patrick You're receiving this because you subscribed to Models & Agents on nerranetwork.com. |
| Issue #186 · Models & Agents · Sep 27, 2026 |
