InsiderLLM

Archives
Log in
Subscribe
September 28, 2026

A 27B for your 12 GB card. It lost one field.


InsiderLLM Weekly issue 23 -- September 28, 2026

A 27B now fits a 12 GB card and decodes 1.8x faster than the Q4 it was squeezed from. It also lost one field on my routing test โ€” the same way every time. That's the main piece. First, a reader caught a stale 3090 price, and it wasn't the only one.


New This Week

  • Bonsai 2 vs Qwen3.8-27B on an RTX 3090: Speed Doubled, Scope Broke. Bonsai 2 27B, PrismML's ternary Qwen3.8, on a 3090: 1.8x the decode of Q4 from a 7.2 GB file, then 12 of 47 against 19 on a routing split. ๐Ÿ“– https://insiderllm.com/guides/bonsai-2-27b-vs-qwen3-8-27b-rtx-3090/
  • What Is Jev, and Can You Run One on Your Own GPU? Jev picks from options you supply instead of writing, in one pass, with a probability per answer. ๐Ÿ“– https://insiderllm.com/guides/what-is-jev-typesafe-explained-local/

Updated:

  • 23 articles corrected for used GPU prices after a reader caught a stale 3090 figure. Each carries a dated correction line saying what it used to say.
  • Best Used GPUs for Local AI: the fair-price ladder. It had been calling fair 3060 listings overpriced. Against 107 sold listings the median is $300, so fair is now $275-325, a great price is under $250, and overpriced starts above $350. ๐Ÿ“– https://insiderllm.com/guides/best-used-gpus-local-ai-2026/
  • Best Way to Run Qwen 3.6 35B MoE Locally. "A separate drafter doesn't help" was a 3060 offload result. On a resident 3090, z-lab's DFlash drafter gave 1.09x overall, from 1.64x on Python down to 0.41x on a short translation. ๐Ÿ“– https://insiderllm.com/guides/best-way-run-qwen-3-6-35b-moe-locally/

Quick Hits

  • Qwen Image 2.1. 7B generator under the Qwen Research License, not Apache. gguf-org/qwen-image-2.1-gguf landed next day: NVFP4 transformer, Q4 text encoder, VAE, 10 GB, so 12 GB cards fit. ๐Ÿ“– https://huggingface.co/gguf-org/qwen-image-2.1-gguf ยท the model: https://huggingface.co/Qwen/Qwen-Image-2.1
  • Qwen Audio 3.1. Five audio models, API prices cut up to 95 percent. No open weights on Hugging Face, GitHub or Qwen's post. Local voice stays June's Qwen3-ASR-1.7B. ๐Ÿ“– https://the-decoder.com/alibaba-launches-qwen-audio-3-1-with-five-new-models-and-slashes-ai-audio-prices-by-up-to-95-percent/ ยท the local option: https://huggingface.co/Qwen/Qwen3-ASR-1.7B-hf
  • llama.cpp v0.5.0. Twelve builds and a version tag: CUDA conv2d speedup, Metal MoE fusion, MiMo-V2.6 support, no breaking changes. We stay on v0.4.0 until it's re-gated on both rigs. ๐Ÿ“– https://github.com/ggml-org/llama.cpp/releases/tag/v0.5.0 ยท ours: https://insiderllm.com/benchmarks/
  • Opus 5.5 and GPT-6. Both closed, both cheaper: Opus 5.5 with stricter cyber safeguards, GPT-6 pitched on cost and fewer mistakes. The API column of our local-vs-API break-even moved. ๐Ÿ“– https://www.theverge.com/ai-artificial-intelligence/998868/anthropic-claude-opus-5-5-cybersecurity ยท https://openai.com/index/introducing-gpt-6-sol-and-luna ยท ours: https://insiderllm.com/guides/local-ai-vs-cloud-api-cost/
  • Buried injections. 629 AgentDojo injections buried in tool output: regex catches 0 percent, Prompt Guard 2 catches 1 percent out of the box, 99 percent once its threshold tuned. ๐Ÿ“– https://github.com/rudratoshs/buried-injections ยท ours: https://insiderllm.com/guides/ai-agent-coordination-incidents-2026/
  • Dettmers's 1.5-bit claim. Qwen 3.6 35B-A3B at 450 tok/s at 1.5 bits per weight on Metal. His lab's first release is a private beta of bitsandbytes2, so there's still nothing public to download; when the code ships, we'll run it on a 3090. ๐Ÿ“– https://timdettmers.com/2026/09/21/dlab-open-source-week/ ยท the beta: https://x.com/Tim_Dettmers/status/2102418118568018281
  • Open Jev-style decision models. Jev-Style-Qwen3.5-2B picks among 2 to 26 options with a probability each, Apache 2.0, on plain llama.cpp. Take the 2 GB Q8: the Q4 agrees with full precision on only 91 percent of choices. Ollaya serves the same class on ONNX Runtime. ๐Ÿ“– https://huggingface.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF ยท https://ollaya.dev/ ยท ours: https://insiderllm.com/guides/what-is-jev-typesafe-explained-local/
  • 2.2x from one setting on an Intel Arc iGPU. Qwen3.6-35B on a laptop's integrated Arc B390 went from 14.7 to 32.5 tok/s by moving the MoE experts off the CPU. An old README priced that at 10 percent โ€” it cost 55 percent of decode. If your model fits, check you aren't offloading experts by habit. ๐Ÿ“– https://grigio.org/how-i-got-2-2x-more-tokens-per-second-from-llama-cpp-on-intel-arc/ ยท ours: https://insiderllm.com/guides/best-way-run-qwen-3-6-35b-moe-locally/

A 27B on a 12 GB Card, Minus One Field

If you have a 12 GB card and want a 27B for chat, try Bonsai 2 27B. It's PrismML's ternary Qwen3.8-27B โ€” every weight stored as โˆ’1, 0 or +1 โ€” and the whole model lands in a 7.2 GB file. On my 3090 it peaked at 8,118 MiB with 8k of context, which leaves a 12 GB card room to spare. It decodes at 77 tok/s against 42 for the Q4 it came from, and still holds 1.8x at 8k deep. PrismML's table claims parity on code; I haven't measured that. An 8 GB card will need a short context, and I haven't measured how short either.

Then the bill. I ran both through the 47-item routing split this site has scored since August. The Q4 got 19 right on all three fields. Bonsai got 12, and the whole gap is one field. On tool, they tied at 28. On mode, Bonsai won, 27 to 24. On scope โ€” the field that asks whether a message points back at something said earlier โ€” it fell from 28 to 18, and 23 of its wrong scopes were the same call: facts where the answer was session or all.

PrismML says it keeps 98.2 percent of the full model's benchmark score. Both numbers can stand. Theirs is a 14-benchmark average in thinking mode against FP16. Mine is one task, thinking off, against the Q4 a consumer card can hold. An average like theirs can't show you one field breaking.

So if you route on it, measure first. With a 24 GB card and no VRAM pressure, the Q4 is still the safer file.

๐Ÿ“– We wrote a full breakdown here: https://insiderllm.com/guides/bonsai-2-27b-vs-qwen3-8-27b-rtx-3090/


That's the week. If you have a 12 GB card, Bonsai 2 is worth a download for chat โ€” just don't hand it your router without scoring it on your own data first. And if you spot a price on this site that looks wrong, tell me. The last one led to corrections on 24 pages.

โ€” Mark, InsiderLLM


Forwarded this and want your own copy?

Tried Bonsai 2 on a 12 GB or 8 GB card? Got a context length that fits? Reply, or hit me at [email protected]. I read everything.

Read this issue on the web: https://insiderllm.com/blog/newsletter-2026-09-28/

Don't miss what's next. Subscribe to InsiderLLM:
Older โ†’ Stop writing the answer. Pick it. 47 items waiting.
Powered by Buttondown, the easiest way to start and grow your newsletter.