<nezhar/>

Archives
Log in
Subscribe
September 28, 2026

September 2026.4

This week the focus is on open alternatives to Jev and new model releases: Ollaya runs the open decision models locally, while Opus 5.5 and GPT-6 Sol and Luna arrive at lower prices.


๐Ÿ“– Story 1: Ollaya โ€“ Ollama for open-source, Jev-style decision models

ollaya.dev ยท Read

Ollaya is to decision models like Jev, TypeSafe's text-free model from last week, what Ollama is to LLMs: an independent, Apache-2.0 runtime that pulls open weights and serves them locally behind TypeSafe's /v1/systemone API, so the official TypeSafe SDK works unchanged against localhost.

The catalogue gathers the open Jev alternatives: winnow, a Gemma 4 fine-tune on llama.cpp, decider and kev on Qwen3.5 bases, the small laya models and zero-shot classifiers like nli. On the project's own typed-decisions test, winnow:e4b scores 0.722 against hosted Jev's 0.738 and answers five questions in 89 ms on an RTX 4090; laya answers in about 10 ms but scores 0.361. Everything runs on CPU, it ships as desktop app, CLI and Docker image, weights are pinned by commit and sha256, and a Modelfile can recalibrate a model on your own labelled data.

Ollaya makes the open decision models easy to compare, and its own numbers show only the larger ones come close to Jev.

๐Ÿ’ฌ HN Discussion

The developer conceded that small models like laya fall well short of Jev on harder queries and that the open models that come close are much bigger. Users agreed: a semantic-grep experiment found laya competitive on easy queries before falling apart, another tester called the gap not close, while one commenter swapped laya in for small jobs on a 4 GB GTX 970.

A community leaderboard ranks decider-4b v2 marginally above Jev and Laya 421M 41st. But its top five, Jev included, score 82 to 88 percent on public questions and 33 to 37 on sealed ones, which one commenter read as weak generalisation. The rest relitigated whether Jev is innovation or a well-marketed zero-shot classifier.

โ†’ Discuss on Hacker News


๐Ÿ“– Story 2: Introducing Claude Opus 5.5

anthropic.com ยท Read

Anthropic released Claude Opus 5.5, claiming Fable 5.1-level work at 40% lower cost than Opus 5. It scores 66.4% on Terminal-Bench 4.0, against 55.8% for Fable 5.1 and 57.9% for GPT-6 Astra, and 1846 Elo on GDPval-AA against Fable's 1735. Astra still leads Terminal-Bench-Science (64.6% vs 58.7%), and Anthropic itself says the real gap to Fable is narrower than the scores suggest.

Prices drop to $4/$20 per million input/output tokens from $5/$25, and cache reads to $0.20 from $0.50, which matters most in long agentic sessions. A fast mode runs up to 2.5x quicker at $8/$40. Thinking is always on, and Fable-class safeguards route most cybersecurity work to Opus 4.8.

The price cut is real; whether the claimed token savings follow depends on the effort setting you run.

๐Ÿ’ฌ HN Discussion

The main dispute was whether the 40% saving holds up outside Anthropic's charts. Artificial Analysis data shows Opus 5.5 using about 21% fewer tokens than Opus 5 at high effort but about 38% more at max. Against Astra the gap shrinks: about $1.34 per task for Opus 5.5 medium and $1.82 at high, versus $1.76 for Astra high. In one patch-review test, Opus 5.5 found 8 of 14 issues for $15.40, while Fable 5.1 found 7 for $66.34.

The second split was over writing style. Some found it readable; others posted samples still full of Opus 5 tics. Several hit false cyber-classifier flags on embedded and routine shell work.

โ†’ Discuss on Hacker News


๐Ÿ“– Story 3: Introducing GPT-6 Sol and Luna

openai.com ยท Read

OpenAI released GPT-6 Sol and Luna, the tiers below Astra, and roughly halved their prices: Sol drops from $4/$20 to $2/$10 per million input/output tokens, Luna from $0.20/$1.20 to $0.10/$0.50. Cached input reads now get a 90% discount.

The benchmarks are framed as cost per task. On AutomationBench, Sol at xhigh scores 33.2% at $0.27 per task, ahead of Opus 5 at max (26.9% at 11x the cost). On DeepSWE v1.1, Sol reaches 68.8%, 1.1 points behind Fable 5, at about 80% lower cost. Competitor numbers come from public reports, with Fable 5 standing in where 5.1 scores were missing. There are no speed figures.

Artificial Analysis is more modest: Sol two points up on its Coding Agent Index at about half the cost per task, Luna two points down at about 60% less. Less a generational jump than a large price cut.

๐Ÿ’ฌ HN Discussion

The main question was whether this is an upgrade or a price cut on a smaller model. Artificial Analysis found GPT-6 Sol ahead of GPT-5.6 Sol on Terminal-Bench 4.0 (43% vs 37%) while Luna regressed (SWE-Atlas-QnA: 44% vs 49%). Some suspect the new Sol is a Terra-sized model; several went back to 5.6 Sol.

The other dispute was whether the per-token cut survives real workloads. One team found GPT-6 Luna needs enough extra tokens to erase the savings; one benchmark site measured Sol reasoning about twice as long, making it slower and roughly 25% more expensive in practice. Against Anthropic the gap is narrower than it looks: Sol's cached reads cost $0.20 per million, the same as Opus 5.5.

โ†’ Discuss on Hacker News


๐Ÿ’ฌ Community Moment

Felt the need to update this meme

https://www.reddit.com/r/ChatGPT/comments/1wq3yzp/felt_the_need_to_update_this_meme/

๐Ÿ› ๏ธ Projects Worth Checking Out

  • GitHub - ollaya-dev/ollaya: Run open decision models locally
  • GitHub - badrisnarayanan/antigravity-claude-proxy: Proxy that exposes Antigravity provided claude / gemini models
  • GitHub - fuergaosi233/claude-code-proxy: Claude Code to OpenAI API Proxy
  • GitHub - The-PR-Agent/pr-agent: ๐Ÿš€ PR Agent: The Original Open-Source PR Reviewer.
  • GitHub - audacity/audacity: Audio Editor
Don't miss what's next. Subscribe to <nezhar/>:
Older โ†’ September 2026.3
GitHub
LinkedIn
nezhar.com
www.flickr.com
Twitter
Powered by Buttondown, the easiest way to start and grow your newsletter.