Agentailor logo

Agentailor

Archives
Subscribe
October 3, 2026

I replayed 193 real eval verdicts to see if Jev can replace my LLM judges

Hi,

Last month was about writing evals. This month I tested the thing grading them, and turned it into a new format.

The first Agentailor Report

An Agentailor Report answers one question with data I collected myself. The method is fixed before any results exist, and every number is reproducible from a public repo.

The first question: can Jev replace the LLM judges in my agent's eval suite? Jev is TypeSafe's new model that returns a typed probability instead of text. The launch coverage said it matched a human reviewer on every decision, from "one test on one agent".

So I took 193 inputs my suite's judges had already graded, labeled them by hand (blind, with a second blind pass that corrected 8 labels), and replayed them on Jev, Claude Sonnet 4.6, Claude Haiku 4.5 and Gemini 3.5 Flash-Lite, five times each.

What came out:

  • Not as a straight swap. Jev scored 90.6%, Sonnet 93.2%, Flash-Lite 94.8%.
  • But it's about 180x cheaper than Sonnet ($0.07 per 1,000 verdicts against $13.21) and about 21x faster.
  • Its probability is the real product. On the 125 items where Jev was confident, it made zero errors. Almost every mistake sat near 0.5.
  • So a cascade works. Let Jev decide the confident items and send the uncertain 13–16% to Sonnet: Sonnet-level accuracy, no real defect missed, about a sixth of the cost. (Those widths are in-sample, so validate them on your own data.)
  • The worst errors weren't the judge's fault. On rubrics where my harness hides context from the judge, every judge failed the same way. A better judge can't fix that.

The full report, with the method and every limitation, is at AR-001: Can Jev Replace Our LLM Judges? If you'd rather read it offline or pass it to your team, there's a PDF version.

The harness is public: judge-replay. pnpm verify recomputes every number with no API keys, and it's built so you can bring your own suite's data. If your evals gate merges on an LLM judge, this is the cheapest experiment you'll run this quarter.

Cameron v3: skills

Cameron v3 moves know-how out of the system prompt. A skill is a folder with a SKILL.md. Only its name and description sit in context, and the agent loads the body when it decides it needs it. The first skill teaches Cameron to answer with a chart, and when not to.

The design rule held: reading a skill approves nothing. If its instructions lead to a tool that writes, that tool hits the same approval gate as always.

The guide, with TypeScript and Python examples: How to Add Agent Skills Support to a Custom Agent Harness.

Devlog 4 covers a bug in evals themselves: a scripted multi-turn case replays fixed user turns, so when the agent asks a sensible question, the script answers something else and the case fails for the wrong reason. The fix is a simulated user. When Your Eval Script Answers a Question Nobody Asked.

Two smaller things

create-mcp-server v0.8.0 drops the greet example. New projects start with a small notes server that signals truncation instead of hiding it, a passing test suite, an AGENTS.md, and the tool-design skill.

The hub and the blog now expose their tools to agents running in your browser via WebMCP. It's Chrome and Edge only for now. agentailor.com/for-agents explains how, and the site shows each tool call live as the agent works.

— Ali

If you'd rather not get these, unsubscribe here. No hard feelings.

Don't miss what's next. Subscribe to Agentailor:
Older → My agent was wrong and never said so. Evals are how I found out.
GitHub
Twitter
LinkedIn
Powered by Buttondown, the easiest way to start and grow your newsletter.