Logo
Agentailor

Agentailor Reports

One question, answered with data.

Independent tests of claims about AI agents, run on a production agent. One question per report, the method fixed before the results, and every number reproducible from a public repository.

  • One questionEach report answers a single, stated question.
  • Method fixed firstDataset, thresholds and runs are set before any result is seen.
  • ReproducibleEvery number recomputes from a public repository.
  • Corrections loggedErrors are fixed in place and dated, never silently.

All reports

AR-001On a real agent's eval rubrics, can Jev replace the LLM judges?Not as a straight swap. As a confident first pass in front of Sonnet, it matches Sonnet's accuracy at about a sixth of the cost.EvalsAI Agents193 items × 5 runs3,995 recorded verdictsagentailor/judge-replay

AGENT BRIEFINGS

Get the next report by email.

What actually matters for building and scaling AI agents in production — and what's just hype. Straight from the work, no filler.