GET THE AUTOPSY ➔

Frontier agents solved the Rails tasks. Then the graders checked whether they knew Rails existed.

An unusually honest indie benchmark: one frozen harness, 504 runs, receipts published. Models ace the work while reaching for the framework as little as 8% of the time, and 91 cents beats models costing far more.

01THE CLAIM
"The Agents on Rails benchmark, published on the official Rails blog, finds frontier agents solve most atomic Rails tasks (up to 92%) while mostly hand-rolling code, with Rails API recall running from 8% to 35%, and price stops predicting score: Luna's full run cost 91 cents while Opus charged 132x the price for nineteen points." [SOURCE ↗]
VERIFIED7 SOURCES · LIVE 2026-08-25
AGENTS ON RAILS TEAM TRACK RECORD1 CLAIM · 0/100 BS RATE →
8%RAILS API RECALL FLOOR, MODELS MOSTLY HAND-ROLL
91CENTS FOR LUNA'S ENTIRE 63-RUN BENCHMARK
132xOPUS COST PREMIUM FOR NINETEEN MORE POINTS
Frontier agents solved the Rails tasks. Then the graders checked whether they knew Rails existed.
02THE CHECK

THE CLAIM. across 21 atomic Rails tasks and 504 runs, frontier agents solve most of the work, Opus 5 at 92%, but mostly by hand-rolling code: Rails API recall runs from 8% (DeepSeek) to 35% (Fable), and cost decouples from quality, with GPT-5.6 Luna clearing 73% for 91 cents total while Opus costs 132x the price for nineteen extra points.

THE CHECK. it holds, within its stated scope. The methodology is the strongest this desk has seen from an indie benchmark: one frozen harness, one bash tool, default settings, hidden behavior tests, three runs per model per task for $491, and the corpus plus harness going open source. The honest limiters are in the report itself: tests check behavior, so hand-rolled fixes pass like idiomatic ones, and six of 21 tasks are solved by every run of every model. Also on the record: Fable 5 would lead at ~95% but went zero for three on the one task worded like a pen-test report.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"Finally a benchmark with receipts: frozen harness, published runs, stated limits. The models solve Rails tasks while mostly not using Rails, and after the first dollar, price stops predicting score."

The first report from Agents on Rails landed on the official Rails blog, a benchmark built by Svyatoslav Kryukov and Artur Petrov. Independent of the model vendors, not of Rails: the framework has obvious skin in the finding that agents do not know it, which the recall metric structurally flatters:

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 7 sources with quotes and screenshots, and our on-record call.

This story is a stable, citable object. If you can falsify a verdict, tell us. Corrections are loud here.