Frontier agents solved the Rails tasks. Then the graders checked whether they knew Rails existed.
An unusually honest indie benchmark: one frozen harness, 504 runs, receipts published. Models ace the work while reaching for the framework as little as 8% of the time, and 91 cents beats models costing far more.
"The Agents on Rails benchmark, published on the official Rails blog, finds frontier agents solve most atomic Rails tasks (up to 92%) while mostly hand-rolling code, with Rails API recall running from 8% to 35%, and price stops predicting score: Luna's full run cost 91 cents while Opus charged 132x the price for nineteen points." [SOURCE ↗]

THE CLAIM. across 21 atomic Rails tasks and 504 runs, frontier agents solve most of the work, Opus 5 at 92%, but mostly by hand-rolling code: Rails API recall runs from 8% (DeepSeek) to 35% (Fable), and cost decouples from quality, with GPT-5.6 Luna clearing 73% for 91 cents total while Opus costs 132x the price for nineteen extra points.
THE CHECK. it holds, within its stated scope. The methodology is the strongest this desk has seen from an indie benchmark: one frozen harness, one bash tool, default settings, hidden behavior tests, three runs per model per task for $491, and the corpus plus harness going open source. The honest limiters are in the report itself: tests check behavior, so hand-rolled fixes pass like idiomatic ones, and six of 21 tasks are solved by every run of every model. Also on the record: Fable 5 would lead at ~95% but went zero for three on the one task worded like a pen-test report.
The first report from Agents on Rails landed on the official Rails blog, a benchmark built by Svyatoslav Kryukov and Artur Petrov. Independent of the model vendors, not of Rails: the framework has obvious skin in the finding that agents do not know it, which the recall metric structurally flatters:
🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNTYou just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 7 sources with quotes and screenshots, and our on-record call.
Couldn't verify your access — this looks like our error, not yours.