The company that grades AI for OpenAI, Anthropic, Google, Meta, and xAI just raised $40M at a $400M valuation. It disclosed a customer relationship with the labs it evaluates.
The benchmark referee has a customer relationship with the labs it grades and is using a May number while August models score 75%. Independence is the product and the conflict.
"Vals AI's independent benchmarks show frontier models correctly complete fewer than 52% of real financial-analyst tasks; a16z valued the evaluator at $400M." [SOURCE ↗]
THE CLAIM. Vals AI raised $40M led by a16z at a $400M valuation on the pitch that it is the independent evaluator of AI. Its benchmarks appear in model cards from OpenAI, Anthropic, Google, Meta, and xAI.
THE CHECK. Artificial Lawyer reported Vals AI disclosed a customer relationship with one or more of the participants it evaluates. The benchmarks that appear in model cards are produced by a company billing the labs whose products carry those scores.
THE TWIST. independence is the product. The customers are the subjects. The conflict is structural, not hidden. Vals disclosed it. But disclosure does not eliminate the conflict; it just puts it on the record.
Vals AI positions itself as the independent evaluator of AI. Its Series A blog opens with: "We started Vals AI to solve this problem as the independent evaluator of artificial intelligence." Its benchmarks appear in model cards from five of the six major frontier labs.
🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNTYou just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 4 sources with quotes and screenshots, and our on-record call.
Couldn't verify your access — this looks like our error, not yours.