Start free trial

The trick: Self-Marked

OpenAI dropped hundreds of AI math papers at once.

Its own README says about 42% of the headline results are machine checked.

Issue 388 October 202617 receipts3 min

OpenAI says it is releasing a broad range of new math results from an internal model, with many proofs checked in Lean.

Before you read on. Your call?

Its own README says about 42% of top-line results are formalized and some unformalized results could have issues.

The twist

The 372 families came out of about 4,000 problems posed, and a statement Terence Tao reposted calls the 700-file drop a demonstration of power.

~42%share of top-line results formalized in Lean
4,000approximate number of problems the model was posed
372result families in the catalogue
719manuscripts in the current catalogue

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Review my 30-day free trial →

30 days free. Payment card required. One introductory trial per customer.

Your trial ends on 30 days after you start. Unless you cancel before then in Account → Manage membership, we charge A$89 for the first year. It renews automatically at A$89/yr until cancelled.

You can ask for a full refund within 14 days after any annual payment, renewals included, with no reason needed, by emailing hello@bskiller.com. This voluntary refund does not limit your rights under the Australian Consumer Law.

By starting your trial, you agree to the Terms.

BS Killer is published by Inferno Tech Pty Ltd, ABN 27 647 413 474.

Already a member? Sign in

The trick has a name

We call it Self-Marked: graded by the party that benefits from the grade. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“OpenAI's math release is real work, but its own README says only about 42% of the top-line results are Lean checked, and it picked them from about 4,000 problems.”

Receipts

  1. Supports web.archive.org: a broad range of new mathematical results produced by an internal frontier model
  2. Context web.archive.org: we are sharing formalizations of many of the proofs in Lean
  3. Context web.archive.org: The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking.
  4. Refutes github.com: The repository has ~42% top-line results formalized.
  5. Refutes github.com: Some of the unformalized results could have issues.
  6. Context github.com: Over the course of the evaluation, the model was posed approximately 4,000 problems.
  7. Context github.com: A family groups related papers, which may include a principal result, companion arguments, consequences, or alternative proofs.
  8. Context github.com: The current catalogue contains 719 manuscripts organized into 372 families.
  9. Context github.com: the writeup for the Re(s) > 11/12 zero-free region for the Riemann zeta function was human edited for readability.
  10. Context terrytao.wordpress.com: This is a guest post by the Association for Human Mathematics
  11. Refutes terrytao.wordpress.com: Releasing over 700 files at once is not a demonstration of scholarship, but a demonstration of power.
  12. Refutes terrytao.wordpress.com: Mathematicians did not ask for this work to be done.
  13. Context fortune.com: OpenAI published AI-generated full or partial solutions Tuesday to more than 370 outstanding mathematical problems
  14. Context fortune.com: made progress on three other Millennium Prize problems but had not fully solved them
  15. Refutes fortune.com: It followed some, but not all, of the steps the advisory group had recommended.
  16. Refutes fortune.com: do not understand the AI output well enough to answer questions on the result
  17. Supports fortune.com: My view is that this is great for mathematics

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

Get tomorrow's check in your inbox.

One AI claim a night, checked against independent sources, receipts attached. Free.

Free email updates. Unsubscribe any time.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.