Subscribe

The trick: Self-Marked

OpenAI built a mental health test, had its own model grade it, and its own model came top.

Answers written by licensed clinicians scored 38.5%. GPT-6 Astra scored 57.3%.

Issue 2525 September 202613 receipts3 min

OpenAI says results on MentalHealthBench show the steady improvement of AI systems in helping people navigate mental health situations.

Before you read on. Your call?

the benchmark uses synthetic conversations, and each reply is graded against expert-written criteria by an automated grader, OpenAI's own GPT-5.6 Sol. OpenAI's GPT-6 Astra tops the reported results at 57.3%. Answers written by clinicians scored 38.5%, which the authors attribute largely to clinicians writing short, in-person style replies, and completions written with the rubric in view scored 99.0%.

The twist

as NxCode puts it, a high score measures coverage of the rubric, not a proven benefit to a person. Nothing in this record measures what happened to anyone after a conversation, so 'helping people' is a step beyond what the test observes.

38.5%score of answers written by licensed clinicians on MentalHealthBench's task-clipped measur
57.3%score of OpenAI's GPT-6 Astra on the same measure
99.0%score of rubric-aware completions written with the grading rubrics in view

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

The trick has a name

We call it Self-Marked: graded by the party that benefits from the grade. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“On OpenAI's own mental health benchmark, graded by OpenAI's own model, real clinicians scored 38.5% and OpenAI's GPT-6 Astra scored 57.3%. The authors put that gap largely down to clinicians answering briefly. The score does not show AI helps people more than therapists do.”

Receipts

  1. Supports web.archive.org: Results on MentalHealthBench show the steady improvement of AI systems in helping people navigate mental health situations.
  2. Context web.archive.org: For each conversation, we use an automated grader, GPT‑5.6 Sol, to assess model responses against the expert-written criteria.
  3. Context web.archive.org: Using privacy-preserving techniques, we created synthetic mental health conversations that accurately reflect real-world usage patterns of AI for mental health.
  4. Context web.archive.org: While ChatGPT is not a substitute for therapy or professional care, expert-informed insights help us measure the progress towards AI models
  5. Supports unite.ai: In OpenAI’s reported results, GPT-6 Astra scored highest at 57.3% task-clipped, followed by GPT-6 Sol at 53.9%, Claude Opus 5.5 at 52.4%, and GPT-6 Luna at 50.2%.
  6. Refutes unite.ai: Expert-authored completions written by clinicians scored 38.5%, which the authors attribute largely to clinicians writing short responses as if in an in-person conversation, often asking a single question or making a simple statement.
  7. Context unite.ai: Rubric-aware completions, written with the grading rubrics provided, scored 99.0%, which the paper describes as a sanity check on the evaluation’s noise ceiling.
  8. Refutes nxcode.io: In a reference comparison, expert-authored completions scored 38.5% on the task-clipped measure, while GPT-6 Astra scored 57.3% and GPT-6 Sol 53.9%
  9. Context web.archive.org: MentalHealthBench was co-created with a global cohort of more than 80 licensed mental health experts from 22 countries.
  10. Context unite.ai: The authors state that no benchmark captures everything that matters in a personal conversation, and they position MentalHealthBench as an auditable diagnostic tool rather than a definitive leaderboard.
  11. Context unite.ai: Scores are reported as task-clipped rubric scores and can be decomposed across ten expert-defined behavioral axes, including context seeking, empathy, urgency calibration, and reality testing.
  12. Refutes nxcode.io: A high score therefore measures coverage of the rubric, not a proven benefit to a person.
  13. Context nxcode.io: It does not observe what happens to the person after the conversation.

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

Next letterPerplexity let nine AI models loose inside its own sandbox with root access and told them to break out. In 108 runs, none got through the wall. Four found a way under the fence.

This story is a stable, citable object. If you can falsify a verdict, tell us. Corrections are loud here.