The trick: Self-Marked
OpenAI built a mental health test, had its own model grade it, and its own model came top.
Answers written by licensed clinicians scored 38.5%. GPT-6 Astra scored 57.3%.
OpenAI says results on MentalHealthBench show the steady improvement of AI systems in helping people navigate mental health situations.
Before you read on. Your call?
TRUE, BUT
38.5%
the benchmark uses synthetic conversations, and each reply is graded against expert-written criteria by an automated grader, OpenAI's own GPT-5.6 Sol. OpenAI's GPT-6 Astra tops the reported results at 57.3%. Answers written by clinicians scored 38.5%, which the authors attribute largely to clinicians writing short, in-person style replies, and completions written with the rubric in view scored 99.0%.
The twist
as NxCode puts it, a high score measures coverage of the rubric, not a proven benefit to a person. Nothing in this record measures what happened to anyone after a conversation, so 'helping people' is a step beyond what the test observes.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
The trick has a name
We call it Self-Marked: graded by the party that benefits from the grade. You'll see it again. Learn to spot it →
Receipts
- Supports web.archive.org:
Results on MentalHealthBench show the steady improvement of AI systems in helping people navigate mental health situations.
- Context web.archive.org:
For each conversation, we use an automated grader, GPT‑5.6 Sol, to assess model responses against the expert-written criteria.
- Context web.archive.org:
Using privacy-preserving techniques, we created synthetic mental health conversations that accurately reflect real-world usage patterns of AI for mental health.
- Context web.archive.org:
While ChatGPT is not a substitute for therapy or professional care, expert-informed insights help us measure the progress towards AI models
- Supports unite.ai:
In OpenAI’s reported results, GPT-6 Astra scored highest at 57.3% task-clipped, followed by GPT-6 Sol at 53.9%, Claude Opus 5.5 at 52.4%, and GPT-6 Luna at 50.2%.
- Refutes unite.ai:
Expert-authored completions written by clinicians scored 38.5%, which the authors attribute largely to clinicians writing short responses as if in an in-person conversation, often asking a single question or making a simple statement.
- Context unite.ai:
Rubric-aware completions, written with the grading rubrics provided, scored 99.0%, which the paper describes as a sanity check on the evaluation’s noise ceiling.
- Refutes nxcode.io:
In a reference comparison, expert-authored completions scored 38.5% on the task-clipped measure, while GPT-6 Astra scored 57.3% and GPT-6 Sol 53.9%
- Context web.archive.org:
MentalHealthBench was co-created with a global cohort of more than 80 licensed mental health experts from 22 countries.
- Context unite.ai:
The authors state that no benchmark captures everything that matters in a personal conversation, and they position MentalHealthBench as an auditable diagnostic tool rather than a definitive leaderboard.
- Context unite.ai:
Scores are reported as task-clipped rubric scores and can be decomposed across ten expert-defined behavioral axes, including context seeking, empathy, urgency calibration, and reality testing.
- Refutes nxcode.io:
A high score therefore measures coverage of the rubric, not a proven benefit to a person.
- Context nxcode.io:
It does not observe what happens to the person after the conversation.
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.