Trending: On-device modelsSearch
iHeartGeek
iTECH

OpenAI opens MentalHealthBench to test AI empathy

OpenAI has published an open benchmark built with more than 80 clinicians in 22 countries to measure how AI handles everyday mental health conversations, not just emergencies.

OpenAI's announcement card for MentalHealthBench: a small speech-bubble icon connected by dotted lines to a larger panel holding plus and minus symbols, above the words Introducing MentalHealthBench on a pale blue background.

OpenAI has published MentalHealthBench, an open benchmark for measuring how AI systems answer people in realistic mental health conversations. It was co-created with more than 80 licensed mental health experts and released publicly on 23 September, methods and all, so that other researchers can run their own evaluations against it.

The gap it is trying to close

Most evaluations of AI in this space have concentrated on emergency scenarios and judged success against broad, predefined criteria. OpenAI argues that has left a blind spot: very little measurement of how models perform across the ordinary range of mental health conversations, or of whether a reply matches what a clinician would have advised beyond simply avoiding a disallowed answer.

MentalHealthBench splits that range into three tiers. Non-acute conversations are everyday exchanges that may carry some emotional weight. High-acuity conversations point to serious distress or significant mental health concerns without an immediate emergency. The third tier covers emergencies, where the right answer includes steering someone towards real-world support. Scenarios are written for four kinds of user: adults, teenagers aged 13 to 17, caregivers and clinicians, across multiple languages and regions. The teen conversations were reviewed by clinicians with youth mental health expertise, and OpenAI notes that this testing approach may not capture every safeguard built into a specific product.

Written by hand, graded by a model

The substance of the benchmark is its rubrics. Experts read each synthetic conversation and wrote criteria for judging the response to the final user turn, each targeting one thing: asking the right question, seeking context, preserving the user's agency or offering practical guidance. Every criterion carries a weight from -10 to +10, where positive points reward helpful behaviour and negative points penalise harmful behaviour, and larger values signal greater clinical importance in that conversation.

Those criteria are not one clinician's opinion. Each conversation was read by at least three experts, and only criteria agreed by at least two and not contradicted by a third survived into the final set. Scoring is then automated: OpenAI uses GPT-5.6 Sol as the grader, assessing model responses against the expert-written criteria. The company says results so far show steady improvement across models, while stating plainly that ChatGPT is not a substitute for therapy or professional care.

What the experts said

Dr. Arthur Evans, chief executive of the American Psychological Association, framed the design of the benchmark around the range it covers: "Mental health exists on a continuum, from flourishing to everyday stress to acute crisis. AI systems that engage people across that range need to be grounded in both clinical science and lived experience."

Dr. Kim A. Baranowski put the case for clinician involvement bluntly, saying AI has the potential to support user well-being "but only if clinicians are meaningfully and actively involved throughout the research, development, and evaluation of emerging technologies." Dr. Kevin La Moureaux, a practising psychiatrist, said patients already bring these tools into their care, and that contributing clinical experience helps ensure the technology treats mental health with the care it deserves.

What users wanted that experts did not

OpenAI ran a second study alongside the benchmark, asking 44 adults who had used AI for mental health or emotional support, drawn from 16 countries and 14 languages, to rate model responses and write their own criteria. The comparison is the most interesting detail in the release. Users put more weight on tone and on practical next steps; experts put more weight on gathering relevant context and interpreting ambiguous situations carefully. The final scoring criteria stayed with expert consensus, and the user study was deliberately limited to non-acute conversations so that participants were not shown distressing material.

Our opinion

There is a quiet admission buried in MentalHealthBench that matters more than its leaderboard: nobody involved claims a rubric can tell you whether someone was actually helped. A benchmark can check that a model asked a good follow-up question, avoided projecting feelings onto the user and pointed towards real support, and those are genuine, testable behaviours that previous safety evaluations mostly ignored. But the 44 adults in the companion study wanted warmth and something to do next, and neither of those fits neatly into a criterion worth plus three. The scoring arrangement is worth watching for a different reason. OpenAI is grading model responses with GPT-5.6 Sol, its own model, against criteria its own paid cohort wrote, and publishing the method so that rivals can attack it. That is a mature way to release a benchmark, and it is also the reason third-party reruns will matter more than the launch numbers. The most useful thing in the release is not the scores, it is the divergence between what clinicians reward and what people say helps, because that gap is where an AI that sounds competent and feels cold gets built.