
OpenAI releases MentalHealthBench for evaluating AI mental-health conversations
The AMW Read
An incremental update from an established foundation-model provider introduces an open evaluation resource relevant across the segment, but reports no model results or demonstrated deployment outcomes.
OpenAI releases MentalHealthBench for evaluating AI mental-health conversations
OpenAI has released MentalHealthBench, an open benchmark developed with more than 80 licensed mental-health professionals from 22 countries and regions. It covers 19 languages and synthetic conversations involving adults, adolescents, caregivers, and clinical professionals, spanning non-acute, high-risk, and emergency situations. Experts created detailed scoring criteria with weights ranging from -10 to +10, rewarding helpful behavior and penalizing harmful responses. At least three experts reviewed each conversation; only criteria they unanimously approved entered the final benchmark.
For foundation-model providers, the release makes sensitive conversational behavior a more explicit dimension of product evaluation. General-purpose assistants encounter contexts where a fluent answer can still be unsafe or unhelpful, so expert-defined criteria offer a way to examine behavior beyond broad capability scores. The multilingual scope also creates room to assess whether responses meet those criteria across languages. However, synthetic conversations and expert agreement on a rubric do not establish clinical effectiveness or demonstrate safety in live deployment. The report provides no comparative model results.
Builders evaluating mental-health interactions can use the benchmark to look for specific helpful and harmful response behaviors, while separately testing their own deployment conditions. OpenAI also says it has strengthened ChatGPT's responses to sensitive conversations, expanded access to crisis resources, and added a trusted-contact feature. For investors, the concrete diligence question is whether benchmark findings translate into measurable improvements in deployed behavior; this announcement alone does not answer it.

