OpenAI has released MentalHealthBench, a benchmark designed to evaluate AI performance in mental health conversations. Built with input from licensed clinicians, it tests models against realistic scenarios where both helpfulness and safety are required simultaneously. This is not a general-purpose safety rubric. It targets the specific failure modes that emerge when users discuss crisis, self-harm, and psychological distress with AI systems.
The benchmark matters because most existing evaluations treat safety and helpfulness as a tradeoff. MentalHealthBench treats them as co-requirements. A model that refuses to engage is not passing. A model that engages recklessly is not passing either. The expert-informed design means the scoring reflects clinical judgment, not just policy compliance.
The full paper is worth reading for how the benchmark was constructed: which clinical experts were consulted, how scenarios were sourced and validated, and where current frontier models actually fail. The gap between a model sounding empathetic and a model behaving safely is larger than most benchmarks reveal. This one tries to measure it.
[READ ORIGINAL →]