OpenAI Unveils MentalHealthBench to Evaluate AI Safety in Mental-Health Conversations
Summary
OpenAI unveils MentalHealthBench, an open safety benchmark for AI mental-health conversations that tests context gathering, user agency and actionable guidance across synthetic emotional, high-acuity and emergency scenarios reviewed by 80 professionals from 22 countries.
Key Points
- OpenAI unveils MentalHealthBench, an open benchmark that evaluates AI responses to mental-health conversations for safety, context gathering, user agency and actionable guidance.
- Astra leads the tested models with a 57.8% overall score, ahead of GPT-6 Sol at 54%, Luna at 50.3% and Claude Opus 5 at 48.1%.
- Eighty mental-health professionals from 22 countries, representing 19 languages and 20 subspecialties, assess synthetic conversations spanning emotional, high-acuity and emergency scenarios.