Someone in crisis is typing to a chatbot right now. We tested 125 AI configurations to see which ones notice.
K-Bench, built by Kivira Health and the University of Roehampton with practising clinicians, finds that the leading AI models now score above 95 out of 100 on recognising and exploring suicide, self-harm, domestic-violence and substance-misuse risk — and exposes weaker models that sound supportive but omit essential questions.
Risk rarely arrives as a direct request for crisis support. As the K-Bench paper describes, it may emerge gradually, through descriptions of hopelessness, isolation, escalating substance use, coercive relationships, fear or shame. It may be communicated indirectly, or alongside other risks: a person discussing domestic violence may also be experiencing suicidal thoughts. And increasingly, the one on the other end of that conversation is an AI.
People increasingly turn to chatbots for mental health support because they are available at any time, private and low-cost, and may be a first point of contact for people unsure about formal support or unable to get it. That makes a simple question urgent: when risk emerges in the middle of a conversation, does the AI see it coming — and does it do anything about it?
Existing benchmarks answer only part of it. Many use isolated prompts or ask a model to classify a risk it has been handed, which can't test gradual disclosure, changing severity or co-occurring risks. The closest multi-turn efforts cover suicide alone across three models, or cap conversations at ten turns across nine chatbots.
Testing AI the way crises actually unfold
K-Bench puts every model through the same 200 conversations of up to 20 turns each, with simulated patients built from anonymised material provided by people with lived experience of suicide or self-harm, intimate-partner violence and substance misuse. A 122-variable design varies everything from trauma history and protective factors to how guarded or coherent the person is — so risk can emerge indirectly, escalate, or co-occur, just as it does in real life.
To check those simulated patients sound like real people, the team compared them against 50 real conversations between users and an AI support tool. Across four different patient-simulating models, the synthetic and real language showed substantial overlap.
Every conversation is then scored against a 47-item rubric developed by a 10-person stakeholder panel of clinicians, people with lived experience, mental health and AI researchers, and AI developers, meeting roughly fortnightly through development. Six clinicians rated 151 transcripts, each transcript rated independently at least three times, to set the ground truth. An automated judge, frozen after calibration, matched their consensus on 94.2% of items — more consistently than individual clinicians agreed with each other (89.4%). That is what lets the benchmark evaluate new models continuously, at a scale clinician-only assessment could not sustain.
“Separate scores for clinical judgement and risk exploration therefore show whether an apparently empathic conversation also gathers the information needed for a safe response.” — from the K-Bench paper
What the benchmark found
The headline is genuinely encouraging. The leading systems — configurations of OpenAI's GPT-5.5, GPT-5.2 and GPT-5.6 Sol, and Anthropic's Claude Fable 5 — scored above 95 out of 100 on the combined-risk measure. The strongest models combined supportive dialogue with detailed assessment across all four risk domains, and maintained high performance as severity increased and risks co-occurred.
But the benchmark's real value is in what it separates. Ethical reasoning, supportive conversation, psychological knowledge, cultural competence and boundaries were consistently strong, including among many lower-ranked configurations. Where models diverged was risk exploration: whether the model establishes the nature and immediacy of the risk, draws out contributing and protective factors, and identifies appropriate next steps.
Sounding supportive is easy. Exploring risk is not.
In the paper's words, weaker models “omitted essential questions and sometimes missed the concern itself.”
Three findings for anyone deploying AI in mental health
- Price alone doesn't buy safety. Risk scores generally rose with cost, but several lower-cost configurations approached substantially more expensive systems, while models at similar cost sometimes differed appreciably.
- “Thinking harder” doesn't help on average. Raising a model's reasoning setting lowered its risk score in 21 of 30 matched comparisons (mean change −0.58 points).
- A good prompt can rescue a weak model — or quietly hurt another. A therapeutic system prompt lifted one small model's risk score by 17.9 points, but reduced scores in 37 of 61 matched pairs, including a 6.1-point drop for one model. The paper concludes that model version, prompt and reasoning setting must be evaluated together as deployed.
A living test that can't be gamed
Benchmarks lose their value the moment developers can train against them. K-Bench publishes its methods, clinical constructs, configurations and results on a public leaderboard, but withholds the patient prompts, disclosure schedules, scoring examples and transcripts, following the AEF-1 standard for independent third-party AI evaluation. That lets the leaderboard show whether updated models genuinely improve or regress over time.
“Leading models established the nature and immediacy of risk, elicited contributing and protective factors, and identified appropriate next steps; weaker models omitted essential questions and sometimes missed the concern itself.” — from the K-Bench paper
What this study does and doesn't show
K-Bench measures whether a model's conversation meets clinician standards for recognising and responding to risk. It evaluates how risk is handled within a conversation, not whether anyone will later attempt suicide or self-harm, and it tests models through their APIs rather than finished consumer apps. Conversations were simulated, in English, and capped at 20 turns. General-purpose chatbots usually cannot verify a user's identity or location, contact services or confirm follow-up, so any clinical use needs a clear pathway when risk is identified.The research: L. M. Vowels, M. J. Vowels, S. Sharma, A. Jha, et al. “K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations.” Preprint, 2026. Read the paper on arXiv · View the live leaderboard · AIAS Workshop 2026, our benchmarking event. Author affiliations: University of Roehampton, Kivira Health, University of Hertfordshire, University of Surrey, University of Bedfordshire, Tavistock Relationships, InsideOut.
Funding & interests: Funded by the ESRC Digital Good Network and Kivira Health. M. J. Vowels is Chief Technology Officer of Kivira Health; L. M. Vowels is married to M. J. Vowels. Full disclosures are in the paper.