KKivira Research & Impact
Sustainable AI · Clinical Safety

Safer mental health AI shouldn’t have to cost the earth. We measured the trade-off across 47 AI configurations.

A study from Kivira Health, the University of Roehampton and the University of Isfahan, accepted to ACM SusMod, puts clinical safety and environmental cost side by side for therapeutic AI — two things rarely evaluated together. The final few points of safety came at a steep price in energy, carbon and water.

Kivira HealthSeptember 20265 min readRead the paper

AI models are increasingly being explored as conversational agents for mental health support. To be safe, such a system has to sustain a multi-turn conversation, pick up clinically relevant risk signals, and respond appropriately to issues such as suicidal ideation or self-harm — without saying anything that could cause harm.

A common response is to reach for the biggest frontier models, on the assumption that larger, more capable systems handle clinical nuance better. They often do score well on safety benchmarks. But serving large models at scale takes substantial computing power, with real consequences for energy use and emissions.

Today, those two questions are usually asked separately. Safety benchmarks rarely report environmental cost, and sustainability estimates rarely consider clinical safety. That makes it hard to tell whether a small gain in safety is worth a large increase in energy, carbon or water.

Putting safety and footprint on the same chart

The team combined two sources. For safety, they used the public K-Bench leaderboard, which scores AI models in simulated multi-turn mental health conversations involving suicide, self-harm, domestic violence and substance misuse, using a rubric calibrated against clinician ratings. The primary outcome was K-Bench's combined risk score, which pools recognising risk and exploring it.

For environmental cost, they used EcoLogits, an open-source tool that estimates the life-cycle footprint of AI inference from model, hardware and data-centre assumptions. Each model was scored on four indicators — energy use, global warming potential, water consumption and abiotic resource depletion — all standardised to one million generated output tokens.

Of 90 leaderboard configurations, 47 across 13 base models had enough public information to be estimated. The team then mapped the Pareto frontier: the models that no other model beats on both safety and footprint at once.

“Relying solely on larger models or additional inference-time computation may be an inefficient strategy for improving safety in therapeutic AI systems.” — from the paper

The last few points are the most expensive

The highest-scoring model, gpt-5.5, reached a risk score of 96.02 at an estimated 6.44 kWh per million output tokens. claude-haiku-4.5 scored 93.41 using an estimated 0.11 kWh — 98.3% less energy for a 2.61-point lower score. Put the other way, that final 2.61-point gain in safety corresponded to an approximately 60-fold increase in estimated energy use. Across the four indicators, gpt-5.5 had roughly 57 times the global warming potential, 56 times the water consumption and 39 times the abiotic depletion of claude-haiku-4.5.

The relationship was non-linear: across all four indicators, the largest jumps in estimated impact came among the models with the highest safety scores. Higher cost did not guarantee higher safety either — gemini-3.5-flash carried one of the larger energy estimates in the table but a lower score than claude-haiku-4.5.

Thinking harder didn't reliably make models safer

Match the model to the risk

The authors argue that therapeutic AI deployment should be treated as a multi-objective problem, weighing clinical safety, environmental impact, cost and operational risk together rather than simply picking the highest safety score. One practical route is dynamic model selection, or model cascading: smaller, efficient models handle lower-risk interactions, and larger models are brought in where the additional performance is clinically relevant.

That fits the wider Kivira approach of measuring what matters and matching care to need. Alongside dimension- and case-level analysis, the same K-Bench scores that show which models handle risk well can help show when a larger model's extra performance is worth its cost.

“Aggregate K-Bench scores are comparative, not universal safety thresholds.” — from the paper

What this study does and doesn't show

Environmental figures are EcoLogits estimates based on modelled hardware and data-centre assumptions, not direct measurements from providers, and the four indicators share those assumptions, so they are not fully independent. Fifteen base models, including Claude Fable 5, were excluded because the information needed to estimate them wasn't available, and EcoLogits gives the same estimate for every configuration of a base model. Efficient models did not reach the highest safety scores; whether a smaller model is acceptable depends on the use case, risk profile and monitoring, and small score differences are not necessarily clinically meaningful.

The research: A. A. Safaei (University of Isfahan), L. M. Vowels (University of Roehampton), M. J. Vowels (Kivira Health), A. Jha (Kivira Health) & S. Rahimi (University of Roehampton). “Quantifying the Relationship Between Clinical Safety and Environmental Impact in Therapeutic LLMs.” Accepted to ACM SusMod, MODELS Companion 2026, Málaga, Spain. doi:10.1145/3837062.3839341 · arXiv:2608.11830.

Funding & interests: Supported by the Digital Good Network and Kivira Health. M. J. Vowels and A. Jha are affiliated with Kivira Health, which also developed K-Bench.

More from Kivira Research
Someone in crisis is typing to a chatbot right now125 AI configurations, clinician-calibrated Mental illness can take up to 23 years to reach a clinicianA reference standard for mental health care Before a treatment reaches you, someone should check the study behind itAn AI audit of 2,448 clinical papers An established face set in mental health research shows only White facesAI-generated faces, 1,047 participants