KKivira Research & Impact
Social Cognition · Generative AI

An established face set used in mental health research shows only White faces. We tested an AI alternative.

Difficulty reading emotion in faces is one of the best-established findings in schizophrenia-spectrum disorders. A new study from Kivira Health and the University of Roehampton shows that AI-generated faces spanning a range of backgrounds are recognised clearly and pick up symptom-related differences comparable to those found with an established photo set.

Kivira Health & University of RoehamptonSeptember 20265 min readRead the paper

A glance across a room tells us whether someone needs comfort, feels threatened, or wants to be left alone. Errors at that first stage can accumulate into misunderstandings, reduced social participation and poorer everyday functioning. The largest, most established recognition impairments occur in schizophrenia-spectrum disorders, and difficulties are also reported in bipolar disorder, depression, OCD, social anxiety, ADHD and insomnia.

That makes the emotion-recognition task — look at a face, name the feeling — a transdiagnostic outcome in mental health research. But the test is only as good as its faces. FACES, one of the most established photo sets, contains only White models. A review of 36 facial-expression databases found that 21 either included a single racial group or didn't report race at all.

A White-only set gives participants from different backgrounds systematically different measurement experiences, and leaves open whether apparent group differences reflect emotion-recognition ability, familiarity with the faces, or both. Newer, more diverse photo sets exist, but each identity must be recruited, photographed, directed, edited, consented and normed — slow, expensive, and hard to extend.

Faces designed, not photographed

Laura Vowels and Matthew Vowels generated 161 photorealistic portraits: 23 identities of different ages and genders, labelled Asian, Black, Indian, Latin and White, each showing anger, disgust, fear, happiness, sadness, surprise and a neutral expression. Because the people don't exist, appropriately documented synthetic images can reduce the privacy and consent constraints of sharing real photographs, and new identities or conditions can be added without recruiting new actors.

Then came the real test. More than a thousand adults across the UK and US — recruited so that many reported mental health conditions — completed the task online, seeing both the AI faces and the FACES photographs in random order. Together they produced over 83,000 valid trials, alongside standard questionnaires for psychotic-like experiences, obsessive-compulsive symptoms, mania, ADHD, anxiety, depression and more.

“The images can measure real differences between participants, which is a higher bar than looking convincing.” — from the paper

Clearer, faster, and just as revealing

People recognised the AI expressions more accurately and faster (2.29 vs 2.64 seconds). Crucially, performance on the diverse AI set and the White-only AI subset did not differ significantly.

Higher accuracy isn't automatically better: a task that's too easy compresses differences between people. That's where the study's key result lies. People with more psychotic-like experiences and more obsessive-compulsive symptoms made more errors on the AI faces — just as they did on the photographs, with no significant difference in the strength of the link. The size of the psychosis association was in line with earlier research. Analysed together, each test explained its own share of symptom variation, and the AI faces also tracked mania and ADHD symptoms.

In other words, the AI task didn't make everyone perform alike: people with greater symptom burden continued to make more errors even when the expressions were comparatively clear.

Why this matters beyond the lab

Measurement that looks like the people it serves

For Kivira, this is a concrete example of a larger aim: making mental health measurable in a way that works for everyone who takes the test. The emotion-recognition task ran as one of four browser-based Kivira tasks in a session of about 40 minutes.

“AI generation changed the demographic composition and overall difficulty of the task while preserving the symptom-related differences for which facial emotion recognition is used in clinical research.” — from the paper

What this study does and doesn't show

The links between face-reading and symptoms are small — useful for research across many people, but not strong enough to diagnose anyone. Symptoms were self-reported, with no clinical interviews, and the task was done once, so stability over time isn't yet known. The AI images differed from the photographs in more than diversity (they were sharper and more uniform), and some emotions, like happiness, were near ceiling. Ethnicity was compared only as White versus non-White.

The research: L. M. Vowels (University of Roehampton) & M. J. Vowels (Kivira Health). “Emotion Recognition From AI-Generated and Human Faces: Performance, Mental Health, and Demographic Diversity.” Preprint, 21 September 2026. Read the paper on PsyArXiv.

Funding & interests: Funded by Kivira Health, which paid participant reimbursement and provided the digital study platform. M. J. Vowels is employed by Kivira Health; L. M. Vowels is married to M. J. Vowels.

More from Kivira Research
Someone in crisis is typing to a chatbot right now125 AI configurations, clinician-calibrated Before a treatment reaches you, someone should check the study behind itAn AI audit of 2,448 clinical papers Mental illness can take up to 23 years to reach a clinicianA reference standard for mental health care Safer mental health AI shouldn’t have to cost the earthClinical safety vs environmental cost, 47 configurations