An established face set used in mental health research shows only White faces. We tested an AI alternative.
Difficulty reading emotion in faces is one of the best-established findings in schizophrenia-spectrum disorders. A new study from Kivira Health and the University of Roehampton shows that AI-generated faces spanning a range of backgrounds are recognised clearly and pick up symptom-related differences comparable to those found with an established photo set.
A glance across a room tells us whether someone needs comfort, feels threatened, or wants to be left alone. Errors at that first stage can accumulate into misunderstandings, reduced social participation and poorer everyday functioning. The largest, most established recognition impairments occur in schizophrenia-spectrum disorders, and difficulties are also reported in bipolar disorder, depression, OCD, social anxiety, ADHD and insomnia.
That makes the emotion-recognition task — look at a face, name the feeling — a transdiagnostic outcome in mental health research. But the test is only as good as its faces. FACES, one of the most established photo sets, contains only White models. A review of 36 facial-expression databases found that 21 either included a single racial group or didn't report race at all.
A White-only set gives participants from different backgrounds systematically different measurement experiences, and leaves open whether apparent group differences reflect emotion-recognition ability, familiarity with the faces, or both. Newer, more diverse photo sets exist, but each identity must be recruited, photographed, directed, edited, consented and normed — slow, expensive, and hard to extend.
Faces designed, not photographed
Laura Vowels and Matthew Vowels generated 161 photorealistic portraits: 23 identities of different ages and genders, labelled Asian, Black, Indian, Latin and White, each showing anger, disgust, fear, happiness, sadness, surprise and a neutral expression. Because the people don't exist, appropriately documented synthetic images can reduce the privacy and consent constraints of sharing real photographs, and new identities or conditions can be added without recruiting new actors.
Then came the real test. More than a thousand adults across the UK and US — recruited so that many reported mental health conditions — completed the task online, seeing both the AI faces and the FACES photographs in random order. Together they produced over 83,000 valid trials, alongside standard questionnaires for psychotic-like experiences, obsessive-compulsive symptoms, mania, ADHD, anxiety, depression and more.
“The images can measure real differences between participants, which is a higher bar than looking convincing.” — from the paper
Clearer, faster, and just as revealing
People recognised the AI expressions more accurately and faster (2.29 vs 2.64 seconds). Crucially, performance on the diverse AI set and the White-only AI subset did not differ significantly.
Emotions correctly identified
Higher accuracy isn't automatically better: a task that's too easy compresses differences between people. That's where the study's key result lies. People with more psychotic-like experiences and more obsessive-compulsive symptoms made more errors on the AI faces — just as they did on the photographs, with no significant difference in the strength of the link. The size of the psychosis association was in line with earlier research. Analysed together, each test explained its own share of symptom variation, and the AI faces also tracked mania and ADHD symptoms.
In other words, the AI task didn't make everyone perform alike: people with greater symptom burden continued to make more errors even when the expressions were comparatively clear.
Why this matters beyond the lab
- Fairer measurement. White and non-White participants did not differ significantly on any face set, and there was no broad same-ethnicity advantage.
- Tunable difficulty. The same identities could be generated at subtler intensities to detect milder difficulties, or to track change during treatment.
- Shareable by design. The images will be shared so other researchers can reproduce, challenge and extend the work.
Measurement that looks like the people it serves
For Kivira, this is a concrete example of a larger aim: making mental health measurable in a way that works for everyone who takes the test. The emotion-recognition task ran as one of four browser-based Kivira tasks in a session of about 40 minutes.
“AI generation changed the demographic composition and overall difficulty of the task while preserving the symptom-related differences for which facial emotion recognition is used in clinical research.” — from the paper
What this study does and doesn't show
The links between face-reading and symptoms are small — useful for research across many people, but not strong enough to diagnose anyone. Symptoms were self-reported, with no clinical interviews, and the task was done once, so stability over time isn't yet known. The AI images differed from the photographs in more than diversity (they were sharper and more uniform), and some emotions, like happiness, were near ceiling. Ethnicity was compared only as White versus non-White.The research: L. M. Vowels (University of Roehampton) & M. J. Vowels (Kivira Health). “Emotion Recognition From AI-Generated and Human Faces: Performance, Mental Health, and Demographic Diversity.” Preprint, 21 September 2026. Read the paper on PsyArXiv.
Funding & interests: Funded by Kivira Health, which paid participant reimbursement and provided the digital study platform. M. J. Vowels is employed by Kivira Health; L. M. Vowels is married to M. J. Vowels.