Before a treatment reaches you, someone should check the study behind it. We checked 2,448.
Kurate, an AI-assisted system built by Kivira Health, turns each paper into a structured, source-traceable quality assessment, comparing it with its own registration, protocol and funding records. In its first large audit, only one in five papers was both graded A or B and retained as supporting evidence.
Clinical decisions rest on published studies. Scientific search tools can find the relevant papers, but they typically assume peer review has already done the work of separating strong evidence from weak.
That assumption doesn't always hold. The clinical literature has well-documented distortions: incomplete registration and reporting, gaps between the outcomes a trial planned and the ones it published, selective publication, and small samples that can exaggerate how strong the evidence looks. Even visible preregistrations are often not checked during peer review.
The traditional safeguard — systematic review with risk-of-bias and GRADE-style certainty ratings — works, but it needs trained reviewers, months of effort, and repeated updating as the evidence moves. Publication volume has outpaced review capacity.
Reading a trial like an auditor
Kurate treats each paper as the end of a paper trail. It ingests the published article and attaches whatever external context is available: the trial's registry entry, its protocol, funding records, citations and integrity signals. Then it compares what was planned with what was reported. Did the trial reach the sample size its power calculation required? Do the reported outcomes match the registered ones? Was the analysis prespecified? Who funded it?
Each paper is scored on eight dimensions — statistical power, causal identification, preregistration, selective reporting, measurement validity, analysis prespecification, reporting transparency and conflict of interest — and each judgement is linked to the source evidence behind it. The result is a letter grade from A to F, plus a separate decision on whether the paper should be used as supporting evidence at all.
That matters because an audit is only as good as its reading. So the team tested Kurate against an independent benchmark in which human experts had marked, item by item, what 200 clinical trial documents actually say. On trial protocols, Kurate matched the experts' consensus with an agreement score of 0.94 — above the published automated system (0.80) and statistically indistinguishable from an individual human expert (0.89). On results papers it scored 0.81, on par with that automated system and below the human figure.
“A paper may be peer-reviewed, published, and directly relevant to the question being asked, and may nonetheless be underpowered, selectively reported, weakly prespecified, or poorly documented.” — from the Kurate paper
What 2,448 papers look like, graded
Across the corpus, grades spread widely. About three in ten papers earned an A or B. Its separate evidence-use decision was more conservative still. Nearly half the papers — 1,143 — were excluded as supporting evidence for downstream synthesis (they stay in the database, but aren't relied on). Only 20% were both graded A or B and retained.
How the 2,448 papers graded
The most common problems were exactly the ones that are hardest to spot from the paper alone. Selective reporting was flagged as weak in 65% of papers, statistical power in 61%, and analysis prespecification in 55%. Each is a criterion where appraisal depends on connecting several research documents — the paper and its registry entry or protocol.
Mental health stood out. The 912 mental-health papers — the largest domain in the corpus — averaged a quality score of 0.53, below the corpus average of 0.57 and behind cancer (0.63) and infectious disease/HIV (0.63). 52.9% graded C or better, against 62.7% overall.
There is good news in the timeline. Before 2000, just 4.4% of papers graded C or better; for 2000–2009 it was 21.9%, and for 2020–2026, 73.2% — a pattern the authors note is consistent with better registration and reporting infrastructure.
One trial, fully traced
- The team walked through SPRINT, the landmark blood-pressure trial, as a worked example. Kurate gave it an A (0.92).
- It linked the paper to a registration first submitted before the study start date, found the key reported endpoints matched the preregistered outcomes, and found the early stop was disclosed rather than hidden.
- It scored two dimensions lower where the paper-level record was less explicit — allocation concealment and blinding, and author-level conflict disclosures routed to external forms.
From one paper to a whole field
Because every paper is ingested once and stored as a structured record, the same assessments can support study-level reports, quality-aware search, and corpus-scale summaries across domains, interventions, funders, institutions or portfolios. A query interface lets users ask a clinical question in PICO format and returns a summary of the evidence across the matching studies.
For Kivira, this is the evidence layer that its reference standard for mental health care calls for: an evidence library with strength-of-evidence grading and provenance tracking, from primary study to system output.
What this study does and doesn't show
Kurate's reading of papers was checked against human experts; its grades and keep-or-exclude decisions have not yet been, because no comparable expert-labelled benchmark exists. The benchmark comparison covered only the items Kurate models, which may flatter it. The corpus is a targeted set of open-access trials, not a random sample of medicine, so its percentages describe this corpus. Grades are Kurate outputs, not verified properties of each study, and are meant to support expert judgement rather than replace it.The research: “Kurate: Scalable Scientific Quality Analysis.” Manuscript, 2026; preprint forthcoming. Try the Kurate demo.
Interests: Kurate is developed by Kivira Health.