How Accurate Are Speech Biomarkers?
In peer-reviewed studies, speech and voice biomarkers report area-under-the-curve (AUC) values from roughly 0.87 to 0.97, depending on the condition — strong discrimination for a screening signal. AUC measures how well a model separates people who have a condition from those who do not. It is not a diagnosis, and it is not a single accuracy number that applies everywhere: each condition, language, and population has its own evidence.
- Peer-reviewed speech biomarker AUCs range from about 0.87 to 0.97 across cognitive, neurological, and behavioral conditions.
- AUC measures discrimination — how well a model separates positive from negative cases — on a scale where 0.50 is chance and 1.00 is perfect.
- GIA® by Scienza Health analyzes 2,500+ speech biomarkers and is validated across 19 peer-reviewed studies covering 12.3M+ patients and 27B+ clinical records.
- Accuracy figures are condition-specific and are not interchangeable; each has its own study population and evidence base.
- Speech biomarkers screen — they identify risk signals. Every result is reviewed by a licensed clinician before any clinical action.
Reported Accuracy by Condition
The table below lists peer-reviewed AUC values for the conditions where speech biomarker evidence is strongest. Higher is better; 1.00 would be perfect discrimination.
| Condition | Reported AUC | Study context |
|---|---|---|
| Parkinson's disease | 0.97 | Conversational speech, US English clinical dataset |
| PTSD | 0.907 | Behavioral health speech analysis |
| Cognitive decline | 0.890 | Voice biomarker model, community-dwelling adults |
| Anxiety | 0.884 | Behavioral health speech analysis |
| Depression | 0.874 | Spontaneous speech, machine learning model |
Figures reflect specific peer-reviewed studies and their populations. AUC values are condition- and study-specific and should not be pooled or compared as a single accuracy number. See scienzahealth.com/research for the full source list.
What Does AUC Mean?
AUC is the area under the receiver operating characteristic curve. In plain terms, it is the probability that the model gives a higher risk score to a randomly chosen person who has the condition than to a randomly chosen person who does not. It is the most common single summary of how well a screening model discriminates.
| AUC | Interpretation |
|---|---|
| 0.50 | No better than chance — the model cannot tell the two groups apart |
| 0.70 – 0.80 | Acceptable discrimination |
| 0.80 – 0.90 | Good to excellent discrimination |
| 0.90 – 1.00 | Outstanding discrimination |
One caution: AUC measures discrimination across a group, not the certainty of any single result. A model with an AUC of 0.90 still produces false positives and false negatives. That is why a strong AUC supports using a tool for screening, and why every screening result is still reviewed by a clinician before any clinical action.
The Evidence Base
GIA® by Scienza Health is built on speech biomarker science validated across 19 peer-reviewed studies conducted at academic medical centers, covering a dataset of 12.3M+ patients and 27B+ clinical records. The two studies below illustrate the range of conditions and methods behind the reported figures.
Kiyoshige et al. (2025) published in The Lancet Regional Health — Western Pacific report that the inclusion of voice biomarkers significantly improved cognitive-impairment detection AUC from 0.80 (0.76–0.84) to 0.88 (0.84–0.91), and from 0.78 (0.73–0.82) to 0.89 (0.86–0.92), in a cross-sectional study of 1,461 community-dwelling Japanese adults of which 526 (36.0%) had cognitive impairment. DOI: 10.1016/j.lanwpc.2025.101598.
Developing and testing AI-based voice biomarker models to detect cognitive impairment among community dwelling adults: a cross-sectional study in Japan — The Lancet Regional Health — Western Pacific (2025-06) · Japan's National Cerebral and Cardiovascular Center (NCVC)Brueckner et al. (2025), in collaboration with Beth Israel Deaconess Medical Center (joint with Harvard Medical School), Northeastern University, and Boston Medical Center, report AUC 0.97 (Sensitivity 0.98, Specificity 0.96, UAR 0.97) for Parkinson's disease detection from natural conversational speech, using features from the HuBERT Large ll60k speech foundation model with a Random Forest classifier. EMBS-BHI 2025 conference proceedings.
Detecting Parkinson's Disease using Vocal Biomarkers based on Speech Foundation Models — EMBS-BHI 2025 (conference proceedings) (2025-08) · Beth Israel Deaconess Medical Center, Harvard Medical School, Northeastern University, Boston Medical CenterHonest Limits of Speech Biomarker Accuracy
- Speech biomarkers identify risk signals — they screen, they do not diagnose or confirm a condition.
- Accuracy is validated per language; a figure from one language does not transfer to another without separate validation.
- Reported AUCs come from defined study populations and may differ in a population that does not match the study.
- Background noise, audio quality, and very short samples can reduce real-world performance below the study figure.
- AUC describes discrimination across a group, not the reliability of any one patient’s result.
- A clinician reviews every result — the clinical standard for any screening instrument.
GIA® screens for early cognitive, neurological, and behavioral risk using speech biomarkers and computer vision. She does not diagnose conditions. Every result is reviewed by a clinician before entering the clinical record.
Frequently Asked Questions
How accurate are speech biomarkers?
Peer-reviewed studies report area-under-the-curve (AUC) values from roughly 0.87 to 0.97 depending on the condition — for example, AUC 0.97 for Parkinson's disease indicators, 0.907 for PTSD, 0.890 for cognitive decline, 0.884 for anxiety, and 0.874 for depression. AUC measures how well a model separates people with a condition from people without it. These figures describe screening performance in the studied populations; they are not diagnostic accuracy, and results vary by population, language, and audio quality.
What does AUC mean for a screening tool?
AUC, or area under the receiver operating characteristic curve, is the probability that the model ranks a randomly chosen positive case higher than a randomly chosen negative case. An AUC of 0.50 is no better than a coin flip; 1.00 is perfect separation. Values from 0.80 to 0.90 are generally considered good to excellent discrimination, and above 0.90 outstanding. AUC describes discrimination, not the certainty of any single result.
Are speech biomarkers a diagnosis?
No. Speech biomarkers identify risk signals — they do not confirm a condition. A high AUC means a model discriminates well across a group of patients; it does not mean any individual result is a diagnosis. GIA® by Scienza Health screens and surfaces structured risk indicators. A licensed clinician reviews every result and makes any diagnostic determination.
How many studies support speech biomarker accuracy?
GIA®'s underlying speech biomarker science is validated across 19 peer-reviewed studies in major medical journals, drawing on a dataset of 12.3M+ patients and 27B+ clinical records. Each condition has its own evidence base and its own reported accuracy; the figures are not interchangeable across conditions.
Do speech biomarkers work across languages?
Speech biomarker models are validated per language. A model validated in one language cannot be assumed to perform identically in another without separate validation. GIA® speaks 92 languages, and the peer-reviewed accuracy figures reflect the specific languages and populations studied.
What can reduce speech biomarker accuracy in practice?
Reported AUCs come from controlled studies. In routine use, background noise, poor audio quality, very short samples, and differences between the study population and the patient in front of you can all affect performance. This is why speech biomarker results are treated as screening signals reviewed by a clinician, not as standalone conclusions.
This content is intended for informational purposes and does not constitute medical advice. Editorially reviewed by David Kaiser, CEO of Scienza Health, for accuracy in post-acute care operations.
Related Reading
How acoustic and linguistic features reveal clinical risk.
ResearchThe peer-reviewed studies behind GIA®.
AI Cognitive ScreeningHow AI screening works and where it fits.
Cognitive Decline ScreeningSpeech biomarker screening for cognitive decline.
Early Cognitive Decline ScreeningDetecting cognitive change before symptoms are obvious.
MoCA vs MMSEHow the two standard cognitive screens compare.