Speech and Language Features as a Digital Biomarker

Speech carries motor, cognitive and affective information at once, which makes it a rich signal and an unusually sensitive one to collect.

Status
Exploratory
Unit
acoustic and linguistic feature sets
Data type
Composite
Sensor
Microphone
Worn
Phone

Evidence maturity

Graded with the V3 framework: whether the sensor measures accurately, whether the algorithm has been validated against a reference standard, and whether the measure has been shown to matter clinically.

Verification
Established
Analytical validation
Emerging
Clinical validation
Limited

Acoustic measures separate early untreated Parkinson's disease from controls and linguistic features identify Alzheimer's disease in narrative speech, but most evidence is cross sectional and features are sensitive to recording conditions, language and accent.

What are Speech and Language Features

Speech and language features are quantitative properties extracted from recorded speech. They fall into two families. Acoustic features describe how the voice sounds: pitch variability, loudness, voice quality, timing, pauses and articulation precision. Linguistic features describe what is said: vocabulary richness, syntactic complexity, use of pronouns and repetition.

What makes speech attractive as a measure is that producing it requires precise motor control, language retrieval, working memory and executive planning simultaneously. A change in any of those can surface in speech, which is why the same recording can carry information about Parkinson's disease and about cognitive decline.

That richness is also the problem. A signal that reflects many things at once is hard to attribute to one, and speech is among the most identifying and most sensitive data types a study can collect.

How it is measured

Collection uses structured tasks rather than ambient recording: sustained vowel phonation, reading a standard passage, describing a picture, or answering a prompt. Standardising the task matters because speech features depend heavily on what the person is asked to do.

Acoustic analysis extracts measures of voice quality and timing from the waveform. Work in early untreated Parkinson's disease established that quantitative acoustic measurements can characterise speech and voice disorders before treatment confounds them, and dysphonia measures have been evaluated for telemonitoring. Linguistic analysis applies natural language processing to transcripts, and linguistic features have been shown to identify Alzheimer's disease in narrative speech.

Review and recommendation papers now exist specifically for evaluating speech based digital biomarkers, which is a sign the field is beginning to standardise.

Clinical use

Two research streams dominate. In movement disorders, speech is used as an early and sensitive motor marker, since articulation involves fine, rapid, highly coordinated movement. In neurodegeneration and psychiatry, linguistic and paralinguistic features are studied as markers of cognitive decline, depression severity and thought disorder.

The practical attraction is that a speech sample takes under a minute, needs no hardware beyond a phone, and can be collected remotely at high frequency. Measures are reported alongside the Consensus Auditory-Perceptual Evaluation of Voice, which is the direct clinical counterpart in which trained listeners rate the same properties by ear, and alongside voice related quality of life instruments that capture what the impairment costs.

Regulatory status

No regulatory qualification as an endpoint. Speech based measures appear in trials as exploratory outcomes, and recommendation papers for evaluating them have only recently been published.

Limitations

Consent and data governance are heavier here than for any other measure in this library. A voice recording is biometric, identifying, and may contain content the participant did not intend to share. Protocols must state what is recorded, whether raw audio is retained or discarded after feature extraction, who can access it, and how identification risk is managed.

Technically, features depend on the recording device, microphone, room acoustics and background noise, so cross study comparison is unreliable without controlled capture. They also depend on language, dialect, accent and education, and models trained on one population perform worse on another, which is a fairness problem as well as an accuracy one.

Evidence is mostly cross sectional. Longitudinal demonstration that speech features track progression or respond to treatment remains limited.

References

  • Robin J, et al. Evaluation of speech-based digital biomarkers: review and recommendations. Digit Biomark. 2020. pubmed.ncbi.nlm.nih.gov
  • Rusz J, et al. Quantitative acoustic measurements for characterization of speech and voice disorders in early untreated Parkinson's disease. J Acoust Soc Am. 2011. pubmed.ncbi.nlm.nih.gov
  • Fraser KC, et al. Linguistic features identify Alzheimer's disease in narrative speech. J Alzheimers Dis. 2016. pubmed.ncbi.nlm.nih.gov
  • Little MA, et al. Suitability of dysphonia measurements for telemonitoring of Parkinson's disease. IEEE Trans Biomed Eng. 2009. pubmed.ncbi.nlm.nih.gov
Devices that capture it
Related instruments

Direct counterpart of the Consensus Auditory-Perceptual Evaluation of Voice, in which trained listeners rate by ear the same voice properties an algorithm extracts from the waveform. The voice handicap and quality of life instruments capture what the impairment costs the person.

Use case
Diagnostic · Monitoring