The Complete Guide to Voice Biomarkers: How the Voice Reveals the Mind
Not what you say, but how you say it. The voice carries many signals that can connect to a person's state.
Minsoo · Founder & CEO
Voice AIRead 14 min
A voice biomarker is an acoustic feature extracted from speech that may relate to a person’s physical or psychological state. Just as blood values or imaging reflect the body’s condition, the voice can be a signal reflecting one’s state.
When we speak, the voice carries more than the content of our words. How we speak — pitch, speed, where we pause, the tremor and roughness of the voice — comes along with it. This article covers what voice biomarkers are, which features are examined, why they change with state, and what to be careful about.

01. Which features we examine
Features commonly used in voice research fall into three groups.
Prosodic features- Fundamental frequency (F0): the pitch of the voice. Mean, range, and intonation contours are analyzed.
- Speaking rate and pauses: how fast, and how often/long one pauses while speaking.
- Vocal energy/intensity: loudness and its variation.
- Jitter: tiny variation in the period of vocal fold vibration.
- Shimmer: tiny variation in the amplitude of vibration.
- HNR (Harmonics-to-Noise Ratio): the ratio of regular to noise components in the voice.
- MFCCs, formants, and other measures reflecting timbre and articulation.
These features have long been used in clinical and engineering research, and are now being extended with deep-learning representation learning.
02. Why the voice changes with state
Phonation is not simply making sound. It’s produced through precise coordination of breathing and the vocal and articulatory organs, and this process is affected by autonomic nervous system activity, muscle tension, and cognitive load.
- Under tension or stress, muscle tone and breathing patterns shift, subtly changing pitch and voice quality.
- In fatigue or exhaustion, speaking rate, energy, and pause patterns can change.
- Under high cognitive load, planning speech becomes harder, so pauses and hesitations can increase.
In other words, the voice is one channel through which the state of mind and body leaks out. That’s why research on how voice connects to stress, emotion, and depression has continued steadily. The specific link between voice and stress is detailed in Voice biomarkers and stress.
03. How far the research has come
Voice-based state estimation has been studied especially in depression and suicide-risk assessment, with several systematic reviews accumulated (see References). Reviews summarizing voice-based assessment across psychiatric disorders also exist.
The current trends can be summarized as:
- Statistically significant associations between voice features and emotion, stress, and depression are reported repeatedly.
- However, effect sizes and reproducibility vary widely by dataset, language, and task.
- There is a performance gap between lab conditions and real, in-the-wild environments.
04. Limits you must know
The most important thing when discussing voice biomarkers is the limits.
- Many confounders. Colds, sleep deprivation, caffeine, ambient noise, microphone quality, sex, age, language, and personal speaking habits all affect the voice.
- A single sample cannot confirm. Judging burnout or depression from one voice sample is risky.
- It is not a medical diagnosis. Association findings do not equal a diagnostic tool.
- Bias risk. Training on data skewed toward certain populations can make performance uneven.
So a trustworthy approach looks not at one absolute value, but at within-person change across multiple time points.
05. SYMPLE’s approach: change relative to a personal baseline
What SYMPLE focuses on is not “is this person’s F0 higher than others’,” but “how is it changing compared to this person’s usual?”
People naturally speak fast or slow, high or low. Applying one standard to everyone risks mistaking individual differences for state changes. So we build a personal baseline and observe longitudinal change. This principle is also explained as the core of early detection in Early detection of employee burnout.
Voice is one signal used in that process, interpreted together with other signals like language and check-ins. And to protect privacy, organizations receive only group-level signals that cannot identify individuals.
06. Summary
- Voice biomarker = an acoustic feature extracted from speech that may relate to one’s state.
- Key features: F0, jitter/shimmer, speaking rate/pauses, energy, spectral features.
- Phonation is affected by the autonomic nervous system, muscle tension, and cognitive load, so it changes with state.
- Don’t overtrust a single sample or treat it as a medical diagnosis. Change relative to a personal baseline is the key.
Learn what SYMPLE builds with this technology in About SYMPLE.
References
- Cummins N, Scherer S, Krajewski J, et al. A review of depression and suicide risk assessment using speech analysis. Speech Communication. 2015;71:10–49.
- Low DM, Bentley KH, Ghosh SS. Automated assessment of psychiatric disorders using speech: A systematic review. Laryngoscope Investigative Otolaryngology. 2020;5(1):96–116.
- Mundt JC, Vogel AP, Feltner DE, Lenderking WR. Vocal acoustic biomarkers of depression severity and treatment response. Biological Psychiatry. 2012;72(7):580–587.
- Teixeira JP, Oliveira C, Lopes C. Vocal acoustic analysis – Jitter, Shimmer and HNR parameters. Procedia Technology. 2013;9:1112–1122.
- Insel TR. Digital phenotyping: technology for a new science of behavior. JAMA. 2017;318(13):1215–1216.
Frequently asked questions
- What is a voice biomarker?
- It's an acoustic feature extractable from speech that may relate to a person's physical or psychological state. Features studied include pitch (F0), speaking rate, pauses, voice-quality measures like jitter and shimmer, and spectral features. Just as blood tests or imaging serve as traditional biomarkers, the voice can be seen as a signal reflecting one's state.
- What do jitter and shimmer mean?
- Jitter is the tiny cycle-to-cycle variation in vocal fold vibration period; shimmer is the tiny variation in its amplitude. Both relate to voice-quality stability and can change with tension, fatigue, and emotional state, which is why they are frequently used features in voice research.
- Can voice detect depression or stress?
- Many studies have analyzed the relationship between voice and depression, stress, and emotional states, with systematic reviews especially in depression and suicide-risk assessment. However, research findings are not the same as diagnosis. Because voice is strongly affected by individual and environmental factors, it's best used as a supporting signal for observing long-term change rather than confirming a state from one sample.
- Why do voice biomarkers change with a person's state?
- Phonation is produced through the coordination of breathing and the vocal and articulatory organs, and this process is affected by autonomic nervous system activity, muscle tension, and cognitive load. For example, tension can alter muscle tone and breathing patterns, subtly changing pitch and voice quality.
- Is voice data safe from a privacy standpoint?
- Voice is sensitive data, so design choices like data minimization, purpose limitation, on-device or de-identified processing, and access control matter. SYMPLE does not provide an individual's raw data to the organization; it aims to provide only group-level signals that cannot identify a specific person.