Clinical Summary

Clinical Summary: Using Speech to Develop Multivariable Prediction Models for Major Depressive Disorder and Generalized Anxiety Disorder During Pregnancy

Depressive and anxiety disorders during pregnancy are common, consequential, and still hard to screen for reliably in routine care when time, training, and workflow are limited. This study tests whether speech collected during diagnostic interviews can identify major depressive disorder and generalized anxiety disorder better than pregnancy and sociodemographic factors alone.

Design Prediction Model Development and Evaluation Study
N the total sample size included 146 participants
Population pregnant, 19 years and older, and a resident of BC
Setting phase 3

Key Findings

  • Full interview speech models achieved the best overall performance, with F1-score of 77.48% (±5.59) for MDD and 79.57% (±14.81) for GAD.
  • Verified-segment patient-only models performed substantially worse for MDD, with best F1-score of 45.05% (±24.84), and lower for GAD, with F1-score of 64.89% (±16.68).
  • Models trained on solely pregnancy/sociodemographic features showed poor performance for both disorders, with best F1-score of 39.33% (±5.36) for MDD and 33.04% (±15.50) for GAD.
  • Adding pregnancy/sociodemographic characteristics to the best speech models did not improve classification: GAD remained 79.57%±14.81, while MDD decreased to 75.00%±17.09; unpaired t tests showed no statistical differences for MDD (P value=~.82) or GAD (P value=1.0).
  • In wav2vec2 models, the best F1-scores were 43.89% (±8.15) for MDD and 51.62% (±14.09) for GAD, both below the best OpenSMILE-based full interview models.
Clinical Bottom Line

In this dataset, speech-based models from full clinical interviews outperformed pregnancy and sociodemographic predictors for both major depressive disorder and generalized anxiety disorder during pregnancy. Adding conventional risk factors did not improve performance over speech alone.

Practice Implications

  • If speech-based screening is developed for prenatal mental health care, longer conversational samples may be more clinically useful than brief patient-only excerpts, given the gap between full interview and verified-segment F1-scores for MDD (77.48% [±5.59] vs 45.05% [±24.84]) and GAD (79.57% [±14.81] vs 64.89% [±16.68]).
  • Do not rely on pregnancy and sociodemographic characteristics alone to identify clinically diagnosed MDD or GAD during pregnancy, because these models performed poorly with F1-scores of 39.33% (±5.36) and 33.04% (±15.50).
  • When interpreting future speech-based tools, expect diagnostically relevant signal in pause duration, F0, and shimmer for MDD, and in HNR, jitter, and F0 for GAD, as these were among the most influential features.
  • Apply these findings cautiously outside similar populations: the sample included 146 participants, only 19 had MDD and 28 had GAD, and 72% identified their ethnic background as White.
Read full article
Physicians Postgraduate Press, Inc. (PPP) makes no warranties about the accuracy or completeness of any information published in The Journal of Clinical Psychiatry or other PPP materials, and disclaims liability for any use or non-use of that information. Clinicians should not rely solely on these materials and should exercise their own professional judgment when making patient care decisions on an individualized basis.