HOW-TO GUIDES 1 guide
Frequently Asked Questions
10 questions-
Yes. In this community sample of pregnant participants, speech-based models built from full clinical interviews showed good performance for both disorders, with a best F1-score of 77.48% (±5.59) for major depressive disorder and 79.57% (±14.81) for generalized anxiety disorder. The diagnoses were based on SCID-5 interviews rather than symptom screening scales.
-
Yes. Full-interview models outperformed verified patient-only speech segments for both conditions, especially for major depressive disorder. For major depressive disorder, the best F1-score was 77.48% (±5.59) with full interviews versus 45.05% (±24.84) with verified segments; for generalized anxiety disorder, the best F1-score was 79.57% (±14.81) versus 64.89% (±16.68).
The authors noted that full interviews were much longer on average than verified segments, which may have contributed to better performance. Verified segments averaged 1.47 ± 1.17 minutes, while full interviews averaged 40.74 ± 16.81 minutes.
-
Yes. Models trained only on pregnancy and sociodemographic characteristics performed poorly compared with speech-based models. The best F1-score for pregnancy/sociodemographic-only models was 39.33% (±5.36) for major depressive disorder and 33.04% (±15.50) for generalized anxiety disorder, versus 77.48% (±5.59) and 79.57% (±14.81), respectively, for the best full-interview speech models.
-
No. Adding pregnancy and sociodemographic characteristics to the best speech pipelines did not improve model performance for generalized anxiety disorder and slightly reduced performance for major depressive disorder. The best combined-model F1-score was 79.57% (±14.81) for generalized anxiety disorder, unchanged from speech alone, and 75.00% (±17.09) for major depressive disorder, compared with 77.48% (±5.59) for speech alone.
Unpaired t tests showed no statistical difference between speech-only and combined models for major depressive disorder (P≈.82) or generalized anxiety disorder (P=1.0).
-
For major depressive disorder, pause duration, F0, and shimmer were among the most influential features. For generalized anxiety disorder, harmonics-to-noise ratio, jitter, and F0 were among the most influential features.
The authors noted that these patterns were similar to findings reported in the general population. They also explained that low F0 reflects reduced pitch variation and more monotonous speech, while jitter and shimmer reflect altered voice quality.
-
No. In this dataset, OpenSMILE-based models performed better than wav2vec2-based models. The best wav2vec2 F1-scores were 43.89% (±8.15) for major depressive disorder and 51.62% (±14.09) for generalized anxiety disorder, both below the best full-interview OpenSMILE models.
-
Major depressive disorder and generalized anxiety disorder were ascertained using the Structured Clinical Interview for DSM-5 (SCID-5). The control group included participants without any current or past known psychiatric disorders, which the authors selected to reduce classification bias given that speech features can overlap across psychiatric conditions.
-
The total sample included 146 pregnant participants from British Columbia, Canada, all aged 19 years or older. Within the full dataset, 19 participants had major depressive disorder, 28 had generalized anxiety disorder, and 102 were included as controls without current or past known psychiatric disorders; 3 participants had comorbid major depressive disorder and generalized anxiety disorder.
-
Not in this sample. The authors used error dependence plots and association rules and did not observe relevant error regions suggesting discrimination based on demographic or pregnancy characteristics. However, they cautioned that some subgroups were very small.
-
The main limitations were small MDD and GAD sample sizes, limited diversity, and the use of full-interview recordings that included both participant and interviewer speech. The authors also emphasized that the majority of participants (72%) identified their ethnic background as White, so the resulting models may be mainly applicable to that population.
They further noted that findings from the full-interview models should be interpreted cautiously because the recordings contained both voices, even though all interviews were conducted by the same interviewer. The study was a secondary data analysis, and the authors said primary prospective data collection is needed to confirm the findings.