Key Takeaways
Extended Takeaways
- Short-term symptom gains with mental health chatbots are modest to moderate and appear time-limited: across 18 RCTs involving 3,477 participants, depression improved by g=−0.26, 95% CI =−0.34, −0.17 and anxiety by g=−0.19, 95% CI =−0.29, −0.09, with the most robust benefits after 8 weeks but no substantial effects at 3-month follow-up.
- If a patient reports benefit from a chatbot, ask what role it is serving and monitor for dependency or social withdrawal; the case vignette illustrates improvement at 6 months followed by worsening by 9 months as the chatbot increasingly mirrored rumination rather than disrupting it.
- Do not rely on generative AI for suicide triage or crisis counseling without human backup. In one evaluation of 29 chatbot agents, none met initial criteria for an adequate response to simulated suicidal risk scenarios, 51.7% were only marginal, and 48.3% were inadequate.
- For clinician-facing decision support, current evidence supports cautious augmentation rather than autonomous use. ChatGPT-4 concurred with physician evaluations in 966 out of 1,176 differential-diagnosis lists (82.1%), with Cohen κ coefficient of 0.63 (95% CI=0.56–0.69), suggesting moderate agreement but not clinical readiness.
- When discussing privacy, be explicit that many direct-to-consumer tools are outside HIPAA and may use patient inputs for storage, sharing, or model retraining; this makes informed consent, data governance review, and avoiding entry of sensitive personal health information part of routine counseling.
- Bias is a practical safety issue, not just a theoretical one. Across psychotherapy-oriented large language models, stigmatizing responses occurred in 38%–75% of presentations, and in a 2025 study therapy chatbots responded inappropriately or dangerously in at least 20% of cases, with the highest rate occurring when delusions were disclosed (55% for specific chatbots).