HOW-TO GUIDES 3 guides
Frequently Asked Questions
12 questions-
They may provide short-term reductions in anxiety and depression symptoms, but the durability of benefit is uncertain. A meta-analysis of 18 randomized controlled trials involving 3,477 participants found improvement in depression symptoms (g=-0.26, 95% CI=-0.34 to -0.17) and anxiety symptoms (g=-0.19, 95% CI=-0.29 to -0.09), with the most robust benefits after 8 weeks of treatment. However, no substantial effects persisted at 3-month follow-up for either condition.
The article also notes that randomized trials of generative AI chatbots have usually compared them with waitlist, information-only controls, or app-based psychoeducation rather than expert clinician-delivered psychotherapy, which limits how far the findings can be interpreted clinically.
-
Patients most commonly use generative AI chatbots for emotional support, help with coping, and mental health advice. The article states that users value these tools because they provide easy access to a “nonjudgmental ear,” an immediate sense of therapeutic alliance, and help with reframing negative thoughts.
In a recent cross-sectional study of 5.4 million individuals, 13.1% of US youths used generative AI for mental health advice, rising to 22.2% among those aged 18 years or older. Among these users, 65.5% engaged with generative AI at least monthly and 92.7% reported finding the advice helpful, particularly because of its immediacy and perceived privacy.
-
Many users do perceive generative AI chatbot responses as highly empathic. In one systematic review of 15 studies, about 73% of users judged large language model responses to be as perceptive as or more empathic than those of a human practitioner in head-to-head comparisons.
The article cautions, however, that this likely reflects the chatbot's ability to simulate empathy through consistent validation and pattern recognition rather than authentic emotional understanding or lived human experience.
-
No. The article advises that generative AI tools should be viewed as adjuncts to care rather than substitutes for professional mental health treatment, especially when a patient is at significant risk of harm.
Although some trials show short-term symptom improvement, the evidence base is limited by comparisons with nonexpert interventions, uncertain long-term benefit, and concerns about bias, safety, and crisis management. The authors specifically recommend avoiding generative AI as a substitute for care in situations such as active suicidal thoughts, psychosis, or severe substance use, where clinician assessment and crisis management are required.
-
The main risks are inaccurate or inappropriate advice, poor handling of complex psychiatric presentations, bias, privacy problems, and overreliance. The article explains that chatbot output can vary with how a prompt is phrased, may become illogical in some contexts, and may fail to incorporate nuanced clinical information needed for personalized psychiatric care.
It also notes that chatbots may mirror rather than challenge distorted thoughts, which can worsen rumination or distorted belief content, and that realistic “therapy-like” interactions may foster dependency or social withdrawal. In addition, many direct-to-consumer tools are not covered by HIPAA, so entering sensitive personal health information may jeopardize privacy and confidentiality.
-
Current evidence suggests they are not reliable enough to be used alone for suicide triage or crisis counseling. In one study of 29 generative AI-powered chatbot agents responding to simulated suicidal risk scenarios, none met the initial criteria for an adequate response; 51.7% met relaxed criteria for only a marginal response, and 48.3% were judged inadequate.
Common problems included failure to provide emergency contact information and lack of contextual understanding. The article also reports that large language model chatbots aligned with expert clinicians on very low- or very high-risk suicide queries but were inconsistent for intermediate-risk cases.
-
Natural language processing may help identify dynamic suicide risk signals that structured assessments miss between visits, but it is best understood as an adjunct rather than a replacement for validated suicide assessment tools. The article explains that NLP can analyze free-text clinical narratives and other longitudinal real-world data to model changing risk trajectories and detect contextual or linguistic markers of distress.
Early studies suggest NLP-augmented models can outperform clinician checklists and traditional scales in short-term risk stratification. In a cohort of more than 120,000 adult patient encounters, suicide risk detection was most effective when face-to-face Columbia Suicide Severity Rating Scale screening was combined with real-time electronic health record-based machine learning prediction, using the complementary strengths of clinician assessment and data-driven modeling.
-
It may help as decision support, but the article says the evidence for real-world clinical utility, safety, and readiness remains limited. Potential clinician-facing uses include triage, remote diagnosis and monitoring, personalized treatment planning, and predictive analytics.
In one exploratory study, ChatGPT-4 generated 1,176 differential-diagnosis lists from 392 case descriptions, and physician evaluators concurred with 966 of 1,176 lists (82.1%), with a Cohen kappa of 0.63 (95% CI=0.56-0.69), which the article interprets as fair to good agreement. The authors emphasize that this suggests possible value for decision support, not clinical readiness.
-
Many direct-to-consumer generative AI tools operate outside HIPAA and outside traditional clinical oversight, which raises concerns about privacy, confidentiality, and informed consent. The article notes that data collection and sharing practices are often opaque, and user inputs may be stored, shared, or used for model retraining without meaningful clinical supervision.
For that reason, the article advises clinicians to discuss where and how a tool stores data, whether data are encrypted, whether personal data may be used to train models, which legal jurisdictions govern data processing, and whether users can delete their data.
-
Primary care providers should counsel patients to use generative AI tools as adjuncts for self-management, symptom tracking, or skills practice, not as replacements for professional care. The article recommends first clarifying what the patient wants the tool to do, such as CBT support, motivational interviewing, monitoring, emotional support, or advice, and then discussing whether there is scientific evidence supporting that function.
Clinicians should also review limitations directly, including lack of nuance, factual errors, inability to manage crises, privacy risks, algorithmic bias, and the possibility of dependence. The article further recommends documenting these discussions in the medical record and, when possible, steering patients toward tools that meet local standards or have been independently evaluated without industry bias.
-
The article says safe use requires human oversight, clear crisis escalation pathways, documentation, auditing, and defined accountability across clinicians, health systems, developers, and regulators. Clinicians remain responsible for patient care decisions even when AI informs care, while health systems are responsible for implementation and monitoring, developers for transparency and validation, and regulators for broader standards and enforcement.
The authors also stress that market availability does not guarantee safety or clinical reliability, including for FDA-cleared products, because postmarket surveillance for model drift, accuracy, and equitable performance remains limited. Clinicians using these tools should understand their capabilities, limitations, and failure points before integrating them into care.
-
In the case vignette, the patient initially reported improvement but later developed problematic reliance on the chatbot. At 6 months, her symptoms had improved to PHQ-9=6 and GAD-7=8, which she attributed to help reframing negative thoughts with chatbot support.
By 9 months, however, she felt the chatbot was mirroring and amplifying her ruminative loops and that she had become more isolated from peers. Her scores worsened to PHQ-9=14 and GAD-7=18, after which she and her primary care provider reviewed both the benefits and limitations of the chatbot, she agreed to increase her citalopram dose, and she began seriously pursuing a CBT-proficient psychotherapist.