At a daylong Digital Health Advisory Committee meeting, FDA, clinicians, payers, and developers wrestled with how generative AI mental health devices can expand access without compromising safety. Below is a recap of the daylong discussion. (Watch the replay here.)
AI Disclosure: This report was written in collaboration with generative AI, with specific direction, resources, editing, and review provided by me.
A high-stakes test case for AI in medicine
On November 6, the FDA’s Digital Health Advisory Committee (DHAC) devoted eight hours to a single topic: generative AI-enabled digital mental health medical devices.
The brief was narrow, but the implications are broad. DHAC examined “AI therapists” and related tools that diagnose, treat, or otherwise meaningfully guide care for psychiatric conditions. Unlike wellness apps, these products look and behave like clinical tools. Some are fully automated, some supervised by clinicians. Most could be deployed at scale to solve a deep and concerning health problem.
“Across the country now, we’re experiencing a mental health crisis that is affecting people of all ages,” said Michelle Tarver, MD, PhD, director of the Center for Devices and Radiological Health (CDRH). Yet to date, no generative-AI mental-health device has been authorized for use in the United States. For medtech leaders, the DHAC discussion is an early roadmap for what FDA will likely expect from companies building AI-driven mental-health products: evidence of clinical benefit, robust trial design, lifecycle performance monitoring, and clear human-oversight models — all tuned to the unique risks of AI that talks back.
Opening Remarks: Responsible innovation under crisis conditions
Tarver used her opening remarks to anchor the discussion in both urgency and restraint. She underscored CDRH’s mission to “protect and promote public health by assuring that patients have safe, effective, and high-quality medical devices,” while simultaneously supporting innovation via predictable, efficient regulatory pathways.
Tarver linked generative AI directly to unmet mental health needs. Citing CDC data, she noted the large and rising share of Americans diagnosed with mental, emotional, or behavioral conditions, and the critical need for access to diagnosis and treatment. Generative AI-enabled digital tools, she argued, could help reach people in rural or underserved communities who lack access to clinicians — if they can be kept safe as they evolve in real-world use.
She framed the day’s purpose succinctly: to better understand “the benefits, the risks to health, and the risk mitigations that might be considered” for these devices, and to illuminate “a path forward for these very important technologies.”
For executives, the signal is clear: FDA is not trying to slow this category down; it’s trying to define how to say yes responsibly.
FDA Perspective: How CDRH is thinking about generative AI mental-health devices
Gioia Guerrieri, DO, DFAPA, FASAM, digital behavioral health lead in the Digital Health Center of Excellence, laid out the strategic and policy scaffolding that will underpin FDA decisions for AI mental-health products.
“Americans are facing a mental health crisis,” said Guerrieri. More patients of all ages are being diagnosed with psychiatric conditions while access to providers shrinks; at the same time, advances in digital mental-health technology — especially generative AI chatbots — are accelerating.
Guerrieri highlighted three pillars medtech teams should assume from the outset:
- Risk-based regulation and function-level analysis
Not all mental health software is regulated as a medical device. Guerrieri reiterated FDA’s long-standing approach: some software functions are not devices at all (e.g., stress-reduction tips), some are devices but low-risk enough for enforcement discretion (e.g., a “skill of the day” app for diagnosed anxiety), and others are clearly devices requiring premarket review (e.g., an app providing computerized behavioral therapy to treat a psychiatric disorder).
For generative AI, the same logic applies: claims, intended use, and level of clinical reliance will determine scrutiny. - Total product lifecycle (TPLC) and performance monitoring
Guerrieri pointed to recent draft guidance on lifecycle management and marketing submissions for AI-enabled devices. While still draft, it signals FDA’s expectation that companies design for design-development-deployment-maintenance as an integrated continuum — not a one-time submission.
Because AI systems are “sensitive to changes in input data, which may lead to changes in performance that create a risk to patients over time,” FDA is explicitly interested in post-market performance monitoring plans as part of the initial submission, especially for generative models that may adapt or be updated over time. - Real-world performance as a core evidentiary theme
Guerrieri also flagged FDA’s open request for public comment on best practices to measure and evaluate real-world performance of AI-enabled devices — a direct outgrowth of the 2024 DHAC meeting. For developers, this is a preview: real-world evidence strategies will not be optional for meaningful AI products.
FDA Perspective: Diagnostics — data, bias, and drift
Jay Gupta, assistant director of the Neurodiagnostic Devices Team in the Division of Neurosurgical, Neurointerventional, and Neurodiagnostic Devices (DHT5A), walked through the diagnostic side of digital mental health.
Gupta defined “digital mental health diagnostics” as devices that use software to monitor, assess, or support the diagnosis or prognosis of mental-health conditions — often by analyzing physiologic signals, behavior, or performance on tasks. Examples include assessment aids for ADHD and autism spectrum disorder, some of which have already gone through de novo or 510(k) pathways.
Key regulatory questions Gupta flagged — all highly relevant to generative-AI diagnostic tools — included:
- Training-data representativeness: Does the model see enough diversity across age, sex, race, comorbidity, and care setting to be reliable in real-world populations?
- Bias mitigation and clinical impact: If performance differs across subgroups, how will the sponsor detect, quantify, and mitigate that — and what does that mean for labeling and use?
- Performance drift: How will sponsors monitor whether diagnostic accuracy degrades as practice patterns, populations, or upstream LLMs change?
In practical terms, Gupta’s remarks telegraphed that diagnostic accuracy alone will not be enough. Sponsors should be prepared to show who the device works for, how drift will be detected and managed, and what happens when it fails.
FDA Perspective: Therapeutics — classification, change control, and harm prevention
Pamela Scott, assistant director for neuromodulation psychiatry devices in CDRH’s OHT5, focused on therapeutic applications: generative-AI tools that actually deliver interventions — for example, structured CBT-style chat, coping skills, or relapse-prevention plans.
Scott’s remarks centered on three design-time decisions that will shape any regulatory discussion:
- Device classification: Depending on intended use, risk profile, and degree of autonomy, AI mental-health therapeutics could fall into different device classes. A low-risk adjunct to clinician-led therapy looks very different from a fully autonomous intervention for high-risk patients.
- Software change management and Predetermined Change Control Plans (PCCPs): With Congress now explicitly authorizing PCCPs for devices, Scott underscored that sponsors can propose a structured plan for future model updates — but must clearly specify the scope, methods, and monitoring that will keep those updates safe.
- Unintended harm and crisis handling: Because these systems interact directly with vulnerable users, their failure modes include not just incorrect advice but escalation of risk (e.g., reinforcing suicidal ideation). Scott’s emphasis was on evidence of efficacy plus evidence that safeguards actually work in realistic usage.
Takeaway for medtech leaders: think about your PCCP, safety red-teaming, and crisis-response logic as first-class design artifacts, not afterthoughts.
Anthony Becker: AI support in extreme settings, and the hallucination problem
Lieutenant Commander Anthony Becker, a U.S. Navy psychiatrist, offered a front-line perspective on generative AI in high-stakes, resource-constrained environments.
He described settings where service members or sailors may have limited or no immediate access to mental health professionals, but do have secure digital channels. In those environments, AI-based tools could provide screening, psychoeducation, and basic coping support between human contacts — potentially reducing stigma and encouraging earlier help-seeking.
At the same time, Becker underscored the risk of hallucinations: plausible but false or dangerous outputs from generative models. In a mental-health context, a hallucinated therapeutic rationale, misinterpretation of risk, or invalid reassurance could directly harm patients.
His core recommendation was that generative AI systems be validated and deployed as adjuncts to clinicians, not replacements — especially for high-severity conditions — and that their crisis-handling behavior be tested under adversarial conditions before deployment.
Vaile Wright: Ethics, human factors, and the empathy gap
Vaile Wright, PhD, senior director in the Office of Health Care Innovation at the American Psychological Association, focused on ethical and human-factors dimensions of generative AI in mental health.
Wright acknowledged the appeal: AI chatbots can be available 24/7, speak in plain language, and may feel less stigmatizing than presenting to a clinic. They could extend reach to people who would never otherwise engage.
But she highlighted a few dangers:
- Misinformation or non-evidence-based advice masquerading as therapy
- Data-privacy and confidentiality risks in highly sensitive conversations
- Erosion of therapeutic empathy and alliance, particularly if users mistake a non-sentient system for a human relationship
Wright argued for standards that ensure evidence-based therapeutic content, transparent communication that the agent is an AI, and design that respects cultural and contextual differences. Without that, she warned, generative tools could amplify existing inequities rather than reduce them.
Open Public Hearing: Optimism, fear, and a plea for guardrails
During the two-hour Open Public Hearing session, clinicians, developers, and patient advocates echoed a common theme: demand for innovation is already here, but safety guardrails are not.
Participants described patients using general-purpose LLMs for mental-health support today, without any clinical oversight, and urged FDA to avoid a regulatory approach that only captures “good actors” while leaving the broader universe of unregulated tools untouched.
At the same time, multiple speakers raised concerns about biased outputs, lack of transparency about training data, and weak or absent crisis-response logic in many consumer-facing apps. The message to FDA was less “don’t regulate” and more “regulate in a way that rewards evidence-based tools and discourages low-quality imitators.”
Nicholas Jacobson: Building Therabot and why data quality matters
Nicholas Jacobson, PhD, associate professor at Dartmouth Geisel School of Medicine, offered an academic case study in developing Therabot, a generative-AI psychotherapy system.
“Mental health is exceedingly common, occurring in about one in three people annually in any given year,” said Jacobson. Yet most with a mental-health disorder do not receive minimally adequate treatment — a gap his team set out to address with digital therapeutics and, ultimately, generative AI.
Jacobson described three iterative phases:
- Naïve training on peer-to-peer data
Early models trained on peer-support conversations generated responses that reinforced pathological sentiments instead of delivering therapy — “very much not safe to actually deploy,” as Jacobson put it. - Training on psychotherapy transcripts
A second attempt using psychotherapy training videos emulated “the bad habits of psychotherapists,” producing stereotyped, unhelpful responses. The team concluded that widely available training corpora “contained systematic problems that are not really suitable for clinical use.” - Handcrafted, clinician-authored dialogues
The team then built a bespoke training set: thousands of dialogues written and peer-reviewed by clinicians, grounded in evidence-based CBT and related therapies. That intensive effort — more than 100,000 hours of expert fine-tuning — yielded models with high rates of clinically appropriate, personalized responses and strong symptom reduction in randomized controlled trials.
Jacobson urged FDA to focus not just on end-product performance but on development processes and data curation practices, warning that “product-centric approval” could lock in obsolete models in a fast-moving field. He also cautioned that heavy regulation of purpose-built clinical tools could create a “streetlight effect,” pushing users toward unregulated general-purpose LLMs instead.
For medtech leaders, his message was this: safety and efficacy in this space are fundamentally a data-quality and process-quality problem, not just a model-architecture problem.
Patricia Areán: What clinical-trial rigor looks like for AI chatbots
Patricia Areán, PhD, founder of the University of Washington CREATIV Lab and former division director at the National Institute of Mental Health’s Division of Services and Interventions Research, walked through what credible trials must look like for therapeutic chatbots.
“Before we get to the efficacy stage, intervention design is critical,” said Areán. “Don’t just develop the AI because you can, think clearly about the problem you are trying to solve.”
Her key points:
- User-centered design and context of use: Sponsors should define where and how the chatbot will be used (primary care, self-help, stepped care, etc.), and design interfaces accordingly.
- Appropriate comparators: If the intent is to substitute for therapy, non-inferiority trials against standard therapy may be appropriate; if the intent is to improve on waitlists or usual care, superiority designs are more relevant.
- Sample verification and supervision: Fully remote recruitment via social platforms raises serious concerns about participant authenticity. Areán recommended human-in-the-loop oversight in trials to ensure safety and verify that the chatbot’s responses remain within protocol.
- Follow-up duration: Short follow-ups may miss relapse or delayed adverse effects. She encouraged at least multi-month follow-up in early studies.
She also introduced a pragmatic idea: hybrid efficacy-effectiveness designs with phased funding or deployment, allowing tools that meet early safety and target-engagement benchmarks to scale while more data accumulates.
Andrew Trister: Lessons from other AI fields
Andrew Trister, MD, PhD, chief medical and scientific officer at Verily, zoomed out to compare mental-health AI to more mature fields like imaging and digital biomarkers.
He noted that the “Overton window” for generative AI use in everyday life has shifted rapidly since 2022, with people using open models for everything from essays to health questions — often outside any regulatory perimeter. In other domains, Trister argued, progress came from continuous real-world evidence programs, robust post-market surveillance, and clear notions of acceptable failure modes.
For mental health, he advocated:
- Continuous learning systems with guardrails, where performance monitoring feeds back into model updates via PCCPs.
- Interoperability and data-standard alignment, so outputs can integrate into EHRs and clinician workflows without breaking safety chains.
- Human-in-the-loop architectures, especially where stakes are high or ambiguity is significant.
He warned against premature large-scale deployment of unvalidated generative models in vulnerable populations, arguing instead for adaptive regulation proportional to both risk and evidentiary maturity — a point likely to resonate with regulatory and product-strategy leaders alike.
Brooke Trainum: Liability, consent, and scope of practice
Brooke Trainum, JD, senior director of practice policy at the American Psychiatric Association, focused on legal and ethical obligations when AI mediates care.
Trainum raised questions that any medtech legal or compliance team will recognize:
- Who is liable when an AI chatbot gives harmful advice — the manufacturer, the prescribing clinician, the institution?
- How should informed consent describe AI involvement so that patients truly understand what they are interacting with?
- When does reliance on AI cross into unlicensed practice or extend beyond a clinician’s scope?
She emphasized preserving clinician accountability and transparent communication about AI use, and reiterated APA’s stance: supportive of innovation, but only with human supervision, traceable outputs, and alignment with existing ethical frameworks around confidentiality, professional boundaries, and duty of care.
Bradley Karlin: Why payers will demand strong evidence
Bradley Karlin, PhD, ABPP, MSCP, MBA, program manager in the Resilient Systems Office at ARPA-H, brought the payer and value lens.
Karlin argued that enthusiasm alone will not unlock coverage. Payers will look for:
- Robust clinical-effectiveness data in relevant populations
- Durable impact on symptoms and functioning, not just short-term engagement metrics
- Cost-effectiveness, demonstrated through reductions in higher-cost services (e.g., ED visits, hospitalizations) or improved productivity
Without that level of evidence, he cautioned, reimbursement pathways for generative-AI mental-health devices will remain limited, even if patients and providers are eager to use them.
Caroline Pearson: Independent assessments show uneven performance
Caroline Pearson, executive director of the Peterson Health Technology Institute, summarized PHTI’s evaluations of virtual solutions for depression, anxiety, and opioid-use disorder.
Her findings will matter to anyone building or buying digital therapeutics: relatively few products demonstrated strong, durable clinical benefit when evaluated independently, and many relied on limited or low-quality evidence bases. Claims often outpaced supporting data.
Pearson urged FDA and payers to coordinate around standards for:
- Real-world validation methodologies
- Transparent labeling about what is and is not supported by evidence
- Substantiation requirements for marketing claims, particularly for AI-driven tools that may evolve over time
Her argument was straightforward: aligning regulatory and reimbursement expectations will reward robust products and drive out low-value offerings, improving trust in the entire category.
Committee Discussion – How to define benefit, manage risk, and monitor over time
In its first question group, the committee dug into the core design principles for generative-AI mental-health devices.
Themes included:
- Data-quality dependence: Panelists repeatedly returned to Jacobson’s point that bad training data yields harmful behavior. Sponsors should be ready to show not just aggregate performance metrics but also how training corpora were curated and audited.
- Bias and subgroup performance: Given historic inequities in mental-health diagnosis and treatment, several members stressed the need for pre-specified analyses by demographic and clinical subgroups, and for sponsors to address performance gaps explicitly.
- Dynamic monitoring and updating: Members supported the idea that for generative systems, post-market monitoring is part of the core safety case, not an afterthought. That includes monitoring for new kinds of failure modes as usage patterns evolve.
There was broad agreement on transparent labeling, explicit human-oversight expectations in indications for use, and staged rollouts that expand to broader or higher-risk populations only after early real-world performance is confirmed.
Committee Discussion – Over-the-counter AI for depression?
A later discussion examined a hypothetical: an over-the-counter generative-AI device, without clinician involvement, for people with major depressive disorder.
The mood was cautious to skeptical. One member noted that the proposal “would make me very nervous,” particularly for anything beyond mild depression. Another, Omer Liran, MD, said, “With today’s technology, I’m very uncomfortable with the idea that an AI chatbot will be used in the absence of a human provider…”
When FDA staff pressed on what evidence would be needed, panelists pointed to:
- Prior demonstration of safety and effectiveness under clinician supervision
- Large-scale real-world data showing favorable safety profiles, especially around suicidality and crisis events
- Clear mechanisms to verify diagnosis and ensure the tool is used by appropriate patients, even if OTC
Ray Dorsey, MD, suggested that any move to OTC would require “years” of safety data across thousands of individuals, given the scale of potential use.
Committee Discussion – Multi-condition, autonomous devices
The committee also considered whether a single generative-AI device could autonomously diagnose and treat multiple mental-health conditions related to “sadness,” including users without any formal diagnosis.
Members highlighted several concerns:
- Comorbidity and diagnostic overlap make it difficult to validate performance for each condition independently.
- A model tuned for one condition could behave unpredictably in another, especially without clear guardrails.
- Users without a diagnosable disorder could still be harmed by inappropriate or pathologizing responses.
As a result, panelists leaned toward risk-tiered regulation (with more autonomy allowed at lower risk), requirements to demonstrate non-harm in off-label populations, and rigorous real-world performance monitoring to detect unexpected patterns of use or harm.
What this means for medtech leaders
By the time DHAC Chair Ami Bhatt, MD, closed the meeting, one theme was unmistakable: FDA is open to generative-AI mental-health devices, but only on the back of serious, discipline-specific rigor.
For senior leaders building or evaluating these technologies, several implications stand out:
- Design for regulation from day one: Your training data strategy, PCCP, post-market monitoring plan, and crisis-response logic are central to your value proposition, not appendices.
- Assume evidence expectations comparable to drugs or high-impact devices, including randomized trials, subgroup analyses, and multi-month (or longer) follow-up.
- Expect human-in-the-loop models to be favored, at least initially, with clear role definitions and accountability.
- Plan for payer scrutiny: Clinical effectiveness and cost-effectiveness will ultimately determine adoption, not just regulatory clearance.
The DHAC meeting gave attendees a clear take on FDA’s thinking on AI-enabled mental health software. With it, medtech leaders can better understand the design constraints and regulatory parameters as they build the next generation of AI-enabled mental health devices.

