Skip to content
  • clinical AI

What AI can and cannot do in psychotherapy

Gesell Team10 min read

“Will AI replace your therapist?” The two fashionable answers — total enthusiasm and total dismissal — fail for the same reason: they talk about “AI” as if it were one thing. In clinical practice today, two categories coexist that almost every conversation conflates: consumer chatbots that a client uses on their own, and clinician-side tools that help with documentation. The evidence says very different things about each. What follows is a capability map with the sources in hand: what AI does well today, what it does only if you verify it, and what remains — by evidence, by ethics, and by law — human work.

What AI does well today: drafting, structuring, summarizing

Start with the firm ground. When the WHO published its January 2024 guidance on large multi-modal models — the technology behind ChatGPT-style tools — it listed clerical work among its five broad health applications: tasks “such as documenting and summarizing patient visits within electronic health records.” Drafting a note from your session record, structuring it into a clinical format like SOAP, summarizing a long history: that is the use case where the evidence is accumulating.

And it is measured evidence now, not promise. At Stanford Health Care, a study in JAMIA followed 45 physicians across eight ambulatory specialties for three months and 17,428 encounters with an ambient AI scribe: median time per note fell 0.57 minutes, daily documentation time fell 6.89 minutes, and total daily EHR time fell 19.95 minutes versus baseline. The companion study measured the same users’ experience: significant reductions in task load and burnout, with 65% reporting more efficient documentation and 52% better note quality. A 2025 study in JAMA Network Open, spanning 263 clinicians across six US health systems, found the share experiencing burnout dropped from 51.9% to 38.8% after adopting an ambient scribe — with no control group, so read it as promising pre-post evidence, not proof.

The caveats come from the authors themselves, and they are worth copying: absolute time saved was “modest,” in the Stanford team’s words; the tool was used in just over half of encounters; and some physicians saw their documentation time go up, because writing time turned into editing time. AI does not eliminate documentation work: it shifts it toward review. And the caveat that matters most to you: this time-savings evidence comes from general ambulatory medicine. In psychotherapy, the studies are still pilots — and what they show is exactly the next section.

What only works with your verification

Language models hallucinate: they generate fluent, plausible-sounding content that is factually incorrect, unsupported by the source data, or entirely fabricated. This is not an anecdote. A blinded evaluation of a commercial ambient scribe — peer-reviewed, with the vendor (Suki AI) involved in the research, a fact worth knowing as you read it — found hallucinations in 31% of AI-generated notes, against 20% of physician-written ones. The same study shows why the risk is treacherous: blinded reviewers preferred the AI notes overall (47% versus 39%) and rated them more thorough. Capability and error live in the same output; a note can read flawlessly and contain something nobody said.

In mental health, the signal points the same way. The only psychiatric-interview documentation study we could verify — a proof-of-concept from Charité in Berlin, six simulated interviews processed through Whisper and GPT-4 — found human reports more accurate than AI reports (mean accuracy 0.94 versus 0.78), with the AI omitting far more clinical findings. The authors’ own conclusion: human supervision remains crucial. And errors persist even with mitigation: in a 2025 adversarial experiment that planted fake clinical details, mitigation prompting reduced hallucinations without bringing them anywhere near zero. That scenario is not routine note drafting, but the lesson travels: review is not optional.

The WHO names the behavioral risk this creates: “automation bias” — letting errors pass that you would otherwise have caught, because you trust the system. And the APA’s 2025 ethical guidance on AI — which states it does not represent official APA policy — draws the responsibility line: “AI should augment, not replace, human decision-making,” psychologists remain responsible for final decisions, and negligent reliance on AI without validation or oversight could create liability. Wherever you practice, that has a concrete legal echo: clinical notes are attributed and signed. In Mexico, for example, the clinical-record norm requires every note to carry the full name and signature of its author. The note a system drafted and you signed is, legally and ethically, your note. The operating rule fits in one line: AI drafts; you review, correct, and sign every note. Never without you.

What stays human

The therapeutic alliance. The evidence consistently associates alliance quality with treatment outcome: in the meta-analysis by Flückiger and colleagues (2018) — 295 studies, more than 30,000 patients — the correlation was r = .278 (equivalent to d = .579), and nearly identical for internet-delivered therapy (r = .275). It is an association, not proof of causation, but it holds across treatment approaches, measures, patient characteristics, and countries. Building that alliance is the work of the relationship, not of software — and clients seem to know it: in a pilot where GPT-4 generated psychological reports rated comparable in quality by expert reviewers, 50% of clients still preferred human providers and only 10% preferred the AI.

Informed consent and the limits of confidentiality. The APA Ethics Code assigns these duties to the psychologist personally: informing clients as early as feasible about the nature and course of therapy, fees, and the limits of confidentiality (Standard 10.01), and discussing those limits at the outset (4.02). They are non-delegable; no tool assumes them for you.

Diagnosis and final clinical decisions. The first principle of the WHO’s 2021 guidance is protecting human autonomy: humans should remain in control of health-care systems and medical decisions. The APA’s 2025 guidance says the same in clinical terms: the final decision belongs to the professional.

Risk assessment. A risk signal in session — a mention of ideation, a positive PHQ-9 item 9 — calls for your immediate clinical risk assessment, under your own protocol and training. No system output substitutes for it, postpones it, or documents it for you.

Accountability. The WHO’s responsibility principle sums it up: AI performs tasks, but ensuring the tools are used under appropriate conditions and by appropriately trained people is the responsibility of people. The signature at the bottom of the note is yours.

A consumer chatbot is not psychotherapy

In November 2025, the American Psychological Association published a health advisory on generative AI chatbots and wellness apps. Its first recommendation is blunt: “Do not rely on GenAI chatbots and wellness apps to deliver psychotherapy or psychological treatment” — at best they may be “a supportive adjunct, not substitute, to an ongoing therapeutic relationship.” The underlying diagnosis is about design, not malice: these systems were not created to deliver mental health care, and most lack scientific validation, adequate safety protocols, and regulatory approval, even though they are widely used for exactly that.

Two of the advisory’s cautions carry particular weight in the consulting room. On crises: “The ability of these tools to consistently and safely manage a user in crisis is limited and unpredictable,” and relying solely on an app during a mental health emergency can be dangerous. Your frame should get ahead of that: your client should know where to turn in an emergency — the crisis lines and emergency services where you practice, agreed on explicitly at intake — and any risk disclosure in session triggers your immediate clinical assessment, not a system’s. On vulnerable users: the advisory documents unsafe interactions that have already occurred — chatbots encouraging self-harm, eating-disorder behavior, or delusional thinking — and notes that many of these products are engineered to maximize engagement rather than any healthy outcome.

None of this is moral panic: the APA itself acknowledges preliminary research suggesting that some apps built specifically for mental health may offer supportive benefits in some contexts. And the most actionable recommendation is simple: the advisory strongly recommends that users tell their health providers which tools they use, so providers can spot guidance that is unhelpful or inconsistent with the treatment plan. Turn it into a routine question:

“Do you use any app or chatbot for your mood between sessions? It is not a gotcha — knowing helps me integrate it, or flag it if something it tells you runs against what we are working on.”

One honest note on jurisdiction: an equivalent official advisory may simply not exist where you practice. In Mexico, for example, there is no AI-specific health regulation we could verify on official sources — the general data-protection and clinical-record frameworks are what apply. Verify the rules where you practice.

How to evaluate a tool: questions, not promises

For clinician-side tools, the honest standard is not the vendor’s promises but your questions. In 2024, APA Services published a thirteen-step guide for evaluating AI tools and a companion checklist whose questions work in any country: does the company use your data to train its underlying model? How long is data retained, and where is it stored? Does it attest compliance with the data-privacy laws “in the jurisdiction in which you practice”? What clinical evidence backs the tool? The guide closes with a disclaimer worth imitating: the APA does not endorse any specific AI tools. Neither do we ask for your trust: put these questions to any vendor — our own tool included — and demand the answers in writing. The full interrogation — legal roles, retention, deletion, breach notification — is developed in our guide on AI and patient data safety, the companion piece to this article.

Transparency with your client is not optional

Running AI over session content — audio, transcript, or written record — is processing of sensitive health information, and the converging standard is express, documented consent, obtained before you start and never assumed by default; in Mexico, for example, the 2025 data-protection law requires it expressly and in writing for health data. The APA’s 2025 guidance says where to document it: in the written informed consent, specifying when, how, and what type of tool you use, with a genuine option to decline. And peer-reviewed mental-health-specific guidelines from 2025 state it as a principle: disclose to the client whenever AI is used in their treatment, including its capabilities and limitations. How to hold that conversation — and draft that document — is covered in our informed consent guide.

The division of labor, in one sentence

AI drafts, structures, and summarizes; you listen, decide, assess risk, and answer for every note. That division is exactly the one gesell.ai is built around: it drafts your notes in SOAP, DAP, BIRP, or GIRP format from your session records, generates treatment plans with SMART goals, and keeps each client’s chart structured — and at every step, you review, adjust, and own every output. The human column of the map we do not touch, because it cannot be done. Nor should it be.

References

Share this article

About the author

Gesell Team

Clinical and product content written by the gesell.ai team together with certified clinical psychologists.

← Back to the blog