Skip to content
  • treatment planning
  • assessment

How to measure patient progress session by session

Gesell Team11 min read

“How is your client doing?” Most of us answer that question from clinical impression: how they look, what they tell us, how the session feels. Impression is indispensable, but measuring is something else — and the data on how many of us measure is uncomfortable. According to a review published in JAMA Psychiatry, fewer than 20% of behavioral health practitioners use measurement-based care — 17.9% of psychiatrists, 11.1% of psychologists, and 13.9% of masters-level practitioners — and as little as 5% use it on the schedule the evidence supports: every session. This guide covers what measuring seriously means, which instruments to use, how much change counts as real change, and where progress belongs in your record.

What measurement-based care actually is

Scott and Lewis define it as the systematic collection of data to monitor client progress and directly inform care decisions, and report that, used as a framework to guide practice, it produces superior outcomes compared with usual care. The definition has two halves, and the second is the one that usually goes missing: administering scales is not measurement-based care if the result never enters a clinical decision.

The JAMA Psychiatry review breaks the practice into four components: a routinely administered standardized measure, ideally before each clinical encounter; practitioner review of the result; patient review of the result; and collaborative reevaluation of the treatment plan in light of the data. Components three and four are what separate real practice from a filing cabinet of scores: the number gets discussed in session, and it can change the course of treatment.

The same review summarizes 21 randomized clinical trials comparing this practice with usual care: outcomes improved significantly, particularly for patients who were not responding to treatment, with medium-to-large effect sizes (0.22 to 0.70). Read twice where the benefit concentrates: in nonresponders — exactly the cases where clinical impression takes longest to accept that something is not working. The review also documents barriers at four levels — patient (confidentiality concerns), practitioner (the belief that measures add nothing beyond clinical judgment), organization, and system — and the nonresponder data answers the practitioner-level one directly.

Every session, not every quarter

Measuring at intake and at termination is not monitoring: it is a photo at the entrance and one at the exit, with no film in between. The Lewis review distinguishes measurement-based care from low-frequency outcome monitoring — done every few months or once a year — which is too spaced out to steer an individual treatment. In Scott and Lewis’s worked case example, the PHQ-9 and GAD-7 are administered before each session: by the time the session starts, the data is already on the table.

Is every-session measurement feasible outside a research trial? England’s public system has been demonstrating it at national scale for years. The official NHS Talking Therapies manual (v7, updated March 2024) asks patients to complete a depression measure and an anxiety measure at every session, because most report significant levels of both. The practical rationale is elegant: if a patient finishes earlier than expected, or a measure is missed once, there is always a last available score — and that habit is how the programme obtains outcome data on 98.5% of patients who have a course of treatment. Without it, missing scores come disproportionately from the people who improved less, and a service ends up grading itself better than it is. These are the operational standards of one national programme, not a universal norm — but they prove that session-by-session measurement works at scale.

The minimal kit: PHQ-9 and GAD-7

For depression and anxiety, the usual pair is also the best documented, and both scales are free on their official site, translations included. The PHQ-9 scores 0 to 27, and its original validation set thresholds of 5, 10, 15, and 20 for mild, moderate, moderately severe, and severe symptoms. The GAD-7 scores 0 to 21, and in its original validation cut points of 5, 10, and 15 are interpreted as mild, moderate, and severe anxiety. The official instruction manual explicitly endorses the double use this guide proposes: the PHQ-9 grades severity for initial treatment decisions and serves as an outcome tool to determine treatment response; even in the mild band (5-9), the suggested action is watchful waiting and repeating the instrument at follow-up.

The bands are what turn a number into a clinical reading: a PHQ-9 that moves from 17 to 9 did not just “drop eight points” — it crossed from moderately severe to mild. And for the far end of the process there is a citable threshold: CMS quality measure #370 in the United States defines depression remission as a PHQ-9 below 5, in patients who entered care scoring above 9. It is a US pay-for-quality definition, not a universal clinical criterion, but it gives the word “remission” a concrete reference point.

“Session 8. PHQ-9: 9 (session 1: 17; session 4: 15). GAD-7: 7 (previous: 11). Reports resuming exercise and sleeping better; avoidance of work meetings persists. Charts reviewed with the client: agrees to sustain behavioral activation and prioritize workplace exposure.”

How much change is real change

Scores wobble: one point up or down between weeks is expected noise, not news. To avoid over-celebrating or over-reacting you need change thresholds — and here the honest, complete answer is that several exist, they are estimated with different methods, and they do not coincide exactly.

The English operational standard is the most explicit. Table 9 of the NHS Talking Therapies manual sets clinical caseness at PHQ-9 ≥ 10 and GAD-7 ≥ 8, and reliable change at ≥ 6 points on the PHQ-9 and ≥ 4 on the GAD-7; a patient shows reliable improvement when one measure drops by that amount and neither shows a reliable increase, and the same apparatus defines reliable deterioration. A 2024 evaluation in BMJ Open restates those definitions and adds recovery: moving from caseness (PHQ-9 ≥ 10 or GAD-7 ≥ 8) to no caseness on both measures. In parallel, the GAD-7 literature commonly takes a 4-point change as the minimal clinically important difference, attributed to Toussaint and colleagues’ sensitivity-to-change analysis, as restated in the peer-reviewed literature — a commonly used threshold, not a universal constant.

The practical reading for your practice: movements of 1 or 2 points do not get interpreted on their own; a sustained drop in the 4-to-6-point range, depending on the scale, is a genuine signal of progress; and a reliable increase is a signal to review the plan this week, not at the quarterly review. That is the point of measuring often: deterioration shows up on the chart before it shows up in the narrative.

Symptoms are not the whole story

The English system does not measure symptoms alone: its main disability measure is the WSAS (Work and Social Adjustment Scale), which assesses how much the mental health problem interferes with work, home management, leisure, social life, and family. The manual warns that disability often decreases as symptoms improve — but not always — and explicitly contemplates continuing treatment when PHQ-9 and GAD-7 scores are already low but interference with everyday functioning remains significant. A PHQ-9 of 6 in a client who has not returned to work or to their social life is not a closed case.

Two more options, with WHO backing. WHODAS 2.0 is a generic health-and-disability instrument usable across all diseases — mental disorders included — and independent of diagnosis; it covers six domains (cognition, mobility, self-care, getting along, life activities, and participation) and comes in 36-item and 12-item versions, with the short form explaining 81% of the variance of the full one. And if you prefer a positively framed measure, the WHO-5 Well-Being Index consists of five positively phrased items producing a 0-100 score (the raw 0-25 is multiplied by 4): a 10-point change is considered clinically relevant, a score of 50 or below is the recommended cue to screen for depression, and it exists in more than 30 languages.

Progress is also measured against the plan

Scales tell you whether symptoms are moving; your treatment plan tells you whether the client is getting closer to what they came for. When goals are written in verifiable form — “reduce the PHQ-9 from 16 to below 10 in 12 weeks,” “return to full work attendance in 6 weeks” — every session can answer in one line how far there is to go, and the plan review stops being a ritual and becomes a comparison between what was agreed and what was measured.

Reviewing the data with the client — components three and four — has its own framework: shared decision making, which Slade defines as the process in which clinicians and patients work together to select treatments based on clinical evidence and the patient’s informed preferences. The evidence on its effect on clinical outcomes remains inconclusive, so present it as what it is: good practice. In the room it translates into simple gestures: show the chart, ask “does this match how you have been feeling?”, and decide together which goal comes next.

Where progress lives in the record

Measured progress that never reaches the record might as well not exist. The APA’s Record Keeping Guidelines (US guidance) name client progress as one of the things records exist to document, alongside treatment plans and services provided — and appropriate records protect both client and psychologist if anything ever reaches a legal or ethical proceeding. Some jurisdictions go further and make a per-session progress note a normative requirement: in Mexico, for example, the clinical-record norm requires an evolution note at every outpatient attention, covering clinical evolution, relevant results, diagnoses, prognosis, and treatment — a score with its previous comparison is exactly that. Verify what applies where you practice.

In format terms, the score goes in the O of your SOAP note — a measurable observation, next to the mental status exam — and the client’s response to the work lives in the R of a GIRP note. The format matters less than the habit: the number belongs in the signed note of the session, with its date and its comparison, not on a loose sheet no one will find again.

What a score does not do

  1. It does not diagnose. The PHQ-9 and GAD-7 are screening and monitoring tools; diagnosis integrates your clinical judgment, the history, and the context. A 12 is not a disorder: it is a signal you interpret.
  2. It does not decide for you. The data informs the decision; it does not replace it. A score that contradicts everything else you know about the case is a question to explore in session, not a verdict.
  3. It does not get filed unread. If you administer the PHQ-9 every session, sooner or later there will be a positive item 9, and an endorsed item 9 calls for an immediate clinical risk assessment under your own protocol and training — the PHQ-9 guide covers that point — plus its documentation in the note.
  4. It is not an anonymous worksheet. Every score you store is health information about your client, and many jurisdictions treat it as sensitive data: in Mexico, for example, the data-protection law in force requires express written consent to process health data. Getting informed consent right from the first session covers this too — and, as always, verify the rules where you practice.

Make measuring cost minutes, not sessions

The remaining objection is time, and it is the easiest to solve: a brief scale takes minutes to complete before the session; what actually consumes time is scoring, charting, and comparing by hand, week after week. gesell.ai includes built-in validated scales with automatic scoring — PHQ-9 and GAD-7 among them — tracks progress across sessions, and surfaces recurring themes in the process, always as decision support that you evaluate. The platform produces the chart; the clinical reading — and the conversation with your client — remain yours.

References

Share this article

About the author

Gesell Team

Clinical and product content written by the gesell.ai team together with certified clinical psychologists.

← Back to the blog