Skip to content
  • assessment tools
  • assessment

PHQ-9: administration and interpretation in practice

Gesell Team10 min read

The PHQ-9 is one of the most widely used depression screening instruments, and yet in everyday practice it tends to get used at half capacity: administered once, the number written down, the form filed away. Used well, it does three different jobs: it screens, it grades severity, and it tracks progress from session to session. This guide covers how to administer it, how to read the score without turning it into a verdict, what exactly a positive item 9 requires of you, and what evidence stands behind the free official translations — including the Spanish versions many clinicians need.

What the PHQ-9 is and what backs it

The PHQ-9 is the depression module of the Patient Health Questionnaire, the self-administered version of PRIME-MD. Each of its nine items maps to a DSM-IV criterion for depression and is scored 0 to 3 by how frequently the symptom occurred over the last two weeks; the total score runs from 0 to 27.

The original validation study (Kroenke, Spitzer & Williams, 2001) administered it to 6,000 patients across 8 primary care and 7 obstetrics-gynecology clinics, and assessed criterion validity against an independent structured interview by mental health professionals in 580 of those patients. Internal consistency was excellent: Cronbach’s alpha of 0.89 in the primary care study and 0.86 in the ob-gyn study. The paper’s own conclusion captures the instrument’s two jobs: beyond making criteria-based diagnoses, it is a reliable and valid measure of depression severity. Hold on to that sentence — everything below depends on keeping those two functions distinct, and on remembering that neither one replaces your clinical assessment.

Administering it in practice

The PHQ-9 is self-administered: the client completes it alone, rating each symptom on four frequency anchors from “not at all” to “nearly every day.” In practice, most people finish it in a few minutes, in the waiting room or at the start of the session — which leaves the clinical hour for discussing the result rather than collecting it.

Three administration details that prevent errors later:

  1. The tenth item does not count. The closing question about how much the symptoms interfere with work and daily life is clinically valuable, but the official instruction manual is explicit: that difficulty item is not used to calculate any PHQ score or diagnosis.
  2. If you use the PHQ-2 as a gateway, know the rule. For the ultra-brief versions, the manual states that a score of 3 or greater should prompt administration of the full PHQ-9, plus a clinical interview to determine whether a disorder is present.
  3. Use the official forms. The instrument is in the public domain: the official site states that the questionnaires and their translations can be downloaded, reproduced and distributed with no permission required. If you serve Spanish-speaking clients, there are official Spanish versions, including a Spanish-for-Mexico PHQ-9 whose footer states the no-permission terms in Spanish. With a free, verified original available, there is no reason to administer copies of unclear provenance.

When introducing it, a brief framing beats a technical explanation:

“Before we start, I’d like you to answer this short questionnaire about the last two weeks. There are no right or wrong answers; it helps me see how you’ve been doing and compare it with previous weeks.”

Remember, too, that the scores you record are health information in the client’s record — treat them with the same care as your notes, make sure your informed consent covers how records are stored and accessed, and verify the data-protection rules that apply where you practice.

Interpreting the score

The severity table nearly everyone quotes actually combines two sources, and it helps to know which says what. The 2001 paper established the cutpoints: scores of 5, 10, 15 and 20 represent mild, moderate, moderately severe and severe depression. The bottom band and the suggested actions come from the official manual, which credits them to Kroenke & Spitzer (2002):

Score Severity Action proposed by the manual
0–4 None–minimal None
5–9 Mild Watchful waiting; repeat PHQ-9 at follow-up
10–14 Moderate Treatment plan: counseling, follow-up and/or pharmacotherapy
15–19 Moderately severe Active treatment with pharmacotherapy and/or psychotherapy
20–27 Severe Active treatment

Note that those actions were written for primary care, with pharmacotherapy as a first-line option; in a psychotherapy practice they work as general orientation, not as prescriptions.

On the classic screening threshold: in the original validation, a score ≥ 10 had 88% sensitivity and 88% specificity for major depression against the structured interview. The manual adds a useful image: ≥ 10 is a “yellow flag” pointing to a possible clinically significant condition, and ≥ 15 a “red flag” marking someone in whom active treatment is probably warranted.

Now the nuance most summaries skip: those figures come from US primary care, and optimal cutoffs shift with the population. In a Colombian study of 243 adult primary-care users in Bucaramanga, the optimal cutoff was ≥ 7 (sensitivity 90.38%, specificity 81.68%, alpha 0.80). The lesson is not to change your cutoff on your own authority — it is that the cutoff guides screening while your clinical judgment decides what happens with each case.

Because the PHQ-9 screens and monitors; it does not diagnose by itself. The manual calls the diagnoses it produces “provisional,” and its own worked example shows the clinician confirming major depression only after ruling out a history of mania, physical or medication causes and normal bereavement — and after questioning the suicidal ideation the patient had endorsed. The scale proposes; the clinician decides.

Item 9 is never just filed — it gets assessed

Item 9 asks about thoughts of being better off dead or of hurting oneself in some way. It is the only item the diagnostic algorithm counts at any frequency: while other symptoms must be present “more than half the days,” the manual specifies that suicidal ideation is counted whenever it is present at all.

That special status has a practical translation worth stating plainly: any answer other than “not at all” on item 9 requires an immediate clinical risk assessment, in that same session, under your own protocol and training. The manual backs this up: the final decision about actual risk of self-harm requires a clinical interview, not a score. This guide will not hand you a step-by-step protocol — that comes from your training in suicide risk and from the professional guidelines that apply where you practice — but here are three verifiable supports:

  • A structured complement. The Columbia-Suicide Severity Rating Scale (C-SSRS) organizes the assessment through plain-language questions about ideation, preparatory acts and attempts; it is available in more than 100 country-specific languages and is free for community and healthcare settings.
  • Crisis resources for your client. Make sure the client leaves the session knowing whom to contact between sessions — the crisis line serving your client’s location, plus local emergency services. That does not replace your clinical plan; it completes it.
  • Documentation. Record the score, the assessment you conducted, what you found, and what your decision rests on. A chart showing a positive item 9 with no follow-up assessment is exactly the gap you cannot explain later.

From screening to monitoring

The PHQ-9’s second job is the one that delivers the most value and gets practiced the least: repeating it. The manual says so expressly — the PHQ-9’s responsiveness to change is well established, and the instrument is used both to grade initial severity and as an outcome tool to determine treatment response.

That is the heart of measurement-based care: systematic data collection to monitor client progress and directly inform care decisions, as Scott & Lewis (2015) define it. Their case example is replicable in any practice: the therapist administers the PHQ-9 and GAD-7 before each session and shares graphs of the client’s scores with her, turning the number into clinical conversation.

The severity bands make that tracking legible. A client who moves from 12 to 8 did not just “drop four points”: they crossed from moderate to mild symptomatology, and that is a genuine progress data point that belongs in the O section of your SOAP note and can anchor verifiable goals in your treatment plan — “PHQ-9 below 10 sustained for a month” can be evaluated; “improved mood” cannot.

Common PHQ-9 mistakes

  1. Treating the score as a diagnosis. A 16 is not major depressive disorder; it is a red flag your clinical assessment confirms or rules out.
  2. Filing a positive item 9. The one mistake on this list that never tolerates “I’ll look at it next session.”
  3. Administering it once and never again. Without repeated measurement there is no progress curve, only an intake snapshot.
  4. Using copies or translations of unclear provenance when the official versions are free from the instrument’s own site.
  5. Not sharing the result with the client. A score discussed and graphed is therapeutic material; a score in a drawer is paperwork.
  6. Forgetting the population context. Cutoffs are conventions derived from specific samples, not universal thresholds.

The Spanish versions have real data

The official manual includes a caveat few people quote: unlike the English versions, few of the PHQ translations have been psychometrically validated against an independent structured psychiatric interview. That is one more reason to use the official forms and to know the local evidence rather than assume it.

For Spanish, that evidence exists. In the Mexican Teachers’ Cohort (Familiar et al., 2015), with 55,555 participants, the Spanish PHQ-9 showed a one-factor structure with loadings from 0.71 to 0.90 and a Cronbach’s alpha of 0.89, and the authors endorse it for research and clinical use; 12.6% of that sample showed moderate-to-high depressive symptoms. In rural Chiapas (Arrieta et al., 2017), with 215 adults analyzed, internal consistency held at alpha ≥ 0.8 even across gender, literacy and age subgroups. If you work with Spanish-speaking clients, the version you download from the official site is not an improvised translation: it is an instrument with documented performance in Spanish-speaking populations.

Make the instrument work for you

Everything above — administering before the session, scoring, graphing, comparing bands, documenting — is easy with one client and heavy with thirty. That is where a platform helps: gesell.ai includes built-in validated scales with automatic scoring, PHQ-9 and GAD-7 among them, integrated into each client’s structured chart, so today’s score sits next to the previous ones and is ready for your note. The output is always decision support you evaluate: the scale contributes the data point; the clinical judgment remains yours.

References

Share this article

About the author

Gesell Team

Clinical and product content written by the gesell.ai team together with certified clinical psychologists.

← Back to the blog