Skip to content
  • assessment tools
  • assessment

GAD-7: administration and interpretation in practice

Gesell Team11 min read

The GAD-7 sits next to the PHQ-9 in almost every practice that measures anything: seven items, a few minutes to complete, and an anxiety score you can compare across sessions. It also shares the PHQ-9’s most common fate — administered once, the number written down, the form filed away. And it carries a confusion of its own that almost no summary clears up: the famous “10” on the GAD-7 means three different things depending on what you are using it for. This guide covers how to administer it, how to read the score without turning it into a verdict, what evidence actually stands behind the Spanish versions many clinicians need, and how to turn the number into real treatment monitoring.

What the GAD-7 is and what backs it

The GAD-7 was developed by Spitzer, Kroenke, Williams and Löwe as a brief measure of generalized anxiety disorder. It has seven items covering the last two weeks, each scored 0 to 3 by symptom frequency — “not at all,” “several days,” “more than half the days,” “nearly every day” — for a total score of 0 to 21.

The original validation study (2006) was run in primary care: 2,740 adult patients completed the questionnaire, and 965 of them had a telephone interview with a mental health professional within the following week, which served as the diagnostic criterion. Internal consistency was excellent (Cronbach’s alpha of .92) and test-retest reliability good (intraclass correlation of 0.83). The authors also wrote down the caveat that organizes everything below: the GAD-7 provides only probable diagnoses, which should be confirmed by further evaluation. No score replaces your clinical assessment; it feeds it.

Administering it in practice

The GAD-7 is self-administered: the client completes it alone, and it is genuinely brief — in the Spanish validation by García-Campayo and colleagues, mean completion time was 2 minutes 30 seconds. It fits in the waiting room or at the start of the session, which leaves the clinical hour for discussing the result rather than collecting it.

Three administration details that prevent problems later:

  1. Use the official version for your client’s language and country, complete. The instrument is in the public domain: per the official instruction manual, no permission is required to reproduce, translate, display or distribute it. If you serve Spanish-speaking clients, note that the official Spanish versions are not interchangeable: the Spanish-for-Mexico form and the Spanish-for-Spain form word the stem and several items differently. Pick the version that matches your client’s population and never mix items across versions.
  2. If you use the GAD-2 as a gateway, know the rule. For the ultra-brief version (the first two items, range 0 to 6), the manual states that a score of 3 or greater should prompt administration of the full GAD-7, plus a clinical interview to determine whether a disorder is present.
  3. The score you record is health information. Treat it with the same care as your notes: make sure your informed consent covers how assessment data is stored and accessed, and verify the data-protection rules that apply where you practice — many jurisdictions treat health data as a specially protected category.

When introducing it, a brief framing beats a technical explanation:

“I’d like you to answer this short questionnaire about the last two weeks. There are no right or wrong answers; it helps me see how your anxiety has been and compare it with previous weeks.”

Interpreting the score: the three “10s”

Start with the severity bands, which come from the original study and are confirmed in the official manual: cutpoints of 5, 10 and 15 mark mild, moderate and severe anxiety.

Score Severity
0–4 Minimal
5–9 Mild
10–14 Moderate
15–21 Severe

Now the three distinct jobs the same number does:

The first “10” grades severity. It is the boundary where symptoms stop being mild. The manual attaches a memorable image: a score of 10 or greater is a “yellow flag” pointing to a possible clinically significant condition, and 15 is a “red flag” marking someone in whom active treatment is probably warranted.

The second “10” screens for generalized anxiety disorder. In the original validation, a cutoff of 10 or greater optimized detection — 89% sensitivity and 82% specificity against the criterion interview — and the authors propose it as a reasonable cut point for identifying probable cases of GAD.

The third “10” screens beyond GAD. According to the official manual, the GAD-7 also has moderately good operating characteristics for three other common anxiety disorders — panic disorder, social anxiety disorder and PTSD — and it recommends the same cutpoint of 10 or greater for further evaluation when screening for anxiety disorders in general (the study behind that statement is Kroenke et al., 2007). The nuance: that same study proposed a lower cutoff of 8 for the broad screen, as restated by Johnson and colleagues (2019).

And one rule governs all three: optimal cutoffs are population-dependent. In Johnson and colleagues’ heterogeneous psychiatric sample (1,201 patients), the optimal cutoff was 8, with sensitivity of 0.92 and specificity of 0.70; in Spanish primary care, a cutoff of 10 yielded 86.8% and 93.4%. For context on how high “high” is: in a representative German general-population sample, the mean score was 2.18 and a score of 10 fell between the 96th and 99th percentiles — Germany-specific reference data, not clinical norms for your setting. The practical lesson is the same as with the PHQ-9: the cutoff guides screening; your clinical judgment decides each case.

What a high score does not tell you

Four cautions, all documented in the primary sources:

  1. It is a probable diagnosis, not a confirmed one. The phrase belongs to the authors themselves: GAD-7 diagnoses should be confirmed by further evaluation. A 16 is not GAD; it is a red flag your interview confirms or rules out.
  2. Two weeks are not six months. The GAD-7 asks about the last two weeks, but the GAD diagnostic criterion used in the validation required at least six months of symptoms. A high score cannot, by itself, establish the duration criterion.
  3. Anxiety can also be medical. The manual requires ruling out a physical disorder, a medication or another substance as the biological cause of the symptoms before concluding an anxiety syndrome.
  4. Anxiety and depression travel together. In the original study the GAD-7 correlated 0.75 with the depression measure (PHQ-8); even so, the authors concluded that measuring both is complementary rather than duplicative. If the presentation suggests both, administer both instruments and read them side by side.

A note on risk, even though the GAD-7 demands less protocol than its sibling: unlike the PHQ-9, it contains no item about suicidal ideation. That does not delegate the topic. If any risk disclosure surfaces during administration or conversation, it calls for an immediate clinical risk assessment under your own protocol and training, documented in the record — and make sure the client leaves knowing which local crisis line and emergency services serve their location. That completes your clinical plan; it does not replace it.

The Spanish versions: what the evidence says and what is missing

The official manual carries a caveat few people quote: most translations are linguistically valid, but few have been psychometrically validated against an independent structured psychiatric interview. It helps to keep two things separate that usually get blurred: the official translation you download, and the published evidence on the GAD-7 in Spanish.

The most-cited Spanish validation is the cultural adaptation by García-Campayo and colleagues (2010), in Spanish primary care with 212 patients: Cronbach’s alpha of 0.936, test-retest reliability (intraclass correlation) of 0.926, a one-dimensional structure, and a cutoff of 10 with 86.8% sensitivity and 93.4% specificity. One precision worth keeping: that adaptation was published on its own track, and it has not been established that it is word-for-word identical to the Spanish-for-Spain form on the official site — cite each for what it is.

What about Latin America? Evidence exists, with a scope worth stating honestly. Galindo-Vázquez and colleagues (2023) evaluated the GAD-7 in 163 Mexican patients in oncology genetic counseling: Cronbach’s alpha of 0.899, a one-factor structure explaining 62.3% of the variance, and the conclusion that the instrument is valid and reliable in a Mexican population. That is reliability and factor-structure evidence in a specific clinical sample; what we could not locate is a Mexican diagnostic-accuracy validation — sensitivity and specificity against a structured interview. The cutoffs you use come from other populations: treat them as orientation, not as locally established thresholds. The largest regional dataset comes from Peru: 4,431 people from the general population (ages 12 to 65), a one-factor structure, satisfactory internal consistency (omega of 0.90) and measurement invariance across sex and age.

From screening to monitoring

The GAD-7’s second job is the one that delivers the most value and gets practiced the least: repeating it. That is the heart of measurement-based care as Scott & Lewis (2015) define it — systematic data collection to monitor progress and inform care decisions, a framework that, per the preliminary research they cite, produces better outcomes than usual care. Their case example is replicable in any practice: the therapist administers the PHQ-9 and GAD-7 before each session and shares the score graphs with the client.

How often should you repeat it? There is no official cadence standard, and you should distrust anyone who cites one. The instrument itself sets the natural floor: it asks about two weeks, so readministering sooner measures the same period twice. What matters is consistency — same instrument, same moment (ideally before the session), scores graphed over time. One honest note from the manual itself: the GAD-7 has demonstrated sensitivity to change as a secondary outcome in depression trials, though as of that document it had not been studied as a primary outcome in anxiety trials.

The bands make the tracking legible. A client who moves from 13 to 8 did not just “drop five points”: they crossed from moderate to mild symptomatology, and that is a progress data point that belongs in the O section of your SOAP note and can anchor verifiable goals in your treatment plan — “GAD-7 below 10 sustained for a month” can be evaluated; “less anxious” cannot.

“GAD-7: 8 (previous: 13). Crossed from the moderate to the mild band, consistent with the reduced avoidance she reports. Continue graded exposure; readminister before next session.”

Common GAD-7 mistakes

  1. Treating the cutoff as a diagnosis. The instrument screens and grades severity; diagnosis integrates the interview, the duration criterion and the differential.
  2. Using the same “10” for everything. Band boundary, GAD screening cutoff and broad anxiety-screening cutoff are three different uses with different evidence.
  3. Administering it once and never again. Without repeated measurement there is no progress curve, only an intake snapshot.
  4. Mixing country versions or using copies of unclear provenance, when official versions are free from the instrument’s own site.
  5. Measuring only anxiety when the presentation suggests both poles. With a 0.75 correlation with the depression measure, overlap is the rule, not the exception.
  6. Forgetting the medical differential. An anxiety syndrome with a physical or pharmacological cause is not solved with more psychotherapy.

Turning the number into clinical conversation

Everything above — administering before the session, scoring, graphing, comparing bands, documenting — is easy with one client and heavy with thirty. That is where the platform helps: gesell.ai includes built-in validated scales with automatic scoring, the GAD-7 and PHQ-9 among them, integrated into each client’s structured chart, so today’s score sits next to the previous ones, ready for your note and for the conversation in session. Always as decision support you evaluate: the scale contributes the data point; the clinical judgment remains yours.

References

Share this article

About the author

Gesell Team

Clinical and product content written by the gesell.ai team together with certified clinical psychologists.

← Back to the blog