The PHQ-9 is the most frequently used depression screener in the world (Carroll et al., 2020): a brief self-report measure whose items map one-to-one onto the DSM symptom criteria for major depressive disorder, so the same responses serve as both a screening result and a severity score.
It was developed by Robert Spitzer, Janet Williams, Kurt Kroenke and colleagues as the depression module of the Patient Health Questionnaire and validated as a standalone severity measure in 2001; it is free to use, with no permission required to reproduce, translate, display, or distribute it.
A large meta-analytic literature supports its internal consistency and its screening accuracy across settings, languages, and reference standards, and shows that the summed-score cutoff outperforms the DSM-based algorithm when the goal is screening.
A high score is not a diagnosis: confirmation requires a clinical interview, and any endorsement of the suicidal-ideation item warrants clinical follow-up regardless of the total score.
At a Glance
Items
9 symptom items, one per DSM criterion for major depressive disorder, plus one unscored functional-impairment question
Administration time
A few minutes
Response format
4-point frequency scale, 0 = not at all to 3 = nearly every day, rated over the last 2 weeks
Scores
Single total score, 0-27; severity bands at 5, 10, 15, and 20; optional DSM-based diagnostic algorithm
Validated populations
Adults in primary care and obstetrics-gynecology settings (Kroenke et al., 2001); adolescents aged 13-17 (Richardson et al., 2010)
License
Free to use; the official form states “No permission required to reproduce, translate, display or distribute”
Original citation
Kroenke, Spitzer, & Williams (2001), Journal of General Internal Medicine
Introduction
The Patient Health Questionnaire-9 (PHQ-9) is the most frequently used depression screener globally (Carroll et al., 2020). It originated as the depression module of the Patient Health Questionnaire (PHQ), the self-report version of PRIME-MD, developed by Robert L. Spitzer, Janet B.W. Williams, Kurt Kroenke and colleagues with an educational grant from Pfizer Inc. (Spitzer, Kroenke, & Williams, 1999), and was validated as a standalone depression severity measure by Kroenke, Spitzer, and Williams (2001). Its nine items correspond one-to-one to the nine DSM diagnostic criteria for major depressive disorder (criteria that are unchanged from DSM-IV, against which the instrument was developed, to DSM-5), which allows the same nine responses to serve both screening and severity measurement.
The PHQ-9 changed routine depression assessment by packaging diagnostic-criterion coverage into a brief, freely reproducible self-report form. Before instruments of this kind, depression assessment in medical settings often relied on lengthy clinical interviews or on scales that measured distress without clear diagnostic relevance.
Understanding Depression as a Clinical Syndrome
Major depressive disorder is one of the most common mental health conditions worldwide, affecting approximately 5-7% of adults in any given year. It is characterized by persistent low mood or loss of interest, together with cognitive, behavioral, and physical symptoms that impair functioning.
The PHQ-9 measures depression as the DSM defines it (criteria unchanged from DSM-IV to DSM-5): at least five symptoms present during the same two-week period, at least one of which is depressed mood or anhedonia. This criterion-based construction is what lets the instrument function as a screening tool, an aid to diagnostic assessment, and a treatment-monitoring measure at once.
Theoretical Foundation
The PHQ-9 is grounded in the DSM diagnostic framework (the DSM-IV criteria on which it was built, unchanged in DSM-5), which defines major depressive disorder by symptom criteria. Unlike depression scales that sample general distress, the PHQ-9 systematically covers each of the nine criteria required for diagnosis:
Anhedonia (loss of interest or pleasure)
Depressed mood (feeling down or hopeless)
Sleep disturbances (insomnia or hypersomnia)
Fatigue or loss of energy
Appetite changes (increase or decrease)
Feelings of worthlessness or excessive guilt
Diminished ability to concentrate
Psychomotor agitation or retardation
Recurrent thoughts of death or suicidal ideation
The two-week recall window matches the DSM duration requirement, and the frequency-based response scale (not at all, several days, more than half the days, nearly every day) captures symptom persistence, which is what separates clinical depression from transient mood change.
🏥 Key insight: Because each PHQ-9 item corresponds to one DSM symptom criterion, the same nine responses can be read two ways, summed as a severity score or checked against a diagnostic algorithm; the instrument is among those recommended by the US Preventive Services Task Force for routine depression screening in adults (Siu et al., 2016).
Key Features
Assessment Characteristics
9 items corresponding exactly to the DSM depression criteria, plus one unscored functional-impairment question
4-point frequency scale (0-3) covering the last 2 weeks
A few minutes to complete
Dual functionality as screening tool and severity measure
Validated in adult primary care and obstetrics-gynecology samples (Kroenke et al., 2001) and in adolescents aged 13-17 (Richardson et al., 2010)
Free to use – no permission required to reproduce, translate, display, or distribute
Depression Dimensions Assessed
Anhedonia – Loss of interest or pleasure in doing things
Depressed mood – Feeling down, depressed, or hopeless
Sleep disturbance – Trouble sleeping or sleeping too much
Fatigue – Feeling tired or having little energy
Appetite changes – Poor appetite or overeating
Guilt/worthlessness – Negative self-evaluation and self-blame
Concentration problems – Difficulty focusing on activities
Psychomotor changes – Moving/speaking slowly or being restless
Suicidal ideation – Thoughts of death or self-harm
Versions & Adaptations
Full PHQ – the PHQ-9 is the depression module of the complete Patient Health Questionnaire, the self-report version of PRIME-MD (Spitzer, Kroenke, & Williams, 1999)
PHQ-2 – the first two items (anhedonia and depressed mood) as an ultra-brief initial screener (Kroenke, Spitzer, & Williams, 2003)
PHQ-8 – omits the ninth (suicidal-ideation) item, developed for general-population and telephone research (Kroenke et al., 2009)
PHQ-A – adolescent modification of the PHQ (Johnson et al., 2002)
Translations – over 70 languages and dialects, though only a subset have been formally psychometrically validated (Carroll et al., 2020)
Research and Clinical Applications
Primary care screening – Standard depression detection in medical settings
Mental health assessment – Initial evaluation in psychiatric settings
Treatment monitoring – Tracking symptom change during therapy
Clinical trials – Outcome measure in depression research
Healthcare quality – Performance measurement and quality improvement
Population health – Community mental health surveillance
Collaborative care – Communication tool across care teams
Assess depression symptoms experienced over the past 2 weeks.
Scoring and Interpretation
Response Format
Participants rate how often they have been bothered by each problem over the last 2 weeks using a 4-point frequency scale:
0 = Not at all
1 = Several days
2 = More than half the days
3 = Nearly every day
Complete PHQ-9 Items
The PHQ-9 is free to reproduce, so the full item set can be shown. The stem reads: “Over the last 2 weeks, how often have you been bothered by any of the following problems?”
Little interest or pleasure in doing things
Feeling down, depressed, or hopeless
Trouble falling or staying asleep, or sleeping too much
Feeling tired or having little energy
Poor appetite or overeating
Feeling bad about yourself — or that you are a failure or have let yourself or your family down
Trouble concentrating on things, such as reading the newspaper or watching television
Moving or speaking so slowly that other people could have noticed. Or the opposite — being so fidgety or restless that you have been moving around a lot more than usual
Thoughts that you would be better off dead, or of hurting yourself in some way
Functional Impairment Question
After the 9 items, participants answer:
“If you checked off any problems, how difficult have these problems made it for you to do your work, take care of things at home, or get along with other people?”
Not difficult at all
Somewhat difficult
Very difficult
Extremely difficult
(This question is not included in the total score but provides important clinical context.)
Scoring Procedure
Sum all 9 item responses (range: 0-27); no items are reverse-keyed
Active treatment with pharmacotherapy and/or psychotherapy
20-27
Severe
Immediate initiation of pharmacotherapy and/or psychotherapy
Severity thresholds of 5, 10, 15, and 20 come from the original validation (Kroenke et al., 2001); the proposed treatment actions follow Kroenke and Spitzer (2002). The same ordinal bands (0-4, 5-9, 10-14, 15-19, 20-27) are used throughout the meta-analytic literature (Manea et al., 2015).
Provisional diagnosis of major depressive disorder requires:
5 or more items scored as 2 (more than half the days) or 3 (nearly every day)
PLUS:
Must include Item 1 (anhedonia) OR Item 2 (depressed mood)
Item 9 (suicidal ideation) counts if present at any frequency (score ≥1)
This DSM-based algorithm is the originally proposed categorical scoring method, verified against the official Patient Health Questionnaire office-coding instructions. Its sensitivity is substantially lower than that of the summed-score cutoff (see Research Evidence), so the summed ≥10 cutoff is generally preferred for screening (Manea et al., 2015).
Suicide Risk Assessment
Any positive response to Item 9 warrants clinical attention. The following tiers reflect common clinical practice rather than a single validated protocol:
Score 2-3 (“more than half the days” or “nearly every day”): Comprehensive suicide risk evaluation
A final decision about the actual risk of self-harm requires a clinical interview.
Cutoffs and Change Benchmarks
Optimal screening cutoff: ≥10 (sensitivity 88%, specificity 88%) (Kroenke et al., 2001, as reported in Manea et al., 2015)
Cutoff trade-offs: lower cutoffs raise sensitivity (at ≥8: sensitivity 0.95, specificity 0.75) and higher cutoffs raise specificity (at ≥12: sensitivity 0.79, specificity 0.91), but ≥10 maximized combined sensitivity and specificity overall and across subgroups (Levis et al., 2019)
Adolescents (ages 13-17): a slightly higher cutoff (≥11) was optimal (Richardson et al., 2010)
Meaningful change: a ≥5-point reduction, the minimal clinically important difference estimated as two standard errors of measurement, indicates clinically significant improvement (Löwe, Unützer, Callahan, Perkins, & Kroenke, 2004)
The PHQ-9 is interpreted against these published severity bands and cutoffs rather than against population norms; no normative means and standard deviations are presented here because none were verified for this page.
Research Evidence and Psychometric Properties
Reliability Evidence
Internal consistency: α = 0.86 (95% CI [0.85, 0.87]) in a reliability-generalization meta-analysis of 60 studies with 232,147 participants (Ajele & Idemudia, 2025); in the original validation samples, α = 0.89 (primary care) and 0.86 (obstetrics-gynecology) (Kroenke et al., 2001)
Test-retest reliability: r = 0.84 between self-administered and telephone-administered forms within 48 hours (Kroenke et al., 2001); pooled test-retest reliability of 0.82 in the meta-analysis (Ajele & Idemudia, 2025)
Diagnostic Accuracy
Summed-score cutoff (≥10):
Original validation: sensitivity 88% and specificity 88% for major depression at the ≥10 cutoff (Kroenke et al., 2001, as reported in Manea et al., 2015); area under the ROC curve 0.95 (Kroenke et al., 2001)
Individual participant data meta-analysis: 58 studies (17,357 participants); at ≥10, sensitivity 0.88 (95% CI 0.83-0.92) and specificity 0.85 (95% CI 0.82-0.88) against semistructured diagnostic interviews, with ≥10 optimal across subgroups (Levis et al., 2019)
Predictive values: positive predictive value 24-66% across assumed prevalences of 5-25%; negative predictive value 96-99% (Levis et al., 2019)
Algorithm (DSM-based) scoring method:
In a diagnostic meta-analysis of 27 validation studies, the algorithm scoring method showed pooled sensitivity 0.58 (95% CI 0.50-0.66) and pooled specificity 0.94 (95% CI 0.92-0.96) (Manea et al., 2015)
Among studies reporting both methods, the summed score at ≥10 achieved higher sensitivity (0.77, with specificity 0.85) than the algorithm (sensitivity 0.53); the summed ≥10 cutoff has better diagnostic performance for screening purposes or where high sensitivity is needed (Manea et al., 2015)
Validity Evidence
Criterion validity against diagnostic interviews: good screening accuracy against structured and semistructured diagnostic interviews, with accuracy strongest when semistructured (SCID-type) interviews served as the reference standard (Levis et al., 2019)
Treatment Sensitivity
Minimal clinically important difference: a 5-point change, estimated as two standard errors of measurement (Löwe, Unützer, Callahan, Perkins, & Kroenke, 2004)
Sensitivity to change: standardized effect sizes of -1.33 in patients rated improved, -0.21 in unchanged patients, and +0.47 in deteriorated patients (Löwe, Kroenke, Herzog, & Gräfe, 2004)
Response and remission conventions: the original validation proposed a score below 10 combined with a ≥50% decline as an improvement criterion, to be verified clinically (Kroenke et al., 2001); scores below 5 are commonly used as a remission benchmark in the collaborative-care literature (Löwe, Unützer, Callahan, Perkins, & Kroenke, 2004)
Cross-Cultural Validation
Global validation: 49 validation studies across low- and middle-income countries (Carroll et al., 2020)
Language versions: translated into over 70 languages and dialects, though only a subset have been formally psychometrically validated (Carroll et al., 2020)
Special Populations
Adolescents (ages 13-17):
Validated for detecting major depression among adolescents aged 13-17: an optimal cutoff of ≥11 gave sensitivity 89.5%, specificity 77.5%, and an area under the ROC curve of 0.88 (Richardson et al., 2010)
Usage Guidelines and Applications
Primary Clinical Applications
Depression screening in primary care settings; the USPSTF issues a Grade B recommendation to screen adults for depression, without specifying a screening interval (Siu et al., 2016)
Initial mental health assessment in psychiatric and counseling settings
Treatment progress monitoring during active therapy
Outcome measurement in healthcare quality improvement programs
Collaborative care models for systematic tracking across care teams
Clinical Decision Support by Severity
Minimal depression (0-4):
No treatment indicated; continue routine screening
Mild depression (5-9):
Watchful waiting with repeat PHQ-9 at follow-up
Consider lifestyle interventions, brief counseling, or psychoeducation
Moderate depression (10-14):
Active treatment planning: psychotherapy, medication, or their combination
Discuss treatment preferences and establish a follow-up schedule
Moderately severe depression (15-19):
Active treatment with pharmacotherapy and/or evidence-based psychotherapy
More frequent monitoring; consider specialist referral; assess suicide risk
Severe depression (20-27):
Immediate treatment initiation, with strong consideration for combination therapy
Expedited psychiatric referral, frequent monitoring, and comprehensive suicide risk assessment
Suicide Risk Protocol
Any endorsement of Item 9 (score ≥1) suggests the following clinical practice:
Prompt clinical follow-up
Comprehensive suicide risk assessment
Safety planning and lethal means counseling
Increased monitoring frequency
Consider referral to emergency/crisis services if high risk
Treatment Monitoring Guidelines
Frequency (customary collaborative-care practice):
Administer at baseline, then every 2-4 weeks during active treatment
Continue monthly once stabilized; resume frequent monitoring if relapse is suspected
A score below 5 is a common treatment target/remission benchmark; a ≥50% reduction from baseline is the customary responder criterion, consistent with the improvement metric proposed in the original validation (Kroenke et al., 2001)
Cultural and Special-Population Considerations
Medical comorbidity: somatic items (fatigue, sleep, appetite) may overlap with medical illness; clinical judgment is essential for interpretation
Cultural expression: depression expression varies across cultures, and some populations may somaticize emotional distress; consider cultural context when interpreting results
Age: for adolescents aged 13-17 a higher screening cutoff (≥11) was optimal (Richardson et al., 2010); older adults may underreport mood symptoms, so assess cognitive factors alongside the questionnaire
Limitations and Cautions
Not diagnostic alone: a clinical interview is required for definitive diagnosis
Cannot detect bipolar disorder: the item set contains no mania/hypomania items, so additional screening is needed
Symptom overlap: medical illness may inflate scores
Self-report limitations: depends on insight and willingness to disclose
Two-week timeframe: may miss episodic or rapidly cycling symptoms
Copyright and Usage Responsibility: Check that you have the proper rights and permissions to use this assessment tool in your research. This may include purchasing appropriate licenses, obtaining permissions from authors/copyright holders, or ensuring your usage falls within fair use guidelines.
The PHQ-9 is in the public domain and freely available for all uses worldwide. The official PHQ-9 form states: “No permission required to reproduce, translate, display or distribute.” The PHQ Screeners terms of use further state that the content “is expressly exempted from Pfizer’s general copyright restrictions” and “is free for download and use.” The PHQ was developed with an educational grant from Pfizer Inc.
Proper Attribution: When using or referencing this scale, cite the original development:
Kroenke, K., Spitzer, R. L., & Williams, J. B. W. (2001). The PHQ-9: Validity of a brief depression severity measure. Journal of General Internal Medicine, 16(9), 606-613. https://doi.org/10.1046/j.1525-1497.2001.016009606.x
Spitzer, R. L., Kroenke, K., & Williams, J. B. W. (1999). Validation and utility of a self-report version of PRIME-MD: The PHQ Primary Care Study. JAMA, 282(18), 1737-1744. https://doi.org/10.1001/jama.282.18.1737
Kroenke, K., Spitzer, R. L., & Williams, J. B. W. (2001). The PHQ-9: Validity of a brief depression severity measure. Journal of General Internal Medicine, 16(9), 606-613. https://doi.org/10.1046/j.1525-1497.2001.016009606.x
Kroenke, K., & Spitzer, R. L. (2002). The PHQ-9: A new depression diagnostic and severity measure. Psychiatric Annals, 32(9), 509-515. https://doi.org/10.3928/0048-5713-20020901-06
Short Forms and Adaptations:
Kroenke, K., Spitzer, R. L., & Williams, J. B. W. (2003). The Patient Health Questionnaire-2: Validity of a two-item depression screener. Medical Care, 41(11), 1284-1292. https://doi.org/10.1097/01.MLR.0000093487.78664.3C
Kroenke, K., Strine, T. W., Spitzer, R. L., Williams, J. B. W., Berry, J. T., & Mokdad, A. H. (2009). The PHQ-8 as a measure of current depression in the general population. Journal of Affective Disorders, 114(1-3), 163-173. https://doi.org/10.1016/j.jad.2008.06.026
Johnson, J. G., Harris, E. S., Spitzer, R. L., & Williams, J. B. W. (2002). The patient health questionnaire for adolescents: Validation of an instrument for the assessment of mental disorders among adolescent primary care patients. Journal of Adolescent Health, 30(3), 196-204. https://doi.org/10.1016/S1054-139X(01)00333-0
Diagnostic Accuracy Meta-Analyses:
Levis, B., Benedetti, A., Thombs, B. D., & DEPRESsion Screening Data (DEPRESSD) Collaboration. (2019). Accuracy of Patient Health Questionnaire-9 (PHQ-9) for screening to detect major depression: Individual participant data meta-analysis. BMJ, 365, l1476. https://doi.org/10.1136/bmj.l1476
Manea, L., Gilbody, S., & McMillan, D. (2015). A diagnostic meta-analysis of the Patient Health Questionnaire-9 (PHQ-9) algorithm scoring method as a screen for depression. General Hospital Psychiatry, 37(1), 67-75. https://doi.org/10.1016/j.genhosppsych.2014.09.009
Reliability Generalization:
Ajele, K. W., & Idemudia, E. S. (2025). Charting the course of depression care: A meta-analysis of reliability generalization of the patient health questionnaire (PHQ-9) as the measure. Discover Mental Health, 5(1), 50. https://doi.org/10.1007/s44192-025-00181-x
Cross-Cultural Validation:
Carroll, H. A., Hook, K., Rojas Perez, O. F., Denckla, C., Cooper Vince, C., Ghebrehiwet, S., Ando, K., Touma, M., Borba, C. P. C., Fricchione, G. L., & Henderson, D. C. (2020). Establishing reliability and validity for mental health screening instruments in resource-constrained settings: Systematic review of the PHQ-9 and key recommendations. Psychiatry Research, 291, 113236. https://doi.org/10.1016/j.psychres.2020.113236
Treatment Monitoring:
Löwe, B., Kroenke, K., Herzog, W., & Gräfe, K. (2004). Measuring depression outcome with a brief self-report instrument: Sensitivity to change of the Patient Health Questionnaire (PHQ-9). Journal of Affective Disorders, 81(1), 61-66. https://doi.org/10.1016/S0165-0327(03)00198-8
Löwe, B., Unützer, J., Callahan, C. M., Perkins, A. J., & Kroenke, K. (2004). Monitoring depression treatment outcomes with the Patient Health Questionnaire-9. Medical Care, 42(12), 1194-1201. https://doi.org/10.1097/00005650-200412000-00006
Clinical Guidelines:
Siu, A. L., & US Preventive Services Task Force. (2016). Screening for depression in adults: US Preventive Services Task Force recommendation statement. JAMA, 315(4), 380-387. https://doi.org/10.1001/jama.2015.18392
Special Populations:
Richardson, L. P., McCauley, E., Grossman, D. C., McCarty, C. A., Richards, J., Russo, J. E., Rockhill, C., & Katon, W. (2010). Evaluation of the Patient Health Questionnaire-9 Item for detecting major depression among adolescents. Pediatrics, 126(6), 1117-1123. https://doi.org/10.1542/peds.2010-0852
A burdened elephant struggling under the weight of a broken heart, storm cloud, and heavy rock, powerfully illustrating the emotional load measured by the PHQ-9 (Patient Health Questionnaire)
Frequently Asked Questions
What does the PHQ-9 measure?
The PHQ-9 measures depression severity based on the nine DSM diagnostic criteria for major depressive disorder (criteria unchanged from DSM-IV, on which the instrument was built, to DSM-5). It covers depressed mood, anhedonia, sleep disturbance, fatigue, appetite changes, guilt or worthlessness, concentration problems, psychomotor changes, and suicidal ideation, each rated for frequency over the past two weeks.
How is the PHQ-9 scored?
Sum all nine item responses (each rated 0-3) for a total score of 0-27; no items are reverse-keyed. Severity bands are 0-4 (minimal), 5-9 (mild), 10-14 (moderate), 15-19 (moderately severe), and 20-27 (severe). A summed score of 10 or above is the standard screening cutoff, with sensitivity and specificity of 88% in the original validation (Kroenke et al., 2001, as reported in Manea et al., 2015); in adolescents aged 13-17 a cutoff of 11 was optimal (Richardson et al., 2010).
Is the PHQ-9 free to use?
Yes. The official form states that no permission is required to reproduce, translate, display, or distribute it, and the PHQ Screeners terms of use exempt the content from Pfizer's general copyright restrictions, making it free for download and use. The PHQ was developed with an educational grant from Pfizer Inc.
Can the PHQ-9 diagnose depression?
No. A DSM-based algorithm exists for a provisional diagnosis (five or more symptoms at least 'more than half the days', including depressed mood or anhedonia, with the suicidal-ideation item counting at any frequency), but its sensitivity is substantially lower than the summed cutoff of 10 or above (Manea et al., 2015), and a definitive diagnosis always requires a clinical interview.
What short forms and adaptations of the PHQ-9 exist?
The PHQ-2 uses the first two items as an ultra-brief screener, the PHQ-8 omits the suicidal-ideation item for general-population and telephone research, and the PHQ-A is the adolescent modification. The PHQ-9 itself is the depression module of the full Patient Health Questionnaire, and it has been translated into over 70 languages and dialects, a subset of which have been formally validated.
How does the PHQ-9 differ from the Beck Depression Inventory-II?
The PHQ-9 maps its nine items directly onto the DSM diagnostic criteria, while the BDI-II has 21 items sampling broader depressive symptomatology. The PHQ-9 is briefer, free to use (the BDI-II requires purchase), and designed for medical settings and routine screening.
Frequently Asked Questions
What does the PHQ-9 measure?
The PHQ-9 measures depression severity based on the nine DSM diagnostic criteria for major depressive disorder (criteria unchanged from DSM-IV, on which the instrument was built, to DSM-5). It covers depressed mood, anhedonia, sleep disturbance, fatigue, appetite changes, guilt or worthlessness, concentration problems, psychomotor changes, and suicidal ideation, each rated for frequency over the past two weeks.
How is the PHQ-9 scored?
Sum all nine item responses (each rated 0-3) for a total score of 0-27; no items are reverse-keyed. Severity bands are 0-4 (minimal), 5-9 (mild), 10-14 (moderate), 15-19 (moderately severe), and 20-27 (severe). A summed score of 10 or above is the standard screening cutoff, with sensitivity and specificity of 88% in the original validation (Kroenke et al., 2001, as reported in Manea et al., 2015); in adolescents aged 13-17 a cutoff of 11 was optimal (Richardson et al., 2010).
Is the PHQ-9 free to use?
Yes. The official form states that no permission is required to reproduce, translate, display, or distribute it, and the PHQ Screeners terms of use exempt the content from Pfizer's general copyright restrictions, making it free for download and use. The PHQ was developed with an educational grant from Pfizer Inc.
Can the PHQ-9 diagnose depression?
No. A DSM-based algorithm exists for a provisional diagnosis (five or more symptoms at least 'more than half the days', including depressed mood or anhedonia, with the suicidal-ideation item counting at any frequency), but its sensitivity is substantially lower than the summed cutoff of 10 or above (Manea et al., 2015), and a definitive diagnosis always requires a clinical interview.
What short forms and adaptations of the PHQ-9 exist?
The PHQ-2 uses the first two items as an ultra-brief screener, the PHQ-8 omits the suicidal-ideation item for general-population and telephone research, and the PHQ-A is the adolescent modification. The PHQ-9 itself is the depression module of the full Patient Health Questionnaire, and it has been translated into over 70 languages and dialects, a subset of which have been formally validated.
How does the PHQ-9 differ from the Beck Depression Inventory-II?
The PHQ-9 maps its nine items directly onto the DSM diagnostic criteria, while the BDI-II has 21 items sampling broader depressive symptomatology. The PHQ-9 is briefer, free to use (the BDI-II requires purchase), and designed for medical settings and routine screening.