BDI-II: Beck Depression Inventory-II

Reviewed by: Constantin Rezlescu | Associate Professor | UCL Psychology

TL;DR

  • The BDI-II is a self-report measure of the severity of depressive symptoms in adolescents and adults, asking respondents to describe how they have felt over the past two weeks across cognitive, affective, and somatic symptoms.
  • Aaron Beck, Robert Steer, and Gregory Brown published it in 1996 as a revision of the original 1961 Beck Depression Inventory, updating item content and extending the reference period to align with DSM-IV depression criteria while keeping its grounding in Beck's cognitive model.
  • Reliability is high across psychiatric, college, and adolescent samples, scores correspond closely to the original inventory, and a large review literature supports its use in many languages, though optimal screening cutoffs vary widely by setting.
  • It is a commercial Pearson instrument that must be licensed before use, and it measures severity rather than diagnosing depression; a clinical interview remains necessary for diagnosis.

At a Glance

Items 21; each item presents four graded statements scored 0-3 (items 16 and 18, sleep and appetite, present seven options distinguishing increases from decreases, still scored 0-3)
Administration time 5-10 minutes (the publisher lists 5 minutes)
Response format For each item, respondents select the one statement that best describes how they have felt during the past two weeks, including today
Scores Single total score, 0-63; cognitive-affective and somatic factors are reported in psychometric studies
Validated ages 13 through 80 years (publisher specification)
License Commercial. © Pearson (formerly The Psychological Corporation); purchase and user qualification (level B) required
Original citation Beck, Steer, & Brown (1996), Manual for the Beck Depression Inventory-II

Introduction

The Beck Depression Inventory-II (BDI-II) is a 21-item self-report measure of the severity of depressive symptoms in adolescents and adults aged 13 through 80. It was developed by Aaron T. Beck, Robert A. Steer, and Gregory K. Brown (1996) as a revision of the original Beck Depression Inventory (Beck, Ward, Mendelson, Mock, & Erbaugh, 1961), by way of the intermediate amended version known as the BDI-IA, and it remains among the most widely used and cited depression measures in clinical practice and research. Unlike brief screening tools built primarily for case detection, the BDI-II was designed as a severity measure: it quantifies how intensely the depressive syndrome has been experienced over the preceding two weeks.

Understanding Depression Severity Through the Cognitive Model

Beck, who developed cognitive therapy, framed depression as arising from and maintained by characteristic patterns of negative thinking: the cognitive triad of negative views about the self, the world, and the future. The BDI-II operationalizes the depressive syndrome that follows from this model as 21 symptoms, each rated on graded statements. These span cognitive-affective symptoms, such as sadness, pessimism, self-dislike, and worthlessness, and somatic symptoms, such as loss of energy, changes in sleep and appetite, and fatigue.

Theoretical Foundation

The 1996 revision aligned the inventory with DSM-IV criteria for major depressive disorder (criteria essentially unchanged in DSM-5) while retaining its cognitive orientation. Items of the earlier version covering body image, work difficulty, weight loss, and somatic preoccupation were replaced by items on agitation, worthlessness, concentration difficulty, and loss of energy; coverage of atypical symptoms such as increased sleep and increased appetite was added; and the reference period was extended from one week to two weeks to match the diagnostic timeframe (Dozois, Dobson, & Ahnberg, 1998; Wang & Gorenstein, 2013).

Reviews of the original inventory document both strengths, including high internal consistency and reliable differentiation of depressed from nondepressed groups, and weaknesses, including instability of scores over short intervals and doubtful objectivity of interpretation (Richter, Werner, Heerlein, Kraus, & Sauer, 1998). In psychiatric outpatients who completed both versions, the BDI-II showed slightly higher internal consistency than the BDI-IA (α = .91 versus .89) and yielded total scores about two points higher (Beck, Steer, Ball, & Ranieri, 1996).

🎯 Key insight: The BDI-II is a severity index, not a diagnostic test; it grades the intensity of the depressive syndrome through Beck’s cognitive lens, which is why the same instrument serves trials as an outcome measure and cognitive therapists as a map of treatment targets.

Key Features

Assessment Characteristics

  • 21 items, each a set of graded statements describing one symptom at increasing severity
  • Scored 0-3 per item; total score 0-63, higher scores indicating greater severity
  • 5-10 minutes to complete (the publisher lists 5 minutes)
  • Ages 13 through 80, per the publisher’s specification
  • Two-week reference period, including today, matching the DSM diagnostic timeframe
  • Commercially licensed through Pearson; purchase and user qualification required

Dimensions Assessed

Cognitive-affective symptoms: sadness, pessimism, past failure, loss of pleasure, guilty feelings, punishment feelings, self-dislike, self-criticalness, suicidal thoughts or wishes, crying, agitation, loss of interest, indecisiveness, worthlessness

Somatic symptoms: loss of energy, changes in sleeping pattern, irritability, changes in appetite, concentration difficulty, tiredness or fatigue, loss of interest in sex

This grouping reflects the two-factor (cognitive-affective and somatic) solutions reported in the psychometric literature; the exact assignment of items to factors varies by study and sample (Dozois et al., 1998; Wang & Gorenstein, 2013), and the instrument itself is scored as a single total.

Versions & Adaptations

  • BDI (Beck et al., 1961), the original inventory
  • BDI-IA, the amended intermediate version (1979/1987), compared directly with the BDI-II by Beck, Steer, Ball, and Ranieri (1996)
  • BDI-II (Beck, Steer, & Brown, 1996), the current 21-item version
  • BDI-FastScreen for Medical Patients (BDI-FS) (Beck, Steer, & Brown, 2000), a Pearson-published short form for medical settings, where somatic item overlap complicates the full scale
  • The English BDI-II has been translated into 17 languages (Wang & Gorenstein, 2013); validation quality varies by language and sample

Research Applications

  • Treatment outcome measurement – change scores, responder rates, and remission rates in psychotherapy and medication research
  • Severity assessment and symptom profiling – detailed measurement beyond simple case detection
  • Screening research – with cutoffs selected for the target population rather than a single universal threshold
  • Cross-cultural measurement research – examining the structure and comparability of depressive symptoms across languages and settings

View Testable Demo

â–º Click here to try the Testable implementation

Assess depression severity across cognitive, affective, and somatic symptom domains.

Scoring and Interpretation

Response Format

Each of the 21 items presents four statements describing one symptom at increasing levels of severity, scored 0 to 3. Respondents select the single statement that best describes how they have felt during the past two weeks, including today. Items 16 (changes in sleeping pattern) and 18 (changes in appetite) instead present seven options that distinguish increases from decreases (for example, sleeping more versus sleeping less); one option is selected and the item is still scored 0-3.

Item Content

The BDI-II is a commercial instrument, so its statements are not reproduced here. Each item is identified by a symptom label; the 21 labels are: sadness, pessimism, past failure, loss of pleasure, guilty feelings, punishment feelings, self-dislike, self-criticalness, suicidal thoughts or wishes, crying, agitation, loss of interest, indecisiveness, worthlessness, loss of energy, changes in sleeping pattern, irritability, changes in appetite, concentration difficulty, tiredness or fatigue, and loss of interest in sex. As an illustration of the graded format, the sadness item runs from a denial of sadness at 0 to persistent sadness the respondent feels unable to bear at 3.

Scoring Procedure

  1. Sum the 21 item scores (each 0-3) for a total score of 0-63.
  2. Higher scores indicate greater depression severity.
  3. Individual items may be inspected for symptom patterns; any endorsement of item 9 (suicidal thoughts or wishes) should prompt a suicide risk assessment.

Severity Classification

Total score Severity level
0-13 Minimal depression
14-19 Mild depression
20-28 Moderate depression
29-63 Severe depression

These are the bands published in the manual (Beck, Steer, & Brown, 1996) and restated by Wang and Gorenstein (2013). Scores of 14 and above fall outside the minimal range and are commonly treated as indicating clinically significant symptoms. For screening decisions, however, optimal cutoffs vary widely by sample and setting: roughly 10-16 in non-clinical, 7-20 in medical, and 19-31 in psychiatric samples (Wang & Gorenstein, 2013). A single universal cutoff should not be assumed.

Meaningful Change

  • Minimal clinically important difference (MCID): a 17.5% reduction in BDI-II score from baseline is the smallest change that patients themselves report as feeling “better”; the threshold rises to about 32% in longer-duration, treatment-resistant depression (Button et al., 2015, p. 3269)
  • Fixed point rules lack support: Button et al. (2015, p. 3270) note that rule-of-thumb absolute differences of 2-3 points have no empirical basis, which is why the ratio-based MCID is preferred
  • Remission: scores in the minimal range (0-13) are often taken to indicate remission, though trial criteria vary
  • Response: a reduction of 50% or more from baseline is a standard responder convention in depression treatment research

Descriptive Sample Statistics

The mean below is descriptive of one published clinical sample, not normative; representative population norms are provided with the licensed Pearson materials.

Sample N M SD Source
Adolescent psychiatric outpatients, ages 12-18 210 18.23 12.74 Steer, Kumar, Ranieri, & Beck (1998)

In that sample, girls scored on average about five points higher than boys (Steer et al., 1998).

Research Evidence and Psychometric Properties

Reliability Evidence

  • Internal consistency (development samples): α = .92 in psychiatric outpatients and α = .93 in college students (Beck, Steer, & Brown, 1996; both values reported in Dozois et al., 1998, p. 84)
  • Test-retest reliability: r = .93 over a one-week interval in an outpatient sample (Beck, Steer, & Brown, 1996; reported in Steer, Kumar, Ranieri, & Beck, 1998, p. 135)
  • Independent samples: α = .91 in 1,022 undergraduates (Dozois et al., 1998, p. 85); α = .89 in 160 college students (Steer & Clark, 1997, p. 133); α = .92 in 210 adolescent psychiatric outpatients aged 12-18 (Steer et al., 1998, p. 127)
  • Reliability generalization: across the 118 studies reviewed, coefficient alpha averaged around 0.90, ranging from 0.83 to 0.96 (Wang & Gorenstein, 2013)

Validity Evidence

Convergent validity:

  • Original BDI: r = .93 between the original BDI and the BDI-II administered to the same respondents, supporting the convergent validity of the revision (Dozois et al., 1998, p. 85)
  • BDI-IA: in psychiatric outpatients, the BDI-II showed α = .91 versus .89 for the BDI-IA and scored about two points higher on average (Beck, Steer, Ball, & Ranieri, 1996)
  • Hamilton Depression Rating Scale: r = .71 with the clinician-rated measure (Beck, Steer, & Brown, 1996)
  • Known-groups: significant differences between depressed and non-depressed groups (Beck, Steer, & Brown, 1996)

Discriminant validity:

  • Anxiety measures: r = .47-.60, showing overlap but distinctiveness (Beck, Steer, & Brown, 1996)
  • Beck Hopelessness Scale: r = .68, an expected relationship between distinct constructs (Beck, Steer, & Brown, 1996)

Factor structure:

  • Two-factor model most common: a cognitive-affective and a somatic(-vegetative) factor; in undergraduates the two-factor solution accounted for 46% of the variance, with the factors correlated r = .60 (Dozois et al., 1998, pp. 86-87)
  • Alternative models: single general-factor and three-factor solutions are also reported in the literature reviewed by Wang and Gorenstein (2013)
  • Total-score use supported: between-factor correlations of 0.49-0.87 across studies support interpreting the single total score (Wang & Gorenstein, 2013)

Screening Performance

  • Discrimination: area under the ROC curve of approximately 0.75 and higher across studies (Wang & Gorenstein, 2013)
  • Cutoffs are sample-dependent: optimal values range 10-16 in non-clinical, 7-20 in medical, and 19-31 in psychiatric samples (Wang & Gorenstein, 2013)

Cross-Cultural Evidence

  • Translations: the English BDI-II has been translated into 17 languages and is used in Europe, the Middle East, Asia, and Latin America (Wang & Gorenstein, 2013)
  • Translation equivalence: psychometric properties are broadly consistent across language versions (Wang & Gorenstein, 2013)

Special Populations

  • Adolescents: internal consistency was high (α = .92) in 210 adolescent psychiatric outpatients aged 12-18 (Steer, Kumar, Ranieri, & Beck, 1998, p. 127); adult cutoffs are often applied with clinical judgment, though adolescent-specific cutoff studies vary
  • Medical populations and older adults: somatic items can be elevated for reasons other than depression, which contributes to the wide variation in optimal cutoffs across samples (Wang & Gorenstein, 2013); the BDI-FastScreen (Beck, Steer, & Brown, 2000) was developed for medical patients for precisely this reason

Usage Guidelines and Applications

Primary Applications

  • Depression severity assessment in mental health and research settings, beyond simple case detection
  • Treatment outcome monitoring as one of the reference-standard self-report measures for tracking therapy and medication response
  • Clinical trials as a primary or secondary endpoint, with change scores, responder rates (50% or greater reduction), or minimal-range remission as outcomes
  • Cognitive therapy practice, where item-level responses identify cognitive targets consistent with Beck’s therapeutic model

Research Design Considerations

  • It is a severity measure, not a diagnostic instrument: pair it with a clinical interview when diagnosis matters
  • Choose and report the cutoff for the target population: published optima span 7-31 depending on sample type (Wang & Gorenstein, 2013)
  • Interpret change against the empirically derived MCID (a 17.5% reduction from baseline; Button et al., 2015) rather than ad hoc point rules
  • Budget for Pearson licensing before data collection; the instrument is not free to reproduce or administer

Cultural Considerations

  • Use a validated translation where one exists (17 languages; Wang & Gorenstein, 2013) and check local validation work before use
  • Cutoffs established in one setting or culture do not transfer automatically; optimal values vary widely by sample (Wang & Gorenstein, 2013)

Limitations and Cautions

  • Not diagnostic alone: a clinical interview is required for a diagnosis of major depressive disorder
  • Commercial licensing: materials must be purchased from Pearson; unauthorized reproduction is infringement
  • Somatic symptom overlap: scores can be inflated in medical illness or chronic pain (Wang & Gorenstein, 2013)
  • Response bias: as a face-valid self-report, it can be influenced by motivation to appear more or less depressed
  • Suicidality item: any endorsement of item 9 should prompt a suicide risk assessment; plan this workflow before administering the scale

Import & Customize Testable Template

â–º Import scale to your Testable account – Add this scale. Modify instructions, edit questions, adjust presentation. Test anyone (including yourself)

â–º Try Testable version – View the full implementation of this scale in Testable.

â–º View detailed implementation guide in Testable – Step by step instructions for complete customization.

â–º Browse other tests and scales in Testable Library – The largest collection of ready-made psychological tests and scales.

Copyright and Usage Responsibility: Check that you have the proper rights and permissions to use this assessment tool in your research. This may include purchasing appropriate licenses, obtaining permissions from authors/copyright holders, or ensuring your usage falls within fair use guidelines.

The BDI-II is copyrighted by Pearson (formerly The Psychological Corporation) and must be purchased for legal use. It requires proper licensing for administration in clinical practice and research settings, including user qualification; Pearson provides the official test materials, scoring keys, normative data, and interpretation guidelines. Unauthorized reproduction or administration constitutes copyright infringement, which is why item statements are not reproduced on this page.

Proper Attribution: When using or referencing this scale, cite the development manual:

  • Beck, A. T., Steer, R. A., & Brown, G. K. (1996). Manual for the Beck Depression Inventory-II. San Antonio, TX: The Psychological Corporation. https://doi.org/10.1037/t00742-000

Academic and Research Use: Educational institutions and research organizations can obtain appropriate licensing for academic use, including student training and research projects. Contact Pearson for specific academic pricing and usage agreements.

References

Primary Development:

  • Beck, A. T., Steer, R. A., & Brown, G. K. (1996). Manual for the Beck Depression Inventory-II. San Antonio, TX: The Psychological Corporation. https://doi.org/10.1037/t00742-000

Original BDI and Lineage:

Psychometric Evaluation:

  • Beck, A. T., Steer, R. A., Ball, R., & Ranieri, W. F. (1996). Comparison of Beck Depression Inventories-IA and -II in psychiatric outpatients. Journal of Personality Assessment, 67(3), 588-597. https://doi.org/10.1207/s15327752jpa6703_13
  • Dozois, D. J. A., Dobson, K. S., & Ahnberg, J. L. (1998). A psychometric evaluation of the Beck Depression Inventory-II. Psychological Assessment, 10(2), 83-89. https://doi.org/10.1037/1040-3590.10.2.83
  • Steer, R. A., & Clark, D. A. (1997). Psychometric characteristics of the Beck Depression Inventory-II with college students. Measurement and Evaluation in Counseling and Development, 30(3), 128-136. https://doi.org/10.1080/07481756.1997.12068933
  • Steer, R. A., Kumar, G., Ranieri, W. F., & Beck, A. T. (1998). Use of the Beck Depression Inventory-II with adolescent psychiatric outpatients. Journal of Psychopathology and Behavioral Assessment, 20(2), 127-137. https://doi.org/10.1023/A:1023091529735
  • Button, K. S., Kounali, D., Thomas, L., Wiles, N. J., Peters, T. J., Welton, N. J., Ades, A. E., & Lewis, G. (2015). Minimal clinically important difference on the Beck Depression Inventory-II according to the patient’s perspective. Psychological Medicine, 45(15), 3269-3279. https://doi.org/10.1017/S0033291715001270

Short Form:

  • Beck, A. T., Steer, R. A., & Brown, G. K. (2000). BDI-FastScreen for Medical Patients. Pearson. (no DOI)

Comprehensive Reviews:

  • Wang, Y. P., & Gorenstein, C. (2013). Psychometric properties of the Beck Depression Inventory-II: A comprehensive review. Revista Brasileira de Psiquiatria, 35(4), 416-431. https://doi.org/10.1590/1516-4446-2012-1048
Illustration of a sad panda sitting hunched over in a misty bamboo forest with head lowered and eyes downcast, surrounded by gray fog, with the Testable logo and text "BDI-II Beck Depression Inventory-II"
A melancholy panda sitting alone in the fog — embodying sadness, hopelessness, and loss of interest measured by the BDI-II (Beck Depression Inventory-II)

Frequently Asked Questions

What does the BDI-II measure?

The BDI-II measures the severity of depressive symptoms in people aged 13 through 80. Its 21 items cover cognitive-affective symptoms (such as sadness, pessimism, guilt, self-dislike, and worthlessness) and somatic symptoms (such as loss of energy, changes in sleep and appetite, and fatigue), each rated for the past two weeks including today, yielding a total score from 0 to 63.

How long does the BDI-II take to complete?

About 5 to 10 minutes; the publisher lists 5 minutes. Each of the 21 items asks the respondent to choose the one statement, from a graded set, that best describes how they have felt during the past two weeks.

Is the BDI-II free to use?

No. The BDI-II is copyrighted by Pearson (formerly The Psychological Corporation) and must be purchased, with user qualification requirements. Unauthorized reproduction or administration constitutes copyright infringement, which is also why the item statements are not reproduced on this page.

How is the BDI-II scored?

Each item is scored 0 to 3 and the 21 items are summed to a total of 0 to 63. The manual's severity bands are 0-13 minimal, 14-19 mild, 20-28 moderate, and 29-63 severe. Any endorsement of item 9 (suicidal thoughts or wishes) should prompt a suicide risk assessment. For screening decisions, optimal cutoffs vary widely by sample type, so a single universal threshold should not be assumed.

How does the BDI-II differ from the original BDI?

The 1996 revision replaced items on body image, work difficulty, weight loss, and somatic preoccupation with items on agitation, worthlessness, concentration difficulty, and loss of energy, added coverage of increased sleep and appetite, and extended the reference period from one week to two weeks to match DSM-IV criteria. Scores on the two versions correlate at .93, and the BDI-II runs about two points higher than the BDI-IA in psychiatric outpatients.

Does the BDI-II have subscales?

The instrument is scored as a single total. Psychometric studies most often report a two-factor structure, with a cognitive-affective and a somatic factor, but the assignment of items to factors varies by study and sample, and high between-factor correlations support interpreting the total score.

What counts as meaningful change on the BDI-II?

The empirically derived minimal clinically important difference is a 17.5% reduction in score from baseline, rising to about 32% in longer-duration, treatment-resistant depression (Button et al., 2015). A reduction of 50% or more is a standard responder convention, and scores in the minimal range (0-13) are often taken to indicate remission, though trial criteria vary. Fixed point rules of 2-3 points have no empirical support.

Frequently Asked Questions

What does the BDI-II measure?

The BDI-II measures the severity of depressive symptoms in people aged 13 through 80. Its 21 items cover cognitive-affective symptoms (such as sadness, pessimism, guilt, self-dislike, and worthlessness) and somatic symptoms (such as loss of energy, changes in sleep and appetite, and fatigue), each rated for the past two weeks including today, yielding a total score from 0 to 63.

How long does the BDI-II take to complete?

About 5 to 10 minutes; the publisher lists 5 minutes. Each of the 21 items asks the respondent to choose the one statement, from a graded set, that best describes how they have felt during the past two weeks.

Is the BDI-II free to use?

No. The BDI-II is copyrighted by Pearson (formerly The Psychological Corporation) and must be purchased, with user qualification requirements. Unauthorized reproduction or administration constitutes copyright infringement, which is also why the item statements are not reproduced on this page.

How is the BDI-II scored?

Each item is scored 0 to 3 and the 21 items are summed to a total of 0 to 63. The manual's severity bands are 0-13 minimal, 14-19 mild, 20-28 moderate, and 29-63 severe. Any endorsement of item 9 (suicidal thoughts or wishes) should prompt a suicide risk assessment. For screening decisions, optimal cutoffs vary widely by sample type, so a single universal threshold should not be assumed.

How does the BDI-II differ from the original BDI?

The 1996 revision replaced items on body image, work difficulty, weight loss, and somatic preoccupation with items on agitation, worthlessness, concentration difficulty, and loss of energy, added coverage of increased sleep and appetite, and extended the reference period from one week to two weeks to match DSM-IV criteria. Scores on the two versions correlate at .93, and the BDI-II runs about two points higher than the BDI-IA in psychiatric outpatients.

Does the BDI-II have subscales?

The instrument is scored as a single total. Psychometric studies most often report a two-factor structure, with a cognitive-affective and a somatic factor, but the assignment of items to factors varies by study and sample, and high between-factor correlations support interpreting the total score.

What counts as meaningful change on the BDI-II?

The empirically derived minimal clinically important difference is a 17.5% reduction in score from baseline, rising to about 32% in longer-duration, treatment-resistant depression (Button et al., 2015). A reduction of 50% or more is a standard responder convention, and scores in the minimal range (0-13) are often taken to indicate remission, though trial criteria vary. Fixed point rules of 2-3 points have no empirical support.
Last Updated: