HDRS/HAM-D: Hamilton Depression Rating Scale

Reviewed by: Constantin Rezlescu | Associate Professor | UCL Psychology

TL;DR

  • The Hamilton Depression Rating Scale (HDRS/HAM-D) is a clinician-administered interview measure of depression severity, developed by Max Hamilton in 1960 to quantify change in patients already diagnosed with depression rather than to diagnose it.
  • It became the gold-standard outcome measure of antidepressant clinical trials through a gradual consensus among trialists; regulatory bodies later accepted it as a widely used endpoint rather than mandating it.
  • Total-score reliability is generally high and convergent and discriminant validity are adequate, but item-level reliability, content validity, and the scale's factor structure have drawn sustained, well-documented criticism.
  • Its heavy weighting of somatic and behavioural symptoms can inflate scores in medically ill patients, and structured versions such as the SIGH-D and GRID-HAMD were developed to standardize administration.

At a Glance

Items 17 scored items (HDRS-17, the most widely used form); a 21-item version adds 4 items intended to subtype, not to score severity
Administration time 15-20 minutes, structured clinical interview
Response format Clinician ratings after interview; nine items rated 0-4 and eight items rated 0-2
Scores Single total severity score, range 0-52 (HDRS-17); no formal subscales, though 6- and 7-item core-symptom short forms exist
Validated populations Adults; originally developed for patients already diagnosed with depression
License Public domain; structured derivatives (SIGH-D, GRID-HAMD) may carry separate copyright
Original citation Hamilton (1960), Journal of Neurology, Neurosurgery and Psychiatry

Introduction

The Hamilton Depression Rating Scale (HDRS), also known as the Hamilton Depression Scale (HAM-D), is a clinician-administered measure of depression severity developed by German-born British psychiatrist Max Hamilton (Hamilton, 1960). For over four decades it served as the “gold standard” for measuring depression severity in clinical trials and pharmaceutical research, shaping how the efficacy of depression treatments is evaluated (Bagby et al., 2004; Worboys, 2013).

Understanding Clinician-Rated Depression Severity

Hamilton created the scale to enable psychiatrists to chart changes in already diagnosed patients through particular treatment regimes, converting qualitative clinical judgments into quantitative data (Worboys, 2013, p. 204-205). It was one of the first and most influential clinician-administered tools for measuring treatment response, though it was born into a world of already competing rating scales (Worboys, 2013, p. 204). Its dominance emerged gradually and “from below,” from an emerging consensus among psychiatrists undertaking clinical trials for depression, which from the 1960s were principally trials of psychopharmaceuticals (Worboys, 2013, abstract and p. 217). Unlike self-report questionnaires, the HDRS requires trained clinical judgment and structured interview technique, which is central to its role in pharmaceutical research.

Theoretical Foundation

The HDRS predates modern diagnostic systems such as the DSM-5, emerging from Hamilton’s clinical experience and the prevailing understanding of depression in the late 1950s. Rather than mapping onto specific diagnostic criteria, the scale reflects a broad clinical assessment approach that emphasizes observable signs and symptoms across several descriptive domains:

  • Mood and affect – depressed mood, feelings of guilt, suicidal ideation
  • Psychomotor and cognitive features – agitation, retardation, loss of insight
  • Somatic symptoms – insomnia patterns, appetite and weight changes, loss of energy
  • Anxiety components – psychological anxiety, somatic anxiety, hypochondriasis
  • Behavioral functioning – work and activities, sexual interest, general somatic symptoms

The clinician-administered format was intentional: the HDRS relied mainly on clinician observation of bodily (somatic) and behavioural features, which Hamilton weighted more heavily than the few items based on patients’ own reports of their feelings (Worboys, 2013, p. 207).

🏥 Key insight: The HDRS remains the most widely used primary outcome measure in depression clinical trials, and is a widely accepted, not mandated, primary endpoint in antidepressant studies reviewed by regulatory agencies including the FDA. When the scale was sanctioned “from above” in the 1980s by the WHO, the FDA, and other licensing agencies, this acknowledged its widespread use rather than imposing it “top down” (Worboys, 2013, p. 217-218).

Key Features

Assessment Characteristics

  • 17 scored items in the standard HDRS-17; versions with 17-21 or more items exist
  • 15-20 minutes administration time by a trained clinician
  • Adult populations, originally designed for patients already diagnosed with depression
  • Mixed item scales: nine items rated 0-4 and eight items rated 0-2
  • Past-week timeframe: ratings cover symptoms over the previous week
  • Clinician-administered, requiring structured interview skills

Dimensions Assessed

The following groupings are descriptive; they are not empirically derived subscales, and the scale’s factor structure has not replicated consistently across samples (Bagby et al., 2004, p. 2173-2174).

  • Mood evaluation – core depressive symptoms and guilt
  • Sleep disturbances – three insomnia items covering early, middle, and late sleep difficulties
  • Somatic symptoms – physical manifestations and energy levels
  • Anxiety components – psychological and somatic anxiety features
  • Behavioral functioning – work capacity and activity engagement
  • Suicidal ideation – severity of thoughts of death and suicide

Versions & Adaptations

  • HDRS-17 – the original 17 scored items (Hamilton, 1960); the standard form for severity measurement
  • HDRS-21 – adds 4 items intended to subtype the depression, not to be summed into the severity score
  • HDRS-7 – a seven-item shortened version developed for efficiency
  • Maier subscale – six core depression items extracted from the HDRS-17
  • Bech melancholia subscale – a six-item severity scale focused on core symptoms
  • SIGH-D – the Structured Interview Guide for the HDRS, which standardizes administration and improved item test-retest reliability (Williams, 1988)
  • SIGH-SAD – an extended structured version adding atypical depression symptoms
  • GRID-HAMD – a standardized version with improved anchor points and training materials

Research Applications

  • Pharmaceutical trials – primary efficacy endpoint in antidepressant studies
  • Clinical research – established comparator for validating new depression measures
  • Treatment monitoring – baseline and outcome assessment in clinical settings
  • Regulatory submissions – widely used and accepted endpoint for demonstrating drug efficacy in depression
  • International research – a standardized measure used across global clinical trials

View Testable Demo

â–ş Click here to try the Testable implementation

Conduct standardized clinician-administered depression severity assessment for clinical research participants.

Scoring and Interpretation

Response Format

Each item is scored by the clinician on the basis of a structured interview covering the previous week’s symptoms. Items use either 3-point scales (0-2) or 5-point scales (0-4) depending on the symptom being assessed; the clinician rates severity from patient responses and clinical observation.

Illustrative Items

The HDRS is in the public domain, so the full form may be reproduced freely; three items illustrate the anchor structure.

Item 1: Depressed Mood (rated 0-4)

  • 0: Absent
  • 1: These feeling states indicated only on questioning
  • 2: These feeling states spontaneously reported verbally
  • 3: Communicates feeling states non-verbally through facial expression, posture, voice, tendency to weep
  • 4: Patient reports virtually only these feeling states in spontaneous verbal and non-verbal communication

Item 3: Suicide (rated 0-4)

  • 0: Absent
  • 1: Feels life is not worth living
  • 2: Wishes he were dead or any thoughts of possible death to self
  • 3: Suicidal ideas or gesture
  • 4: Attempts at suicide (any serious attempt rates 4)

Item 4: Insomnia Early (rated 0-2)

  • 0: No difficulty falling asleep
  • 1: Complains of occasional difficulty falling asleep—i.e., more than 1/2 hour
  • 2: Complains of nightly difficulty falling asleep

Scoring Procedure

  1. Rate each of the 17 items from the structured interview covering the previous week’s symptoms.
  2. Apply the item scales: nine items are rated 0-4 and eight items are rated 0-2.
  3. Sum the 17 item ratings for a total severity score of 0-52.
  4. If the 21-item form is used, the 4 additional items serve to subtype the depression and are not added to the severity total.

HDRS-17 Severity Interpretation

Total Score Severity Level Clinical Interpretation
0-7 Normal/Remission No depression or successful treatment response
8-13 Mild depression Mild symptoms requiring monitoring
14-18 Moderate depression Clinically significant depression warranting treatment
19-22 Severe depression Substantial depression requiring active intervention
≥23 Very severe depression Severe depression requiring intensive treatment

These bands are the widely used convention in depression research; they do not derive from a single primary publication and should be reported as conventional rather than empirically anchored.

Clinical Thresholds and Research Standards

  • Entry criteria: a score ≥20 is typically required for clinical trial participation
  • Treatment response: ≥50% reduction from the baseline score (widely used trial convention)
  • Remission: a score ≤7, indicating minimal residual symptoms (widely used convention)

Population Norms

No normative sample means (M, SD) for the HDRS-17 were identified during verification, and none are presented here. Interpretation in practice rests on the conventional severity bands and trial thresholds above, not on population norms.

Research Evidence and Psychometric Properties

Reliability Evidence

Internal consistency:

  • Cronbach’s α range: 0.46-0.97 across the studies reviewed, with 10 studies reporting estimates ≥0.70, indicating variable and often inadequate internal consistency (Bagby et al., 2004, p. 2165, Table 2)
  • Pooled internal consistency: α = 0.79 (pooled estimate 0.789, 95% CI 0.766-0.810) in a meta-analysis spanning 49 years, with substantial heterogeneity across studies (Trajković et al., 2011)

Inter-rater reliability:

  • Total score inter-rater reliability: Pearson’s r ranged from 0.82 to 0.98 and intraclass r from 0.46 to 0.99 across studies for the total 17-item scale (Bagby et al., 2004, p. 2165-2166, Table 2); the pooled meta-analytic inter-rater ICC is 0.937 (95% CI 0.914-0.954) (Trajković et al., 2011)
  • Item-level agreement: poor for many individual items, particularly those requiring subjective clinical judgment; loss of insight showed the lowest inter-rater agreement (Bagby et al., 2004, p. 2166)
  • Structured interview reliability: the SIGH-D structured interview guide improved test-retest reliability of HDRS items, raising the mean item-level retest reliability to 0.54, though only four items then met the criterion for adequate reliability (Williams, 1988, as reviewed in Bagby et al., 2004, p. 2167)

Test-retest reliability:

  • Total-scale stability: retest reliability for the total scale ranged from 0.81 to 0.98 across the studies Bagby et al. reviewed (Bagby et al., 2004, p. 2166, Table 2); the reliability meta-analysis reports retest reliability of 0.65-0.98, declining over longer retest intervals (Trajković et al., 2011)
  • Item-level stability: poor for individual items over time, ranging from 0.00 to 0.85 (Bagby et al., 2004, p. 2166-2167, Table 3)

Validity Evidence

Convergent validity:

  • Correlation with other depression measures: convergent validity is adequate, meeting the r ≥ 0.50 criterion in correlations with all but two comparison scales; across studies, HDRS correlations with the Beck Depression Inventory ranged widely (about 0.27-0.89) and with the Montgomery-Ă…sberg Depression Rating Scale were 0.68-0.88 (Bagby et al., 2004, p. 2171, Table 4)
  • Agreement with clinical ratings: good correspondence with clinician assessments of depression severity (Bagby et al., 2004, p. 2171)

Discriminant validity:

  • Depression vs. non-depression: adequate ability to distinguish depressed from non-depressed patients (mean cutoff 12.6/13.5; mean sensitivity 0.76, specificity 0.91) (Bagby et al., 2004, p. 2171, Table 5)
  • Overlap with anxiety: the HDRS correlates more highly with depression than with anxiety measures, but its saturation with anxiety-related concepts is nonetheless considerable (Bagby et al., 2004, p. 2172)

Predictive validity:

  • Treatment outcomes: predictive validity meets established criteria, though it suffers somewhat due to the scale’s multidimensionality; in meta-analyses the HDRS was more sensitive to treatment change than the Beck Depression Inventory and Zung self-report scales (Bagby et al., 2004, p. 2172-2174)

Content validity concerns:

  • Sleep symptom overemphasis: content validity is poor; the scale includes three insomnia items while burying or omitting important features of contemporary depression such as concentration difficulties and feelings of worthlessness (Bagby et al., 2004, p. 2170-2171)
  • Item discrimination: loss of insight showed the most variable and poorest performance; several items (loss of insight, hypochondriasis, somatic anxiety, general somatic symptoms, suicide, psychomotor agitation) were identified as problematic, and the loss-of-libido item showed high rater disagreement (Bagby et al., 2004, p. 2165-2169)

Factor Structure

  • Not unidimensional: the HDRS is clearly not unidimensional; the original single-factor assumption is not supported by the empirical data (Bagby et al., 2004, p. 2173)
  • Multi-factor solutions: across 15 studies (17 samples) the number of factors identified ranged from two to eight; recurring factors included an insomnia/sleep factor (13 data sets), a general depression factor (depressed mood, guilt, suicide), and an anxiety/agitation factor (Bagby et al., 2004, p. 2173, Table 6)
  • Cross-sample variability: the factor structure fails to replicate a single unifying structure across studies and samples (Bagby et al., 2004, p. 2173-2174)
  • Construct clarity: the exact structure of the scale’s multidimensionality remains unclear, and the meaning of the total score is correspondingly uncertain (Bagby et al., 2004, p. 2173)

Treatment Sensitivity

  • Medication trial sensitivity: sensitive to change in antidepressant trials; in meta-analyses the HDRS detected treatment change as well as or better than the Beck Depression Inventory and Zung self-report scales (Bagby et al., 2004, p. 2172-2173)
  • Clinically meaningful change: detects treatment-related change over multi-week trials, though unidimensional core-symptom subscales were often at least as sensitive to change as the full scale (Bagby et al., 2004, p. 2172-2173)
  • Response criteria: a ≥50% reduction from baseline is the standard trial definition of response (widely used convention)
  • Remission criteria: a score ≤7 is the established remission criterion in depression research (widely used convention)

Psychometric Limitations

  • Variable internal consistency: inconsistent reliability across studies and populations, with many studies showing α < 0.80 (Bagby et al., 2004, p. 2165)
  • Uneven item quality: many items are poor contributors to the measurement of depression severity, and some sum into a total score whose meaning is multidimensional and unclear (Bagby et al., 2004, p. 2173-2174)
  • Factor structure inconsistency: failure to replicate a single unifying factor structure across different samples (Bagby et al., 2004, p. 2173-2174)
  • Somatic symptom emphasis: the scale relies heavily on somatic and behavioural features, which may inflate scores in medically ill populations (Bagby et al., 2004; Worboys, 2013, p. 207)
  • Standardization challenges: some items require clinical inferences that are difficult to standardize across raters, motivating the development of structured interview guides (Williams, 1988)

Despite these limitations, the scale continues to be used due to historical precedent, regulatory acceptance, and its accumulated evidence base (Worboys, 2013).

Usage Guidelines and Applications

Primary Applications

  • Pharmaceutical clinical trials – primary efficacy endpoint in antidepressant development programs
  • Regulatory submissions – widely used and accepted endpoint for demonstrating drug efficacy in depression
  • Academic research – established comparator for validating new depression assessment tools
  • Clinical settings – depression severity assessment and treatment monitoring in psychiatric outpatient and inpatient care
  • Training programs – educational tool for developing clinical interviewing skills

Research Design Considerations

  • Rater training: formal training in HDRS administration and scoring, with ongoing calibration to prevent rater drift, is essential for reliable use
  • Assessment protocol: a 15-20 minute structured clinical interview covering the previous week’s symptoms, with documentation supporting each item score
  • Sample selection: clinical trials typically require a baseline HDRS score ≥20 for entry
  • Outcome definitions: use the established response (≥50% reduction) and remission (score ≤7) conventions and report them explicitly
  • Quality control: double-rating and consensus procedures are recommended for research applications
  • Standardization: structured versions (SIGH-D, GRID-HAMD) reduce administration variability and are preferable where rater consistency is critical (Williams, 1988)

Cultural and Population Considerations

  • Medically ill populations: the scale’s heavy weighting of somatic and behavioural features may inflate scores where physical illness produces overlapping symptoms (Bagby et al., 2004; Worboys, 2013, p. 207)
  • Cross-sample generalization: the factor structure does not replicate consistently across studies and samples, so subscale-style interpretations should not be transported between populations (Bagby et al., 2004, p. 2173-2174)
  • Intended population: the scale was designed for adults already diagnosed with depression; it is a severity measure, not a diagnostic instrument

Limitations and Considerations

  • Training requirements – extensive clinician training is needed for reliable administration
  • Time intensive – the 15-20 minute clinician interview limits routine clinical use
  • Somatic emphasis – may overestimate depression severity in medically ill patients
  • Item limitations – several items show poor psychometric properties and require careful interpretation (Bagby et al., 2004)

Import & Customize Testable Template

â–ş Import scale to your Testable account – Add this scale. Modify instructions, edit questions, adjust presentation. Test anyone (including yourself)

â–ş Try Testable version – View the full implementation of this scale in Testable.

â–ş View detailed implementation guide in Testable – Step by step instructions for complete customization.

â–ş Browse other tests and scales in Testable Library – The largest collection of ready-made psychological tests and scales.

Copyright and Usage Responsibility: Check that you have the proper rights and permissions to use this assessment tool in your research. This may include purchasing appropriate licenses, obtaining permissions from authors/copyright holders, or ensuring your usage falls within fair use guidelines.

The Hamilton Depression Rating Scale was developed by Max Hamilton and first published in 1960. The original scale is in the public domain: public-domain reproductions state explicitly that the HDRS is in the public domain and that no permission is required to use the instrument. It may be reproduced and administered without a license.

Proper Attribution: When using or referencing this scale, cite the original development:

Structured versions: The Structured Interview Guide for the Hamilton Depression Rating Scale (SIGH-D) and other derivatives (such as GRID-HAMD) may carry separate copyright. Users should verify licensing requirements for specific structured versions before reproducing them.

Training materials: Commercial rater-training and certification materials may require separate licensing agreements even though the underlying scale is public domain.

References

Original Development:

Comprehensive Psychometric Review:

  • Bagby, R. M., Ryder, A. G., Schuller, D. R., & Marshall, M. B. (2004). The Hamilton Depression Rating Scale: Has the gold standard become a lead weight? American Journal of Psychiatry, 161(12), 2163-2177. https://doi.org/10.1176/appi.ajp.161.12.2163

Reliability Meta-Analysis:

  • Trajković, G., StarÄŤević, V., Latas, M., Leštarević, M., Ille, T., Bukumirić, Z., & Marinković, J. (2011). Reliability of the Hamilton Rating Scale for Depression: A meta-analysis over a period of 49 years. Psychiatry Research, 189(1), 1-9. https://doi.org/10.1016/j.psychres.2010.12.007

Structured Interview Development:

Historical Analysis:

  • Worboys, M. (2013). The Hamilton Rating Scale for Depression: The making of a “gold standard” and the unmaking of a chronic illness, 1960–1980. Chronic Illness, 9(3), 202-219. https://doi.org/10.1177/1742395312467658

Related Assessments: BDI-II: Beck Depression Inventory-II and MADRS: Montgomery-Ă…sberg Depression Rating Scale

Frequently Asked Questions

What does the HDRS/HAM-D measure?

The Hamilton Depression Rating Scale measures the severity of depression in patients who have already been diagnosed; it is not a diagnostic instrument. A clinician rates 17 items spanning mood, guilt, suicidal ideation, insomnia, work and activities, psychomotor features, anxiety, and somatic symptoms, based on a structured interview covering the previous week.

Who can administer the HDRS/HAM-D, and how long does it take?

The scale is clinician-administered and requires trained clinical judgment and structured interviewing skills; it is not a self-report questionnaire. Administration takes approximately 15-20 minutes, and formal rater training with ongoing calibration is recommended for research use.

Is the HDRS/HAM-D free to use?

Yes. The original scale is in the public domain, and public-domain reproductions state that no permission is required to use it. Structured derivatives such as the SIGH-D and GRID-HAMD may carry separate copyright, so licensing should be verified before reproducing those versions. Cite Hamilton (1960) when using the scale.

How is the HDRS/HAM-D scored?

The HDRS-17 is scored by summing ratings across 17 items; nine items are rated 0-4 and eight items are rated 0-2, giving a total range of 0-52. Conventional severity bands are 0-7 (normal/remission), 8-13 (mild), 14-18 (moderate), 19-22 (severe), and 23 or higher (very severe). Treatment response is conventionally defined as a reduction of at least 50% from baseline, and remission as a score of 7 or lower.

What is the difference between the HDRS-17 and the HDRS-21?

The HDRS-17 contains the 17 scored items used to measure severity. The 21-item version adds four items intended to subtype the depression; those extra items are not added to the severity total, so the HDRS-17 remains the standard form for scoring.

How does the HDRS/HAM-D differ from the Beck Depression Inventory (BDI)?

The HDRS is clinician-administered and emphasizes observable signs and somatic symptoms, whereas the BDI is a self-report questionnaire focused on subjective cognitive and emotional experience. Reported correlations between the two vary widely across studies, from roughly 0.27 to 0.89 (Bagby et al., 2004). The HDRS is the standard clinical-trial outcome measure; the BDI is more common in screening and self-monitoring.

How reliable is the HDRS/HAM-D?

Reliability is strong at the total-score level but weak at the item level. Inter-rater reliability for total scores is generally high (Pearson's r 0.82-0.98 across studies, pooled ICC 0.937 in the Trajković et al., 2011 meta-analysis), while internal consistency varies widely (α = 0.46-0.97 across studies, pooled α = 0.79). Many individual items show poor inter-rater and retest reliability, and the factor structure is inconsistent across samples (Bagby et al., 2004).
Last Updated: