The Hamilton Depression Rating Scale (HDRS/HAM-D) is a clinician-administered interview measure of depression severity, developed by Max Hamilton in 1960 to quantify change in patients already diagnosed with depression rather than to diagnose it.
It became the gold-standard outcome measure of antidepressant clinical trials through a gradual consensus among trialists; regulatory bodies later accepted it as a widely used endpoint rather than mandating it.
Total-score reliability is generally high and convergent and discriminant validity are adequate, but item-level reliability, content validity, and the scale's factor structure have drawn sustained, well-documented criticism.
Its heavy weighting of somatic and behavioural symptoms can inflate scores in medically ill patients, and structured versions such as the SIGH-D and GRID-HAMD were developed to standardize administration.
At a Glance
Items
17 scored items (HDRS-17, the most widely used form); a 21-item version adds 4 items intended to subtype, not to score severity
Administration time
15-20 minutes, structured clinical interview
Response format
Clinician ratings after interview; nine items rated 0-4 and eight items rated 0-2
Scores
Single total severity score, range 0-52 (HDRS-17); no formal subscales, though 6- and 7-item core-symptom short forms exist
Validated populations
Adults; originally developed for patients already diagnosed with depression
License
Public domain; structured derivatives (SIGH-D, GRID-HAMD) may carry separate copyright
Original citation
Hamilton (1960), Journal of Neurology, Neurosurgery and Psychiatry
Introduction
The Hamilton Depression Rating Scale (HDRS), also known as the Hamilton Depression Scale (HAM-D), is a clinician-administered measure of depression severity developed by German-born British psychiatrist Max Hamilton (Hamilton, 1960). For over four decades it served as the “gold standard” for measuring depression severity in clinical trials and pharmaceutical research, shaping how the efficacy of depression treatments is evaluated (Bagby et al., 2004; Worboys, 2013).
Understanding Clinician-Rated Depression Severity
Hamilton created the scale to enable psychiatrists to chart changes in already diagnosed patients through particular treatment regimes, converting qualitative clinical judgments into quantitative data (Worboys, 2013, p. 204-205). It was one of the first and most influential clinician-administered tools for measuring treatment response, though it was born into a world of already competing rating scales (Worboys, 2013, p. 204). Its dominance emerged gradually and “from below,” from an emerging consensus among psychiatrists undertaking clinical trials for depression, which from the 1960s were principally trials of psychopharmaceuticals (Worboys, 2013, abstract and p. 217). Unlike self-report questionnaires, the HDRS requires trained clinical judgment and structured interview technique, which is central to its role in pharmaceutical research.
Theoretical Foundation
The HDRS predates modern diagnostic systems such as the DSM-5, emerging from Hamilton’s clinical experience and the prevailing understanding of depression in the late 1950s. Rather than mapping onto specific diagnostic criteria, the scale reflects a broad clinical assessment approach that emphasizes observable signs and symptoms across several descriptive domains:
Mood and affect – depressed mood, feelings of guilt, suicidal ideation
Psychomotor and cognitive features – agitation, retardation, loss of insight
Somatic symptoms – insomnia patterns, appetite and weight changes, loss of energy
Behavioral functioning – work and activities, sexual interest, general somatic symptoms
The clinician-administered format was intentional: the HDRS relied mainly on clinician observation of bodily (somatic) and behavioural features, which Hamilton weighted more heavily than the few items based on patients’ own reports of their feelings (Worboys, 2013, p. 207).
🏥 Key insight: The HDRS remains the most widely used primary outcome measure in depression clinical trials, and is a widely accepted, not mandated, primary endpoint in antidepressant studies reviewed by regulatory agencies including the FDA. When the scale was sanctioned “from above” in the 1980s by the WHO, the FDA, and other licensing agencies, this acknowledged its widespread use rather than imposing it “top down” (Worboys, 2013, p. 217-218).
Key Features
Assessment Characteristics
17 scored items in the standard HDRS-17; versions with 17-21 or more items exist
15-20 minutes administration time by a trained clinician
Adult populations, originally designed for patients already diagnosed with depression
Mixed item scales: nine items rated 0-4 and eight items rated 0-2
Past-week timeframe: ratings cover symptoms over the previous week
The following groupings are descriptive; they are not empirically derived subscales, and the scale’s factor structure has not replicated consistently across samples (Bagby et al., 2004, p. 2173-2174).
Mood evaluation – core depressive symptoms and guilt
Sleep disturbances – three insomnia items covering early, middle, and late sleep difficulties
Somatic symptoms – physical manifestations and energy levels
Anxiety components – psychological and somatic anxiety features
Behavioral functioning – work capacity and activity engagement
Suicidal ideation – severity of thoughts of death and suicide
Versions & Adaptations
HDRS-17 – the original 17 scored items (Hamilton, 1960); the standard form for severity measurement
HDRS-21 – adds 4 items intended to subtype the depression, not to be summed into the severity score
HDRS-7 – a seven-item shortened version developed for efficiency
Maier subscale – six core depression items extracted from the HDRS-17
Bech melancholia subscale – a six-item severity scale focused on core symptoms
SIGH-D – the Structured Interview Guide for the HDRS, which standardizes administration and improved item test-retest reliability (Williams, 1988)
SIGH-SAD – an extended structured version adding atypical depression symptoms
GRID-HAMD – a standardized version with improved anchor points and training materials
Research Applications
Pharmaceutical trials – primary efficacy endpoint in antidepressant studies
Clinical research – established comparator for validating new depression measures
Treatment monitoring – baseline and outcome assessment in clinical settings
Regulatory submissions – widely used and accepted endpoint for demonstrating drug efficacy in depression
International research – a standardized measure used across global clinical trials
Conduct standardized clinician-administered depression severity assessment for clinical research participants.
Scoring and Interpretation
Response Format
Each item is scored by the clinician on the basis of a structured interview covering the previous week’s symptoms. Items use either 3-point scales (0-2) or 5-point scales (0-4) depending on the symptom being assessed; the clinician rates severity from patient responses and clinical observation.
Illustrative Items
The HDRS is in the public domain, so the full form may be reproduced freely; three items illustrate the anchor structure.
Item 1: Depressed Mood (rated 0-4)
0: Absent
1: These feeling states indicated only on questioning
2: These feeling states spontaneously reported verbally
3: Communicates feeling states non-verbally through facial expression, posture, voice, tendency to weep
4: Patient reports virtually only these feeling states in spontaneous verbal and non-verbal communication
Item 3: Suicide (rated 0-4)
0: Absent
1: Feels life is not worth living
2: Wishes he were dead or any thoughts of possible death to self
3: Suicidal ideas or gesture
4: Attempts at suicide (any serious attempt rates 4)
Item 4: Insomnia Early (rated 0-2)
0: No difficulty falling asleep
1: Complains of occasional difficulty falling asleep—i.e., more than 1/2 hour
2: Complains of nightly difficulty falling asleep
Scoring Procedure
Rate each of the 17 items from the structured interview covering the previous week’s symptoms.
Apply the item scales: nine items are rated 0-4 and eight items are rated 0-2.
Sum the 17 item ratings for a total severity score of 0-52.
If the 21-item form is used, the 4 additional items serve to subtype the depression and are not added to the severity total.
Substantial depression requiring active intervention
≥23
Very severe depression
Severe depression requiring intensive treatment
These bands are the widely used convention in depression research; they do not derive from a single primary publication and should be reported as conventional rather than empirically anchored.
Clinical Thresholds and Research Standards
Entry criteria: a score ≥20 is typically required for clinical trial participation
Treatment response: ≥50% reduction from the baseline score (widely used trial convention)
Remission: a score ≤7, indicating minimal residual symptoms (widely used convention)
Population Norms
No normative sample means (M, SD) for the HDRS-17 were identified during verification, and none are presented here. Interpretation in practice rests on the conventional severity bands and trial thresholds above, not on population norms.
Research Evidence and Psychometric Properties
Reliability Evidence
Internal consistency:
Cronbach’s α range: 0.46-0.97 across the studies reviewed, with 10 studies reporting estimates ≥0.70, indicating variable and often inadequate internal consistency (Bagby et al., 2004, p. 2165, Table 2)
Pooled internal consistency: α = 0.79 (pooled estimate 0.789, 95% CI 0.766-0.810) in a meta-analysis spanning 49 years, with substantial heterogeneity across studies (Trajković et al., 2011)
Inter-rater reliability:
Total score inter-rater reliability: Pearson’s r ranged from 0.82 to 0.98 and intraclass r from 0.46 to 0.99 across studies for the total 17-item scale (Bagby et al., 2004, p. 2165-2166, Table 2); the pooled meta-analytic inter-rater ICC is 0.937 (95% CI 0.914-0.954) (Trajković et al., 2011)
Item-level agreement: poor for many individual items, particularly those requiring subjective clinical judgment; loss of insight showed the lowest inter-rater agreement (Bagby et al., 2004, p. 2166)
Structured interview reliability: the SIGH-D structured interview guide improved test-retest reliability of HDRS items, raising the mean item-level retest reliability to 0.54, though only four items then met the criterion for adequate reliability (Williams, 1988, as reviewed in Bagby et al., 2004, p. 2167)
Test-retest reliability:
Total-scale stability: retest reliability for the total scale ranged from 0.81 to 0.98 across the studies Bagby et al. reviewed (Bagby et al., 2004, p. 2166, Table 2); the reliability meta-analysis reports retest reliability of 0.65-0.98, declining over longer retest intervals (Trajković et al., 2011)
Item-level stability: poor for individual items over time, ranging from 0.00 to 0.85 (Bagby et al., 2004, p. 2166-2167, Table 3)
Validity Evidence
Convergent validity:
Correlation with other depression measures: convergent validity is adequate, meeting the r ≥ 0.50 criterion in correlations with all but two comparison scales; across studies, HDRS correlations with the Beck Depression Inventory ranged widely (about 0.27-0.89) and with the Montgomery-Åsberg Depression Rating Scale were 0.68-0.88 (Bagby et al., 2004, p. 2171, Table 4)
Agreement with clinical ratings: good correspondence with clinician assessments of depression severity (Bagby et al., 2004, p. 2171)
Discriminant validity:
Depression vs. non-depression: adequate ability to distinguish depressed from non-depressed patients (mean cutoff 12.6/13.5; mean sensitivity 0.76, specificity 0.91) (Bagby et al., 2004, p. 2171, Table 5)
Overlap with anxiety: the HDRS correlates more highly with depression than with anxiety measures, but its saturation with anxiety-related concepts is nonetheless considerable (Bagby et al., 2004, p. 2172)
Predictive validity:
Treatment outcomes: predictive validity meets established criteria, though it suffers somewhat due to the scale’s multidimensionality; in meta-analyses the HDRS was more sensitive to treatment change than the Beck Depression Inventory and Zung self-report scales (Bagby et al., 2004, p. 2172-2174)
Content validity concerns:
Sleep symptom overemphasis: content validity is poor; the scale includes three insomnia items while burying or omitting important features of contemporary depression such as concentration difficulties and feelings of worthlessness (Bagby et al., 2004, p. 2170-2171)
Item discrimination: loss of insight showed the most variable and poorest performance; several items (loss of insight, hypochondriasis, somatic anxiety, general somatic symptoms, suicide, psychomotor agitation) were identified as problematic, and the loss-of-libido item showed high rater disagreement (Bagby et al., 2004, p. 2165-2169)
Factor Structure
Not unidimensional: the HDRS is clearly not unidimensional; the original single-factor assumption is not supported by the empirical data (Bagby et al., 2004, p. 2173)
Multi-factor solutions: across 15 studies (17 samples) the number of factors identified ranged from two to eight; recurring factors included an insomnia/sleep factor (13 data sets), a general depression factor (depressed mood, guilt, suicide), and an anxiety/agitation factor (Bagby et al., 2004, p. 2173, Table 6)
Cross-sample variability: the factor structure fails to replicate a single unifying structure across studies and samples (Bagby et al., 2004, p. 2173-2174)
Construct clarity: the exact structure of the scale’s multidimensionality remains unclear, and the meaning of the total score is correspondingly uncertain (Bagby et al., 2004, p. 2173)
Treatment Sensitivity
Medication trial sensitivity: sensitive to change in antidepressant trials; in meta-analyses the HDRS detected treatment change as well as or better than the Beck Depression Inventory and Zung self-report scales (Bagby et al., 2004, p. 2172-2173)
Clinically meaningful change: detects treatment-related change over multi-week trials, though unidimensional core-symptom subscales were often at least as sensitive to change as the full scale (Bagby et al., 2004, p. 2172-2173)
Response criteria: a ≥50% reduction from baseline is the standard trial definition of response (widely used convention)
Remission criteria: a score ≤7 is the established remission criterion in depression research (widely used convention)
Psychometric Limitations
Variable internal consistency: inconsistent reliability across studies and populations, with many studies showing α < 0.80 (Bagby et al., 2004, p. 2165)
Uneven item quality: many items are poor contributors to the measurement of depression severity, and some sum into a total score whose meaning is multidimensional and unclear (Bagby et al., 2004, p. 2173-2174)
Factor structure inconsistency: failure to replicate a single unifying factor structure across different samples (Bagby et al., 2004, p. 2173-2174)
Somatic symptom emphasis: the scale relies heavily on somatic and behavioural features, which may inflate scores in medically ill populations (Bagby et al., 2004; Worboys, 2013, p. 207)
Standardization challenges: some items require clinical inferences that are difficult to standardize across raters, motivating the development of structured interview guides (Williams, 1988)
Despite these limitations, the scale continues to be used due to historical precedent, regulatory acceptance, and its accumulated evidence base (Worboys, 2013).
Usage Guidelines and Applications
Primary Applications
Pharmaceutical clinical trials – primary efficacy endpoint in antidepressant development programs
Regulatory submissions – widely used and accepted endpoint for demonstrating drug efficacy in depression
Academic research – established comparator for validating new depression assessment tools
Clinical settings – depression severity assessment and treatment monitoring in psychiatric outpatient and inpatient care
Training programs – educational tool for developing clinical interviewing skills
Research Design Considerations
Rater training: formal training in HDRS administration and scoring, with ongoing calibration to prevent rater drift, is essential for reliable use
Assessment protocol: a 15-20 minute structured clinical interview covering the previous week’s symptoms, with documentation supporting each item score
Sample selection: clinical trials typically require a baseline HDRS score ≥20 for entry
Outcome definitions: use the established response (≥50% reduction) and remission (score ≤7) conventions and report them explicitly
Quality control: double-rating and consensus procedures are recommended for research applications
Standardization: structured versions (SIGH-D, GRID-HAMD) reduce administration variability and are preferable where rater consistency is critical (Williams, 1988)
Cultural and Population Considerations
Medically ill populations: the scale’s heavy weighting of somatic and behavioural features may inflate scores where physical illness produces overlapping symptoms (Bagby et al., 2004; Worboys, 2013, p. 207)
Cross-sample generalization: the factor structure does not replicate consistently across studies and samples, so subscale-style interpretations should not be transported between populations (Bagby et al., 2004, p. 2173-2174)
Intended population: the scale was designed for adults already diagnosed with depression; it is a severity measure, not a diagnostic instrument
Limitations and Considerations
Training requirements – extensive clinician training is needed for reliable administration
Time intensive – the 15-20 minute clinician interview limits routine clinical use
Somatic emphasis – may overestimate depression severity in medically ill patients
Item limitations – several items show poor psychometric properties and require careful interpretation (Bagby et al., 2004)
Copyright and Usage Responsibility: Check that you have the proper rights and permissions to use this assessment tool in your research. This may include purchasing appropriate licenses, obtaining permissions from authors/copyright holders, or ensuring your usage falls within fair use guidelines.
The Hamilton Depression Rating Scale was developed by Max Hamilton and first published in 1960. The original scale is in the public domain: public-domain reproductions state explicitly that the HDRS is in the public domain and that no permission is required to use the instrument. It may be reproduced and administered without a license.
Proper Attribution: When using or referencing this scale, cite the original development:
Hamilton, M. (1960). A rating scale for depression. Journal of Neurology, Neurosurgery and Psychiatry, 23(1), 56-62. https://doi.org/10.1136/jnnp.23.1.56
Structured versions: The Structured Interview Guide for the Hamilton Depression Rating Scale (SIGH-D) and other derivatives (such as GRID-HAMD) may carry separate copyright. Users should verify licensing requirements for specific structured versions before reproducing them.
Training materials: Commercial rater-training and certification materials may require separate licensing agreements even though the underlying scale is public domain.
Hamilton, M. (1960). A rating scale for depression. Journal of Neurology, Neurosurgery and Psychiatry, 23(1), 56-62. https://doi.org/10.1136/jnnp.23.1.56
Comprehensive Psychometric Review:
Bagby, R. M., Ryder, A. G., Schuller, D. R., & Marshall, M. B. (2004). The Hamilton Depression Rating Scale: Has the gold standard become a lead weight? American Journal of Psychiatry, 161(12), 2163-2177. https://doi.org/10.1176/appi.ajp.161.12.2163
Reliability Meta-Analysis:
Trajković, G., Starčević, V., Latas, M., Leštarević, M., Ille, T., Bukumirić, Z., & Marinković, J. (2011). Reliability of the Hamilton Rating Scale for Depression: A meta-analysis over a period of 49 years. Psychiatry Research, 189(1), 1-9. https://doi.org/10.1016/j.psychres.2010.12.007
Worboys, M. (2013). The Hamilton Rating Scale for Depression: The making of a “gold standard” and the unmaking of a chronic illness, 1960–1980. Chronic Illness, 9(3), 202-219. https://doi.org/10.1177/1742395312467658
The Hamilton Depression Rating Scale measures the severity of depression in patients who have already been diagnosed; it is not a diagnostic instrument. A clinician rates 17 items spanning mood, guilt, suicidal ideation, insomnia, work and activities, psychomotor features, anxiety, and somatic symptoms, based on a structured interview covering the previous week.
Who can administer the HDRS/HAM-D, and how long does it take?
The scale is clinician-administered and requires trained clinical judgment and structured interviewing skills; it is not a self-report questionnaire. Administration takes approximately 15-20 minutes, and formal rater training with ongoing calibration is recommended for research use.
Is the HDRS/HAM-D free to use?
Yes. The original scale is in the public domain, and public-domain reproductions state that no permission is required to use it. Structured derivatives such as the SIGH-D and GRID-HAMD may carry separate copyright, so licensing should be verified before reproducing those versions. Cite Hamilton (1960) when using the scale.
How is the HDRS/HAM-D scored?
The HDRS-17 is scored by summing ratings across 17 items; nine items are rated 0-4 and eight items are rated 0-2, giving a total range of 0-52. Conventional severity bands are 0-7 (normal/remission), 8-13 (mild), 14-18 (moderate), 19-22 (severe), and 23 or higher (very severe). Treatment response is conventionally defined as a reduction of at least 50% from baseline, and remission as a score of 7 or lower.
What is the difference between the HDRS-17 and the HDRS-21?
The HDRS-17 contains the 17 scored items used to measure severity. The 21-item version adds four items intended to subtype the depression; those extra items are not added to the severity total, so the HDRS-17 remains the standard form for scoring.
How does the HDRS/HAM-D differ from the Beck Depression Inventory (BDI)?
The HDRS is clinician-administered and emphasizes observable signs and somatic symptoms, whereas the BDI is a self-report questionnaire focused on subjective cognitive and emotional experience. Reported correlations between the two vary widely across studies, from roughly 0.27 to 0.89 (Bagby et al., 2004). The HDRS is the standard clinical-trial outcome measure; the BDI is more common in screening and self-monitoring.
How reliable is the HDRS/HAM-D?
Reliability is strong at the total-score level but weak at the item level. Inter-rater reliability for total scores is generally high (Pearson's r 0.82-0.98 across studies, pooled ICC 0.937 in the Trajković et al., 2011 meta-analysis), while internal consistency varies widely (α = 0.46-0.97 across studies, pooled α = 0.79). Many individual items show poor inter-rater and retest reliability, and the factor structure is inconsistent across samples (Bagby et al., 2004).