Depression: What It Is, and What Treatment Actually Does
Two people can both be diagnosed with major depression and have no symptom in common. That is not a quirk of the criteria. It is the central problem in the science.
This page covers what depression is, how common and how dangerous it is, whether it is a choice, and what treatment does. It contains no description of methods of self-harm and no drug doses, anywhere, including where a cited study involved them. Supporting someone who has it has its own research literature and its own page, linked at the foot. Crisis numbers are at the foot too, checked against each organisation’s own site on 13 September 2026.
What it actually is
A major depressive episode requires five of nine symptoms in the same two-week period, and at least one of them has to be low mood or loss of interest and pleasure. The other seven are appetite or weight change, insomnia or sleeping too much, agitation or slowing that someone else can see, fatigue, worthlessness or misplaced guilt, trouble concentrating or deciding, and recurrent thoughts of death. The symptoms have to be a change from how the person was, and they have to interfere with their life.
That has been the definition since 1994, and DSM-5-TR in 2022 left it alone. One thing did change, in 2013, and people still argue about it.
DSM-IV blocked the diagnosis if the symptoms started within two months of a death, unless they were severe. DSM-5 deleted that exclusion. The American Psychiatric Association's reasoning was that bereavement-related depression resembles other depression in its features, its course and how it responds to treatment, and that no other kind of loss cancelled the diagnosis: not a marriage ending, not a house burning down. Jerome Wakefield and Michael First argued in World Psychiatry that there was no scientific basis for the removal and that it would turn ordinary grief into illness. DSM-5-TR later added prolonged grief disorder as its own diagnosis, after twelve months, which does not put the old exclusion back.
Alongside the episodic form there is persistent depressive disorder: low mood most of the day, more days than not, for at least two years, with two of six further symptoms and no gap longer than two months. It merged what used to be dysthymia and chronic major depression. If someone meets full episode criteria at some point during those two years, the current guidance is to record both.
Two people, no shared symptoms
Five of nine, one of which must be from a fixed pair, allows 227 different ways to qualify. Several of the criteria are two-sided: you can meet the sleep item by sleeping far too little or far too much, and the appetite item by eating far too little or far too much. Count those poles separately and the number goes to 945. Count the sub-parts of the compound items and it reaches 16,400.
So one person can be sleepless, not eating, agitated and racked with guilt. Another can be sleeping fourteen hours, eating constantly, slowed to a crawl and unable to concentrate. Both have major depression. They share nothing.
That is the arithmetic. The question is what happens in practice, and two studies went and counted.
Fried and Nesse took the baseline data from STAR*D, which is the largest treatment study of depression ever run, and found that nearly half of their 3,703 patients had a symptom profile shared with nobody else in the trial. The most common single profile covered 1.8% of them. Controlling for how severe people were did not reduce the spread.
This is not a filing problem. It bleeds into everything downstream. A drug trial that scores everyone on a single total is treating a room of quite different people as one outcome, so a treatment that works well for a third of them and not at all for the rest produces a small average effect that looks like a weak drug. Biomarker studies keep failing to replicate, which is what you would expect if the samples are not the same phenotype twice. Fried has also shown that seven of the common depression rating scales between them contain 52 distinct symptoms, with any two scales overlapping by about a third, so two trials measuring "depression" may be measuring noticeably different things.
The obvious fix is subtypes, and the subtypes do not fix it. Adding the melancholic specifier to the criteria increases the number of possible qualifying profiles rather than cutting it down. Melancholia has the longest claim to being biologically distinct and the weakest modern evidence that it predicts who responds to what. Psychotic depression is the exception: it is more severe, it carries higher mortality, and it changes what treatment is offered.
None of this establishes that depression is several illnesses wearing one name. Family history, course and treatment response still cluster together better than chance, and psychiatry has tried splitting before and failed to make the pieces stick. What it establishes is that the label is doing less work than the research built on top of it assumes.
How common it is
Ask two American national surveys and you get two answers.
Lifetime estimates run from about 13% to about 21% depending on which survey you pick up, and that spread is wider than most of the differences people argue about in this field. Some of it is real definitional change, since NESARC-III used DSM-5 criteria with the bereavement exclusion already gone. Some of it is instrument. Women come out at roughly 1.7 times men in every one of them, which is the most stable finding in depression epidemiology.
Whether depression is becoming more common is not settled, and the disagreement follows methodology rather than politics. Compton and colleagues found one-year prevalence rising from 3.33% to 7.06% between 1991 and 2002 on the same criteria. A Canadian study found a flat 5% across 1952, 1970 and 1992. Ormel and colleagues, reviewing the epidemiological record in 2022, concluded that true prevalence has been roughly stable since the 1980s and that what has risen is diagnosis and treatment. Nobody has a decomposition that separates more illness from more looking.
How dangerous it is
People with depression die at roughly twice the rate of people without it. A review in World Psychiatry pooling 268 cohort studies and 10.8 million people with depression put all-cause mortality at a relative risk of 2.10 (95% CI 1.87 to 2.35). Older and smaller meta-analyses came in lower, around 1.8, so the estimate has moved with study inclusion. It stays elevated at 1.29 even against controls matched for other conditions, which is the comparison that matters most and the one least often quoted.
Where it gets misread is the split between suicide and everything else. Suicide carries by far the highest relative risk, near tenfold. Physical illness kills far more of the people.
A tenfold increase on a rare event is still a rare event. A 60% increase on heart disease is not.
The 15% figure, and why it was wrongTextbooks said for thirty years that 15% of people with depression die by suicide. It came from a 1970 paper by Guze and Robins, and the error is instructive because it is not an error of data.
They took follow-ups of hospitalised patients and calculated suicides as a share of the deaths that had occurred. That is proportionate mortality, and it inflates whenever the follow-up is short, because the people who die early in a young psychiatric cohort are disproportionately the ones who die by suicide; everyone else has not reached the age where heart disease and cancer arrive. Bostwick and Pankratz recalculated it in 2000 as case fatality, meaning suicides as a share of everyone diagnosed, across more than 50,000 patients.
Blair-West and colleagues came at it from the other end and showed that 15% is arithmetically impossible: multiply the population prevalence of depression by 15% and you get more suicides than actually occur. Their own ceiling was about 3.4%. Boardman and Healy, using UK records, found 2.4% for any affective disorder and 1.1% for people who never reached specialist services.
The risk is real, it is concentrated in the period after a diagnosis and in people who have already been hospitalised, and for most people with depression it is several times lower than what the old textbooks said.
Disability and courseDepression ranked second among all causes of years lived with disability worldwide in the 2019 Global Burden of Disease estimates, and thirteenth for total burden. Those rankings are constructed, not measured: they combine prevalence models with disability weights assigned by survey panels rating how bad conditions sound. Change the weights and the rank moves.
Course depends almost entirely on where you recruit. Eaton and colleagues followed 92 people through their first lifetime episode in Baltimore for up to 23 years. Around half recovered and never had another episode. About 35% recovered and later recurred. Roughly 15% had no year free of it across two decades. Median episode length was twelve weeks. In the NIMH Collaborative Depression Study, which followed 318 already-recovered patients who mostly had episode histories behind them, 63.5% recurred within ten years. Both numbers are correct. They describe different people.
Is it a choice?
Nobody in the literature argues that it is. The live questions are how much of it is constitutional, how much is circumstance, and whether effort changes anything.
Twin studies put heritability at 37% (95% CI 31 to 42), from a meta-analysis by Sullivan, Neale and Kendler, with a Swedish national replication at 38%. Shared family environment contributes almost nothing; the rest is environment specific to the individual. Measure it instead from DNA directly, and common genetic variants account for about 9%. That gap is the usual missing-heritability pattern, and it is unusually wide for depression. Part of it is that twin designs capture rare variants and gene-environment interplay that SNP arrays miss. Part of it may be that twin designs overestimate. Nobody has apportioned the two.
The largest genome-wide studies found 44 associated locations, then 102, each with a tiny effect. The practical consequence is worth stating in one number: a polygenic score built from them ranks a case above a control 57% of the time. A coin does it 50% of the time.
The candidate-gene collapseBefore genome-wide methods, the field tested single plausible genes in samples of a few hundred. The most famous result was Caspi and colleagues in 2003: a short variant of the serotonin transporter promoter, 5-HTTLPR, combined with stressful life events, predicted depression. It became one of the most cited findings in psychiatry.
Culverhouse and colleagues tested it in 38,802 people and found no interaction in any pre-specified subgroup. Border and colleagues tested eighteen historic candidate genes in samples ranging from 62,138 to 443,264 people and found no main effects and no gene-by-environment effects, with the candidate set performing no better than randomly chosen genes. The main effect of stress itself was strong in both.
An entire subfield, two decades of it, was noise that small samples could not see through. That is worth holding on to when reading anything else in this article that rests on a few hundred people.
SerotoninMoncrieff and colleagues published an umbrella review in 2022 covering six areas of serotonin research and concluded there is no consistent evidence of an association between serotonin and depression, and no support for the idea that depression is caused by lowered serotonin. It was downloaded over a million times.
Jauhar, Cowen and more than thirty co-authors replied that the review was not a standard umbrella review, applied quality criteria the authors had devised themselves, and attacked a monocausal story the field had abandoned decades earlier in favour of models where serotonin is one system among several. Moncrieff replied that even granting every criticism, none of it establishes the link.
Both sides are serious researchers and this page does not adjudicate it. What both sides accept is that the "chemical imbalance" line, as the public heard it, ran well ahead of the evidence. Where they part company is whether that has any bearing on whether the drugs work, which is a separate question answered by trials rather than by mechanism.
Circumstance, and whether effort moves itAdversity predicts onset about as reliably as anything in this literature does. Kendler's co-twin design found stressful life events predicting new episodes even within identical twin pairs, at an odds ratio of 3.58, which rules out the obvious genetic and family confounds. Brown and Harris, in Camberwell in the 1970s, found 23% of working-class women were cases against 6% of middle-class women.
Behavioural activation is the best evidence that doing things changes how people feel. It is a therapy built entirely on scheduling and re-entering activities the person has dropped, without working on their thoughts. A 2026 review covering 105 trials and 13,933 patients put it at a standardised mean difference of 0.67 against control conditions, with no detectable difference from other therapies and the effect still there a year after randomisation.
That finding gets used in both directions and deserves neither. It shows behaviour is upstream of mood for some people, so total passivity is not a required feature of the illness. It also required a trained therapist, a session structure and external accountability to produce that effect in people who were still quite ill. The distance between that and telling someone to get out more is the entire point.
What treatment does
Cipriani and colleagues pooled 522 double-blind trials and 116,477 adults across 21 antidepressants. All 21 beat placebo. The odds ratios ran from 2.13 for amitriptyline down to 1.37 for reboxetine, and the overall drug-placebo difference came to a standardised mean difference of 0.30. Nine per cent of the trials were rated high risk of bias, 73% moderate, and the authors graded the certainty of the evidence as moderate to very low.
Irving Kirsch, working from the FDA's own submitted trial data, put the same difference at 1.8 points on the 17-item Hamilton scale: drug groups improved by 9.6 points, placebo groups by 7.8. NICE once used a 3-point Hamilton difference as its rule of thumb for clinical significance, and on that rule the average trial does not clear the bar. Reanalyses of Kirsch's own dataset by Horder and by Fountoulakis found he had understated it, and put the difference nearer 2.2 to 2.7 points, with venlafaxine and paroxetine clearing 3 points.
Turner and colleagues showed why the published literature looks better than the drugs are. Of 74 trials registered with the FDA, 31% were never published. By the FDA's reading 51% were positive; in the published record 94% appeared positive. The effect size in print was 0.41 against 0.31 across all trials, submitted and buried alike.
Both camps agree on the arithmetic. They disagree about what a 2-point shift on a rating scale means to a person, which is a judgement about value rather than a finding the trials can settle.
STAR*DThe number everyone quotes from STAR*D is that 67% of patients remitted after up to four sequential treatment steps. That figure was always a thought experiment: it assumed the people who dropped out, and more than half did, would have remitted at the same rate as those who stayed.
Pigott and colleagues reanalysed the patient-level data in 2023 against the trial's own registered protocol, which specified the blinded clinician-rated Hamilton scale rather than the self-report measure the headline used, and which excluded patients the published analysis had included. They got 35.0%, or 41.3% if missing exit scores are filled in from the self-report. Rush and colleagues replied that Pigott's exclusions removed 941 patients with low exit scores and that other analyses of the same public dataset land near 60%. A third reanalysis, handling dropout with inverse-probability weighting, produced 87.5%.
One dataset, three numbers, none of them from a placebo-controlled trial, because STAR*D had no placebo arm.
Psychotherapy, and what happens when you weight for qualityCuijpers and colleagues have spent fifteen years measuring how much the psychotherapy literature overstates itself. Their 2010 meta-analysis found an overall effect of g=0.67, falling to 0.42 once publication bias was accounted for. By 2020, restricting to trials at low risk of bias and correcting for publication bias together, the pooled figure came down to about 0.31. Roughly a quarter of the trials in that literature are at low risk of bias. Waiting-list controls inflate the numbers relative to active controls, and much of the field uses waiting lists.
CBT, behavioural activation, interpersonal therapy and short-term psychodynamic therapy do not separate from each other. CBT has by far the largest evidence base and no demonstrated advantage.
ExerciseA 2024 network meta-analysis in the BMJ covering 218 studies, 495 arms and 14,170 participants found walking or jogging at g=-0.62, yoga at -0.55, strength training at -0.49, and dance at -0.96. The dance figure came from five small trials. The headline that followed it into the press was that exercise rivals or beats antidepressants and therapy.
One of the 218 studies met Cochrane's criteria for low risk of bias. The authors' own confidence rating was low for walking and jogging and very low for every other modality, and they said so in the paper. Exercise helps. The size of the help is not known to within the precision those numbers imply.
Coming off themDavies and Read reviewed the withdrawal literature in 2019 and reported a weighted average incidence of 56%, across a range of 27% to 86%, with 46% of those affected rating it at the most severe level offered. Jauhar and Hayes replied that the review mixed self-selected online surveys of long-term users with discontinuation arms of short trials, which measure different things. Henssler and colleagues, pooling 62 cohorts in 2024, put incidence nearer 31% and severe withdrawal as uncommon, and were in turn criticised.
The numbers are still disputed. The guidance moved anyway: NICE and the Royal College of Psychiatrists both revised their advice after 2019 to say withdrawal can be severe and prolonged and that tapering should often be slower than the two to four weeks previously suggested. On this one, the critics changed the guidelines before the field agreed on the figure.
The part nobody can explain
Treatment for depression in the United States roughly quadrupled between 1987 and 2007. Antidepressant prescribing rose across every high-income country, psychotherapy became far more available, and screening spread into primary care.
Population prevalence did not fall.
Ormel, Hollon, Kessler, Cuijpers and Monroe set this out in 2022 and named it the treatment-prevalence paradox. They checked and rejected the easiest escape, that incidence secretly rose and masked a real gain, because the incidence studies do not show the necessary rise. What they are left with is some combination of three things: published trials overstate what treatment does, treatment delivered in ordinary clinics does less than treatment delivered in trials, and the benefit falls mainly on people with single episodes while prevalence is carried by the chronic and recurrent cases.
This does not mean the treatments are useless. The individual-level trial evidence says they work, modestly, and the section above is the measure of it. It means the individual effect and the population effect do not add up, and no one has produced a clean account of where the difference goes.
A comparison across 29 OECD countries found national antidepressant prescribing rates were not associated with lower population levels of sadness, worry or unhappiness. National income, education and life expectancy were.
The short version
1. Major depression is five of nine symptoms over two weeks, one of which must be low mood or loss of interest. There are 227 ways to qualify, and two people can meet the criteria with no symptom in common.
2. In the largest treatment study ever run, 3,703 patients produced 1,030 distinct symptom profiles, and nearly half of those profiles belonged to exactly one person.
3. Lifetime prevalence in US national surveys runs from about 13% to about 21% depending on the survey. Women are affected at roughly 1.7 times the rate of men.
4. Mortality is about double, and most of the excess is physical illness rather than suicide, even though suicide carries the higher relative risk.
5. The old 15% lifetime suicide figure was a methodological error corrected in 2000. Current estimates run from under 0.5% for people who never reached services to 8.6% for those hospitalised for suicidality.
6. It is not a choice. Twin heritability is 37%, adversity predicts onset even within identical twin pairs, and the behavioural therapy that works best still needs a therapist to deliver it.
7. Antidepressants beat placebo by a standardised mean difference of about 0.30. Whether that is worth having is a judgement, not a finding. Psychotherapy's effect falls from 0.67 to about 0.31 once trial quality and publication bias are accounted for.
8. Treatment quadrupled and population prevalence did not move, and nobody has explained that.
Supporting someone who is depressed has its own research literature and its own page: How to Help Someone With Depression. It covers what to say and what not to, why asking someone directly about suicide does not increase their risk, and what the law allows when someone refuses help.
United States. The 988 Suicide and Crisis Lifeline takes calls and texts at any hour, free and confidential, in English and Spanish with interpretation in more than 240 languages. Veterans press 1. You do not have to be suicidal to use it. The Crisis Text Line answers texts around the clock: text HOME to 741741. In an emergency, 911.
United Kingdom and Ireland. Samaritans answer on 116 123, free from any phone, at any hour of any day, and the number does not appear on your bill. A Welsh-language line runs on 0808 164 0123. If you would rather write than speak, Shout takes texts at any hour: text SHOUT to 85258. CALM is on 0800 58 58 58 from 5pm to midnight. For urgent NHS mental health help call 111 and choose the mental health option, and 999 in an emergency.
Two things that are out of date almost everywhere else. PAPYRUS closed on 8 September 2026 and HOPELINE247 stopped answering that day, so the number 0800 068 4141 still printed on NHS trust pages, school leaflets and council directories no longer reaches anyone. Use Samaritans or Shout instead. And Samaritans retired their text service in February 2020, so any Samaritans text number you find is dead; they point texters to Shout.