Judgment Playbook

Survivorship bias

Look for the invisible failures before copying the visible winner’s playbook.

What it is

Survivorship bias is the error of reading a lesson off the cases that stayed visible, when the cases that failed have already been removed from view. The method that corrects it rebuilds the full population that once attempted the same move, before any outcome was known. That reconstruction gives you two things a study of winners cannot: an honest base rate, and a test of whether the winner's celebrated attribute was actually rare among the losers.

Where it comes from

The idea has no single author and the phrase has no dated first use. Its first rigorous treatment is by the statistician Abraham Wald (1902–1950). During the Second World War Wald worked in the Statistical Research Group at Columbia University, and in 1943 he wrote "A Method of Estimating Plane Vulnerability Based on Damage of Survivors", classified at the time and reprinted by the Center for Naval Analyses in 1980. His problem was that damage data came only from aircraft that made it home. His method estimated the damage distribution across every aircraft that flew, using the distribution observed in those that returned. The familiar retelling, in which Wald tells the air force to armour the parts with no bullet holes, is a later popularisation; the memorandum itself is a set of estimating equations.

The term became standard in finance in the 1990s. Brown, Goetzmann, Ibbotson and Ross showed in 1992 that a sample truncated by survival can produce the appearance of predictable fund returns where none exists.

What it corrects

You are deciding whether to copy a move: a pricing model, a hiring bar, an expansion into a new category. Your evidence is the set of companies that made that move and are still around to be studied. A competent analyst does the honest thing with that evidence. Read the winners closely, extract what they share, treat the frequent attributes as the causes.

The failure is invisible because the sample looks complete. Nothing in the data announces that it has been filtered, and the filter was the outcome you are trying to explain. Ordinary care makes this worse rather than better: closer study of the survivors makes the shared attribute more salient and the conclusion more confident. Care applied inside a selected sample raises confidence without raising accuracy. Only evidence from outside the sample repairs it.

How it works

  1. Restate the claim as a rate: of everyone who tried this, what share reached the outcome?
  2. Define the population that attempted it, counted before any outcome was known.
  3. Name the rule that removed cases from your evidence: bankruptcy, quiet acquisition, delisting, cancellation, or simply being too dull to write about.
  4. Recover what you can of the missing group, even roughly, from registries, cohort counts, or portfolio totals.
  5. Check whether the winning attribute was also common among the failures.
  6. Keep only the conclusions that survive that comparison, and state what you could not reconstruct.

Worked example

In Search of Excellence, by Tom Peters and Robert Waterman, was published in October 1982. The authors asked McKinsey partners and other business people who was doing impressive work, assembled 62 companies, then applied performance screens that cut the list to 43. General Electric was among those dropped. The book's eight principles were read off the attributes those 43 companies shared.

The selection rule was the outcome itself, so the study could not show that the principles caused the performance. No company that followed the same principles through a bad decade was in the sample, because a bad decade was what the screens excluded. Several of the 43 later struggled. Wang Laboratories, with annual revenues near $3 billion at its peak in the 1980s, filed for bankruptcy protection in August 1992.

The correct repair is narrower than dismissal. A 2002 Forbes analysis found that the 43 companies still returned about 14.1% a year on average after publication, against 11.3% for the Dow companies. The list was not worthless. What it could not support was the causal claim, because the missing half of the evidence had been screened out by design.

In a Business Case Weekly case

In the Y Combinator case, the accelerator's record arrives as a small set of famous names against a portfolio value of roughly $600 billion built on about $1.5 billion deployed across 5,668 companies. At the fork where the case weighs whether bigger batches dilute the model, survivorship bias does one job: it forces the argument to run on rates over all 5,668 rather than on the alumni anyone can name, and it asks what a median funded company actually received. That changes what evidence would settle the question, without settling it for you.

In your answer

  • "This evidence is selected on the outcome, because the sample was assembled after the results were known."
  • "The population that attempted the same move was about N; the cases missing from the story are …"
  • "The same attribute shows up in the failures I can find, so I treat it as a requirement rather than the cause."
  • "I cannot reconstruct the missing group here, so I hold this conclusion at low confidence and would test it by …"

Common misuse

The counterfeit version names the bias and carries on: "there is obviously survivorship bias in these examples", followed by the same conclusion the survivors suggested. The mirror image is using the bias as a veto, treating a successful company as evidence of nothing. The test is one question: did your estimate of the base rate change, and by how much? A correction that leaves every number and every conclusion where it was did none of the method's work.

References

  • Abraham Wald, "A Method of Estimating Plane Vulnerability Based on Damage of Survivors" (1943), reprinted by the Center for Naval Analyses, 1980 — the primary source, and heavy going.
  • Stephen Brown, William Goetzmann, Roger Ibbotson and Stephen Ross, "Survivorship Bias in Performance Studies", Review of Financial Studies 5(4), 1992, 553–580.
  • Mark Carhart, Jennifer Carpenter, Anthony Lynch and David Musto, "Mutual Fund Survivorship", Review of Financial Studies 15(5), 2002, 1439–1463. Measures the bias: 0.07% a year in one-year samples, about 1% a year in samples longer than fifteen years.
  • Jordan Ellenberg, How Not to Be Wrong: The Power of Mathematical Thinking, Penguin Press, 2014. The opening chapter tells the Wald story and is an evening's reading.

Map is not the territory

Next method

Treat metrics, models, and case narratives as useful representations, never the business itself.

Open next tool