Survivorship Bias: Why Backtests Only See the Winners

A backtest on the current coin list ignores every dead coin. What Brown, Goetzmann, Ibbotson and Ross showed, with an example and a checklist.

Key takeaways
  • Survivorship bias appears when a test uses only the assets that still exist, because the ones that failed have dropped out of the data.
  • Brown, Goetzmann, Ibbotson and Ross (1992) showed that a sample truncated by survivorship can make past performance look like it predicts future performance.
  • In an illustrative universe of 20 coins where 6 were delisted after a 90% loss, the survivors average +30% while the whole universe averages −6%.
  • The fix is a point-in-time universe: at each test date, use the assets that existed and were tradable on that date, including those that later died.

Pick the 30 most traded coins today, download five years of their prices, and test your strategy. The result will look better than reality, and no amount of careful coding will fix that, because the error is in the question: you asked how the coins that are on the list today behaved, and being on the list today is itself a result of having done well.

What survivorship bias is

Survivorship bias is the distortion that appears when a dataset contains only the assets that survived to the end of the period. Funds that closed, stocks that were delisted and coins that went to zero are missing, and with them go the worst outcomes. Whatever you measure on what is left looks healthier than what an investor who was there at the start actually experienced.

The classic study is Brown, Goetzmann, Ibbotson and Ross (1992), in the Review of Financial Studies. They looked at the evidence that past mutual fund performance predicts future performance, and analysed “the relationship between volatility and returns in a sample that is truncated by survivorship”. Their conclusion is that this truncation can create the appearance of predictability, and that the effect can be strong enough to account for the strength of the evidence in favour of predictability.1 In plain words, some of what looked like skill that persists was an artifact of the sample.

A numeric example

The numbers below are illustrative, not data from any market. Suppose that in 2022 there were 20 coins in a “top by volume” list. By 2026, 6 of them were delisted or fell 90%, and 14 are still on the list and are up 30% on average.

GroupCoinsAverage result
Survivors (the list today)14+30%
Failed (delisted, −90%)6−90%
Whole 2022 universe20−6%

The whole universe is (14 × 30 + 6 × (−90)) ÷ 20 = −6%. A backtest on today’s list reports +30%, a gap of 36 points that comes from nothing but the choice of the list. A strategy that “buys dips” fares even better in such a sample: every dip in the survivors was followed by a recovery by construction, because the coins that did not recover are not there to be bought.

Why crypto traders run into it

Three features make the trap easy to fall into when testing coins:

  • Coins come and go. New coins appear and old ones are delisted or abandoned, so a list made today can differ a lot from the list of three years ago.
  • Selection by recent volume or market cap. Any list of “top coins” is built from information that includes the future relative to the start of your test. This is look-ahead hidden in the choice of the universe.
  • Dead coins are easy to lose from the data. If you download histories from the symbols an exchange lists today, the delisted ones are simply not on the list, so the missing losers are not forgotten on purpose; they never arrive in your dataset.

How to reduce the bias

  1. Build the universe point in time. At each test date use the assets that existed and were tradable then (for example, ranked by volume known at that date), and keep them in the data after they die.
  2. Store delistings. Keep a record of every symbol that was removed, with its last price, and decide in advance how to treat the final period of a failing coin: with the real exit price if one exists, otherwise at a conservative value.
  3. Test on assets selected before the test period, not after it.
  4. Compare with the universe, not with the survivors. Report your strategy against a buy-and-hold of the whole point-in-time list.
  5. Say what the sample is. If you must use survivors, state that the result is optimistic and by roughly how much, as the example shows.

Survivorship is one of several ways in which a backtest overstates results. Overfitting and selection among many trials are described in why most backtests overstate results, where Bailey and his co-authors note how easy it is to fit a strategy to a sample so that it performs well there.2 A strategy that survives all of these checks is rare, and finding out that yours does not is also a result.

Footnotes

  1. Brown, S. J., Goetzmann, W., Ibbotson, R. G., Ross, S. A. (1992). “Survivorship Bias in Performance Studies”. Review of Financial Studies 5(4), 553–580. DOI 10.1093/rfs/5.4.553. The quotation is from the abstract; the abstract states that numerical examples show the effect can be strong enough to account for the evidence of predictability. ↩

  2. Bailey, D. H., Borwein, J. M., López de Prado, M., Zhu, Q. J. (2014). “Pseudo-Mathematics and Financial Charlatanism”. Notices of the AMS 61(5), 458–471. The numeric example in this article (20 coins, 14 survivors at +30%, 6 failures at −90%) is ours and illustrative. ↩

Sources

  1. Survivorship Bias in Performance Studies. Stephen J. Brown, William Goetzmann, Roger G. Ibbotson, Stephen A. Ross. Review of Financial Studies 5(4), 553–580, 1992
  2. Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance. David H. Bailey, Jonathan M. Borwein, Marcos López de Prado, Qiji Jim Zhu. Notices of the AMS 61(5), 458–471, 2014
OrderBlock.net Research

The team behind the OrderBlock.net scanner. We read the primary research and exchange documentation so you do not have to, and cite every source.

This article is research, not investment advice. Results on history do not guarantee future results.