Why Most Backtests Overstate Results

Three landmark papers explain why a great backtest is usually a statistical accident, and what to do: out-of-sample tests and a higher bar.

Key takeaways
  • Bailey, Borwein, López de Prado and Zhu (2014) show that it is relatively simple to overfit a strategy so that it performs well in-sample; a backtest is only realistic when in-sample and out-of-sample results agree.
  • Selection bias makes it worse: researchers tend to report only positive outcomes, and ignoring the number of trials leads to over-optimistic expectations.
  • The Deflated Sharpe Ratio (Bailey and López de Prado, 2014) corrects a Sharpe ratio for the number of strategies tried and for non-normal returns.
  • Harvey, Liu and Zhu (2016) argue that a new factor should clear a t-ratio of about 3.0 rather than 2.0, and that most claimed findings in financial economics are likely false.

A backtest is the most persuasive object in trading: a clean equity curve, a high Sharpe ratio, a table of monthly returns. It is also, according to a substantial body of research, one of the easiest things in finance to fake without meaning to. This article summarizes three papers that every strategy developer should know, and turns them into a practical checklist.

1. Overfitting: the backtest learns the noise

Bailey, Borwein, López de Prado and Zhu define a backtest as a historical simulation of an algorithmic strategy, and separate two readings of its performance: in-sample (IS), measured on the data used to design the strategy, and out-of-sample (OOS), measured on data the design never saw. Their criterion is short:

“A backtest is realistic when the IS performance is consistent with the OOS performance.”

— Bailey, Borwein, López de Prado, Zhu, Notices of the AMS, 20141

The problem is how easy it is to break that consistency. The authors write that, given any financial series, “it is relatively simple to overfit an investment strategy so that it performs well IS”.1 Overfitting, a term borrowed from machine learning, describes a model that “targets particular observations rather than a general structure”. A typical route is what they call data snooping: tuning parameters to remove the specific losing trades you already know about. After a few iterations, the “optimal parameters” profit from features of that particular sample “but may well be rare in the population”.1

The paper opens with an epigraph from Richard Feynman that sums up the danger: “with a little skill any experimental result can be made to look like the expected consequences”.1

2. Selection bias: you only see the winners

Overfitting happens inside one research process. Selection bias happens across many. In their paper on the Deflated Sharpe Ratio, Bailey and López de Prado note that with large data sets and cheap computing, analysts “can backtest millions (if not billions) of alternative investment strategies”.2 Then a second filter is applied by people:

“researchers and investment managers tend to report only positive outcomes, a phenomenon known as selection bias.”

— Bailey, López de Prado, Journal of Portfolio Management, 20142

The consequence is stated plainly in the same abstract: not controlling for the number of trials behind a discovery “leads to over-optimistic performance expectations”.2 If you test a hundred random rules, a handful will look excellent by chance alone. If only those are shown, the reader sees a remarkable strategy where there is only noise.

3. A higher bar: the Deflated Sharpe Ratio and t > 3

The authors’ remedy is the Deflated Sharpe Ratio (DSR). It corrects the observed Sharpe ratio for two sources of inflation: selection bias under multiple testing, and returns that are not normally distributed (fat tails and skew, both common in trading). In doing so, the abstract says, DSR “helps separate legitimate empirical findings from statistical flukes”.2 The practical message: a Sharpe ratio means little unless you also know how many strategies were tried to find it.

Harvey, Liu and Zhu reached a similar conclusion from a different direction. Reviewing hundreds of published factors that claim to explain stock returns, they built a multiple-testing framework and concluded that a newly discovered factor should clear a t-ratio above about 3.0, not the conventional 2.0.3 Their verdict on the field was blunt: “most claimed research findings in financial economics are likely false”.3

Peer-reviewed research at least has referees and replication. A private backtest built over a weekend has neither, so it deserves at least as much scepticism.

What this means for a trader

The papers translate into rules that cost nothing but discipline:

  1. Write the rules before you run the test. Every change after seeing results is another trial.
  2. Keep a trial log. Count every variant you tried, including those you discarded mentally. Without the count, the Sharpe ratio of the winner is not interpretable.
  3. Split by time. Design on the earlier part of the data and judge on the later part. Random splits leak future information into the past.
  4. Demand consistency. If out-of-sample performance is much worse than in-sample, the strategy is overfit, whatever the in-sample statistics say.
  5. Raise the bar. A result that would pass t > 2 once may be a fluke after fifty tries. Use corrections such as the Deflated Sharpe Ratio, or require roughly t > 3.
  6. Prefer simple rules with an economic reason. Fewer parameters mean fewer ways to fit noise.
  7. Forward-test before risking money. Data that did not exist when the rule was written is the only sample that cannot be optimized.

A note on reading other people’s backtests

The same logic applies when you evaluate a signal service, a course or a published strategy. Ask how many variants were tested, what the out-of-sample period was, whether fees and slippage were included, and whether losing configurations were reported. A seller who cannot answer has, at best, shown you the winner of a selection process you cannot see.

Footnotes

  1. Bailey, D. H., Borwein, J. M., López de Prado, M., Zhu, Q. J. (2014). “Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance”. Notices of the American Mathematical Society 61(5), 458–471. DOI 10.1090/noti1105. The Feynman quotation is the article’s epigraph (Feynman, 1964). ↩ ↩2 ↩3 ↩4

  2. Bailey, D. H., López de Prado, M. (2014). “The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality”. Journal of Portfolio Management. SSRN 2460551. ↩ ↩2 ↩3 ↩4

  3. Harvey, C. R., Liu, Y., Zhu, H. (2016). “…and the Cross-Section of Expected Returns”. Review of Financial Studies 29(1), 5–68. NBER Working Paper 20592. ↩ ↩2

Sources

  1. Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance. David H. Bailey, Jonathan M. Borwein, Marcos López de Prado, Qiji Jim Zhu. Notices of the AMS 61(5), 458–471, 2014
  2. The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality. David H. Bailey, Marcos López de Prado. Journal of Portfolio Management, 2014
  3. …and the Cross-Section of Expected Returns. Campbell R. Harvey, Yan Liu, Heqing Zhu. Review of Financial Studies 29(1), 5–68, 2016
OrderBlock.net Research

The team behind the OrderBlock.net scanner. We read the primary research and exchange documentation so you do not have to, and cite every source.

This article is research, not investment advice. Results on history do not guarantee future results.