Why so many horse racing systems look brilliant in backtests
Finding a racing system that would have made money is surprisingly easy. Finding one that makes money from tomorrow is a completely different problem, and most of the difference is statistical rather than equine.
The short answer
If you search a large historical dataset for rules that would have profited, you will find some. You will find some even if there is nothing there to find, because that is what searching does to data.
The strength of a backtest depends on how many things you tried before you got it. That number is almost never reported, and without it the result cannot be interpreted at all.
The seven ways a backtest flatters you
These are not exotic failures. Every one of them is easy to commit while being careful and meaning well.
Searching, then testing on the same data. You look through ten years of racing, notice that a particular kind of horse did well, and then check whether that kind of horse did well. It did. That is not a test, it is the same observation twice. Andrew Lo and Craig MacKinlay set this out formally for asset pricing in 1990: reusing one dataset to both discover and confirm a pattern biases the confirmation.
Trying many rules and reporting one. Twenty filters, each with a one-in-twenty chance of looking good by luck alone, will hand you roughly one winner. Halbert White's 2000 paper in Econometrica gave this a proper treatment: a test applied after a search has to account for the search, or its significance is fiction.
Tuning until it works. Field size of eight or more becomes seven or more becomes seven to twelve. Each nudge is defensible. Together they are a search, and nobody counts them as one.
Look-ahead. Using anything the punter would not have had at the time. Final starting price is the classic. So is a going description updated after the meeting, or an official rating revised the following week.
Survivorship. Testing on the meetings, seasons or tracks that made it into your file, when the ones that did not are missing for a reason.
Silent exclusions. Non-runners, void races, meetings abandoned halfway. Each removal is small and reasonable and the population quietly stops being the population you claim to have tested.
Telling the story afterwards. The rule works, so you explain why it works. The explanation is invented after the result and adds no evidence whatever, but it makes the finding feel earned.
What the research actually says
This is not a racing problem. It is a statistics problem that racing happens to be an excellent place to commit.
In 2014, David Bailey, Jonathan Borwein, Marcos Lopez de Prado and Qiji Jim Zhu published a paper in the Notices of the American Mathematical Society with the unusually blunt title Pseudo-Mathematics and Financial Charlatanism. Their central claim is that high simulated performance is easily achievable after backtesting a relatively small number of alternative strategy configurations, and that the more configurations are tried, the greater the probability the backtest is overfit. Their practical point is sharper still: because the number tried is rarely reported, a reader usually cannot evaluate how overfit a result is.
Campbell Harvey, Yan Liu and Heqing Zhu made a related argument about finance research in the Review of Financial Studies, that when a literature has tested very many candidate effects, a conventional significance threshold is simply too weak to sort the real ones from the lucky ones. John Ioannidis made the general version of the case in PLoS Medicine in 2005: the reliability of a finding depends on how many hypotheses were in play, not only on the finding's own p-value.
None of these papers is about horse racing, and we are not going to pretend a finance result proves a racing result. What transfers is the statistical problem, which does not care what the rows in the dataset represent.
An illustrative example
Invented for illustration, and typical of the shape.
Somebody notices that second favourites in small-field handicaps on soft ground did unusually well over five seasons. The sample is a few hundred races, the return looks strong, and there is a story to go with it: soft ground exaggerates stamina differences and small fields reduce interference.
What is not in the write-up is that soft ground was tried after good ground and heavy ground did nothing. That the field-size cut moved twice. That handicaps were chosen because the first pass across all race types was flat. That second favourite was reached after favourite and third favourite were both disappointing.
Nobody lied. Each decision was made for a reason. But the finding is the survivor of a search that was never counted, and the write-up describes only the survivor.
What Elite Pass found
- 227 racing systems, rules and theories examined.
- 0 have cleared our current validation bar.
- Our own EP Score backtest ran across 1,024 races, and its most attractive components did not survive price control.
What it does not show
It does not show that no racing edge exists. It shows that the ideas we have tested, under the standards we currently apply, have not earned the claim.
It does not mean the systems examined were foolish. Most were reasonable, and several described effects that are genuinely real before price is accounted for.
It is not a comparison with anybody else's record. We have no way to audit other people's testing and we are not implying one.
Why we publish the failures
The number that matters in our audit is not 227. It is 0, and it is the least flattering figure we own.
We publish it because of exactly the problem above. A record that shows only the ideas that worked is uninterpretable: you cannot tell whether it represents one success in three attempts or one in three hundred. The denominator is the evidence. Without it a good result is just a survivor with a story.
A racing idea should have to earn the right to be believed.
That is also why our daily reads are timestamped and frozen before racing. An idea that can be revised after the result is not being tested. It is being narrated.
The system audit lists what we have examined. The evidence room holds the record. What price control did to our best looking ideas is the specific case of our own most promising work dying this way.
Academic sources
- White, H. (2000). A Reality Check for Data Snooping. Econometrica, 68, 1097 to 1126. doi:10.1111/1468-0262.00152
- Lo, A. W., & MacKinlay, A. C. (1990). Data-Snooping Biases in Tests of Financial Asset Pricing Models. Review of Financial Studies, 3, 431 to 467. doi:10.1093/rfs/3.3.431
- Bailey, D. H., Borwein, J. M., Lopez de Prado, M., & Zhu, Q. J. (2014). Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance. Notices of the American Mathematical Society, 61, 458. doi:10.1090/noti1105
- Harvey, C. R., Liu, Y., & Zhu, H. (2016). ... and the Cross-Section of Expected Returns. Review of Financial Studies, 29, 5 to 68. doi:10.1093/rfs/hhv059
- Ioannidis, J. P. A. (2005). Why Most Published Research Findings Are False. PLoS Medicine, 2, e124. doi:10.1371/journal.pmed.0020124
Related research
- Why we timestamp our own pricesA price you read after the race has already been contaminated by the result. So we write ours down before the off, and never touch them again.
- How should a horse racing system actually be tested?Twelve steps, in order, for testing a racing idea properly. Written by someone who has run 227 of them and cleared none.
- Why we publish the ideas that failedThree of the first five beliefs we tested died. Keeping the dead ones in public is the only honest way to ask anyone to trust the ones still standing.
Elite Pass publishes research, not guaranteed outcomes. Findings are measured against the market and remain subject to replication and sample size. The evidence room and the public ledger carry the full record.