Methodology

Why Backtest Results Don't Repeat in Live Investing

Backtest results rarely repeat live because of overfitting, regime change, crowding, and the costs a test omits. Here is what actually causes the gap and what to expect.

Backtest results rarely repeat in live investing, and the honest expectation is that live outcomes land below the tested ones. The gap comes from four separable causes: choices made after seeing the data, market conditions that change underneath the rule, real trading frictions that a test understates, and plain luck. None of these means backtesting is useless. It means the number a backtest produces is not a forecast of what you will earn.

This is worth internalising before you spend months building a testing framework, because most of the disappointment in systematic investing comes from expecting a backtest to be a promise rather than a diagnostic.

Cause one: the result was partly chosen, not discovered

Every backtest involves choices. The universe, the lookback window, the rebalancing frequency, the number of holdings, the entry and exit thresholds, the start date, the treatment of small caps, the handling of corporate events. Each is a knob. Each knob has a setting that, on this particular history, produced a better number.

If you tried several settings and kept the best one, the result you kept is not a clean measurement. It is the maximum of a set of noisy measurements, and the maximum of noisy things is systematically higher than the truth. This is overfitting, and it is far more common than deliberate curve-fitting because it usually happens through ordinary iteration. Nobody sets out to torture the data. They just keep improving the result until it looks presentable.

The subtler version is selection at the idea level. You tested nine ideas and are showing the one that worked. That one looks like a discovery, but out of nine reasonable ideas tested on one history, one of them looking good is close to what you would expect from chance alone.

The practical defences are well known and uncomfortable: keep a genuinely untouched holdout period, count and record how many variants you tried, and prefer rules with fewer parameters. In-sample versus out-of-sample testing is the minimum honest structure here, though it only helps if the holdout stays untouched. Once you have peeked at it and adjusted, it is in-sample.

Cause two: the market that the rule was tuned to has moved

A rule is a bet that some relationship persists. Relationships in markets are not laws of physics. They are the aggregate behaviour of participants, and participants change.

Several kinds of change matter:

Structural change. Market microstructure, index construction, disclosure rules, taxation, and the mix of participants all evolve. A rule tuned on a decade with one set of institutional arrangements can behave differently under another.

Composition change. The sector mix of an index shifts over time. A rule that looked like a general statement about stocks can turn out to have been a statement about whichever sector happened to dominate the tested period.

Crowding. If enough capital pursues the same signal, the signal gets bid away, and the crowd’s simultaneous exits can turn a mild drawdown into a sharp one. This is the mechanism behind factor crowding. Anything published, packaged and widely tradable is a candidate.

Regime. Interest rates, inflation, liquidity conditions and volatility all cycle. Many strategies are implicitly bets on one regime. Their tested period may simply not contain enough of the regimes they will meet.

The uncomfortable implication is that a longer backtest is better not because it produces a more impressive number but because it covers more regimes. A rule tested only on a period of falling rates and rising markets has not been tested at all in the ways that matter.

Cause three: the test was cheaper and easier to trade than reality

A backtest executes perfectly. Reality does not.

Costs are the obvious gap. Brokerage, exchange charges, securities transaction tax, stamp duty and GST all apply, and they scale with turnover. A rule that rebalances weekly and one that rebalances annually can look similar gross and land in completely different places net.

Slippage and market impact are the less obvious gap. A backtest typically assumes you transacted at the close, or at some single reference price. Live, you move the price against yourself, particularly in the less liquid names where many screens find their most attractive candidates.

Then there is capacity. A test does not know how much money is behind it. A rule that trades small, thinly traded stocks may be perfectly implementable with a modest allocation and impossible at scale. The backtest curve looks identical either way.

Finally there is the data itself. If the test ran on today’s restated financials rather than the numbers that were public on each decision date, the rule quietly used information from the future. That is why point-in-time data is a precondition rather than a refinement. The same applies to the universe: testing on the companies that exist today silently excludes the ones that failed or were delisted, which makes the past look kinder than it was.

Cause four: luck, which is larger than people think

Even a rule with a genuine edge produces noisy results over any finite period. The reverse is also true: a rule with no edge produces a perfectly convincing looking equity curve reasonably often.

The intuition most people lack is how many observations are needed to distinguish a modest edge from noise. A strategy that rebalances quarterly over ten years has produced about forty independent decisions. Forty is not many. A run of good quarters inside that sample is entirely unremarkable, whether or not anything real is driving it.

This is why the statistics attached to a backtest deserve scepticism proportionate to how few independent bets sit behind them. A high-turnover rule with thousands of trades has more evidence per year than a concentrated annual-rebalance rule with the same track length, even if the second one shows a nicer chart.

The right question about a backtest is not “how good is this number.” It is “how many genuinely independent observations produced it, and how many alternatives did I look at before choosing it.”

What to expect, honestly

Given all of that, a reasonable set of expectations looks like this.

Expect the live result to be lower than the test. Not because of any single mistake, but because the test was selected for looking good and reality adds friction.

Expect the drawdowns to be at least as bad as the tested worst, and plan as though they will be worse. The tested maximum drawdown is the worst thing that happened to occur in one sample. It is not a ceiling.

Expect long stretches of nothing. Most rules that work over a decade spend meaningful portions of that decade behind a simple benchmark. If the plan cannot survive that psychologically or commercially, the rule is not implementable regardless of what the test says.

Expect the honest baseline to be hard. A simple buy and hold of a broad index, measured on a total return basis and after costs, is a demanding bar. Any strategy that does not clearly clear it after realistic frictions has not earned its complexity.

Using backtests for what they are good at

None of this argues against testing. It argues for a different job description.

A backtest is excellent at elimination. Most ideas fail obviously once they meet costs, turnover and a proper universe. Killing them cheaply is enormous value.

It is excellent at behavioural description. What does this rule do in a crash? How concentrated does it get? What is its turnover? How long is its worst underwater period? These are structural properties of the rule and they transfer to live much better than the return figure does.

It is excellent at cost discovery. Turnover, average holding period and the size distribution of holdings tell you what implementing the rule will actually cost and whether it is implementable at your size at all.

What it is weak at is forecasting return. Treat the return line as the least reliable output on the page, and the risk, turnover and drawdown lines as the most useful. That inversion, more than any statistical technique, is what separates research that survives contact with live capital from research that does not.

This article is educational. Altys Labs is not a registered research analyst or investment adviser, and nothing here is investment advice or a recommendation to buy, sell, or hold any security.

Frequently asked questions

Why don't backtest results repeat in live investing?

Because a backtest is one measurement of one rule on one stretch of history, and several things separate it from live results: choices made after seeing the data, market conditions that shift, costs and liquidity that a test understates, and ordinary luck. Live results are usually more muted than the test, and sometimes materially different.

How much degradation from backtest to live is normal?

There is no reliable universal number, and any figure quoted as one should be treated with suspicion. What is consistent is the direction: live results tend to be worse than the tested version, because the test was chosen partly for looking good and because real trading adds frictions the test did not model.

Does a backtest have any value then?

Yes, but its value is diagnostic rather than predictive. A backtest is good at ruling ideas out, exposing how a rule behaves in bad periods, and showing what it costs to run. It is weak as a forecast of the return you will earn.