How to Read a Backtest Report
Read a backtest report in reverse order: setup and assumptions first, then risk and turnover, and the return figure last. Here is the sequence and the red flags.
The most useful way to read a backtest report is roughly backwards from how it is presented. Reports lead with the equity curve and the annualised return. Read those last. Start with the setup and the assumptions, move to risk and turnover, and only then look at what the rule earned, because the setup decides whether the return figure means anything at all.
This is not cynicism. It is the practical consequence of the fact that the return line is the most fragile output on the page and the mechanics are the most durable. A rule’s turnover, concentration and drawdown behaviour tend to persist. Its tested return does not.
Step one: read the setup before anything else
Four questions, all of them answerable before you look at a single performance number.
What was the universe, and how was it defined on each past date? This is the single most common place a report breaks. If the test ran on the companies that are in the index today, it silently excluded everything that failed, got delisted, or was acquired along the way. That is survivorship bias, and it flatters results in a way that no later statistic will reveal. A credible report states that the universe was reconstituted as it stood on each date.
What was the data basis? Were the financials used the ones that were public on the decision date, or today’s restated versions? Was the price series adjusted for splits, bonuses and dividends on a consistent basis? Were the returns computed against a total return index or a price index? Each of these quietly shifts the result, and the last one shifts it in the strategy’s favour if the benchmark excludes dividends.
What period, and why that period? Check the start date against what you know about market history. A start date that begins just after a major fall is a choice, whether or not it was a deliberate one. Ask whether the window contains at least one genuine bear market, one rate cycle, and one period where the strategy’s style was out of favour.
What costs and frictions were assumed? Look for explicit numbers: brokerage, exchange charges, securities transaction tax, stamp duty, GST, and an allowance for slippage. A report that says “costs of 0.1 percent per trade” is at least stating a testable assumption. A report that says nothing has assumed zero, and zero is wrong. Transaction costs in backtests are not a rounding item for anything that trades often.
If any of these four is missing or unsatisfying, stop. The rest of the report is built on it.
Step two: read the risk and mechanics section
Now go to the parts that describe how the rule behaves, not what it earned.
Maximum drawdown and time underwater. The worst peak to trough fall tells you the depth of the pain. The length of time spent below a previous high tells you the duration of it, which is usually what ends real allocations. A rule with a modest drawdown that took four years to recover is harder to hold than one with a deeper fall that recovered in six months. Maximum drawdown on its own is only half the picture.
Turnover. This is the cost engine. High turnover means costs, taxes and capacity constraints. It also means the tested result is more sensitive to the cost assumptions you checked in step one. If turnover is high and assumed costs are low, the report is fragile in a specific, identifiable way.
Number of holdings and concentration. A rule that ends up with a handful of positions has a return figure driven by a handful of outcomes. Check whether the report shows the contribution of the largest contributors. If removing the best two or three names collapses the result, the rule has not been demonstrated, it has been illustrated.
Liquidity and size profile of the holdings. Where in the market cap spectrum does the rule fish? A screen that consistently lands on small, thinly traded names may be entirely valid at a small allocation and impossible at a larger one. The equity curve looks the same either way, which is exactly the problem.
Exposure profile. What was the average cash level, the average sector weight, and the average sensitivity to the benchmark? A strategy that was half in cash for the tested period has a return figure that is partly a statement about cash, and risk metrics that are flattering for a reason unrelated to selection.
Step three: now read the returns, sceptically
Finally, the headline numbers. Read them in this order.
The benchmark comparison, on a fair benchmark. The absolute return matters less than the return against a defensible yardstick, measured on a total return basis after costs. A mid-cap oriented rule compared against a large-cap index is being graded generously.
Rolling returns rather than the single annualised figure. A compound annual growth rate is one number describing a whole decade, and it hides everything about consistency. Rolling windows show whether the result came from steady performance or from one extraordinary stretch.
Risk-adjusted measures, as a cross-check and not a verdict. Sharpe and its relatives are useful for comparing rules on a common scale, but they are computed from the same sample as everything else and inherit all its flaws. A high ratio on a flawed setup is a precise measurement of nothing.
The worst periods explicitly. Ask for the worst month, the worst quarter, and the behaviour during the specific market falls in the sample. Aggregates smooth away the episodes that determine whether a strategy is holdable.
The red flags, collected
| Red flag | Why it matters |
|---|---|
| Universe defined by today’s index members | Excludes failures and delistings, flattering the result |
| Zero or unstated transaction costs | Turnover then appears free, which it never is |
| Start date just after a major market fall | The worst regime is outside the sample by construction |
| Very smooth equity curve | Usually a sign of overfitting, low liquidity marks, or a short sample |
| Many parameters, few independent decisions | Not enough evidence to support the specificity of the rule |
| No disclosure of how many variants were tested | You cannot judge selection effects you cannot see |
| Return shown gross, benchmark shown ex-dividend | An unfair comparison in the strategy’s favour |
| Results collapse when the best names are removed | The rule was demonstrated by a few outcomes, not by its logic |
None of these is automatically disqualifying, and a report that discloses a weakness is more trustworthy than one that hides it. The signal is disclosure, not perfection.
What a backtest report does not tell you
Even a clean, well-constructed report has hard limits, and a reader who forgets them will still be misled.
It does not tell you the return you will earn. It tells you what one rule did on one sample, under one set of assumptions, after being selected from some unknown number of alternatives. Why backtest results don’t repeat covers the mechanisms in detail.
It does not tell you whether the rule has an economic reason to work. A statistical relationship without a plausible explanation for why it should persist is a relationship that can stop at any time without warning. The report cannot supply that reasoning; you have to.
It does not tell you how the rule behaves in regimes absent from the sample. Every backtest has an implicit blind spot the size of everything that did not happen during the window.
It does not tell you whether you can hold it. Tolerance for a four-year underperformance stretch is a personal and institutional fact, not a statistical one, and it is the constraint that ends most systematic programmes.
And it does not tell you about capacity, taxes and operational reality unless those were explicitly modelled. Most reports model none of them.
Read in this order, with these limits in mind, a backtest report becomes what it should be: a structured description of how a rule behaves and what it costs to run, rather than a projection of what it will pay. That is a less exciting document, and a far more useful one.
Related reading
- Portfolio metrics explained: the hub for the risk and return statistics that appear on these reports.
- Why backtest results don’t repeat: the four mechanisms behind the gap between tested and live outcomes.
- Common backtesting mistakes: the errors that produce the red flags listed above.
- Survivorship bias in backtests: why the index of today cannot define a past universe.
- Transaction costs in backtests: the drag that turns a good gross result into an ordinary net one.
This article is educational. Altys Labs is not a registered research analyst or investment adviser, and nothing here is investment advice or a recommendation to buy, sell, or hold any security.
Frequently asked questions
How should you read a backtest report?
Start with the setup rather than the headline return: the universe, the period, the data basis, and the assumed costs. Then read the risk and turnover section, because those properties transfer to live results better than returns do. Read the return figure last, and read it as a description of one sample rather than as an expectation.
What are the biggest red flags in a backtest report?
A universe defined by the companies that exist today, a start date that conveniently avoids a known market fall, zero or unstated transaction costs, a very smooth equity curve, more parameters than the number of independent decisions can support, and any report that does not disclose how many variants were tested.
Which metric matters most in a backtest?
No single one. The most durable information is in the risk and mechanics: maximum drawdown, time spent underwater, turnover, and the concentration and liquidity of the holdings. These describe how the rule behaves, which carries over to live investing far more reliably than the return figure does.