Why Restatements Break Models and Backtests
Restatements and reclassifications quietly rewrite a company's past, so a model or backtest built on today's numbers acts on figures nobody could have seen at the time.
Restatements break models and backtests because they rewrite a company’s past after the fact, and most models read that rewritten past as if it had always been visible. When a business recasts an earlier year, the number your model depends on quietly changes, so a test run on today’s history ends up acting on figures nobody could have seen on the dates it claims to act. The result looks sharper than reality, and the error is almost invisible because nothing in the data looks broken.
This is one of the most common ways careful-looking work turns out to be quietly wrong. There is no missing file, no obvious typo, just a series that has been tidied up since the events you are modelling. That tidiness is the problem.
What a restatement actually does to your data
When a company reports a period, it does not only print this quarter’s results. It reprints the comparable earlier period so you can see the change. Most of the time that earlier figure matches what was originally filed. Often it does not, and for legitimate reasons.
A company revises its own history when the basis of the numbers changes. The usual causes are routine:
- Accounting standard changes. When a new rule on leases, revenue recognition, or financial instruments arrives, prior years get recast onto the new basis so the trend stays comparable.
- Demergers and spin-offs. When a business is separated into its own listed entity, the parent’s earlier revenue and profit are restated to exclude it, so you compare like with like.
- Discontinued operations. When a division is sold or wound down, its results are pulled out of continuing operations and shown separately, including in the prior-year column.
- Error corrections. Occasionally a genuine mistake is found and the affected periods are corrected.
None of this is a scandal. A serious analyst wants comparable periods, and recasting the past is how the company delivers them. The trouble starts when a data source silently overwrites the old figure with the new one, and your model then treats the revised number as if it had always been on the tape.
Reclassifications are the quieter cousin
Restatements at least change a number, which gives you a chance to notice. Reclassifications are sneakier, because they can leave the headline untouched.
A reclassification moves amounts between lines or segments without changing the totals. Revenue that used to sit in one segment gets folded into another after a reorganisation. A cost that lived in one expense line moves to a different one. A division shifts from continuing operations into discontinued operations. Total revenue and total profit may look identical, so a quick glance sees nothing wrong. But any model keyed to a specific line or segment is now reading a different series than the one that existed at the time.
This matters because most real models are built on the parts, not just the top total. A forecast that keys off a particular segment’s revenue, or a screen that filters on a specific margin, depends on those sub-lines being stable through time. When a reclassification redraws them, the history your model learned from is no longer the history that was published. The break is real even though the headline never moved.
How the break shows up in a backtest
A backtest is a promise: it claims to know only what was public on each decision date. Restated history breaks that promise in the most flattering possible way.
Picture a company that sold a low-margin division. Before the sale, its reported margins carried that drag. After the sale, and in every recast prior year, the division is gone and the historical margins look structurally higher. Now run a naive study that says “companies whose margins were above some level three years ago went on to do well.” If the data is restated, some of those high past margins are an artifact of divisions removed only later. The study has quietly selected for companies that would go on to clean up their portfolios, which is a fact from the future dressed up as a fact from the past. The rule looks predictive. It is partly just remembering the answer.
This is the exact mechanism behind lookahead bias, where a test uses information it could not have had on the day it claims to act. Restated comparables are one of the cleanest ways lookahead sneaks in, precisely because nothing looks off. There is no error to catch, just a number that is a little too good because it was polished after the fact. And the same trap catches ordinary research, not only formal backtests. Any analyst comparing “then” to “now” using today’s recast history is exposed to it, which is why point-in-time data matters well beyond the quant desk.
The most dangerous number is not the one that is wrong. It is the one that is right today and was not knowable then.
Why as-first-reported data is the fix
The remedy is conceptually simple and operationally unglamorous. You keep each number the way it arrived, not the way it ended up.
That means storing the figure as it was originally filed, stamped with the date it first became public, and keeping every later revision as a separate record rather than painting over the old one. When you then ask what a company’s history looked like as of some past date, you get the version that was visible then. When you want the clean, fully recast series for a fair long-run trend, you can have that too. The point is that you can tell the two apart and choose on purpose.
Building and maintaining that kind of history is genuinely hard, for reasons worth understanding in their own right, which is the whole subject of why point-in-time databases are hard. The short version is that vintages, corporate actions, and reclassifications all conspire to make “the number for that quarter” a moving target rather than a fixed fact. But the payoff is exactly what a model needs: the ability to feed it only what was knowable at each moment.
This is also why data discipline outranks model cleverness here. As the broader argument in data quality beats model quality puts it, the smartest model in the world is only as honest as the history it trains and tests on. If that history is the polished, future-informed version, even a sophisticated system is grading itself on an exam it has already seen.
A practical checklist
You do not need special machinery to protect yourself from this. You need a few habits and a healthy suspicion of tidy history.
- Separate the period from the knowledge date. Every value describes some period and became knowable on some later date. A March quarter is not public in March. Keep both facts.
- Never overwrite a revision. When a number changes between two reports, keep the original alongside the new one. The gap between them is often the most interesting thing in the file.
- Ask what sits between two dates. Before comparing a company to its own past, check whether a demerger, a discontinued operation, a rule change, or a segment redraw happened in between. If it did, the raw comparison is not apples to apples.
- Test on as-first-reported data. For any backtest or historical study, use the version of each number that was public on the decision date, not the version that exists now.
- Treat a restatement as a prompt, not a nuisance. A number moving is a signal that the shape of the business changed. The explanation, found in the notes and the management discussion, is where the real insight usually lives.
None of this makes for exciting work. It makes bigger stores of data and more careful queries. But it is the difference between a model you can defend and one that is quietly lying to you. Companies will keep restating and reclassifying their histories, and they should. The work on your side is to remember that the past you can see today is not always the past that was visible then, and to build every model, screen, and backtest on the history that was actually knowable at the time.
This article is educational. Altys Labs is not a registered research analyst or investment adviser, and nothing here is investment advice or a recommendation to buy, sell, or hold any security.
Frequently asked questions
Why do restatements break models and backtests?
Because a restatement rewrites a company's past numbers after the fact. A model or backtest built on today's restated history acts on figures that were not knowable on the dates it claims to act, so it borrows information from the future and looks smarter than it could ever have been in real time.
What is the difference between a restatement and a reclassification?
A restatement changes the value of a past number, usually to correct an error or apply a new accounting basis. A reclassification keeps the totals the same but moves amounts between lines or segments, for example shifting a division into discontinued operations. Both quietly change the series a model reads, even when the headline profit looks unchanged.
What does as-first-reported data mean?
It means the version of each number as it was originally filed, stamped with the date it first became public, kept separate from any later revision. Building on as-first-reported data lets you test an idea using only what was actually knowable at the time, instead of a tidied-up version assembled later.
Are restatements a sign that something is wrong?
Usually not. Most restatements and reclassifications are routine and legitimate, done to keep periods comparable after a rule change, a demerger, or a reorganisation. The risk is not the restatement itself. It is silently feeding the revised history into a model or backtest that claims to have known only what was public at the time.