AI & Finance

Can Your Financial AI Explain Its Assumptions?

A polished forecast is not enough. Investment teams should test whether financial AI can expose its assumptions, sources, sensitivities and review history.

#financial-ai#financial-modelling#assumptions#model-governance#equity-research
Can Your Financial AI Explain Its Assumptions?

An AI can produce a forecast that looks excellent and is still unusable.

Revenue grows neatly. Margins expand. The balance sheet balances. The valuation table has three scenarios and a sensitivity grid. Yet nobody can answer the question that matters: why is this the assumption?

For an investment team, the quality of a financial model is not measured by presentation. It is measured by whether another analyst can trace the model from source fact to calculation to assumption to decision—and challenge each step.

That is the test financial AI now has to pass.

The four layers a model must keep separate

Every serious model contains at least four kinds of information.

1. Reported facts

These are values a company disclosed: revenue, employee cost, gross advances, credit cost, capacity, order book or a sector-specific operating KPI. Each fact needs a source, period, reporting basis and date on which it became available.

2. Deterministic calculations

These include margins, growth rates, capital turns, working-capital days and valuation outputs. A calculation should have an explicit formula and appropriate inputs. It should also know when it does not apply. A generic industrial ROCE formula should not be forced onto a bank.

3. Forecast assumptions

These are beliefs about the future: volume, pricing, mix, cost, reinvestment, asset quality, funding, terminal growth or the multiple an investor may pay. An assumption should have an owner, rationale, effective date and scenario.

4. Investment decisions

The model can inform fair value, risk and expected return. It cannot decide the mandate, liquidity tolerance, correlated exposure or position size on its own. Those are governed portfolio choices.

Weak systems flatten all four layers into one fluent answer. Strong systems make the boundaries obvious.

A simple example: one number, three different meanings

Suppose a model shows next year’s revenue growth at 14%.

That number could be:

  • management’s guidance;
  • the analyst’s base-case assumption;
  • a mechanical extrapolation of the last four quarters;
  • a market-consensus estimate;
  • the output of an AI reading several documents.

They are not interchangeable.

If management guided to 12–15%, the analyst chose 14% because capacity ramps in the second half, and the model previously used 16%, then a useful system should preserve all three facts. “Revenue growth: 14%” throws away the reasoning.

A proper assumption record might contain:

FieldExample
MetricFY28 revenue growth
Approved value14%
Previous value16%
ScenarioBase
RationaleLower first-half volume; unchanged capacity milestone
EvidenceQ1 result, management call, capacity filing
OwnerNamed analyst
Approved onDated model revision
Review triggerCapacity delay or guidance below 12%

The number becomes research only when its lineage is visible.

The most common ways AI models go wrong

The right number on the wrong basis

The model mixes standalone and consolidated figures, quarterly and year-to-date periods, reported and adjusted earnings, or pre- and post-corporate-action share counts. The arithmetic can still look coherent.

A calculation is guessed rather than defined

Financial terms are not universal labels. Return on capital, free cash flow and even revenue may be assembled differently across companies and sectors. The system should apply a defined method, not choose a formula from linguistic familiarity.

A historical relationship is mistaken for a forecast reason

Margins rose with utilisation in the past, so the model extrapolates further expansion. But perhaps the next increment needs a new plant, commodity costs changed, or pricing competition returned. Correlation can suggest a question; it cannot supply the business explanation.

Management language becomes an analyst assumption

AI is good at extracting confident sentences. It can accidentally turn “we aspire to” into a base case. Guidance should be preserved as management’s statement, then compared with the analyst’s assumption and the company’s delivery history.

The model revises itself without a decision trail

An automatic update may change several years of revenue, working capital and value. If the previous version disappears, the investment committee cannot reconstruct why its view moved.

The narrative frames the conclusion

Language models are sensitive to context. A 2026 CFA Institute Research and Policy Center study sent 900 prompts across nine models and ten investment scenarios. The same financial information received different evaluations depending on positive or negative framing; simple reminders about bias were not enough. Human-balanced framing performed better. See the official CFA Institute research on LLM bias in investing.

For model governance, the implication is concrete: do not let a persuasive narrative become the invisible driver of a numerical assumption.

The seven tests to run on any financial AI

Test 1: ask for lineage

Pick five important historical inputs. Can the system show the source document, page or statement, period, basis and availability date? If not, the model cannot be independently reconstructed.

Test 2: change one assumption

Change one operating driver and observe the flow through revenue, profit, cash flow, balance sheet and value. Unrelated outputs should not move. Related outputs should move for an understandable reason.

Test 3: compare sectors

Ask the system to model an industrial company and a bank. It should not use the same schedule, capital definition or valuation logic. Sector awareness is a basic control, not a premium feature.

Test 4: introduce conflicting evidence

Give it management guidance, a weak operating KPI and an older internal assumption. Does it surface the conflict, or average the evidence into a smooth answer? Good research preserves disagreement until a person resolves it.

Test 5: test missing information

Remove a required input. The correct result may be “unavailable.” A guessed zero or plausible substitute is more dangerous than an explicit gap.

Test 6: inspect revisions

Can you see the previous value, proposed value, reason, evidence and approver? A model without versions is a calculator, not an institutional record.

Test 7: export and reproduce

Can the team take core values and analytical outputs into a spreadsheet, inspect them and reproduce the important conclusions? Export does not prove correctness, but a locked answer box prevents meaningful challenge.

Download the financial-AI assumption audit template to run these tests consistently.

A safer update workflow

The best division of labour is explicit:

AI finds the new evidence
          ↓
deterministic rules validate identities, periods and arithmetic
          ↓
AI maps the evidence to potentially affected assumptions
          ↓
the model shows a proposed change and sensitivity
          ↓
a named analyst accepts, edits or rejects it
          ↓
the approved revision becomes the new reference point

This workflow may look slower than unattended automation. In practice, it removes the far more expensive work of discovering later that nobody knows why the model changed.

What a good explanation sounds like

“We reduced FY28 revenue growth from 16% to 14%” is a change log.

“We reduced FY28 revenue growth from 16% to 14% because the Q1 filing moved commissioning from April to September. The base case now assumes three months less contribution; pricing and steady-state utilisation are unchanged. The revision lowers base-case value by 6%. Revisit if commissioning occurs before July or management guides below 12% growth” is an investment explanation.

It contains:

  • the old and new assumptions;
  • the source and causal driver;
  • what did not change;
  • the model impact;
  • the condition for another review.

AI can help draft this explanation. The evidence and approval still need to be real.

Where Excel fits

Excel remains a valuable inspection environment because analysts can expose formulas, run sensitivities and challenge inputs directly. The false choice is “AI or spreadsheets.”

A stronger architecture is:

  • sourced data and documents in a governed research layer;
  • deterministic calculations with explicit definitions;
  • AI assistance for retrieval, comparison and proposed revisions;
  • spreadsheets for independent inspection and custom analysis;
  • a versioned record of what the firm approved.

The spreadsheet should not be the only filing archive, permission system, alert engine and decision memory. But a research platform should not make verification harder merely to keep the user inside its interface.

Where Altys fits

Altys is designed to keep facts, calculations, assumptions and decisions distinct.

The platform’s India-focused research layer connects point-in-time financials, primary documents, management guidance, sector KPIs, strategies, scorecards, GenGrid research and monitoring. Core outputs across these workflows can be exported to Excel where applicable for independent checking.

The Thesis Model is currently an internal pilot. It supports selected operating-company FCFF workflows and standalone-bank schedules with a cited revenue map, visible assumptions, scenarios, DCF and sensitivity, reviewed AI patches and immutable private revisions. It also supports value export for review. This is deliberately narrower than promising that AI can build every financial model without supervision.

For eligible professional pilots, Altys Vault and Company Thesis connect selected firm-owned research versions to a cited, reviewable thesis workflow. Explicit AI actions may process selected excerpts through the configured model provider; firms should evaluate that handling against their own security and governance requirements.

The goal is not a model that sounds certain. It is a model whose uncertainty, evidence and judgement can all be inspected.

The new standard is explainable change

The first generation of financial AI proved that software can produce an impressive answer. The next generation has to prove why the answer changed.

For investment firms, that means assumption lineage, sector-appropriate logic, sensitivity, human approval, version history and portability. A model should make disagreement easier, not hide it behind better prose.

Before trusting the forecast, ask the simplest possible question: Can the system explain every material assumption well enough for a sceptical colleague to reproduce and reject it?

If the answer is no, the financial model is not yet ready to govern capital.

Frequently asked questions

What should financial AI explain about a forecast?

It should show the source period, calculation method, forecast assumptions, rationale, sensitivity and revision history. A user should be able to separate a reported fact from a formula, an estimate and a final investment decision.

Why are AI-generated financial models risky?

The output can look internally consistent while using the wrong reporting basis, period, formula or unsupported assumption. The risk is not only arithmetic error; it is a plausible model whose logic cannot be inspected.

Can AI update a financial model automatically?

AI can propose updates and map new evidence to affected assumptions. Material forecast changes should pass deterministic checks and named human review before becoming an approved model version.

Does exporting to Excel make an AI model reliable?

No. Exportability helps verification, but reliability also requires sourced inputs, explicit definitions, sector-appropriate formulas, version history and review controls. A spreadsheet is an inspection surface, not proof by itself.

How does Altys approach financial-model governance?

Altys separates sourced facts, deterministic calculations, analyst assumptions and approved revisions. Its Thesis Model is currently an internal pilot for selected operating-company and standalone-bank workflows rather than a universal modelling claim.