Methodology

What Is Data-Driven Investing? A Practical Framework for Indian Investors

Data-driven investing uses explicit evidence and repeatable rules, while preserving context, sources and human judgement. Here is a practical Indian framework.

#data-driven-investing#systematic-investing#investment-research#india#portfolio-governance
What Is Data-Driven Investing? A Practical Framework for Indian Investors

Data-driven investing means using explicit evidence, consistent calculations and a repeatable decision process. It does not mean handing a portfolio to a model or treating a large dataset as truth. The useful version combines numbers with source context, time, human judgement and a record that can be checked later.

For Indian investors, the difference is important. There is plenty of financial data. The harder problem is knowing which version was available, what it actually represented and how it entered the decision.

Data is not the same as evidence

A number becomes evidence only after four questions are answered:

  1. What does it measure? Revenue, billings and gross merchandise value are not interchangeable.
  2. Which entity and basis does it cover? Consolidated and standalone statements can describe different economic realities.
  3. Which period does it describe? A quarter, trailing year and financial year answer different questions.
  4. When did it become knowable? The period end is not necessarily the publication date.

Without those labels, more data can create more confidence without creating more knowledge.

Consider a company reporting 20 percent profit growth. The figure may be correct. It may also include a one-time asset sale, reflect a lower tax rate, compare against an unusually weak quarter or exclude a discontinued business. The number is data. The explanation determines whether it supports the thesis.

Data-driven does not mean fully quantitative

Quantitative investing usually represents selection, sizing or trading decisions through mathematical rules. Data-driven investing is broader.

A fundamental analyst can be deeply data-driven while still making a judgement-based decision. The analyst may:

  • compare reported growth with management’s earlier guidance;
  • track unit volumes, realizations, capacity and working capital;
  • score balance-sheet resilience consistently across peers;
  • record the assumptions behind a valuation model;
  • define which operating changes would weaken the thesis;
  • review forecast errors after results are reported.

The final decision may still be human. What changes is that the evidence and reasoning are explicit.

The opposite is also possible. A strategy can be mathematically complicated but weakly data-driven if it uses restated history, ignores missing values, overfits parameters or cannot explain where its inputs came from.

A practical seven-stage framework

1. Begin with a question, not a dataset

“Find something interesting” is not a research question.

A better question is: do companies with improving cash conversion and stable returns on capital behave differently from companies whose reported profit grows without operating cash flow?

The question determines the data needed, the comparison group and what would count as contrary evidence.

2. Define the information set

List every input and its acceptable source. For Indian company research this may include exchange filings, financial statements, annual reports, investor presentations, concalls, shareholding disclosures, corporate actions and prices.

Then define the cut-off. If the research asks what could have been known on 30 June 2024, a restatement published in 2025 cannot enter the answer.

This is the point-in-time discipline described in our point-in-time data guide.

3. Make the calculation reproducible

Write the formula, units, reporting basis, lookback window and missing-data treatment.

“High quality” is not reproducible. Neither is “reasonable leverage.”

A score becomes reproducible when another analyst can use the same inputs and reach the same result. If a metric is unavailable, preserve that absence. Treating missing as zero does not fill a gap. It changes the meaning of the model.

4. Add business context

The calculation narrows the research. It does not finish it.

A falling margin may be caused by a temporary input-cost increase, a deliberate investment cycle, weaker pricing, adverse mix or a structurally changed business. A debt increase may fund productive capacity or cover persistent cash shortfalls.

Read the filing, concall and operating evidence required to distinguish those explanations. Record which explanation is supported and which remains an assumption.

5. Connect company evidence to the portfolio

A company can be attractive in isolation and redundant inside the existing book.

Data-driven portfolio research asks:

  • How much sector, factor, currency or commodity exposure already exists?
  • Does the position increase drawdown risk in the same scenario as other holdings?
  • Is the proposed size consistent with liquidity and uncertainty?
  • Which mandate or concentration limit applies?

The relevant unit of analysis is eventually the portfolio, not the company page.

6. Define what will be monitored

Turn the thesis into a small set of observable conditions.

For a capacity-led growth thesis, those conditions might include commissioning dates, utilization, volume growth, balance-sheet funding and the margin assumed at mature utilization. For a lender, they might include asset quality, credit cost, deposit growth, funding mix and capital adequacy.

The purpose is not to create more alerts. It is to recognize the evidence that changes the investment case.

7. Close the feedback loop

After the outcome is known, compare it with the original expectation.

Do not ask only whether the stock went up. Ask whether revenue, margin, cash flow and the operating driver behaved as expected. Separate thesis quality, forecast quality, valuation, position sizing and timing.

Repeated errors are valuable. A team that consistently overestimates margins in one sector has discovered something about its own process.

A simple example

Suppose an investor wants to find companies with improving business quality.

The undisciplined version screens for high return on capital and recent profit growth, then reads whichever stories look attractive.

The data-driven version is more explicit:

  1. Define the eligible universe and liquidity requirement as of each date.
  2. Specify return on capital, cash conversion, balance-sheet and growth measures.
  3. Preserve the publication date and reporting basis of every input.
  4. Rank companies using a documented score and expose unavailable values.
  5. Test the rule across different periods, sectors and reasonable parameter changes.
  6. Review filings for one-off gains, capitalized costs, customer concentration and business-model changes the score cannot see.
  7. Apply portfolio caps and record any human veto.
  8. Monitor the operating conditions supporting each approved holding.
  9. Compare subsequent evidence with the original thesis.

The second process is not guaranteed to make more money. It is more likely to reveal whether the method was coherent, implemented consistently and improved honestly.

Five traps that make data-backed research look smarter than it is

Hindsight hidden in clean history

Today’s database may contain restated numbers and current index constituents. A backtest using them can know things the investor did not know.

False precision

A score of 83.7 does not imply the underlying evidence is known to one decimal place. Precise arithmetic can sit on uncertain definitions.

Overfitting

If dozens of thresholds were tried and only the best result is shown, the backtest is partly a search through noise. Test the rationale, nearby specifications and held-back data.

Correlation mistaken for explanation

Two variables moving together does not establish a business mechanism. The investment case should explain why the relationship could persist and what would break it.

The spreadsheet that cannot be reconstructed

A large model is not automatically auditable. Hard-coded values, broken links, overwritten assumptions and unclear versions can make a spreadsheet as opaque as any software system.

Verification needs structure, not merely a file extension.

Why Excel still matters

Investment teams already use Excel to model, review and challenge assumptions. Data-driven software should work with that reality.

A useful export should let an analyst inspect:

  • the values and units used;
  • their reporting basis and availability date;
  • the formula or score definition;
  • the version of the rule;
  • any exclusions or missing values;
  • the resulting rank, weight or alert;
  • links or identifiers for the supporting source.

This is more powerful than a screenshot or a polished PDF. The analyst can recompute the result, change an assumption and find exactly where disagreement begins.

Where Altys fits

Altys is an India-first investment research and rule-based portfolio governance platform for PMS, AIF, family-office and research teams.

It is designed to connect the data-driven process rather than stop at a company answer. Source-linked Indian financial and qualitative evidence feeds screens, deterministic factors, scorecards, point-in-time backtests, models, portfolio rules and monitoring conditions. The human team owns the thesis and capital-allocation decision.

Important strategy, scorecard, research and monitoring outputs can be exported to Excel or Excel-ready files for independent validation. The principle is simple: software doing serious financial work should show enough of its work to be challenged.

Altys is not a broker, stock-tip service or substitute for investment judgement.

The conclusion

Data-driven investing is not about collecting the largest possible dataset. It is about reducing ambiguity in the path from evidence to decision.

The valuable questions are ordinary ones: what did we know, where did it come from, what did we calculate, what did we assume, why did we make an exception, what changed and what did we learn?

When those questions have durable answers, data becomes a process rather than decoration.

This article is educational. It does not recommend any security or strategy. Altys Labs is not a SEBI-registered Research Analyst or Investment Adviser.

Frequently asked questions

What is data-driven investing?

Data-driven investing means important decisions are supported by defined evidence, consistent calculations and a repeatable review process. It does not mean blindly buying whatever a model ranks highest. The data proposes and tests; human judgement defines the question, investigates context and owns the decision.

Is data-driven investing the same as quant investing?

Not necessarily. Quant investing usually expresses selection and portfolio decisions mathematically. A fundamental investor can also be data-driven by using consistent financial evidence, explicit scorecards, source-linked research, documented assumptions and measurable monitoring conditions while retaining human judgement.

What data matters for Indian equity research?

The useful set depends on the strategy, but it commonly includes reported financials, filings, concalls, management guidance, company-specific operating metrics, ownership, prices, corporate actions, valuation, liquidity, macro conditions and portfolio exposure. The source, reporting basis and date each item became knowable matter as much as the value.

Why does point-in-time data matter?

Historical databases are often cleaned and restated using information published later. Point-in-time data lets a researcher reconstruct what was actually available on a past decision date, reducing lookahead bias in backtests and making old decisions reviewable.

Can data-driven research be verified in Excel?

Yes. A good process should let an analyst export important inputs, dates, definitions and calculations so a result can be reproduced independently. Excel is useful because investment teams already understand its formulas, review conventions and audit trail.