Tag
#data-quality
19 articles
-
Common Backtesting Mistakes That Make Results Look Better Than Reality
The most common backtesting mistakes are future information leaking into past decisions, survivorship in the universe, ignored costs, and testing so many variations that something looks good by luck.
Read article -
Corporate Actions and Adjusted Prices: Why Price History Must Be Restated
Splits, bonuses, rights issues and dividends change the share count or pay cash out, so raw price history breaks. Here is how adjustment works and where it still misleads.
Read article -
Quant Investing in India: The State of It, Plainly
Quant investing means decisions driven by measured data and explicit models. Here is what that looks like in India today, the data realities, and the honest limits.
Read article -
Survivorship Bias in Backtests: Why Today's Index Lies About the Past
Survivorship bias is testing a strategy on companies that survived to today. Delisted, merged, and dropped names disappear from the sample, so historical results improve for reasons unrelated to the strategy.
Read article -
Revenue Segmentation Is Harder Than It Looks
Mapping a company's revenue to its real business segments sounds like reading a table. In practice, inconsistent disclosure, shifting definitions, and reclassifications make a clean segment history genuinely hard to build.
Read article -
Structured vs Unstructured Financial Data: Why Real Analysis Needs Both
Structured financial data is the neat tables. Unstructured data is the concalls, notes, and filings around them. Real analysis needs both, and the unstructured half is the hard half.
Read article -
Why Point-in-Time Databases Are Hard to Build
A point-in-time database stores every financial number the way it was actually known on each past date. That sounds simple, but restatements, reclassifications, and corporate actions make it one of the hardest things in financial data.
Read article -
Why Restatements Break Models and Backtests
Restatements and reclassifications quietly rewrite a company's past, so a model or backtest built on today's numbers acts on figures nobody could have seen at the time.
Read article -
Why XBRL Isn't Enough for Real Financial Analysis
XBRL turns filings into machine-readable tags, which is genuinely useful, but tagged numbers are not the same as analysis-ready data. Here is the gap and why it matters.
Read article -
Building an AI That Understands Financial Statements
Understanding a financial statement is not reading its words. It means normalising the data, respecting the accounting identities, and cross-checking every number against the other statements.
Read article -
Structuring Decades of Filings So an AI Can Actually Use Them
A language model cannot reason over a messy pile of filings. Labels drift, statements get restated, formats change, and history is not what it looks like today.
Read article -
The Engineering Challenges Behind Institutional AI
Institutional-grade financial AI is hard for five reasons: data quality, point-in-time correctness, citations, deterministic outputs, and coverage at scale. Here is each one.
Read article -
Why Point-in-Time Data Matters in Research and Backtests
Point-in-time data means using the numbers that were actually knowable on a given date, not today's restated version. Skip it and your research quietly looks smarter than it was.
Read article -
The Hardest Part of AI in Finance Is Not the Model. It Is the Data.
In financial AI, the model is fast becoming a commodity. The durable edge lives in disciplined data work: units, restatements, point-in-time correctness.
Read article -
Can AI Actually Read a Balance Sheet? Where LLMs Break on Financial Statements
Language models are strong at prose and weak at accounting. Here is exactly where they break on real filings, and what makes machine reading of statements reliable.
Read article -
Lookahead Bias, Explained: The Silent Killer of Stock Backtests
Lookahead bias is when a backtest uses information it could not have known at the time. It quietly inflates results, and point in time data is the only real fix.
Read article -
Why AI Investing Apps Keep Getting Indian Stocks Wrong
Most AI investing tools are built for clean global data. Indian equities are full of local quirks that make those tools confidently wrong. Here is why.
Read article -
Why ChatGPT Hallucinates Financial Numbers, and How to Catch It
General chatbots predict plausible text, they do not look up facts, so they invent revenue and profit numbers. Here is why, and how to catch it.
Read article -
Data Quality Beats Model Quality: A Year Reading Indian Filings
After a year building AI to read Indian company filings, the biggest gains came from boring data discipline, not from a better model. Here is what actually moved the needle.
Read article