Build vs Buy: How a New Investment Firm Should Set Up Its Research Stack
Which investment-research capabilities should a new Indian PMS, AIF or family office build, buy or keep in Excel? A practical layer-by-layer decision framework.
A new investment firm should rarely build its complete research platform from scratch. It should buy the common infrastructure, preserve its proprietary investment process, and build only where custom work creates an edge that another firm cannot purchase from the same vendor.
The difficult part is knowing which layer is which.
Foundation models are increasingly available to everyone. Public filings are available to everyone. A polished company summary is therefore unlikely to remain proprietary. The defensible layers are usually the firm’s data history, investment rules, models, decision records, portfolio context and feedback from forecasts that were right or wrong.
Separate the stack into layers
| Layer | Default choice | Why |
|---|---|---|
| Foundation language model | Buy | Expensive to train; increasingly commoditised |
| Public-document collection | Buy | High maintenance, little strategic differentiation |
| Financial-data normalisation | Buy, then validate | Difficult and essential, but shared across firms |
| Point-in-time history | Buy | Operationally demanding and easy to get subtly wrong |
| Screens and calculations | Configure | The engine can be shared; the rule should belong to the firm |
| Scorecards and factor models | Configure or build | Investment philosophy starts becoming proprietary here |
| Forecast models | Hybrid | Reuse plumbing; own assumptions and driver logic |
| Decision workflow | Configure | Preserve the firm’s approvals, exceptions and evidence standard |
| Portfolio context | Integrate | Must reflect the actual book and mandate |
| Monitoring rules | Configure | Company-specific guideposts belong to the research process |
| Research memory | Own | Decisions, outcomes and lessons compound over time |
This is a starting point, not a universal answer. A systematic fund may build more of the factor and execution layer. A concentrated family office may configure more and write less software. The principle is to place engineering where it compounds the firm’s distinctive judgement.
Why “we can build a chatbot” is the wrong estimate
A convincing research demo can be assembled quickly: connect a language model to a folder of annual reports, add a chat interface and ask questions.
The prototype becomes a production research system only after it can answer harder operational questions:
- Which document version is authoritative?
- How are company identities resolved across exchanges and subsidiaries?
- Is the number quarterly, year-to-date or annual?
- How are standalone and consolidated results handled?
- What did the historical user actually know on that date?
- Which formulas apply to banks, insurers and industrial companies?
- What happens when a source disappears or changes format?
- Can every firm see only its own portfolios and notes?
- Can the system prove why an alert fired?
- Can an analyst export and reproduce the result?
The chatbot is the visible tip. Data operations, quality controls and workflow state are the larger cost underneath it.
The hidden costs of building
Data collection never ends
Exchange pages change, PDFs break, company identifiers conflict, filings arrive late and historical documents disappear. The work is not a one-time scrape. It is a monitored production pipeline with retries, freshness checks and honest gaps.
Normalisation needs domain rules
The same label can mean different things across companies. Sector-specific economics cannot be reduced to one generic template. Restatements and corporate actions create seams that must be handled consistently.
Point-in-time history is a separate product
Storing the latest value is easy. Preserving every relevant vintage and deciding when a figure became knowable is harder. Without it, a historical screen may quietly use hindsight.
Quality requires independent controls
A system cannot validate a number by comparing it only with the source that supplied it. Data-quality work needs identities, coverage expectations, independent checks and explicit unavailable states.
Maintenance competes with research
Every engineer maintaining parsers, document queues, permissions and alert jobs is an engineer not improving the firm’s genuinely proprietary investment method. That may still be the right trade, but it should be counted honestly.
What a firm should own
The firm’s edge begins where generic infrastructure stops.
The mandate and universe
Which companies are eligible, which risks are unacceptable and what liquidity the strategy requires are firm decisions.
The research questions
One team cares most about working-capital discipline; another focuses on reinvestment runway; a third specialises in management change. The system should encode those questions without pretending every investor wants the same answer.
Scorecards and model logic
Weights, gates, peer groups, forecast drivers and override rules express the philosophy. Even if software evaluates them, the definition belongs to the firm.
Portfolio context
The same company research leads to different decisions in different books. Position size, correlated exposures, mandate constraints and liquidity are proprietary context.
Decision memory
What the firm believed, which exception it allowed, how the forecast missed and what changed afterward becomes more valuable with every cycle. A competitor cannot buy that history on day one.
The hybrid architecture that usually wins
For most new Indian investment teams, the sensible design is:
bought data and document infrastructure
↓
firm-configured screens and scorecards
↓
Excel or Python models with owned assumptions
↓
firm-specific decision and exception workflow
↓
portfolio-aware monitoring rules
↓
owned research memory and feedback
AI sits across this architecture as an interface and a reading layer. It should not be allowed to blur which parts are sourced, calculated, assumed or decided.
When building more is justified
Custom engineering makes sense when at least one condition is true:
- the data is genuinely proprietary;
- the strategy needs a calculation unavailable in standard systems;
- latency or scale changes the investment outcome;
- the workflow is a repeatable source of differentiation;
- integration with internal data creates a meaningful feedback loop;
- the firm has enough technical capacity to maintain the system through market cycles.
“We prefer our own interface” is usually not enough. “Our forecasting and decision history improves a model no outside vendor can reproduce” may be.
When buying is the disciplined choice
Buy when the capability is necessary but not distinctive:
- collecting public filings;
- standardising common financial statements;
- maintaining corporate actions;
- hosting document search;
- handling user permissions;
- running routine data freshness checks;
- exporting common analytical tables.
The vendor still needs scrutiny. Buying weak infrastructure does not create leverage; it creates a dependency. Use an investment research software checklist and insist on source links, point-in-time history, reproducible calculations and exportability.
Where Excel belongs
Excel is neither the old world nor the enemy. It remains one of the best environments for inspecting assumptions, building driver-based models and challenging calculations.
The boundary is important:
- Use Excel to model and verify.
- Do not use one workbook as the only source of truth for documents, permissions, monitoring, decision history and portfolio-wide data.
A good research platform should make Excel more powerful by feeding it clean inputs and exporting transparent logic, not make it disappear.
Where Altys fits
Altys for new investment firms supplies the shared infrastructure while keeping the firm’s investment process configurable.
Altys collects and organises Indian financials, filings, concalls, guidance, ownership, factors, mutual funds, macro and alternative data on a point-in-time foundation. Firms can express their own screens, scorecards, research questions, models and monitoring rules. Core outputs export to Excel so the team can inspect the inputs and calculations outside the application.
The proposition is not “outsource your investment process.” It is “stop rebuilding the common plumbing, and spend the firm’s scarce time on the parts only the firm can own.”
Frequently asked questions
Should a new investment firm build its own research platform?
Usually not from scratch. Most firms should buy common data collection and workflow infrastructure, preserve their own investment rules, models and decision memory, and build only the layers that create a genuine proprietary advantage.
What should an investment firm keep proprietary?
The firm's mandate, factor definitions, scorecards, forecast assumptions, research questions, decision history, portfolio context and feedback loops are usually more proprietary than the foundation model itself.
Can a new investment firm continue using Excel?
Yes. Excel is valuable for modelling, assumptions and independent verification. It becomes risky when it is also asked to be the firm's database, document archive, permission system and continuous monitoring engine.
What is the hidden cost of building an internal research stack?
The largest costs are usually data licensing and normalization, point-in-time history, changing source formats, quality controls, permissions, monitoring operations and ongoing maintenance rather than the first interface or AI prototype.
What is Altys's build-versus-buy position?
Altys supplies India-first data and research infrastructure while letting firms configure and retain their own screens, scorecards, models, monitoring rules and Excel-based verification. It is designed as a system around the firm's process rather than a replacement for it.