Research Workflow

Build vs Buy: How a New Investment Firm Should Set Up Its Research Stack

Which investment-research capabilities should a new Indian PMS, AIF or family office build, buy or keep in Excel? A practical layer-by-layer decision framework.

#build-vs-buy#research-stack#new-firm#investment-technology#india
Build vs Buy: How a New Investment Firm Should Set Up Its Research Stack

A new investment firm should rarely build its complete research platform from scratch. It should buy the common infrastructure, preserve its proprietary investment process, and build only where custom work creates an edge that another firm cannot purchase from the same vendor.

The difficult part is knowing which layer is which.

Foundation models are increasingly available to everyone. Public filings are available to everyone. A polished company summary is therefore unlikely to remain proprietary. The defensible layers are usually the firm’s data history, investment rules, models, decision records, portfolio context and feedback from forecasts that were right or wrong.

Separate the stack into layers

LayerDefault choiceWhy
Foundation language modelBuyExpensive to train; increasingly commoditised
Public-document collectionBuyHigh maintenance, little strategic differentiation
Financial-data normalisationBuy, then validateDifficult and essential, but shared across firms
Point-in-time historyBuyOperationally demanding and easy to get subtly wrong
Screens and calculationsConfigureThe engine can be shared; the rule should belong to the firm
Scorecards and factor modelsConfigure or buildInvestment philosophy starts becoming proprietary here
Forecast modelsHybridReuse plumbing; own assumptions and driver logic
Decision workflowConfigurePreserve the firm’s approvals, exceptions and evidence standard
Portfolio contextIntegrateMust reflect the actual book and mandate
Monitoring rulesConfigureCompany-specific guideposts belong to the research process
Research memoryOwnDecisions, outcomes and lessons compound over time

This is a starting point, not a universal answer. A systematic fund may build more of the factor and execution layer. A concentrated family office may configure more and write less software. The principle is to place engineering where it compounds the firm’s distinctive judgement.

Why “we can build a chatbot” is the wrong estimate

A convincing research demo can be assembled quickly: connect a language model to a folder of annual reports, add a chat interface and ask questions.

The prototype becomes a production research system only after it can answer harder operational questions:

  • Which document version is authoritative?
  • How are company identities resolved across exchanges and subsidiaries?
  • Is the number quarterly, year-to-date or annual?
  • How are standalone and consolidated results handled?
  • What did the historical user actually know on that date?
  • Which formulas apply to banks, insurers and industrial companies?
  • What happens when a source disappears or changes format?
  • Can every firm see only its own portfolios and notes?
  • Can the system prove why an alert fired?
  • Can an analyst export and reproduce the result?

The chatbot is the visible tip. Data operations, quality controls and workflow state are the larger cost underneath it.

The hidden costs of building

Data collection never ends

Exchange pages change, PDFs break, company identifiers conflict, filings arrive late and historical documents disappear. The work is not a one-time scrape. It is a monitored production pipeline with retries, freshness checks and honest gaps.

Normalisation needs domain rules

The same label can mean different things across companies. Sector-specific economics cannot be reduced to one generic template. Restatements and corporate actions create seams that must be handled consistently.

Point-in-time history is a separate product

Storing the latest value is easy. Preserving every relevant vintage and deciding when a figure became knowable is harder. Without it, a historical screen may quietly use hindsight.

Quality requires independent controls

A system cannot validate a number by comparing it only with the source that supplied it. Data-quality work needs identities, coverage expectations, independent checks and explicit unavailable states.

Maintenance competes with research

Every engineer maintaining parsers, document queues, permissions and alert jobs is an engineer not improving the firm’s genuinely proprietary investment method. That may still be the right trade, but it should be counted honestly.

What a firm should own

The firm’s edge begins where generic infrastructure stops.

The mandate and universe

Which companies are eligible, which risks are unacceptable and what liquidity the strategy requires are firm decisions.

The research questions

One team cares most about working-capital discipline; another focuses on reinvestment runway; a third specialises in management change. The system should encode those questions without pretending every investor wants the same answer.

Scorecards and model logic

Weights, gates, peer groups, forecast drivers and override rules express the philosophy. Even if software evaluates them, the definition belongs to the firm.

Portfolio context

The same company research leads to different decisions in different books. Position size, correlated exposures, mandate constraints and liquidity are proprietary context.

Decision memory

What the firm believed, which exception it allowed, how the forecast missed and what changed afterward becomes more valuable with every cycle. A competitor cannot buy that history on day one.

The hybrid architecture that usually wins

For most new Indian investment teams, the sensible design is:

bought data and document infrastructure

firm-configured screens and scorecards

Excel or Python models with owned assumptions

firm-specific decision and exception workflow

portfolio-aware monitoring rules

owned research memory and feedback

AI sits across this architecture as an interface and a reading layer. It should not be allowed to blur which parts are sourced, calculated, assumed or decided.

When building more is justified

Custom engineering makes sense when at least one condition is true:

  • the data is genuinely proprietary;
  • the strategy needs a calculation unavailable in standard systems;
  • latency or scale changes the investment outcome;
  • the workflow is a repeatable source of differentiation;
  • integration with internal data creates a meaningful feedback loop;
  • the firm has enough technical capacity to maintain the system through market cycles.

“We prefer our own interface” is usually not enough. “Our forecasting and decision history improves a model no outside vendor can reproduce” may be.

When buying is the disciplined choice

Buy when the capability is necessary but not distinctive:

  • collecting public filings;
  • standardising common financial statements;
  • maintaining corporate actions;
  • hosting document search;
  • handling user permissions;
  • running routine data freshness checks;
  • exporting common analytical tables.

The vendor still needs scrutiny. Buying weak infrastructure does not create leverage; it creates a dependency. Use an investment research software checklist and insist on source links, point-in-time history, reproducible calculations and exportability.

Where Excel belongs

Excel is neither the old world nor the enemy. It remains one of the best environments for inspecting assumptions, building driver-based models and challenging calculations.

The boundary is important:

  • Use Excel to model and verify.
  • Do not use one workbook as the only source of truth for documents, permissions, monitoring, decision history and portfolio-wide data.

A good research platform should make Excel more powerful by feeding it clean inputs and exporting transparent logic, not make it disappear.

Where Altys fits

Altys for new investment firms supplies the shared infrastructure while keeping the firm’s investment process configurable.

Altys collects and organises Indian financials, filings, concalls, guidance, ownership, factors, mutual funds, macro and alternative data on a point-in-time foundation. Firms can express their own screens, scorecards, research questions, models and monitoring rules. Core outputs export to Excel so the team can inspect the inputs and calculations outside the application.

The proposition is not “outsource your investment process.” It is “stop rebuilding the common plumbing, and spend the firm’s scarce time on the parts only the firm can own.”

Frequently asked questions

Should a new investment firm build its own research platform?

Usually not from scratch. Most firms should buy common data collection and workflow infrastructure, preserve their own investment rules, models and decision memory, and build only the layers that create a genuine proprietary advantage.

What should an investment firm keep proprietary?

The firm's mandate, factor definitions, scorecards, forecast assumptions, research questions, decision history, portfolio context and feedback loops are usually more proprietary than the foundation model itself.

Can a new investment firm continue using Excel?

Yes. Excel is valuable for modelling, assumptions and independent verification. It becomes risky when it is also asked to be the firm's database, document archive, permission system and continuous monitoring engine.

What is the hidden cost of building an internal research stack?

The largest costs are usually data licensing and normalization, point-in-time history, changing source formats, quality controls, permissions, monitoring operations and ongoing maintenance rather than the first interface or AI prototype.

What is Altys's build-versus-buy position?

Altys supplies India-first data and research infrastructure while letting firms configure and retain their own screens, scorecards, models, monitoring rules and Excel-based verification. It is designed as a system around the firm's process rather than a replacement for it.