AI Can Reproduce Investment Research—If the Workflow Keeps Its Receipts

A new LLM research workflow points to a practical investment use: reproduce claims first, then separate verified evidence from model-generated extensions.

Editorial illustration of papers, code notebooks, and an evidence ledger connected by a verification checkpoint

The useful investment application of a large language model may be less glamorous than asking it for a forecast. It is a reproduction gate: can the system recover a published research result, show the data and code it used, and clearly separate the original finding from any extension it proposes?

A 2026 NBER working paper by Matthew Schwartz, Isaiah Andrews, and Jesse M. Shapiro, “An LLM Workflow That Reproduces, Improves, and Extends Published Economics Research”, makes this workflow concrete. The signal is not that a model has discovered a durable trading edge. It is that research assistants can now move from paper to executable replication plan faster—while leaving a larger surface area for audit.

That changes the front end of investment research. The old sequence was: read a paper, interpret its identification strategy, locate data, write code, and decide whether the result survives. The new sequence can put an LLM between each step: extract the claim, map variables to sources, draft code, run checks, and return a structured discrepancy report. The human analyst still owns the research question and the decision to trust the result, but the first pass becomes a testable artifact rather than a confident summary.

The workflow should be designed as a chain of receipts. Start with a paper and freeze its publication date, sample period, tables, and stated outcome. Ask the model for a claim register: one row per hypothesis, variable definition, sample restriction, and reported statistic. Then require every row to point to either a page/table in the paper, a raw data file, or a code cell. Let deterministic tools calculate joins, transformations, summary statistics, and confidence intervals. Use the model for translation between prose and code, error explanations, and candidate robustness checks—not for silently filling missing values or inventing a dataset.

The output is not “the AI agrees with the paper.” It is a three-way ledger: reproduced, not reproduced, and not testable. A fourth field records why: unavailable data, ambiguous definition, coding mismatch, sampling difference, or a genuine result discrepancy. This is the same operational lesson as an auditable research queue: the useful unit is a dated, inspectable task, not a polished paragraph. It also extends the grounding test for AI advice: citations are necessary, but they do not prove that the cited result was actually reproduced.

For an investment team, this can improve triage. A systematic researcher can use the ledger to decide which academic signals deserve a clean-room implementation. A discretionary analyst can distinguish a paper’s measured evidence from an LLM’s suggested economic story. A data or AI builder can measure failure at each handoff: retrieval, variable mapping, code execution, numerical comparison, and explanation. Those measurements are more useful than a single “research copilot accuracy” score because a workflow can summarize perfectly and still misalign a unit or time window.

Try a safe 60-minute research-only exercise. Choose one published table with public data and no live trading implication. In the first 15 minutes, record the paper’s date, sample, outcome, and baseline statistic by hand. In the next 25, have the model produce a claim register and a reproduction plan, but require source links and explicit unknowns. Spend 15 minutes running or checking the calculations with deterministic tools, then 5 minutes classifying each claim. Baseline: the unaided analyst’s manually recorded claim and statistic. Metrics: proportion of claims with valid provenance, exact numerical agreement where reproducible, false “reproduced” labels, and minutes to a reviewable discrepancy list. Stop if the model proposes data not in the paper, changes the sample without notice, or cannot preserve a clear human-readable audit trail. Do not convert the result into a trade or paper portfolio.

The causal gap matters. Faster reproduction does not establish economic value, out-of-sample robustness, or investability. Published studies are selected for attention; replication attempts may inherit survivorship and reporting bias. A model can also overfit the paper’s language, leak information from later versions, or make a plausible coding choice that changes the estimand. Data licensing, compute budgets, and analyst incentives affect which papers get tested. Even a successful reproduction may fail under a new regime, after costs and slippage, or when many teams act on the same signal.

The core judgment is therefore narrow: AI is becoming useful in investment research when it turns claims into inspectable experiments. The adoption test is not whether the model sounds like an analyst. It is whether the workflow preserves provenance, exposes disagreement, and makes it cheap to stop before an unverified result reaches capital.


阅读中文版本 →