Retrieval Helps Investment Research Only When the Question Needs It

A new financial-reasoning benchmark points to a practical workflow: let retrieval rescue context-complete questions, but measure when it changes a correct answer.

Retrieval Helps Investment Research Only When the Question Needs It

The useful change in AI investment research is becoming narrower and more practical: retrieval should be a decision in the workflow, not a reflex attached to every question. A new benchmark, FinExam-10K, tests models on finance-exam questions that combine domain knowledge, calculation, and judgment. Its result is a warning against “just add more context.”

Across 17 models, the best overall result was 85.29%, but the harder full-coverage track was much lower. Retrieval tools rescued hundreds of errors and overturned many correct answers at the same time. A gate that decided when to invoke a graph-based retrieval system called it for 7.9% of held-out questions and improved accuracy from 70.83% to 71.23%. That is a modest, statistically reported improvement—not an investing edge—but it gives researchers a concrete operating principle: first answer, then retrieve selectively, then compare.

That changes the old research workflow. Previously, an analyst might search documents first, paste a large bundle into a model, and trust the fluent synthesis. The newer workflow treats the model’s initial answer as a triage signal. If the question is answerable from the supplied record and the model is confident, preserve the answer. If it requires a definition, a calculation trail, or a missing passage, retrieve only the relevant evidence and run the question again. The important output is not merely an answer; it is a record of whether retrieval helped, harmed, or was unnecessary.

For an investment team, this is useful in tasks such as checking a filing-derived ratio, reconciling a management statement with a footnote, or classifying a risk disclosure. It is less useful as a prompt to produce a security view. The benchmark concerns structured reasoning under an evaluation protocol, not live markets, prices, execution, or returns. Its evidence supports a workflow design, not a claim that retrieval improves portfolio performance.

Here is a reproducible version. Build a small, time-stamped packet from public filings or research notes. Write 20 questions before opening the model: ten that should be answerable from the packet, five that require arithmetic, and five whose answer is deliberately absent. Ask the model to answer with a short rationale and an evidence pointer. Then let a simple gate choose “retrieve” only when the answer lacks a cited passage, contains an unresolved calculation, or expresses uncertainty. Retrieve one or two relevant passages, not the whole archive. Ask for a revised answer and a change log. A human reviewer checks the arithmetic, the date of every source, and whether the cited text actually entails the conclusion.

This approach also fits the lesson from our earlier piece on decision traces for investment AI: an audit trail is part of the research product. A second useful companion is the discussion of temporal integrity in financial foundation models. Retrieval can improve coverage while quietly importing information that was unavailable at the decision date. The timestamp and source boundary therefore belong in the test, not in a footnote.

Try it in 60 minutes as research-only work. In the first 15 minutes, assemble a dated packet and label each passage with its publication time. In the next 20, create and answer the 20-question set without retrieval. Spend 15 minutes on gated retrieval and revision, then 10 minutes reviewing the log. Use the no-retrieval run as the baseline. Track exact-answer accuracy, arithmetic accuracy, citation entailment (does the passage support the answer?), retrieval rate, and harmful reversals—the share of initially correct answers that become wrong after retrieval. Also record abstentions on deliberately missing questions.

Stop the exercise if you cannot preserve the packet and timestamps, if a reviewer cannot reproduce an answer from the cited passage, or if retrieval produces more harmful reversals than genuine rescues. Do not turn the result into a paper trade or live process. A seven-day extension can repeat the same frozen test on new documents, but keep the scoring rubric fixed and separate each day’s questions from the model’s prior outputs.

The central limitation is selection. A gate trained or tuned on public examples may not recognize a new document style, a changed accounting rule, or a market regime. Retrieval quality can also be the bottleneck: a relevant-looking passage may be stale, incomplete, or semantically adjacent rather than probative. The FinExam-10K result itself shows why aggregate accuracy is insufficient: retrieval can help and hurt, and the average gain can hide both. In production, permissions, document licensing, privacy, model-version changes, and review capacity add further friction.

The right lesson for discretionary investors is to measure evidence quality and reversals, not prompt fluency. Systematic researchers should treat retrieval as a conditional component with an out-of-sample gate and a fixed baseline. Builders should expose the decision log, source timestamps, abstention path, and reviewer override as first-class outputs. AI has changed the research workflow when it makes these comparisons cheap and visible. It has not removed the need to decide whether the evidence is complete, timely, and actually relevant.


阅读中文版本 →