The Analyst in the Prompt: Separate Evidence From Investor Preference

A new audit of LLM financial analysis shows why the research workflow should split evidence scoring from personalized interpretation.

An evidence ledger separating neutral filing analysis from personalized investor preferences

The useful change in AI investment research is not that a language model can summarize a 10-K. It is that a research team can now test whether the model's conclusion changes when the evidence stays fixed but the investor persona changes. That turns “prompt quality” from a vague craft problem into an auditable workflow problem.

A September 2026 working paper, The Analyst in the Prompt, tests 3,575 SEC 10-K and 10-Q filings across 12 models. Its central finding is uncomfortable: most user-context spillover came from interpreting the same retrieved evidence differently, rather than merely retrieving different passages. The paper tests persona-conditioned retrieval, neutral retrieval, and memory-framed context, and finds that moving an investor mindset from the assistant's role into a user profile reduces—but does not eliminate—the drift. Read the study.

That changes the old workflow. Previously, an analyst might ask an AI assistant to “act like a cautious value investor” and then treat the resulting risk assessment as evidence-based research. The safer workflow has two outputs: an EvidenceScore, which must be invariant to the user's preferences, and a PersonalizedScore, which may adapt to a mandate. The first answers “what does this filing support?” The second answers “how might this matter for this mandate?” Mixing them creates model risk that citations alone cannot expose.

The workflow change: freeze the evidence, vary the context

An investment-data or research team can implement this without a trading agent. First, create a point-in-time document bundle: the filing, filing date, relevant tables, and retrieval IDs. Next, retrieve passages with a neutral query and freeze those passages. Ask the model for a structured evidence record: claim, quoted passage, page or section, direction of implication, uncertainty, and a neutral score from -1 to +1. Only then run the same passages through several mandate profiles—such as conservative income, growth, or momentum—and store the personalized output in a separate field.

The key comparison is not whether the personalized answers differ; they should. It is whether the neutral score, cited passage IDs, and uncertainty labels move when the profile changes. A change in the neutral field is interpretation spillover. A change in passage IDs is retrieval spillover. The distinction tells a builder where to add controls: retrieval snapshots for the first, schema separation and independent review for the second.

This is a useful complement to WisdomChain's earlier LLM forecasting memory firewall and agentic trading evidence ledger. The common principle is provenance: a fluent answer is not a stable research object until its inputs, time boundary, and transformation are inspectable.

A bounded test you can run safely

Run a 60-minute research-only exercise on three historical filings from different sectors. Use documents available before a fixed cutoff date; do not use current prices, future filings, or any live order. Prepare one neutral prompt and three mandate profiles. For each filing, record the neutral score, personalized score, cited passage IDs, missing-evidence flags, and whether a reviewer can reproduce the score from the frozen passages.

Use the ordinary unprofiled prompt as the baseline. Track four metrics: neutral-score range across profiles, passage-overlap rate, unsupported-claim count, and reviewer agreement on the evidence label. A practical observation window is the same session plus a 24-hour blind re-run if the model or system supports it. Stop the exercise if the system cannot preserve the document cutoff, if citations do not resolve to the frozen passages, or if a reviewer cannot distinguish evidence from preference. The output is a control report—not a portfolio signal.

Do not interpret a low score range as proof of robustness. Three filings are a diagnostic sample, not an estimate of general reliability. The paper's sample is broad, but it is still a working paper; model versions, prompts, retrieval systems, and SEC-document composition can change. Selection effects matter too: public filings are not the full information set used by professional investors, and a reproducible citation does not establish economic importance. Nor does a neutral score predict returns. Leakage, stale context, sector language, non-stationarity, and reviewer anchoring remain live risks.

The practical judgment is therefore narrow but actionable: personalization belongs downstream of an evidence layer. Discretionary investors should inspect whether a preference has silently entered the factual read. Systematic teams should log prompt versions, retrieval snapshots, and score deltas as model-risk fields. AI builders should optimize for invariant evidence records and explicit disagreement—not for a single persuasive verdict. The workflow has improved when a mandate can change the recommendation-shaped output without changing what the filing actually says.

For the native Chinese companion, see 把证据与投资偏好分开:提示词正在改变研究结论.


阅读中文版本 →