LLM Investor Simulators Need a Behavior Audit Before Portfolio Use

A new paper-trading study changes the useful role of LLMs in investment workflows: audit behavioral fidelity before using them to test decisions.

Abstract investment research desk comparing simulated and observed portfolio decisions

The useful question for an LLM investor simulator is not whether it can produce a plausible trade. It is whether it reproduces the decision process that a real investor would have followed at that point in time. A September 2026 arXiv study makes that distinction practical: across 1,239 aligned user-days from 80 participants in a longitudinal paper-trading study, no evaluated LLM reliably beat a simple recent-activity persistence baseline at predicting whether a person would trade. Accuracy deteriorated further when the task moved from “trade or not” to buy/sell structure, asset selection, and portfolio consequences.

That result does not make simulation useless. It changes the workflow. An investment team can use an LLM as a hypothesis generator for behavioral scenarios, but it should first pass a behavior-fidelity audit. A model that sounds like an investor is not automatically a model of that investor.

The workflow changed from generating personas to testing fidelity

The old workflow is familiar: describe an investor’s preferences, give an LLM market context, and ask what the investor might do. The new workflow treats the simulator as an evaluated measurement instrument. The system receives only information available before the next decision, produces a predicted action, and is scored against the participant’s subsequent paper-trading record.

The study’s rolling next-day design matters. It aligns user interactions, simulated transactions, virtual portfolio states, and point-in-time market information instead of letting the model see the future. Its hierarchy also prevents a flattering aggregate score from hiding a broken decision process: first activity, then action structure, then asset selection, then downstream portfolio trajectory. The researchers report that similar activity-level predictions can still lead to substantially different portfolios.

For research teams, this is a concrete shift in quality control. The output is no longer a narrative persona or a single hit rate. It is a dated prediction record with evidence, uncertainty, and failure labels.

A reproducible audit for investment research

Start with a paper-trading or historical replay dataset containing timestamps, available information, portfolio state, and the observed next action. Freeze the information boundary. For every decision date, provide the simulator only the prior history and the documents that were actually available. Ask for a structured response: trade/no trade, buy/sell/hold, asset or assets, confidence, evidence used, and abstention reason.

Run three baselines before comparing models:

  1. Persistence: repeat the participant’s recent activity state.
  2. Population rate: predict the most common action in the evaluation window.
  3. Human or rule benchmark: use a transparent policy with the same information boundary.

Score the layers separately. Use balanced accuracy or calibration for trade occurrence; macro-F1 for action structure; top-k or rank metrics for asset selection; and portfolio-level divergence for downstream consequences. Do not turn the last metric into a return claim. It is a measure of behavioral resemblance, not an investment recommendation.

The paper also reports that recent trading history strongly governs activity prediction, while asset selection is more sensitive to the available evidence. That suggests a useful system design: let a model summarize and stress-test the evidence, but keep the final behavioral label tied to an auditable record. The companion finding—that ticker-specific research predicts imminent trading observationally without a supported causal interpretation—is a warning against treating correlation as intent.

This complements the investment workflow in agentic trading’s evidence ledger, where claims and provenance are separated, and the series’ earlier discussion of why stronger AI models can become more correlated. A simulator adds another possible common failure: many plausible agents may compress a diverse set of human decisions into the same convenient story.

A 60-minute research-only exercise

Use 20–50 dated paper-trading decisions, or a replayable public dataset. Split them chronologically: the first 70% is the baseline window and the last 30% is the observation window. For each observation date, compare persistence, a transparent rule, and one LLM prompt under the same information cutoff. Record trade occurrence, action structure, asset selection, confidence, evidence citations, and abstentions.

The baseline is the persistence model. Metrics are balanced accuracy for activity, macro-F1 for action type, top-k asset overlap, calibration error, and the percentage of outputs that cite information published after the cutoff. Stop if timestamps cannot be verified, if the sample is too small to estimate uncertainty, or if the model begins producing personalized financial advice. Keep all outputs in a research log; do not place orders or infer suitability.

What the evidence does—and does not—say

This is a preprint under review, not proof that every LLM simulator fails or that human decisions are stable. The sample, participant behavior, market period, prompt design, and paper-trading setting create selection and measurement limits. The paper’s observational association does not establish causality. Survivorship and reporting bias can enter if only completed or legible decisions are retained. A model may also learn the dataset’s logging conventions rather than investor reasoning.

The core judgment is therefore narrow: before an LLM is used to stress-test a portfolio workflow, validate what it reproduces at each decision layer and compare it with simple baselines. The changed investment task is not “ask AI what an investor would do.” It is “measure whether the simulator earns the right to stand in for an investor at all.”

Sources: He et al., “Are LLMs Good Financial User Simulators?” (arXiv:2609.15727); Pozniak et al., “Why Better Models Can Create Riskier Systems” (arXiv:2609.04373).


阅读中文版本 →