Agentic Factor Discovery Needs a Backtest of the Research Process
AI agents can search for investment signals, but the process that discovers them needs its own out-of-sample test.
The next failure in AI investing may not come from a bad factor. It may come from testing a good-looking factor while never testing the machine that kept searching until it found one.
A July 2026 paper, Agentic Empirical Asset Pricing: Methodological Foundations, gives this problem a useful name: Agentic Empirical Asset Pricing (AEAP). Its systems let language-model agents propose hypotheses, turn them into executable code, evaluate them, and refine the search. That changes the investment workflow from “run this model” to “operate a research process.” The practical judgment is simple: a static out-of-sample backtest of the final factor is necessary, but it is not enough. The discovery loop itself needs a time-ordered, repeatable audit.
This matters because selection is part of the model. Suppose an agent generates hundreds of candidate factors, rejects most of them, changes its prompts, fixes code, and reports the survivor. The survivor’s test-period result may be clean while the search procedure benefited from repeated chances, hidden leakage, or an accidentally favorable window. A final Sharpe ratio cannot tell those stories apart. The paper’s contribution is therefore less “AI found a factor” than “evaluate the researcher as an adaptive system.”
The proposed workflow separates three questions. First, can the system produce a useful pool of candidates? Second, do those candidates perform under a fair, point-in-time test? Third, are they genuinely novel rather than rediscovering known characteristics under new names? The authors evaluate productivity, performance, and novelty together, noting that no single metric ranks every system consistently. They also propose rolling re-execution: rerun the discovery process at historical decision dates using only information that would then have been available.
That is a meaningful change for an investment-data team. The agent is no longer just a summarizer or coding assistant. It is a participant in hypothesis formation, implementation, and selection. A reproducible implementation should therefore save the prompt and model version, data snapshot, candidate list, rejected candidates, code hash, decision date, validation rule, and reason for promotion. Human review still matters, but review after the fact cannot restore information that the process was not allowed to use.
The second source, a September 2026 review of AI in equity and crypto markets, makes the economic translation explicit. Information must be available at decision time; a representation must add stable information; a rule must become feasible positions; orders must receive plausible fills; and results must survive costs, capacity, and regime changes. The review finds real progress upstream but thinner public evidence for durable, capacity-aware net performance. That is why process evaluation complements, rather than replaces, portfolio and execution tests.
In practice, an agentic research queue could work like this:
- Freeze a point-in-time dataset and an allowed information cutoff.
- Ask the system for a mechanism and a formal, executable specification—not just a narrative.
- Run unit tests for dates, missing values, joins, corporate actions, and forbidden fields.
- Evaluate candidates in a predeclared validation design, recording every attempt.
- Compare promoted candidates with simple baselines and known factors.
- Re-run the entire search at several earlier dates, then inspect stability, novelty, and implementation costs.
The output is not a trade instruction. It is a research dossier: what was proposed, what survived, what failed, why it was selected, and how much of the result depends on choices made during the search.
A bounded research exercise. In 60–90 minutes, use a public, point-in-time dataset and a paper-only workflow. Baseline: a pre-specified linear model or a simple published characteristic, evaluated with a fixed walk-forward split. Let an AI assistant generate no more than 20 candidate transformations, but require each one to include its economic rationale, timestamp assumptions, code, and rejection status. Measure (1) data and timestamp violations, (2) candidate-to-survivor ratio, (3) out-of-sample rank or forecast stability, (4) turnover and a conservative cost haircut, and (5) whether the same candidates recur when the starting date changes. Stop if you cannot reconstruct any candidate from the saved record, if the assistant uses post-cutoff information, or if you find yourself changing the metric after seeing results. Do not place orders or treat the output as advice.
The limitations are material. A rolling re-execution can still be overfit to its own evaluation design. Novelty checks depend on the reference library. Historical data can contain survivorship and reporting biases, while live markets add latency, impact, financing, and crowding. A model provider’s incentives may favor impressive demos over negative findings. And an agent that has access to its own memory may quietly make later runs incomparable. These are causal and measurement gaps, not merely engineering bugs.
For discretionary investors, the lesson is to audit the evidence trail before admiring the chart. For systematic teams, the unit of validation should include the adaptive search procedure. For builders, the valuable product may be a reproducible ledger of attempts and constraints—not an autonomous claim generator. AI can expand the research surface. It cannot make repeated selection, bad timestamps, or execution costs disappear.
Further reading: Agentic Trading Evidence Ledger, LLM Forecasting Needs a Memory Firewall, and Time-Series Foundation Models Need Deployment Diagnostics.
Sources: Pan, Ding, and Giesecke, “Agentic Empirical Asset Pricing: Methodological Foundations” (arXiv:2609.00731); “Artificial Intelligence in Equity and Crypto Markets: Progress, Profitability Evidence, and the Limits of Automated Investing” (arXiv:2609.04917). Both are research papers/preprints; claims here should not be read as audited live-performance evidence.