AI Can Triage a Filing—If Every Alert Has a Time-Stamped Reason
AI is changing the first pass through filings from reading everything to ranking what deserves review. The safeguard is a reproducible, time-stamped alert ledger.
The first AI-shaped change in filing research is not a smarter valuation. It is triage: deciding which passages deserve a human’s scarce attention. That is useful only if each alert can answer three questions: what changed, compared with which earlier filing, and what exact passage supports the alert?
The SEC’s EDGAR APIs make the raw ingredients public: filings, filing dates, company facts, and structured data. FINRA’s guidance on generative AI in financial services supplies the operational warning: existing supervision, books-and-records, privacy, and accuracy obligations do not disappear because a model produced the summary. Together they point to a modest but important workflow change. Let AI rank review queues; keep the evidence ledger and the decision to rely on an alert with a person.
The old process was serial. An analyst opened the latest 10-K or 10-Q, searched familiar headings, and carried notable passages into a spreadsheet. The new process can be comparative. Give a model two point-in-time filing packages and ask it to propose only differences: new risk language, changed accounting definitions, altered commitments, or a missing disclosure. The output is a queue, not an investment conclusion.
That distinction matters. A ranked queue can reduce the chance that an analyst misses an unusual passage, but it can also bury a quiet change under dramatic language, confuse amended and original filings, or import knowledge published after the review date. The practical unit of quality is therefore not “summary fluency.” It is alert precision, citation entailment, timestamp integrity, and reviewer agreement.
Here is a reproducible version using public documents. First freeze the information set: download the two filings, record accession numbers and publication times, and exclude later documents. Second, segment the filings by heading and paragraph. Third, ask the model to emit JSON-like rows with the old passage, new passage, change type, a one-sentence reason, and a confidence label. Fourth, have a separate check compare every row against the source text and mark “supported,” “ambiguous,” or “unsupported.” Finally, let a human decide which supported rows merit deeper research. Never ask the triage prompt for a buy, sell, price target, or expected return.
This makes the model’s role narrow enough to test. It is a difference detector and queue sorter. It is not the owner of materiality, chronology, accounting judgment, or compliance. For systematic builders, the ledger also creates a useful interface: every alert has a source span, an as-of boundary, a model version, and a reviewer outcome. For discretionary investors, it turns an opaque “AI read the filing” claim into an auditable research artifact.
Try it in 75 minutes with research-only data. Spend 15 minutes selecting one issuer’s latest two comparable filings and freezing the packet. Spend 20 minutes building a manual baseline: a reviewer marks all passages they believe deserve follow-up. Spend 20 minutes running the same comparison through the model. Use the final 20 minutes to verify citations and reconcile the two queues. Measure alert precision (reviewer-supported alerts divided by all alerts), recall against the manual baseline, unsupported-citation rate, duplicate-alert rate, and median review time. Also count alerts that cite a later-than-allowed document.
Set the stop condition before starting: stop and discard the run if more than one in ten alerts is unsupported, if any alert uses post-cutoff information, or if the model’s queue cannot be reproduced from the frozen packet. Do not convert the exercise into a paper trade. A seven-day extension may repeat the test on seven new filing pairs, but preserve the same rubric and log model changes; the purpose is workflow measurement, not a performance claim.
The hard cases are exactly where confidence is least informative. A filing may change because of formatting, an amended submission, a new accounting standard, or a genuinely important business event. Text similarity can miss tables; tables can hide unit changes. The SEC data is authoritative for what was filed, not for what is economically material. FINRA’s supervisory framing adds another boundary: prompts, outputs, retention, access controls, and human review need ownership inside a real investment process.
This complements the decision-trace approach for investment AI and the warning about temporal integrity in financial foundation models. The lesson is not that AI has discovered a hidden signal in filings. It has made a formerly expensive comparison cheap enough to measure. That shifts the research question from “Can the model read?” to “Can the team verify, reproduce, and safely act on what it flagged?”
Sources: SEC EDGAR APIs, SEC company facts, and FINRA, Generative Artificial Intelligence. These establish the document substrate and the control obligations; they do not establish an investment edge.