Search papers, labs, and topics across Lattice.
This study investigates the effectiveness of large language models (LLMs) in processing financial disclosures and their impact on investment judgments, revealing a significant retrieval-integration gap. Despite accurate retrieval of information, the influence of risk disclosures on investment decisions diminishes to experimental noise levels when unrelated context is increased, a pattern consistent across various model families and tasks. The research highlights that while more capable models can delay this gap, the architecture of the workflow鈥攕pecifically whether it employs chunk-and-summarize or targeted restatement鈥攑lays a crucial role in ensuring that retrieved information effectively informs judgments.
Retrieval accuracy alone doesn't guarantee that AI analysts will integrate crucial financial disclosures into their investment decisions, revealing a critical flaw in current evaluation methods.
Large language models (LLMs) are increasingly deployed as AI analysts to process financial disclosures and support AI-assisted investment decisions. Yet such systems are usually evaluated by what they can retrieve, not whether retrieved information affects their judgments. We identify a retrieval-integration gap in long-context financial analysis. Holding focal-firm information fixed and varying only unrelated context from 2,000 to 128,000 tokens, we find that a risk disclosure's influence on investment judgments falls to the experimental noise floor even as direct retrieval remains accurate. The pattern replicates across model families and judgment tasks and in experiments removing real disclosures from actual 10-K filings. More capable models postpone but do not eliminate the gap. Causal memory interventions show that compressed summaries and source-text lookup jointly transmit disclosures into judgments. Workflow architecture determines whether this transmission succeeds: chunk-and-summarize pipelines evict relevant information, whereas a targeted, structured restatement adjacent to the decision restores its influence. AI analyst performance is therefore jointly determined by model capability and workflow architecture. Retrieval-based evaluations can certify systems whose investment judgments ignore information they demonstrably retrieved.