Search papers, labs, and topics across Lattice.
This paper introduces the Frontier Financial Judgement benchmark, designed to evaluate AI agents' ability to replicate expert human judgments in assessing the valuation impact of new financial information. Despite the increasing volume of data, the best-performing agent achieved only a 52.4% match with expert labels, highlighting significant challenges in accuracy and reliability. The study reveals a wide range of false-positive rates among different agents, emphasizing the complexities of deploying AI in real-world financial contexts.
Agents struggle to match expert financial judgments, with the best only achieving 52.4% accuracy, exposing critical limitations in AI's ability to process real-time market information.
We introduce Frontier Financial Judgement, a challenging new benchmark developed in collaboration with professional equity analysts to assess agents' ability to replicate expert human judgements. Rapidly identifying new information, evaluating its implications and determining its valuation impact is one of the most time-consuming and challenging aspects of real-world equity coverage. This is becoming ever more difficult and important as AI rapidly increases the quantity of new information to process. The strongest agent we evaluate on Frontier Financial Judgement matches all expert labels in only 52.4% of cases. We also find significant divergence in estimated false-positive rates among frontier agents, ranging from ~1% for GPT-5.6 Sol to ~32% for Claude Sonnet 4.6. To construct the benchmark and make it representative of real-world settings, we combine human-designed and labelled synthetic articles with live news articles and historical documents, creating 656 items for assessment. The resulting task requires agents to distinguish genuinely new, valuation-relevant financial information from stale, immaterial or misleading news under realistic conditions. We find substantial trade-offs among agent accuracy, cost, false positives and reliability that continue to hinder the reliable deployment of news-flow filtering in practice.