Search papers, labs, and topics across Lattice.
This paper introduces FrontierFinance, a comprehensive benchmark designed to evaluate AI agents in the context of the entire investor workflow, addressing the limitations of existing benchmarks that focus narrowly on financial data extraction. The authors conducted evaluations using 220 expert-crafted queries and found that the quality and efficiency of AI responses are significantly influenced by the tool harness rather than the model alone. Notably, Samaya's in-house system outperformed the strongest frontier model, Claude Fable 5, while the best open-weight model, Kimi K3, nearly matched the proprietary model's performance at a fraction of the cost, highlighting the benchmark's rigor and relevance in assessing frontier intelligence in finance.
The FrontierFinance benchmark reveals that the choice of tool harness can dramatically affect AI performance in finance, with significant cost-efficiency implications for model deployment.
AI agents are increasingly deployed for professional investment research, yet no benchmark captures the complexity of the full investor workflow. Existing benchmarks mainly target financial data extraction, a narrow slice that current models have largely saturated, while reference-based metrics and generic LLM-as-a-judge scoring fall short on the open-ended, long-form answers that real analyst queries demand. We introduce FrontierFinance, a fully open benchmark of 220 expert-crafted queries and 11,543 source-attributed rubrics spanning six crucial use cases across the full investor workflow. FrontierFinance is both broader and harder than existing public finance benchmarks. Evaluating frontier models and agent systems under a common harness restricted to publicly available data, we find that the tool harness, not the model alone, strongly shapes quality and efficiency; that Samaya's in-house system leads at 56.0%, ahead of the strongest frontier model (Claude Fable 5, 49.2%) at roughly 2.2x lower cost; and that the best open-weight model (Kimi K3, 46.4%) nearly matches the best proprietary model at 4.5x lower cost. Screening&Discovery and Sector, Industry&Macro remain the hardest use cases across all systems, where even the best systems reach only 33% and 39%. We make the dataset and grading code publicly available.