Search papers, labs, and topics across Lattice.
This paper introduces FinanceComplexQA, a benchmark designed to evaluate agentic reasoning on complex financial documents, addressing the significant performance variability observed among different agents in real-world financial analysis. By leveraging the Finance-LaTeX SKILL, the authors generated 2,000 professional financial documents and 6,000 question-answer pairs, creating a comprehensive dataset that simulates real-world scenarios. The evaluation of leading retrieval-augmented generation (RAG) systems reveals critical insights into their capabilities and limitations in numerical computation, multi-hop reasoning, and content summarization.
Agentic reasoning tools struggle with complex financial documents, revealing substantial performance gaps that could impact decision-making in finance.
Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate large-scale information and generate reliable and accurate content. However, when handling complex real-world problems, different agents still show significant performance variation. In this work, we design Finance-LaTeX SKILL, a skill for synthesizing financial documents with complex layouts based on expert knowledge. Using an agent workflow built on this skill, we generate 2,000 professional financial documents along with 6,000 high-quality question-answer pairs. To evaluate the overall capability of agents, we introduce FinanceComplexQA, a comprehensive open-ended generation benchmark for financial documents that closely resembles real-world scenarios. It contains 2,026 deep research tasks targeting 1009 financial documents. FinanceComplexQA has 8 key features: bilingual support; coverage of six mainstream scenarios and seven tasks; expert-level document reasoning questions; deep research of complex layouts; relatively stable and permanent reference answers; and precise evaluation through an Agent-as-a-Judge with multiple evaluation metrics. Using FinanceComplexQA, we conduct a comprehensive evaluation of leading RAG systems and agentic reasoning tools for financial document QA. Through identifying and analyzing failure cases, we provide an in-depth study of their capabilities in numerical computation, multi-hop reasoning, content summarization, and industry analysis.