Search papers, labs, and topics across Lattice.
This paper introduces the AI Financial Intelligence Benchmark (AFIB) to evaluate LLMs on financial analysis across factual accuracy, completeness, data recency, consistency, and failure patterns. They tested GPT, Gemini, Perplexity, Claude, and SuperInvesting on 95+ structured financial analysis questions derived from real-world equity research. SuperInvesting achieved the highest aggregate performance, while retrieval-oriented systems like Perplexity excelled in data recency but struggled with analytical synthesis.
SuperInvesting, a specialized AI system, significantly outperforms general-purpose LLMs like GPT and Gemini on a new financial intelligence benchmark, suggesting domain-specific architectures are crucial for reliable investment research.
Large language models are increasingly used for financial analysis and investment research, yet systematic evaluation of their financial reasoning capabilities remains limited. In this work, we introduce the AI Financial Intelligence Benchmark (AFIB), a multi-dimensional evaluation framework designed to assess financial analysis capabilities across five dimensions: factual accuracy, analytical completeness, data recency, model consistency, and failure patterns. We evaluate five AI systems: GPT, Gemini, Perplexity, Claude, and SuperInvesting, using a dataset of 95+ structured financial analysis questions derived from real-world equity research tasks. The results reveal substantial differences in performance across models. Within this benchmark setting, SuperInvesting achieves the highest aggregate performance, with an average factual accuracy score of 8.96/10 and the highest completeness score of 56.65/70, while also demonstrating the lowest hallucination rate among evaluated systems. Retrieval-oriented systems such as Perplexity perform strongly on data recency tasks due to live information access but exhibit weaker analytical synthesis and consistency. Overall, the results highlight that financial intelligence in large language models is inherently multi-dimensional, and systems that combine structured financial data access with analytical reasoning capabilities provide the most reliable performance for complex investment research workflows.