Search papers, labs, and topics across Lattice.
This paper addresses the challenges in financial statement fraud detection (FSFD) by proposing a robust framework that integrates structured financial data with unstructured textual information from financial reports using Large Language Models (LLMs). The authors highlight the limitations of existing methods that rely on random data splits, which yield overly optimistic performance estimates, and introduce a novel benchmark task called Company-Isolated FSFD (CI-FSFD) for more realistic evaluation. Their approach not only outperforms existing techniques on the CI-FSFD task but also underscores the importance of utilizing textual data for effective fraud detection in financial contexts.
Textual data from financial reports can significantly enhance fraud detection, as evidenced by our framework's superior performance on a novel benchmark task.
Financial statement fraud detection (FSFD) is crucial for market integrity but faces challenges from increasingly sophisticated schemes and under-utilized textual data in financial reports. Existing methods often rely on random data splits, leading to overoptimistic performance estimates that do not reflect real-world generalization to new companies or future periods. To address this recurring problem with the state of the art, we propose a robust FSFD framework leveraging Large Language Models (LLMs) to integrate both structured financial data and unstructured textual information from financial reports. We provide a more realistic evaluation through a novel and challenging benchmark task called Company-Isolated FSFD (CI-FSFD). We construct and make publicly available a comprehensive U.S. company dataset combining financial statements, summarized MD&A text, and fraud labels. Our approach achieves the best performance on the challenging CI-FSFD task, demonstrating the critical value of textual data and robust evaluation for reliable financial fraud detection.