Search papers, labs, and topics across Lattice.
The paper introduces LAVA, a modular framework designed for the validation and augmentation of financial documents, addressing the challenges posed by heterogeneous layouts and complex business rules. By employing a four-stage process that includes document-rule retrieval and layout-preserving information extraction, LAVA enhances accuracy and consistency in high-stakes auditing tasks. Evaluated against a large benchmark, LAVA significantly outperforms existing methods in managing hallucinations and edge cases while ensuring efficient token usage.
LAVA achieves superior accuracy in financial document auditing by effectively grounding complex business rules and managing edge cases.
Financial document validation in production, such as payroll auditing, tax compliance, and loan underwriting, demands exceptional accuracy, consistency, and reproducibility under strict enterprise constraints. In practice, documents arrive with heterogeneous layouts and formats, semantically rich and context-dependent content, and embedded business rules that current pipelines struggle to process reliably. We introduce LAVA (Logic-Aware Validation and Augmentation), a modular, backbone-agnostic pipeline built on multimodal large language models, that integrates a four-stage design: document-rule retrieval, layout-preserving information extraction, auxiliary metadata enrichment, and auditable symbolic/arithmetic verification. LAVA supports robust rule grounding, fine-grained error attribution, and consistent, traceable end-to-end execution, capabilities essential for high-stakes deployment. Evaluated on a large real-world benchmark with diverse financial documents and dozens of expert-curated validation rules, LAVA outperforms baselines in hallucination control and edge-case handling while maintaining efficient token usage, demonstrating practicality for high-volume, time-critical validation.