Search papers, labs, and topics across Lattice.
The FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering across English, Chinese, Arabic, and Hindi, focusing on the ability of systems to accurately interpret domain-specific terminology and numerical data. With a final test set of 800 questions, the task revealed top accuracies ranging from 92.0% in Hindi to 97.5% in English and Arabic, highlighting consistent performance from leading teams across all languages. The methodologies employed included retrieval augmentation, direct answer-option scoring, and selective self-consistency, showcasing the effectiveness of advanced techniques in multilingual financial contexts.
Systems achieved up to 97.5% accuracy in multilingual financial question answering, revealing the potential for high-performance AI across diverse languages.
FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi. The task tests whether systems can select the correct answer to finance questions involving domain terminology, numerical interpretation, and conceptual financial reasoning across languages and scripts. The final-test set contains 800 questions, with 200 questions per language; gold answers were withheld during submission, and each language was ranked independently by accuracy. The final leaderboards contain 13 English, 11 Chinese, 11 Arabic, and 10 Hindi ranked submissions. Top accuracies range from 92.0% in Hindi to 97.5% in English and Arabic, with the same leading teams appearing near the top across all four languages. The documented systems used retrieval augmentation, direct answer-option scoring, language-specific prompting, selective self-consistency, confidence checks, and LLM-based review stages.