Search papers, labs, and topics across Lattice.
The FinMMEval 2026 Task 2 evaluates multilingual short-answer financial question answering by pairing English questions with financial statements and news in multiple languages, including Chinese, Japanese, Spanish, and Greek. The task features a final-test set of 256 items, categorized into easy and expert tiers, with systems ranked based on macro-averaged ROUGE-1 F1 scores against withheld gold answers. The results reveal a competitive landscape where the top four systems are closely clustered, highlighting advancements in retrieval-augmented generation and cross-lingual evidence handling.
The top-performing systems in multilingual financial question answering are separated by less than one percentage point, showcasing the intense competition and subtlety in model performance.
FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence. Each final-test item pairs an English question with financial statements and news in English, Chinese, Japanese, Spanish, and Greek. Participating systems submit one concise answer per item in JSONL format. The final-test set contains 256 items, split evenly between easy and expert tiers; each tier contains four question templates instantiated over 32 company-report groups. Gold answers were withheld during submission, and systems were ranked by macro-averaged item-level ROUGE-1 F1 against organizer-held reference answers. The final leaderboard includes 12 ranked submissions. The strongest systems are closely clustered, with the top four separated by less than one percentage point in ROUGE-1 F1. The submitted system papers document retrieval-augmented generation, cross-lingual evidence handling, structured prompting, answer compression, and validation strategies.