Search papers, labs, and topics across Lattice.
This paper introduces BEAR-Bench, a bilingual benchmark designed to evaluate the reasoning capabilities of Multimodal Large Language Models (MLLMs) on complex, text-dense documents in both English and Russian. The benchmark consists of 1000 human-annotated questions derived from professional business and scientific texts, addressing the limitations of existing benchmarks that are often language-restricted and focus primarily on information extraction. Evaluation of 16 MLLMs reveals significant performance gaps, highlighting the need for improved reasoning in these models and offering insights into hallucination detection methods based on model outputs.
MLLMs struggle with reasoning on complex documents, with even the best models showing significant performance gaps on BEAR-Bench.
While Multimodal Large Language Models (MLLMs) have made significant strides in visual comprehension, their ability to reason about text-dense, professional documents remains incompletely evaluated. Existing benchmarks emphasize information extraction, require external domain knowledge, or cover professional documents only as one of many settings. They are also largely English- or Chinese-centric, leaving other languages and Russian, in particular, substantially underrepresented. To address these limitations, we introduce BEAR-Bench (Bilingual Enterprise and Academic Reasoning), a self-contained, complex English-and-Russian benchmark comprising 1000 human-annotated questions based on text-rich business and scientific documents. We evaluate 16 proprietary and open-weight MLLMs, including Gemini 3.1 Pro and Qwen3.5-397B, on BEAR-Bench and observe clear headroom even for the strongest systems. Finally, we use the resulting model outputs to compare existing hallucination detection methods, evaluating not only how often models fail on BEAR-Bench but also how reliably those failures can be identified.