Search papers, labs, and topics across Lattice.
This paper introduces OmniHandwritingOCR, a comprehensive benchmark designed to evaluate multimodal large language models (MLLMs) in the context of handwritten optical character recognition (OCR). By encompassing a diverse array of handwritten scenarios鈥攊ncluding multilingual text, writer errors, and complex mathematical expressions鈥攖his benchmark addresses the limitations of existing OCR evaluations that primarily focus on printed text. The evaluation of thirteen systems reveals significant shortcomings, particularly in handling complex multi-line formulas and highlights the tendency of generative models to produce visually unsupported corrections, underscoring the need for improved robustness in MLLMs for handwritten OCR tasks.
Current OCR systems struggle with complex handwritten inputs, with performance plummeting on multi-line formulas and generative models often hallucinating errors.
Multimodal large language models (MLLMs) are increasingly used as OCR systems in document and knowledge-processing pipelines, but their ability to faithfully read real handwriting remains underexplored. Existing OCR benchmarks focus largely on printed text or clean single-line inputs, leaving limited coverage of realistic handwritten OCR scenarios such as multilingual handwriting, writer errors, and structurally complex mathematical expressions. We introduce OmniHandwritingOCR, a diagnostic benchmark for evaluating MLLMs and OCR systems on handwritten OCR. It covers handwritten text recognition and handwritten mathematical expression recognition across six subtasks and twelve subsets, totaling 77.57K labeled images from public datasets and newly collected student writings. A key component is a difficulty-stratified multi-line formula corpus designed to test robustness under increasing structural complexity. We evaluate thirteen open- and closed-source systems with five complementary metrics under a unified protocol. Results show that current systems remain far from faithful transcription: performance drops sharply on complex multi-line formulas, model rankings vary across language and formula settings, and several generative models hallucinate plausible but visually unsupported corrections. OmniHandwritingOCR provides a challenging testbed for diagnosing language, content, structural, and visual-grounding failure modes of multimodal models in handwritten OCR scenarios.