Search papers, labs, and topics across Lattice.
This paper introduces WildHandBench, a comprehensive benchmark designed to evaluate the performance of machine learning models on handwritten text, addressing gaps in existing benchmarks that primarily focus on printed documents. The benchmark consists of 500 handwritten documents across various structures and languages, and it employs a novel Prior-Driven Error (PDE) metric to differentiate between errors stemming from language priors and those from visual evidence. The evaluation reveals that while the best model achieves only 71.85% accuracy, humans outperform models with a narrow margin, highlighting a significant reliance of models on language priors that is not captured by traditional accuracy metrics.
Models struggle with handwritten text, showing a 63-91% reliance on language priors for errors, while humans exhibit a more balanced error profile.
While the top model on OmniDocBench now reaches 96.34% overall on printed-document parsing, the ability of current models to handle challenging handwritten documents remains largely uncharacterized. Existing benchmarks focus on isolated text or formulas, overlook handwritten tables and real-world degradation, and report aggregate accuracy without explaining why models fail. We present WildHandBench, a benchmark containing 500 handwritten documents across three structures (free text, tables, formulas), four languages, and nine real-world scenarios. We introduce a Prior-Driven Error (PDE) metric that quantifies whether errors originate from language priors rather than visual evidence. Evaluating 18 state-of-the-art models together with calibrated human baselines, we find: (1) the best model achieves only 71.85% overall; (2) humans outperform all models yet the gap is narrow (77.09% vs. 71.85%); and (3) model errors are qualitatively different from human errors -- 63-91% of model errors are prior-driven versus only 49% for humans, exposing systematic reliance on language priors that conventional accuracy metrics cannot capture.