Search papers, labs, and topics across Lattice.
This paper introduces ADOPD 2026, an advanced framework for document understanding that integrates spatially grounded reasoning with visual and semantic elements. By enriching existing page anchors with human-cleaned captions and chain-of-thought traces, the framework enables models to better identify document element types and improves the accuracy of dense counting tasks. The findings reveal significant long-tail semantic failures in standard benchmarks, underscoring the necessity of a more holistic approach to document intelligence that connects visual anchors with reasoning capabilities.
Long-tail semantic failures in document understanding are exposed when models are forced to reason with a unified vocabulary of visual anchors rather than treating elements in isolation.
Existing document understanding benchmarks have largely focused on locating page elements, yet real-world document intelligence requires models to reason jointly about region semantics, spatial relations, and visual structure. We present ADOPD 2026, a reasoning-oriented extension of ADOPD that turns page decomposition into spatially grounded document understanding. ADOPD 2026 enriches page anchors inherited from ADOPD 2024 dataset with human-cleaned captions, semantic tags, and generated chain-of-thought (CoT) traces grounded to document regions. Instead of treating boxes, masks, and tags as independent supervision signals, we cast text blocks, visual entities, semantic labels, bounding boxes, and polygon masks as a shared vocabulary of visual anchors. This representation supports three connected capabilities. First, region-level semantic tagging asks models to identify document element types from both page context and local appearance, revealing long-tail semantic failures that standard layout benchmarks often hide. Second, unified vision-language grounding generates text regions and visual entities together with coordinates or polygonal outlines, transforming detection and segmentation outputs into structured anchors that can be reused by downstream reasoning systems. Third, current state-of-the-art models still struggle with dense counting tasks evaluated on DocCount, a benchmark derived from ADOPD 2026, highlighting the need for the Thinking-with-Anchors pipeline in document semantic understanding. By connecting page decomposition to verifiable visual-anchor reasoning, ADOPD 2026 provides a task framework that moves document understanding beyond localization toward anchor-grounded document intelligence.