Search papers, labs, and topics across Lattice.
This study critically evaluates the effectiveness of using LLMs as judges in closed-loop table recognition tasks, specifically through the analysis of deterministic TEDS evaluation with datasets like FinTabNet and OmniDocBench. The findings reveal that judge signals were weak, with rankings often non-reproducible and only one selection policy outperforming random choices, indicating that judge scores alone do not enhance optimization. Additionally, while a structure-preserving instruction mitigated severe losses in candidate generation, it did not improve overall performance, suggesting that effective iterative refinement necessitates a more reliable verification signal beyond mere evaluative feedback.
Judge signals from LLMs may not enhance optimization in table recognition, revealing a critical gap between evaluation ability and practical utility.
LLM-as-a-judge is widely used to provide feedback and selection signals in closedloop regeneration, but this use remains insufficiently validated. We study it in table recognition, where deterministic TEDS evaluation provides a controlled testbed, using FinTabNet and OmniDocBench. Three findings emerge. First, judge signals were weak on both datasets: scores frequently tied, rankings were not reproducible, and the only selection policy that beat random on both datasets depended on an earliest-iteration tie rule, so its advantage cannot be attributed to the judge scores alone. Iteration produced better candidates, but the judge failed to recover them. Second, severe losses occurred even without specific judge feedback. A structurepreserving instruction significantly reduced the severe-loss rate on FinTabNet and was directionally consistent on OmniDocBench. The contrasts support target-preservation failure under unconstrained regeneration as a proximate mechanism of the observed severe losses. Third, the structure-preservation constraint reduced the severe-loss tail but produced no improvement. In an exploratory 2x2 analysis, the same protection was not stably observed when judge feedback was retained. These results do not dispute the value of LLMs as evaluators. Instead, they show that evaluation ability does not imply optimization utility. Iterative refinement requires, at minimum, a verification signal that deterministically detects structural change, rather than judge scores alone.