Search papers, labs, and topics across Lattice.
This paper introduces a protocol-level identifiability audit designed to evaluate the reasoning capabilities of large language models (LLMs) without requiring model calls. By formalizing the relationship between policies, observation supports, and estimands, the authors demonstrate that traditional benchmarks may fail to accurately reflect the underlying behavioral properties of LLMs. The findings reveal significant discrepancies between base accuracy and selective-response fidelity, highlighting the limitations of existing evaluation methods and the importance of structured evaluation design.
Traditional LLM benchmarks can mislead, as they often fail to distinguish between different reasoning capabilities, collapsing multiple policies into a single equivalence class.
LLM benchmark scores can be precise even when the observation protocol does not identify the behavioral property they are intended to measure. In a controlled, solver-grounded setting, we formalize a protocol-level identifiability audit over a finite behavioral policy class: given policies H, observation support O, and estimand $\tau$, we test whether O separates every pair with different $\tau$. The audit requires zero model calls and resolves our diagnostic case: base-only observation collapses seven frozen deterministic policies into one equivalence class; full support yields seven classes and no cross-estimand collisions; every leave-one-out support retains a constructive collision witness. Empirically, both constrained-generation variants have pair-validity 1.0, yet base accuracy and selective-response fidelity diverge - 0.620 versus 0.324 across six balanced oracle-transition directions (cluster-bootstrap 95% CI [0.600, 0.642] vs. [0.304, 0.345]) - and the gap recurs on a second deterministic source (0.646 vs. 0.331). The audit also synthesizes a minimum identifying support $O^*$ for the frozen policy class: two cells instead of the full 36-cell tensor. This case shows how evaluation-design validity can be checked structurally before model inference and why base correctness does not determine intervention-response fidelity.