Search papers, labs, and topics across Lattice.
This paper introduces GAMUT, a two-level meta-rubric framework designed to evaluate the factual completeness of open-ended generations in long-form text. By addressing the limitations of traditional evaluation methods that focus solely on precision, GAMUT provides a structured approach to assess both the presence and organization of necessary information in generated content. The framework demonstrates its effectiveness through rigorous testing, revealing that even state-of-the-art models struggle with factual completeness, achieving a maximum score of only 58.7% on the benchmark.
Open-ended generation models are failing to capture factual completeness, with top performers only achieving 58.7% accuracy on a new benchmark designed to assess this critical aspect.
Evaluating the factuality of long-form generations has focused predominantly on precision, measuring whether the claims a model makes are correct. The dominant decompose-search-verify pipeline catches incorrect claims well but says little about whether a response contains all the information it should. Measuring factual completeness, the missing half of factuality, is harder: it requires enumerating the full set of facts a complete answer should contain, and these facts rarely form a flat list. They often involve open-ended sets where coverage is what matters, ordered processes, and relationships among facts that a list of independent boolean checks fails to capture. We introduce a two-level meta-rubric framework for evaluating open-ended generation, and instantiate it as Gamut (Grounded Assessment of Multimodal Factuality), a benchmark for factual completeness in long-form generation. The framework rests on a two-level rubric representation: a structured meta-rubric captures the organization and importance of the required content, which is then mechanically compiled into a flat checklist of binary, machine-gradable rubrics that an LLM judge scores reliably. We construct 1,813 questions grounded in real wearable imagery across 10 diverse domains, each paired with an evidence-backed rubric verified by expert human annotators. Because the framework is modality-agnostic, we also release a text-only variant. Evaluating 14 frontier and open-weight models, we find the benchmark genuinely challenging (best score 58.7% from Gemini 3.1 Pro), highly discriminative, and robust to the choice of judge.