Search papers, labs, and topics across Lattice.
The paper introduces ADMITBench, a structured framework designed to evaluate the admissibility of recommendations made by industrial large language models (LLMs) based on a safety-governed evaluation contract. This framework emphasizes the importance of aligning LLM advisories with evidence, authority, and specific plant consequences, ensuring that recommendations are rigorously checked against a versioned profile. The key finding highlights that ADMITBench provides a systematic approach to assess LLM outputs, enhancing their reliability in industrial contexts without implying safety certification for the models or processes involved.
A rigorous framework that ensures LLM recommendations are not only evidence-based but also aligned with specific industrial safety protocols.
This white paper presents ADMITBench, a reference framework for evaluating industrial LLM advisories at the level of the proposed action. The framework implements a versioned, safety-governed evaluation contract that checks whether a recommendation is supported by the available evidence, permitted under the stated authority and procedure, and acceptable under the plant-specific consequence checks encoded in the selected evaluation profile. In this report, \emph{safety-governed} means that eligibility is determined through explicit, non-compensatory checks derived from a versioned plant profile; it does not mean that the evaluator, model, or plant has been safety-certified. Release 0.1.0 is a public reference implementation for technical and research evaluation, not an authorisation for physical execution.