Search papers, labs, and topics across Lattice.
This paper introduces a framework for auditing the faithfulness of multimodal large language models (LLMs) in grid diagnosis by assessing their reliance on task-specific evidence. It employs a comparative analysis of self-reported reliance, behavioral reliance derived from interventions, and preregistered engineering importance to identify discrepancies in evidence usage. The framework's effectiveness is validated through case studies on three LLMs, demonstrating its capability to detect and rectify faithfulness failures without compromising performance.
Task-conditional faithfulness failures in multimodal LLMs can be detected and corrected, ensuring models use appropriate evidence for grid diagnosis.
Multimodal large language models (LLMs) can combine topology, measurements, and incident text for grid diagnosis, yet answer accuracy does not establish that task-appropriate evidence was used. This letter proposes a general framework in order to conduct task-conditional faithfulness audit. It compares self-reported reliance, intervention-derived behavioral reliance, and preregistered engineering importance. The framework first registers task-specific evidence requirements and compares them with self-reported reliance and behavioral changes under controlled modality ablations. To resolve detected discrepancies, we design an evidence-gated correction and re-audit mechanism that regenerates failed responses under evidence constraints and independently re-ablates them to verify improved grounding without performance loss. Case studies evaluate three differently scaled LLMs on IEEE 39- and 118-bus scenarios. These results validate the framework ability to detect, diagnose, and correct task-conditional faithfulness failures.