Search papers, labs, and topics across Lattice.
This study investigates the reliability of language models in safety-critical environments, specifically air traffic control (ATC), where traditional semantic metrics like F1 score may misrepresent operational safety. By employing a consequence-aware evaluation framework, the authors reveal a significant semantic-safety gap, where conventional performance metrics yield inflated reliability estimates compared to consequence-aware assessments. The research highlights that while risk-aware fine-tuning can improve model performance, it does not fully bridge the gap, underscoring the necessity of consequence-aware evaluation for safe deployment in critical applications.
Conventional semantic metrics can mislead safety assessments, revealing a critical gap in operational reliability for language models in air traffic control.
Can language models be trusted in safety- critical operations? In such settings, strong per- formance on semantic metrics does not guaran- tee operational reliability: a misread altitude, a dropped execution condition, or a confused call- sign may score well under standard F1 yet carry sharply asymmetric operational consequences. We study this problem in air traffic control (ATC), where controller-pilot communication demands near-zero error tolerance, and use consequence-aware evaluation to test whether semantic scores misstate operational reliabil- ity. The framework is instantiated in a con- trolled diagnostic ATC benchmark grounded in aviation standards and feedback from 40 air traffic controllers across three countries. Evaluating 8 models, we uncover a system- atic semantic-safety gap: conventional scores give substantially higher performance estimates than consequence-aware evaluation, even for models that appear reliable under standard met- rics. Risk-aware fine-tuning narrows but does not close this gap, showing that consequence- aware evaluation is a necessary complement to standard NLP metrics before any real safety- critical deployment claim