Search papers, labs, and topics across Lattice.
This paper introduces MIST, a benchmark designed to evaluate language models' susceptibility to misleading external signals by categorizing reasoning items into four distinct conditions. The authors present SCOPE, an optimization method that enhances model performance by balancing preference pairs across clean, correct, misleading, and irrelevant contexts, rather than solely focusing on misleading items. The results show that SCOPE significantly reduces the rate at which misleading signals corrupt correct answers while maintaining accuracy in other contexts, advocating for a shift in how model robustness is assessed.
Models that ignore context may seem robust, but they can fail spectacularly when the context is actually trustworthy.
Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong. The obvious remedy, training models to resist such signals, hides a failure mode: a model that ignores all context looks robust yet is useless when the context is worth trusting. We recast the problem as selective trust and introduce MIST, a human-annotated benchmark that renders each reasoning item under four matched conditions (clean, misleading, correct-context, and irrelevant-context), together with SC2W, a paired metric counting how often a misleading signal flips a clean-correct answer to wrong. Across a comprehensive benchmark study, we observe that such a susceptibility is universal. We then propose SCOPE, which mines clean-correct/misleading-wrong failures and optimizes a standard Direct Preference Optimization (DPO) objective over matched preference pairs balanced equally across all four conditions, rather than over misleading items alone. Our approach substantially reduces SC2W on popular open-sourced models while preserving accuracy when the added context is clean, correct, or irrelevant. With this work, we argue that models should be judged on selective trust, not on resistance alone.