Search papers, labs, and topics across Lattice.
To quantify how external source claims destabilize language model outputs, the authors evaluated four instruction-tuned LLMs across 220,000 trials on MMLU-Pro and IndicMMLU-Pro spanning English and four Indic languages. They track the neutral-conditioned misleading cue adoption rate (NC-MCAR), which isolates instances where an unverified cue causes a model to abandon an answer it previously got correct under a neutral prompt. Strikingly, attributing a fixed wrong option to an "expert" caused models to flip away from gold answers 41.1% of the time, compared to just 12.5% when attributed to a "majority" opinion.
Instruction-tuned models abandon their own correct answers over 41% of the time when an incorrect option is merely prefaced with a bare, unverified "expert" claim.
Language models often receive a question together with a claim about what another source answered. We audit whether such claims destabilize answers in multiple-choice question answering. For each item, we hold one wrong option fixed across misleading conditions and vary the cue template attached to it. We introduce \emph{neutral-conditioned misleading cue adoption rate} (NC-MCAR), which measures switches to that option only on valid cued trials where the same model first selected the gold answer under a neutral prompt. This is a measure of answer instability, not proof that the model knew the answer or that all deference is irrational. We evaluate four instruction-following models on MMLU-Pro and IndicMMLU-Pro in English, Hindi, Bengali, Tamil, and Telugu. Across 220{,}000 outputs, the expert template yields 41.1\% aggregate NC-MCAR, compared with 12.5\% for the majority template. These two conditions use the same wrong option and final instruction. Filler accuracy remains well above expert-wrong accuracy, while correct-cue prompts have high valid-response accuracy. The audit documents answer instability relevant to grounding under the tested forced-choice prompts: a bare, unverified source claim can outweigh an answer that was previously consistent with the task evidence.