Search papers, labs, and topics across Lattice.
This study investigates the effects of instruction tuning on the confidence and lexical diversity of language models in question-answering tasks. The authors find that while instruction tuning increases verbalized confidence, it does not significantly improve predictive accuracy and can lead to decreased likelihood-based calibration. Additionally, the research reveals a complex relationship between instruction tuning and rationale diversity, with cross-rationale diversity consistently decreasing and surface-level lexical diversity showing variable changes across different models and benchmarks.
Instruction tuning boosts model confidence but often at the cost of rationale diversity and calibration accuracy.
Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recently been shown to exhibit verbalized overconfidence. In question answering, verbalized model overconfidence may be associated with the consistency of the generated supporting rationales. In this paper, we study whether corresponding changes in the lexical diversity of generated answer rationales accompany changes in model confidence induced by instruction tuning. We evaluate three matched base and instruction-tuned models across question-answering benchmarks and find that instruction tuning consistently alters answer confidence, despite limited changes in predictive accuracy and decreases in likelihood-based calibration. Secondly, we observe a non-uniform effect of instruction tuning on rationale diversity: cross-rationale diversity consistently decreases, whereas surface-level lexical diversity varies in both direction and magnitude across models and benchmarks. Finally, we find that these differences persist after controlling for answer selection and rationale length, confirming that confidence and rationale diversity capture distinct effects of instruction tuning.