Search papers, labs, and topics across Lattice.
This paper investigates the memorization patterns of discriminatively trained reward models (RMs) by analyzing their performance on two human preference datasets. The authors find that RMs tend to misallocate memorization towards easier preference pairs, memorize specific shortcuts related to dataset characteristics, and overgeneralize heuristic correlates when faced with novel preference pairs. These insights reveal that current training methods lead to biased RMs that struggle with context-dependent quality assessments, highlighting significant limitations in their applicability.
Reward models are biased by their training data, leading to misjudgments in context-dependent scenarios that could undermine their effectiveness in real-world applications.
This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference datasets. We show that RMs 1) misallocate memorization to easy, high margin preference pairs, 2) memorize dataset-specific shortcuts (e.g., model identity, user sampling strategy), and 3) overgeneralize simple heuristic correlates of human preference (e.g., length, compliance) when confronted with unseen preference pairs. Overall, our findings indicate that discriminative training of RMs from human preference data results in biased RMs not yet capable of judging response quality in context-dependent scenarios.