Search papers, labs, and topics across Lattice.
This paper introduces VIBE-Bench, a benchmark designed to evaluate Personalized Large Language Models (PLLMs) in scenarios where user profiles do not align with their preferences, a situation termed profile-preference conceptual misalignment (PRCM). The authors demonstrate that existing PLLMs struggle to effectively reason about user preferences when faced with this misalignment, relying instead on shallow semantic correlations that fail to capture deeper, cross-concept mappings. The findings highlight PRCM as a critical failure mode for PLLMs and establish VIBE-Bench as a valuable tool for advancing the understanding of preference reasoning in AI systems.
Current Personalized Large Language Models falter when user profiles and preferences diverge, revealing a critical gap in their reasoning capabilities.
Personalized Large Language Models (PLLMs) aim to tailor responses to individual users, where a central challenge is preference reasoning: inferring query-relevant preferences from user-related history. Existing benchmarks, however, largely assume that such preference can be retrieved from semantically related history. We study an underexplored but practically important regime, profile-preference conceptual misalignment (PRCM), where observable profile cues and query-specific preferences lie in different concept spaces, making semantic retrieval inconsistent for personalization. We introduce VIBE-Bench, a benchmark with two psychology-grounded tasks, 3,504 personas and 12,239 dialogues, including a manually verified gold test set, and requires cross-concept preference reasoning beyond surface semantic overlap. Experiments with several personalization methods show that current PLLMs largely rely on shallow semantic correlations and fail to acquire robust cross-concept mappings. These findings establish PRCM as a distinct failure regime in PLLMs and position VIBE-Bench as a focused testbed for advancing preference reasoning beyond semantic matching.