Search papers, labs, and topics across Lattice.
This paper introduces IBA-Bench, a novel benchmark designed to evaluate implicit behavioral alignment in personalized large language model (LLM) agents by utilizing longitudinal interaction histories that capture evolving user preferences and implicit cues. The authors highlight the limitations of existing benchmarks that rely on static user profiles and demonstrate that these approaches overlook the knowledge-to-action gap in preference-conditioned task execution. Experimental results reveal that while personalization remains a significant challenge for current LLM agents, the proposed IBA-Agent framework significantly enhances behavioral alignment across nine diverse application domains.
Personalization in LLM agents is more complex than previously thought, with existing benchmarks failing to capture the dynamic nature of user preferences and their impact on task execution.
Large Language Models have enabled increasingly capable autonomous agents, yet personalization remains critical for making such agents practically useful. Recent benchmarks have begun evaluating personalization in agents, but they largely rely on static preference snapshots, fixed interaction logs, or question answering over predefined user profiles. Such designs fail to capture the complexity of evolving user preferences and neglect preference-conditioned task execution-a discrepancy we term as the knowledge-to-action gap. To address this challenge, we introduce IBA-Bench, a benchmark for implicit behavioral alignment constructed from longitudinal interaction histories that contain noise, implicit cues, and temporal inconsistencies. Unlike prior work, IBA-Bench evaluates whether an agent can execute tasks while satisfying implicit user constraints inferred from historical interactions. We further propose IBA-Agent, an agent framework that reconciles conflicting priorities through broad retrieval and trajectory-level alignment. Experiment results on IBA-Bench show that effective personalization remains a significant challenge for state-of-the-art LLM agents, and the proposed IBA-Agent substantially improves behavioral alignment in complex scenarios across nine application domains.