Search papers, labs, and topics across Lattice.
This paper introduces test-time adaptation through human-agent interaction (TAHI), which leverages iterative feedback from users to refine AI agent performance on open-ended tasks. By integrating personalized criteria into the agent's context and weights, TAHI significantly enhances task success rates for individual users across writing and visual creation domains. The approach not only improves individual task success by 4.5-20.9% but also generates evolving rubrics that capture 16.0-22.3% more failures than traditional methods, demonstrating both personalized and generalized performance gains.
Personalized AI agents can achieve up to 20.9% better task success by learning from user feedback in real-time, reshaping the landscape of human-AI collaboration.
AI agents are trained on population-scale data to encode broad capabilities spanning those of many practitioners. Yet the artifacts they produce rarely meet the personal bar professionals need to stake their reputation on. On realistic, open-ended tasks where success criteria are heterogeneous and insufficiently documented, individual expertise lives precisely in the elevation and departure from the average. In practice, iterative human-agent interaction surfaces criteria that users cannot fully specify up front, yet apply repeatedly across tasks. We argue this cross-session interaction data is a rich, underused signal for closing the gap to individual expertise. In this work, we propose test-time adaptation through human-agent interaction (TAHI), which integrates these signals into agent context and weights, and crystallizes each user's training and evaluation criteria via an evolving rubric module. We adapt agents to 30 individuals in two high-utility domains, writing and visual creation, on a total of 600 tasks. Our agents improve solo task success by 4.5-20.9% within only tens of tasks. Meanwhile, our evolving rubric module serves as a scalable annotation tool, creating evaluation rubrics that catch 16.0-22.3% more failures than those from LMs or humans alone. While agents are adapted towards individuals, we show these personalized agents also produce improvements in success of up to 8.8% that generalize across users.