Search papers, labs, and topics across Lattice.
This study reveals that user feedback, often dismissed as noisy, is a valuable signal for improving Large Language Models (LLMs) when evaluated correctly. By constructing both synthetic and naturalistic datasets, the authors demonstrate that model revisions informed by user feedback resolve targeted issues more effectively than those based solely on baseline outputs. The research also uncovers a systematic bias in evaluation paradigms, where judges fail to recognize improvements made through feedback, favoring less effective baseline responses instead.
User feedback can significantly enhance LLM performance, but current evaluation methods often overlook its benefits, leading to misguided assessments of model improvements.
Harnessing naturally occurring feedback from user interactions offers a promising learning signal for Large Language Models (LLMs). However, recent studies suggest this feedback is inherently noisy and difficult to leverage effectively. We challenge this conception by demonstrating that user feedback is a highly actionable signal for improvement, and that its perceived ineffectiveness stems from a systematic bias in current evaluation paradigms. To isolate the usefulness of feedback, we construct synthetic data with a definitive ground truth, alongside naturalistic data to validate that our findings hold in real-world scenarios. By comparing model revisions generated with and without access to feedback across both settings, we show that feedback-informed revisions resolve targeted issues at significantly higher rates than baseline revisions. Finally, we expose the root of the evaluation bias: when a model successfully fixes an issue exclusively due to feedback, LLM judges frequently fail to identify the genuinely corrected response, systematically preferring inferior baseline outputs instead.