Search papers, labs, and topics across Lattice.
This paper introduces Verbal Reinforcement Learning (VRL) as a novel paradigm for enhancing language agents through natural language feedback, which can articulate intent, preferences, and causal relationships. The authors categorize VRL into three pillars: using language as a grounding signal for defining tasks, as deliberative feedback for guiding reasoning without model updates, and as a learning signal for parameter adjustment during training. By synthesizing existing work and outlining the implications of each pillar, the study highlights how verbal reinforcement is transforming agent development and presents both challenges and opportunities for creating more aligned AI systems.
Verbal feedback is not just a communication tool; it fundamentally reshapes how language agents learn and operate across their lifecycle.
Natural language is emerging as a primary feedback channel for improving language agents, capable of conveying intent, preferences, and causal structure in forms interpretable by both humans and modern language models. We call this paradigm Verbal Reinforcement Learning (VRL) and offer the first unified account of it. We organize the field around a single axis, \textit{when} verbal feedback takes effect in an agent's lifecycle and \textit{what} it modifies, yielding three pillars: (1) \textbf{Language as Grounding Signal}, where language defines the task itself by specifying goals, states, and reward structures; (2) \textbf{Language as Deliberative Feedback}, where natural language guides reasoning at test time without the need to update model parameters; (3) \textbf{Language as Learning Signal}, where language-based feedback shapes model parameters through training. Within each pillar, we synthesize representative work, distinguish key subcategories of approaches, and outline the distinct role language plays in shaping agent behavior. Together, this taxonomy shows how verbal reinforcement is reshaping agent development, while also defining the challenges and opportunities for building more capable and aligned agents.