Search papers, labs, and topics across Lattice.
3
0
4
0
Achieving better sample efficiency and scalability in reward learning, PreferenceEKF redefines how we approach uncertainty in reinforcement learning from human feedback.
Grounding reward learning in natural language rationales makes policies 2x more robust to spurious correlations and distribution shifts.
Learning robotic reward functions from a million trajectories reveals that comparing entire trajectories, not just individual frames, unlocks better generalization and learning from suboptimal data.