Search papers, labs, and topics across Lattice.
Affiliation:
2
0
4
SAPO achieves a 15.1 percentage point improvement over PPO while slashing memory costs and runtime by a third, revolutionizing how we optimize agentic RL.
Robots can now learn fine-grained manipulation skills directly from human demonstrations, thanks to a new imitation learning framework that automatically figures out the right rewards and handles noisy real-world data.