Search papers, labs, and topics across Lattice.
2
5
6
4
Continuous scoring from LLM-as-a-Verifier leads to state-of-the-art verification accuracy and improved sample efficiency in reinforcement learning tasks.
Verification at test time can be a surprisingly effective alternative to scaling policy learning for vision-language-action alignment, yielding substantial gains in both simulated and real-world robotic tasks.