Search papers, labs, and topics across Lattice.
Affiliation:
4
0
6
Achieving reliable robot policy adaptation from a single demonstration could revolutionize how robots learn and interact with their environments.
Different views of the same problem reveal hidden reasoning paths, enabling VLMs to achieve unprecedented accuracy in multimodal reasoning tasks.
Allowing language models to explore unsafe reasoning can actually enhance their ability to discern harmful from harmless prompts, reducing over-refusal without sacrificing safety.
ExpRL outperforms traditional reinforcement learning methods by effectively rewarding intermediate reasoning steps, leading to better LLM performance on complex tasks.