Search papers, labs, and topics across Lattice.
Affiliation:
6
0
8
7
Free-form language reasoning can dramatically enhance robotic manipulation, outperforming traditional instruction-based methods in complex tasks.
Achieving reliable robot policy adaptation from a single demonstration could revolutionize how robots learn and interact with their environments.
Different views of the same problem reveal hidden reasoning paths, enabling VLMs to achieve unprecedented accuracy in multimodal reasoning tasks.
Allowing language models to explore unsafe reasoning can actually enhance their ability to discern harmful from harmless prompts, reducing over-refusal without sacrificing safety.
ExpRL outperforms traditional reinforcement learning methods by effectively rewarding intermediate reasoning steps, leading to better LLM performance on complex tasks.
Forget simple scaling laws: the compute-optimal number of parallel rollouts in LLM RL plateaus, revealing distinct mechanisms for easy vs. hard problems.