Search papers, labs, and topics across Lattice.
5
0
7
13
Simply adding more multimodal environments can hinder agent performance, but targeted diversity and structured difficulty can transform training outcomes.
SFT leads to task conflicts that can cripple multi-task learning, while RL's variance-limited updates enable seamless task coexistence.
Suboptimal trajectories can amplify errors in long-horizon planning, revealing critical insights into the limitations of current training methods.
Catastrophic collapses in tool-use performance can be mitigated by strategically interleaving supervised fine-tuning with reinforcement learning.
MLLMs can now reason about streaming video with significantly improved accuracy and reduced output length thanks to a novel memory-anchored framework that overlaps watching and thinking.