Search papers, labs, and topics across Lattice.
3
0
7
3
Simply adding more multimodal environments can hinder agent performance, but targeted diversity and structured difficulty can transform training outcomes.
SFT leads to task conflicts that can cripple multi-task learning, while RL's variance-limited updates enable seamless task coexistence.
A single model now rivals specialized vision-language models in understanding, while also generating and editing images, thanks to a unified discrete diffusion framework.