Search papers, labs, and topics across Lattice.
3
0
6
4
Current MLLMs falter in following complex video instructions, revealing a significant oversight in their evaluation metrics.
A unified decision process for multi-modal reasoning reveals that joint optimization of text and image generation can dramatically enhance performance in complex reasoning tasks.
Stabilizing RL training is now possible by modulating importance ratios with hazard-aware penalties, preventing both mode collapse and policy erosion.