Search papers, labs, and topics across Lattice.
3
0
5
7
Generation-guided training can significantly enhance multimodal understanding in MLLMs without any inference overhead.
Optimal data ratios discovered through DecoupleMix yield competitive VLM performance with 80B fewer training tokens than traditional methods.
MLLMs struggle with video temporal-logical reasoning, showing a substantial performance gap compared to human capabilities, especially as complexity increases.