Search papers, labs, and topics across Lattice.
3
0
5
2
Generation-guided training can significantly enhance multimodal understanding in MLLMs without any inference overhead.
Optimal data ratios discovered through DecoupleMix yield competitive VLM performance with 80B fewer training tokens than traditional methods.
LLMs and VLMs can encode viewpoint, but they still hallucinate observations when reasoning about viewpoint rotation, revealing a critical gap in their spatial intelligence.