Search papers, labs, and topics across Lattice.
2
0
4
A unified visual representation can achieve high-fidelity image reconstruction and multimodal understanding without the need for separate generative pathways.
Even top-tier MLLMs can falter dramatically in counting tasks that require complex reasoning, revealing a critical gap in multimodal intelligence.