Search papers, labs, and topics across Lattice.
3
0
7
6
Current VLMs struggle with specialized domains, failing to adapt effectively in both zero-shot and ICL scenarios, revealing critical gaps in their spatio-temporal reasoning abilities.
SpatialClaw enables agents to dynamically compose and adapt their reasoning strategies, achieving a remarkable 11.2-point accuracy boost over traditional spatial agents.
Encoder-decoders are out: a decoder-only architecture for large view synthesis not only slashes parameters but also beats the state-of-the-art, even outperforming per-scene optimized 3DGS in some cases.