Search papers, labs, and topics across Lattice.
2
0
4
4
Video language models falter dramatically in counting transient events, with less than 0.2% accuracy in high-frequency scenarios.
Spatially grounded visual memory can dramatically enhance the performance of EQA agents, breaking the accuracy-efficiency tradeoff in continuous settings.