Search papers, labs, and topics across Lattice.
3
0
5
3
Spatially grounded visual memory can dramatically enhance the performance of EQA agents, breaking the accuracy-efficiency tradeoff in continuous settings.
Even top-performing Vision Language Models miss 30% of logical anomalies, revealing critical gaps in their reasoning abilities.
SOTA audio QA models are getting punked by trivia questions a toddler could answer, revealing a stark gap between current capabilities and true audio understanding.