Search papers, labs, and topics across Lattice.
2
0
5
3
Achieving a 25% performance improvement in visual reasoning tasks, UniVR reveals the untapped potential of reasoning directly from visual data.
Despite impressive headline scores, today's best video MLLMs can't reliably ground their answers in space and time, achieving <1% accuracy when required to identify the spatio-temporal evidence supporting their predictions.