Search papers, labs, and topics across Lattice.
5
0
5
2
GenEvA achieves up to 10.1 points higher accuracy in long-video understanding while using less than 0.4% additional video tokens.
Achieving a 14-point boost in grounding accuracy, VistaRef redefines how we approach spatial orientation in AR and human-robot interaction.
PointVG-R achieves a groundbreaking 15.86-point boost in mIoU by integrating geometric reasoning into visual grounding tasks, reshaping our approach to spatial interpretation in models.
Shifting from answer-centric to evidence-centric reasoning, CoVER-7B outperforms even leading closed-source models in long-video understanding tasks.
Mimicking how clinicians review capsule endoscopy videos鈥攆irst screening, then weaving context, and finally converging evidence鈥攜ields surprisingly effective summarization of these ultra-long videos.