Search papers, labs, and topics across Lattice.
Aalto University, Finland
4
0
7
1
A single-vector visual document retrieval system that is 15.6 times smaller and an order of magnitude faster than existing multi-vector models while retaining high accuracy.
Quoting evidence directly from documents boosts evidence recall by up to 39 points while halving hallucination rates, challenging the reliance on coordinate-based methods.
Transforming human motion into structured language allows LLMs to achieve unprecedented accuracy in motion understanding without the constraints of traditional encoding methods.
VLMs struggle to align assembly diagrams and videos because they occupy disjoint visual representation spaces, revealing a fundamental limitation in cross-modal understanding.