Search papers, labs, and topics across Lattice.
Aalto University, Finland
5
0
5
A single-vector visual document retrieval system that is 15.6 times smaller and an order of magnitude faster than existing multi-vector models while retaining high accuracy.
Quoting evidence directly from documents boosts evidence recall by up to 39 points while halving hallucination rates, challenging the reliance on coordinate-based methods.
VLMs struggle to align assembly diagrams and videos because they occupy disjoint visual representation spaces, revealing a fundamental limitation in cross-modal understanding.
Shrinking a 2B vision-language retriever to a 70M text-only model achieves 95% of the original quality and outperforms a 2B baseline, while slashing query latency by 50x.
Ditch global embeddings for text-motion retrieval: this method uses joint-angle motion images and token-patch late interaction to achieve state-of-the-art accuracy and interpretability.