Search papers, labs, and topics across Lattice.
Affiliation:
1
0
3
Throwing away learned attention in favor of pure 3D spatial coverage allows VLMs to retain over 93% of their multi-view 3D reasoning performance while slashing visual token budgets by ~92%.