Search papers, labs, and topics across Lattice.
Affiliation:
3
0
4
Pruning visual tokens based on head alignment can retain nearly all performance while drastically reducing computational costs.
DB-Nav redefines object navigation by leveraging relational biases to filter out unreliable semantic cues, achieving superior performance without the overhead of complex vision-language reasoning.
Forget hand-tuning: VisPCO automatically finds optimal visual token pruning configurations in VLMs, outperforming predefined strategies across diverse benchmarks.