Search papers, labs, and topics across Lattice.
2
0
4
CapProbe reveals that many vision-language models exhibit substantial gaps in coverage, challenging the reliability of existing caption evaluation metrics.
SlimVLM achieves unprecedented efficiency in Vision-Language Models by intelligently pruning redundant visual tokens without sacrificing performance.