Search papers, labs, and topics across Lattice.
5
0
5
0
Role-aware matching in image retrieval can boost performance by up to 23% without the need for task-specific training.
AdaptiveEmbed reveals that tailoring representation capacity to individual sample needs can dramatically enhance multimodal retrieval performance.
Dynamic visual evidence retrieval can accelerate multimodal decoding by over 2x while enhancing draft acceptance rates.
Hallucinations in LVLMs can be cut by over 43% without sacrificing grounded object coverage, thanks to a novel verifier-guided approach.
LVLMs are better at spotting their own mistakes than generating correct answers in the first place, and this self-awareness can be exploited to reduce hallucinations.