Search papers, labs, and topics across Lattice.
3
0
6
14
Captions selected with VEGAS align significantly better with human attention, boosting retrieval performance and challenging the status quo of video captioning metrics.
Text, combined with learned image embeddings, can compress maps by 2x while preserving localization accuracy, offering a practical solution to the growing memory demands of robotic mapping.
Ditch BLEU and ROUGE: ViSIL offers a unified metric for multimodal video captioning that actually correlates with VQA performance and human judgment by measuring information loss via VLM inference.