Search papers, labs, and topics across Lattice.
2
0
3
4
Captions selected with VEGAS align significantly better with human attention, boosting retrieval performance and challenging the status quo of video captioning metrics.
Ditch BLEU and ROUGE: ViSIL offers a unified metric for multimodal video captioning that actually correlates with VQA performance and human judgment by measuring information loss via VLM inference.