Search papers, labs, and topics across Lattice.
4
0
4
5
RefCaptioner not only outperforms existing models in video captioning but also enables precise grounding of visual elements to multiple reference images, enhancing factual accuracy.
Current models falter in executing cross-modal editing instructions, revealing significant gaps in audio-visual consistency and fidelity.
MultiRef-Compass reveals that current MR2AV systems have substantial performance gaps, highlighting the urgent need for a standardized evaluation framework in this novel domain.
Current video generation models face a critical trade-off between faithfully executing keyframes and producing natural-looking videos, with performance degrading under increased keyframe density.