Search papers, labs, and topics across Lattice.
4
0
4
2
RefCaptioner not only outperforms existing models in video captioning but also enables precise grounding of visual elements to multiple reference images, enhancing factual accuracy.
Recovering the original speaker's identity from voice-converted audio is now possible with 90.99% accuracy, even under adverse conditions.
MultiRef-Compass reveals that current MR2AV systems have substantial performance gaps, highlighting the urgent need for a standardized evaluation framework in this novel domain.
Current video generation models face a critical trade-off between faithfully executing keyframes and producing natural-looking videos, with performance degrading under increased keyframe density.