Search papers, labs, and topics across Lattice.
HKUST(GZ)
4
0
6
3
RefCaptioner not only outperforms existing models in video captioning but also enables precise grounding of visual elements to multiple reference images, enhancing factual accuracy.
Existing models mismanage tool use, but Beacon achieves a balance that enhances performance on complex tasks while preserving accuracy on simpler ones.
MultiRef-Compass reveals that current MR2AV systems have substantial performance gaps, highlighting the urgent need for a standardized evaluation framework in this novel domain.
Current video generation models face a critical trade-off between faithfully executing keyframes and producing natural-looking videos, with performance degrading under increased keyframe density.