Search papers, labs, and topics across Lattice.
7
0
5
19
Embedding reference tokens at semantic positions allows for unprecedented precision in multi-reference video editing, setting a new benchmark for instruction quality.
Current video generation models face a critical trade-off between faithfully executing keyframes and producing natural-looking videos, with performance degrading under increased keyframe density.
MultiRef-Compass reveals that current MR2AV systems have substantial performance gaps, highlighting the urgent need for a standardized evaluation framework in this novel domain.
AVSCap-7B achieves superior audio-visual synergy, outperforming existing models by effectively linking non-speech sounds to visual actions.
Dynamic identity memory and large-scale counterfactual self-supervision enable Argus to outperform previous methods in subject-preserving video generation, achieving unprecedented robustness against occlusion and viewpoint changes.
Current video editing models falter under the weight of complex user instructions, often omitting critical edits and introducing artifacts.
Today's best multimodal models can only solve half of compositional visual tool-use tasks, revealing a critical gap in their ability to plan and execute complex, multi-step visual reasoning.