Search papers, labs, and topics across Lattice.
Affiliation:
14
0
10
19
Manifold drift can lead to substantial misalignment in generative models, but ThermoDPO offers a powerful solution that anchors preference optimization to the pretrained data manifold.
OmniPack achieves a remarkable 98% performance retention with a staggering 83.3% reduction in computational load, revolutionizing token compression for omni-modal models.
Evolving contexts can transform the stability and effectiveness of on-policy distillation, leading to superior performance in open-ended tasks.
Existing models mismanage tool use, but Beacon achieves a balance that enhances performance on complex tasks while preserving accuracy on simpler ones.
RefCaptioner not only outperforms existing models in video captioning but also enables precise grounding of visual elements to multiple reference images, enhancing factual accuracy.
Captions generated by PercepCap are not only more accurate but also grounded in explicit spatio-temporal perception, revealing the underlying reasoning behind each description.
Current models falter in executing cross-modal editing instructions, revealing significant gaps in audio-visual consistency and fidelity.
Embedding reference tokens at semantic positions allows for unprecedented precision in multi-reference video editing, setting a new benchmark for instruction quality.
Current video generation models face a critical trade-off between faithfully executing keyframes and producing natural-looking videos, with performance degrading under increased keyframe density.
MultiRef-Compass reveals that current MR2AV systems have substantial performance gaps, highlighting the urgent need for a standardized evaluation framework in this novel domain.
AVSCap-7B achieves superior audio-visual synergy, outperforming existing models by effectively linking non-speech sounds to visual actions.
Dynamic identity memory and large-scale counterfactual self-supervision enable Argus to outperform previous methods in subject-preserving video generation, achieving unprecedented robustness against occlusion and viewpoint changes.
Current video editing models falter under the weight of complex user instructions, often omitting critical edits and introducing artifacts.
Today's best multimodal models can only solve half of compositional visual tool-use tasks, revealing a critical gap in their ability to plan and execute complex, multi-step visual reasoning.