Search papers, labs, and topics across Lattice.
Qwen Applications
13
0
8
7
Joint audio-video models often stay in sync with each other while drifting entirely from the script; explicitly routing text guidance onto the shared temporal axis cuts shot boundary error by 96% down to 42 milliseconds.
Fine-grained metrics reveal that robots can recover from failures more effectively than previously thought, reshaping our understanding of their capabilities.
Achieving seamless audio-visual identity swapping in talking videos while preserving original dynamics could revolutionize content creation and personalization.
ATDEdit achieves unprecedented image preservation during semantic edits, setting a new benchmark in diffusion-based image editing.
Memory systems that intelligently filter and prioritize information can drastically improve GUI agent performance, as shown by FocusMem's superior results across multiple benchmarks.
OmniPack achieves a remarkable 98% performance retention with a staggering 83.3% reduction in computational load, revolutionizing token compression for omni-modal models.
RefCaptioner not only outperforms existing models in video captioning but also enables precise grounding of visual elements to multiple reference images, enhancing factual accuracy.
Existing models mismanage tool use, but Beacon achieves a balance that enhances performance on complex tasks while preserving accuracy on simpler ones.
GeoLens outperforms traditional single-tool approaches by effectively integrating multiple visual reasoning tools, achieving superior accuracy and efficiency in complex remote sensing tasks.
Naive matching in diffusion distillation can inadvertently amplify errors due to hidden information in teacher models, leading to a surprising failure mode called Negative Branch Asymmetry.
Embedding reference tokens at semantic positions allows for unprecedented precision in multi-reference video editing, setting a new benchmark for instruction quality.
Current video generation models face a critical trade-off between faithfully executing keyframes and producing natural-looking videos, with performance degrading under increased keyframe density.
MultiRef-Compass reveals that current MR2AV systems have substantial performance gaps, highlighting the urgent need for a standardized evaluation framework in this novel domain.