Search papers, labs, and topics across Lattice.
WeChat Vision, Tencent Inc.
9
0
12
2
Runtime load balancing in FVAttn slashes attention processing time by over four times, transforming video generation efficiency.
Surpassing larger models, this agent achieves 91.4% retrieval accuracy in long-horizon multimodal dialogues by leveraging episodic memory for efficient context management.
Visual reranking and active rejection in MMAgent-R$^2$ significantly boost retrieval accuracy in challenging KB-VQA tasks, outperforming traditional methods.
Shifting from bounding boxes to pixel-level segmentation in MLLMs leads to significant gains in visual reasoning accuracy and segmentation performance.
Existing video world models struggle with long-term memory retention, and MBench exposes their critical limitations while providing a structured path for future improvements.
Forget brute-force scaling: REVERSE shows that teaching an agent *how* to search and verify evidence lets a smaller model beat giants at image geo-localization.
Aligning diffusion models with human preferences just got a fidelity upgrade: DRM leverages the generative backbone itself for rewards, unlocking step-wise guidance that boosts image quality.
Current video understanding models struggle with long-horizon robustness and non-speech audio, as revealed by the new OmniPro benchmark designed for comprehensive omni-modal proactive evaluation.
Removing objects from video now means removing their shadows and reflections too, thanks to a new method that teaches diffusion models to "understand" object-scene physics.