Search papers, labs, and topics across Lattice.
Harbin Institute of Technology
7
0
9
MLLMs can identify broad user interests from social media, but they falter on fine-grained preferences, revealing a critical gap in personalization capabilities.
ActWorld bridges the navigation-interaction gap in interactive world models, enabling rich object interactions that were previously unattainable.
Fusing event data with global illumination unlocks significantly better low-light image enhancement, outperforming prior art by up to 1.24dB PSNR and 0.069 SSIM.
Linear transport flows between degraded and clean image domains enable fast, adaptable image restoration that outperforms existing methods in distortion-perception balance.
Unified benchmarks reveal the state-of-the-art in simultaneously addressing multiple real-world image degradations like blur, low-light, and rain.
EQA agents can now handle dynamic, human-populated scenes better thanks to a training-free method that selectively remembers only the most informative visual evidence.
Forget generic CoT: Embed-RL uses reinforcement learning to generate reasoning traces that are explicitly optimized for multimodal embedding tasks, leading to significant performance gains.