Search papers, labs, and topics across Lattice.
The Hong Kong University of Science and Technology (Guangzhou)
11
0
14
13
MLLMs can now achieve over 8% improvement in spatial reasoning for egocentric scenes by leveraging a novel Ego-element Graph for enhanced perception.
Front-running can be effectively neutralized by a power-weighted randomized lottery that reshapes transaction incentives and preserves user rewards.
OmniCoT recalibrates the challenges of panoramic reasoning, enabling MLLMs to leverage global evidence for complex multi-step inference.
Routing SQL queries based on complexity allows DecoSearch to achieve unprecedented execution accuracy while using an order of magnitude fewer tokens than traditional methods.
LLM-generated rewards in RL can be misleading early in training, but RHyVE dynamically selects the best reward signal based on policy competence, leading to improved performance.
Existing affordance prediction models fall flat when confronted with the wide-angle, distorted reality of panoramic vision, but a new training-free pipeline called PAP rises to the challenge.
Achieve state-of-the-art panoramic segmentation by training on local perspective views and generalizing to full 360° images, even with geometric distortions and unseen classes.
Pre-trained video diffusion models can be deterministically adapted into state-of-the-art zero-shot depth estimators, sidestepping the need for massive labeled datasets.
Current MLLMs are surprisingly bad at understanding human intent in egocentric videos at a step-by-step level, achieving only 33% accuracy on a new benchmark designed to prevent future-frame leakage.
Forget short looping animations – this new diffusion model generates hour-long, real-time human animations with lip-sync accuracy and emotional expressiveness, all while running on just two GPUs.
The first comprehensive survey of Visual Document Retrieval reveals how MLLMs are reshaping the field, highlighting the shift towards RAG and agentic systems for complex document understanding.