Search papers, labs, and topics across Lattice.
The Hong Kong University of Science and Technology (Guangzhou)
5
0
6
13
MLLMs can now achieve over 8% improvement in spatial reasoning for egocentric scenes by leveraging a novel Ego-element Graph for enhanced perception.
OmniCoT recalibrates the challenges of panoramic reasoning, enabling MLLMs to leverage global evidence for complex multi-step inference.
Modality competition in federated learning can be mitigated by structuring training into dedicated phases, leading to significant performance gains and reduced communication overhead.
Existing affordance prediction models fall flat when confronted with the wide-angle, distorted reality of panoramic vision, but a new training-free pipeline called PAP rises to the challenge.
Pre-trained video diffusion models can be deterministically adapted into state-of-the-art zero-shot depth estimators, sidestepping the need for massive labeled datasets.