Search papers, labs, and topics across Lattice.
Peking University
7
0
7
Training-free HOI detection achieves superior performance by harnessing the multimodal reasoning of foundation models, challenging the need for dataset-specific supervision.
MLLMs struggle to match human accuracy in fine-grained perception, with a striking performance gap revealed by the new DiCoBench benchmark.
I2V models not only excel at dynamic editing but also provide a unique lens for diagnosing errors in Human-Object Interaction tasks.
Current LVLMs are inadequate at fine-grained image recognition, revealing critical bottlenecks in visual and semantic processing that need urgent attention.
Superficial reasoning in video temporal grounding can be transformed into high-quality, time-aware insights with the right optimization framework.
You can now automatically transform structurally flawed photos into aesthetically pleasing images, thanks to a new framework that plans and executes edits based on photographic principles.
MLLMs are better at understanding videos than directly grounding text queries within them, and a self-correction training loop can close the gap.