Search papers, labs, and topics across Lattice.
Affiliation:
6
0
6
Reasoning segmentation models stumble by cramming semantic interpretation and spatial localization into a single prompt token; decoupling "what" from "where" hits state-of-the-art grounding while fine-tuning less than 0.4% of the base MLLM.
Achieving a top rank in the MeViS-Audio challenge, this system showcases the power of agreement-based segmentation in accurately identifying objects described by spoken expressions.
Over 1,500 submissions revealed stark differences in model performance across diverse domains, highlighting the challenges of generalizing egocentric video understanding.
Over 1,100 submissions reveal groundbreaking advancements in sports video understanding, with new methods pushing the boundaries of action prediction and localization.
Current video generation benchmarks overlook crucial aspects of physical plausibility and temporal coherence, highlighting the need for holistic evaluation metrics like PhyScore.
Maritime computer vision is making strides in real-time feasibility, as evidenced by the MaCVi 2026 challenge results.