Search papers, labs, and topics across Lattice.
6
0
9
2
MUSE decouples image annotation from question generation, enabling diverse tasks with controllable difficulty while reducing annotation effort, and it is hoped MUSE will serve as a standardized benchmark for advancing multi-modal understanding in situated educational applications.
Experiments show improved depth accuracy and temporal consistency over scale-only calibration, demonstrating how localization geometry can support both pose recovery and dense robot perception.
UniMPA, a Unified Memory-Prediction-Action model that addresses transition ambiguity by modeling the intended future state evolution through a shared action-grounded transition interface, is proposed.
Masking compositional concepts in one modality while leveraging contextual cues from another can dramatically enhance the compositionality of vision-language models.