Search papers, labs, and topics across Lattice.
5
0
9
0
Forecasting future states from video prefixes reveals a significant gap in current models' retrieval capabilities, with LFTR leading the way in bridging this divide.
TIGER achieves remarkable speedup in multimodal generation by intelligently routing visual tokens based on textual context, outpacing traditional methods.
Conveying depth as text rather than images can significantly boost spatial reasoning in vision-language models, challenging conventional approaches.
Tactile feedback from WT-UMI enables humanoid robots to achieve superior manipulation performance by effectively bridging the gap between human intuition and robotic execution.
Weak-to-strong reward models can ace the test but still fail in the real world, revealing a hidden brittleness in current preference learning approaches.