Search papers, labs, and topics across Lattice.
University of Chinese Academy of Sciences
5
1
6
4
A simple logistic regression can match the performance of advanced models in evaluating emotion descriptions, raising questions about the validity of current multimodal benchmarks.
Agents can now seamlessly transition between seeking and following tasks in dynamic environments, setting a new benchmark for embodied AI performance.
FlowWAM achieves a remarkable 92.94% success rate in manipulation tasks by harnessing optical flow as a video-native action representation.
Structured supervision can boost VLA model performance by over 50% in complex robotic tasks, transforming how we approach fine-tuning in manipulation.
Ditch the multi-camera setup: VERM leverages foundation models to synthesize a single, task-optimized "virtual eye" view for robots, slashing training time by 1.89x.