Search papers, labs, and topics across Lattice.
The Hong Kong University of Science and Technology (Guangzhou)
4
0
4
By integrating semantic foresight from vision-language models, SG-WAM ensures that robotic actions are not just visually accurate but also linguistically aligned, transforming how robots interpret and execute tasks.
Robust-WAM achieves superior out-of-distribution generalization in robot control by seamlessly integrating semantic foresight into action predictions while leveraging extensive VGM pretraining.
Achieving 98.0% success in cross-embodiment manipulation without manual action alignment could redefine how we approach robot control across diverse platforms.
Ditch slow, multi-step video generation: S-VAM distills the structured generative priors of multi-step denoising into a single forward pass for real-time robot action prediction.