Search papers, labs, and topics across Lattice.
15
0
10
3
VA-Judger transforms reward modeling by integrating human preference feedback, leading to coherent and high-quality video-audio generation that outperforms traditional metrics.
Robots can now learn from their experiences and adapt in real-time, paving the way for a new era of general-purpose robotic agents.
MLLMs misallocate attention, leading to hallucinations, but SPAR reclaims focus on meaningful visual information with minimal computational cost.
Human demonstrations can yield over 10x the recovery data for robots, dramatically enhancing their ability to recover from failures in real-world tasks.
SegDiff achieves real-time adaptability in robot manipulation by seamlessly combining continuous trajectory prediction with keypose segmentation, outperforming traditional methods.
Fine-grained contact states can be distinguished through the dynamic correlation of tactile motion, transforming how we approach contact-rich manipulation in robotics.
Semantic visual-action tokenization in RepWAM significantly enhances robotic manipulation performance, outperforming traditional reconstruction-based approaches.
Reinforcement learning boosts multimodal performance, raising task scores and creating unexpected synergies between image generation and editing.
UniDexTok slashes reconstruction errors by over 98% for dexterous hands, achieving unprecedented accuracy without relying on retargeting.
Aligning shallow and deep features in representation autoencoders leads to a dramatic improvement in image reconstruction quality, setting new benchmarks in the field.
OmniGen-AR can seamlessly generate images from a wide array of conditions, outperforming existing methods that are limited to single-modality inputs.
Visual features can be the game-changer in recommendation systems, but they鈥檙e often overlooked鈥擱EVEAL flips the script by making them a focal point.
Gradual bridging with embodied trajectory-coupled data transforms VLMs into robust robot control policies, overcoming significant transfer challenges.
ActiveMimic reveals that leveraging active perception from egocentric videos can close the performance gap with robot-pretrained models, transforming how we approach robot learning.
Language models are increasingly doing their real work in the "invisible" latent space, not the tokens we see.