Search papers, labs, and topics across Lattice.
Affiliation:
6
0
10
4
Experiments show that WholeBodyWAM consistently benefits from increased motion-pretraining scale, improves future-motion prediction and downstream task performance, and transfers effectively to real-world humanoid manipulation.
The authors', a persistent language field for structured Gaussian scenes that requires persistent semantic ownership, conserved evidence, and a hierarchy that balances stability, detail, and representation cost, is introduced.
Clip-level captions compress away the continuous visual dynamics needed for true long-horizon reasoning; Kairos restores this signal with dense, time-resolved annotations tracking actions, entities, and attributes across 10-to-30-minute videos.
Achieve human-like dexterity in humanoid robots by unifying visual-language cues with learned whole-body proprioceptive dynamics, outperforming prior methods in complex manipulation tasks.