Search papers, labs, and topics across Lattice.
6
0
8
46
Semantic visual-action tokenization in RepWAM significantly enhances robotic manipulation performance, outperforming traditional reconstruction-based approaches.
ActiveMimic reveals that leveraging active perception from egocentric videos can close the performance gap with robot-pretrained models, transforming how we approach robot learning.
Forget fine-tuning: VLA-Pro dynamically fuses task-specific LoRA adapters retrieved from memory to achieve state-of-the-art cross-task generalization in robotic manipulation.
Today's best language models can barely make sense of your messy group chats and fragmented digital life, achieving only 19% accuracy on a new benchmark of real-world reasoning.
Even high-performing Vid-LLMs can be easily misled into retracting correct judgments and fabricating justifications under adversarial feedback.
CaTok achieves state-of-the-art ImageNet reconstruction with fewer training epochs by learning causal 1D image representations, outperforming existing visual tokenizers.