Search papers, labs, and topics across Lattice.
Department of Foundation Model, 2012 Labs
9
1
7
3
A single-layer action head on a 30-layer video backbone enables Faster-WAM to achieve a 3.2x speedup in inference latency while maintaining competitive performance and robust generalization.
RoboHarness achieves remarkable improvements in long-horizon planning by seamlessly orchestrating diverse robot policies without retraining, even in uncertain environments.
Future visual cues can dramatically enhance navigation performance, even when not accessible during actual deployment.
Extracting interaction cues from a frozen video model enables robots to achieve up to 90.6% success in manipulation tasks without costly rollout processes.
SVP-IL boosts success rates on ambiguous language tasks by over 60% with minimal training data, revolutionizing data efficiency in robotic manipulation.
Overcoming perceptual uncertainty in vision-language navigation is now possible by explicitly modeling geometric, semantic, and appearance uncertainty with a novel Uncertainty-Aware Gaussian Map.
Policies trained with view augmentation from a novel feed-forward 3D Gaussian Splatting framework maintain robust execution under severe spatial perturbations where baselines fail.
Web-scale video pretraining lets robots handle real-world chaos better than vision-language models trained on curated robotics datasets.
A $14K bimanual robot with a Python-first control framework could democratize embodied AI research by lowering the barrier to entry for complex manipulation tasks.