Search papers, labs, and topics across Lattice.
University of California, Santa Barbara
5
0
10
1
Current action-conditioned world models fail to reliably follow diverse off-expert actions, risking the effectiveness of policy learning in real-world applications.
Pruning 77.8% of visual tokens without losing performance could revolutionize the efficiency of multimodal large language models.
FORCE achieves a remarkable 79% increase in success rates for VLA models while eliminating the need for costly human interventions during training.
Achieving a 6.7x speedup in 3D scene reconstruction without sacrificing quality could redefine efficiency benchmarks in visual geometry tasks.
Just because your agent can write and store memories well doesn't mean it can actually *use* them effectively in a dynamic, multimodal world.