Search papers, labs, and topics across Lattice.
3
0
7
2
This work introduces LIFD (Look, Imagine, Focus, and Do), a framework for persistent, 3D-aware scene memory that learns a scene-token representation from multi-view agreement and completes it from a single RGB view and recurrent memory.
Overcoming the 2D-to-4D spatial bottleneck in VLMs does not require native 3D architectures; factorizing visual projections into verifiable planar, depth, and temporal RL objectives delivers immediate 4-6% benchmark gains.
Hyper-redundant robots get a 75% accuracy boost thanks to a neural network that adaptively blends learned behavior with kinematic priors.