Search papers, labs, and topics across Lattice.
11
0
12
16
Orthogonal JEPA achieves superior predictive performance by factorizing latent states, allowing for more efficient learning in complex systems.
Leading MLLMs falter on the new VideoGAIA benchmark, scoring under 60% accuracy in complex, multi-turn video understanding tasks.
Operation laundering in vision encoders can be effectively mitigated, revealing hidden boundaries in learned assignments that traditional methods obscure.
Executable Blender code transforms text-to-video generation, enabling unprecedented control over scene dynamics and visual fidelity.
Unsupervised reward optimization allows protein language models to self-improve, achieving near-oracle performance without the need for labeled data.
OmniAgent not only outperforms larger models but also scales performance with reasoning turns, revolutionizing how we approach video understanding.
Surpassing existing methods, Orchestra-o1 achieves a 10.3% accuracy improvement on the OmniGAIA benchmark by enabling seamless collaboration across multiple modalities.
Forget static imitation learning: LaST-R1 unlocks near-perfect robotic manipulation (99.8% success) by adaptively reasoning about physical dynamics *before* acting, then refining with RL.
Generating realistic human-object interaction videos from text, images, audio, *and* pose is now possible, opening the door to automated content creation workflows.
Finally, a neural interatomic potential that accurately models long-range electrostatic interactions without sacrificing SO(3) equivariance or energy-force consistency.
Visual grounding in VLAs weakens in deeper layers, but injecting multi-level visual features and pruning irrelevant tokens can boost performance by 9% in simulation and 7.5% in the real world.