Search papers, labs, and topics across Lattice.
7
0
4
5
Block3D slashes text-to-3D generation time from over 25 seconds to under 5, revolutionizing efficiency without compromising quality.
Latent spatial memory can accelerate video generation by over 10 times while dramatically reducing memory usage, revolutionizing how we model dynamic scenes.
Long video generation fails not just because of limited context length, but because of *how* that context is allocated – and ReCA's hierarchical approach shows a way to fix it.
Expert-level video aesthetics can be captured and improved using a hierarchical rubric and reward models trained with a progressive learning scheme.
Text-to-video models can now learn geometrically consistent world dynamics via reinforcement learning, without expensive architectural changes.
Counterintuitively, VLMs can achieve higher VQA accuracy by intentionally degrading visual inputs, suggesting that high-resolution details can act as noise that hinders reasoning.
Generative video models can now simulate a continuously evolving world, even when objects are out of sight, thanks to a new framework that maintains persistent global state.