Search papers, labs, and topics across Lattice.
12
0
11
37
ZipTok3D reconstructs complex 3D objects with up to 32 times fewer tokens than traditional methods, revolutionizing the efficiency of 3D generation.
Achieving a $5.15\times$ speedup in text-to-3D generation without sacrificing quality could revolutionize real-time 3D content creation.
State-of-the-art generative models struggle to maintain physical consistency and coherent interactions over time, revealing critical gaps in their world modeling capabilities.
K-Forcing accelerates token generation by 2.4-3.5x without abandoning the autoregressive backbone, making it a game-changer for high-load deployments.
Latent spatial memory can accelerate video generation by over 10 times while dramatically reducing memory usage, revolutionizing how we model dynamic scenes.
Long video generation fails not just because of limited context length, but because of *how* that context is allocated – and ReCA's hierarchical approach shows a way to fix it.
Finally, a feed-forward 3D reconstruction method that spits out meshes ready for physics engines, no expensive post-processing needed.
Text-to-video models can now learn geometrically consistent world dynamics via reinforcement learning, without expensive architectural changes.
Forget representational differences - the secret to better feed-forward 3D scene modeling lies in tackling five core design problems.
LLMs can achieve 2.5x higher throughput and 10.7x KV memory reduction in long-context reasoning by compressing the KV cache using trigonometric functions derived from pre-RoPE query/key vector distributions.
Counterintuitively, VLMs can achieve higher VQA accuracy by intentionally degrading visual inputs, suggesting that high-resolution details can act as noise that hinders reasoning.
The medical imaging AI community is being held back by a fragmented data landscape, but a new metadata-driven fusion paradigm offers a path to unlocking the power of foundation models.