Search papers, labs, and topics across Lattice.
16
0
13
3
Oxygen-TryOn achieves unprecedented realism in virtual try-on by synthesizing images across diverse fashion categories, outperforming both proprietary and open-source models.
SGF bridges the gap in video generation by allowing future losses to inform past latent encodings, resulting in unprecedented long-video extrapolation capabilities.
Reducing sampling steps from 50 to just 8 without sacrificing quality could revolutionize how we approach generative modeling.
DiT-Reward not only outperforms existing models in image evaluation but also accelerates inference by 1.65x without sacrificing quality.
Human raters overwhelmingly prefer JoyAI-VL-Interaction over existing video-call assistants, showcasing a leap in real-time interaction capabilities.
AnchorEdit achieves state-of-the-art performance in multi-turn image editing by maintaining subject identity across 10+ interactions, revolutionizing iterative design workflows.
Raw context outperforms compact memory designs, revealing that memory structure is crucial for effective video generation in action-conditioned models.
Ultra Flash achieves real-time high-resolution video generation at unprecedented frame rates, pushing the boundaries of what’s possible in streaming video AI.
Despite the advancements in multimodal agents, even the best models struggle with interactive spatial reasoning, achieving only a 17.4% success rate in complex real-world tasks.
Bidirectional interaction between enhanced understanding, controllable spatial editing, and novel-view-assisted reasoning enables a unified multimodal model to achieve spatial intelligence beyond general visual competence.
Unlock the full potential of your pretrained video diffusion models with a surprisingly simple four-stage post-training framework that drastically improves visual quality, temporal coherence, and instruction following.
Forget external teachers – the best way to boost your RL model's performance is to learn from its future self.
Bridging the gap between human manipulation and robotic control, JoyAI-RA unlocks enhanced cross-embodiment behavior learning through multi-source pretraining.
EasyVideoR1 achieves a 1.47 times throughput improvement in video understanding tasks by eliminating redundant video decoding and leveraging a comprehensive task-aware reward system.
Spatial reasoning gets a major boost: OpenSpatial-3M, a new dataset, enables models to leapfrog existing benchmarks by 19%.
Achieve real-time, synchronized audio-visual generation at 25 FPS by distilling a bidirectional diffusion model into a fast, autoregressive architecture, overcoming training instability with novel alignment and token handling techniques.