Search papers, labs, and topics across Lattice.
19
0
16
5
Oxygen-TryOn achieves unprecedented realism in virtual try-on by synthesizing images across diverse fashion categories, outperforming both proprietary and open-source models.
Fine-grained hallucination diagnosis can dramatically enhance the reliability of multimodal language models by revealing the types of hallucinations they produce and how to correct them.
Spatial reasoning hallucinations in MLLMs can be drastically reduced by integrating geometric evidence, challenging the effectiveness of traditional mitigation methods.
SGF bridges the gap in video generation by allowing future losses to inform past latent encodings, resulting in unprecedented long-video extrapolation capabilities.
KAT-Coder-V2.5 outperforms existing models in agentic tool-use, showcasing a new paradigm for autonomous coding agents within executable environments.
Reducing sampling steps from 50 to just 8 without sacrificing quality could revolutionize how we approach generative modeling.
DiT-Reward not only outperforms existing models in image evaluation but also accelerates inference by 1.65x without sacrificing quality.
AnchorEdit achieves state-of-the-art performance in multi-turn image editing by maintaining subject identity across 10+ interactions, revolutionizing iterative design workflows.
Human raters overwhelmingly prefer JoyAI-VL-Interaction over existing video-call assistants, showcasing a leap in real-time interaction capabilities.
Despite the advancements in multimodal agents, even the best models struggle with interactive spatial reasoning, achieving only a 17.4% success rate in complex real-world tasks.
Ultra Flash achieves real-time high-resolution video generation at unprecedented frame rates, pushing the boundaries of what’s possible in streaming video AI.
Raw context outperforms compact memory designs, revealing that memory structure is crucial for effective video generation in action-conditioned models.
Real-time infinite video generation is now feasible, achieving over 1.3 million frames in a single 24-hour rollout without sacrificing quality.
Bidirectional interaction between enhanced understanding, controllable spatial editing, and novel-view-assisted reasoning enables a unified multimodal model to achieve spatial intelligence beyond general visual competence.
Unlock the full potential of your pretrained video diffusion models with a surprisingly simple four-stage post-training framework that drastically improves visual quality, temporal coherence, and instruction following.
Bridging the gap between human manipulation and robotic control, JoyAI-RA unlocks enhanced cross-embodiment behavior learning through multi-source pretraining.
Spatial reasoning gets a major boost: OpenSpatial-3M, a new dataset, enables models to leapfrog existing benchmarks by 19%.
Existing image editing models fall short when it comes to precise spatial manipulations, but a new benchmark and dataset reveal the path to closing the gap.
Achieve real-time, synchronized audio-visual generation at 25 FPS by distilling a bidirectional diffusion model into a fast, autoregressive architecture, overcoming training instability with novel alignment and token handling techniques.