Search papers, labs, and topics across Lattice.
12
0
10
4
Current LLMs only achieve 27.3% accuracy in reasoning about scientific lineage, revealing a critical gap in their compositional capabilities.
Generating visually faithful driving simulations just got a boost with a novel framework that stabilizes error accumulation and enhances realism in closed-loop scenarios.
Role-aware training can boost video diffusion models' physical consistency by up to 39.4% without sacrificing visual fidelity.
Large-scale structured academic visual data can transform image generation from mere aesthetic appeal to verifiable knowledge-grounded creation.
Current image editing models struggle with physics-based reasoning, as revealed by the new PhyEditBench benchmark.
SSR not only resolves destructive parameter collisions in LoRA merging but also guarantees mathematical optimality, setting a new standard for efficiency in diffusion model training.
Current video MLLMs struggle to grasp fleeting visual events, with top models barely surpassing 39% accuracy on critical momentary tasks.
NITP achieves a remarkable 5.7% performance boost on MMLU-Pro by transforming how LLMs are trained, moving beyond sparse supervision to dense semantic predictions.
Forget prompt engineering: this new region proposal network spots objects across diverse datasets without *any* text or image prompts.
Current image editing models stumble when domain-specific knowledge is required, as revealed by a new benchmark spanning disciplines from natural science to social science.
Forget billion-scale datasets: EvoTok achieves state-of-the-art image tokenization for both understanding and generation using a residual evolution process trained on just 13M images.
DreamWorld achieves more world-consistent video generation by jointly modeling multiple heterogeneous dimensions of world knowledge, moving beyond surface-level plausibility.