Search papers, labs, and topics across Lattice.
10
0
14
13
WorldScape Policy 2.0 achieves unprecedented long-horizon autonomous planning by integrating reasoning-augmented memory with multimodal instruction processing.
Large-scale structured academic visual data can transform image generation from mere aesthetic appeal to verifiable knowledge-grounded creation.
ACE revolutionizes context management for LLM agents, enabling them to adaptively retain critical information without loss, leading to superior decision-making performance.
Language-action pretraining can lead to VLA policies that are not only more robust but also less dependent on visual cues, achieving up to 45% higher success rates in real-world tasks.
FlowTracer reveals that optimizing token-level rewards based on attention-induced information flow can dramatically enhance reasoning performance in LLMs.
NITP achieves a remarkable 5.7% performance boost on MMLU-Pro by transforming how LLMs are trained, moving beyond sparse supervision to dense semantic predictions.
A billion-scale SAR foundation model closes the domain generalization gap in SAR imagery by 10% mIoU, thanks to a physics-guided MoE architecture.
Hallucinations in RL-based image editing and generation are tamed with FIRM, a new framework that trains robust reward models on curated datasets to provide more accurate guidance.
Current image editing models stumble when domain-specific knowledge is required, as revealed by a new benchmark spanning disciplines from natural science to social science.
Forget billion-scale datasets: EvoTok achieves state-of-the-art image tokenization for both understanding and generation using a residual evolution process trained on just 13M images.