Search papers, labs, and topics across Lattice.
16
0
13
4
WorldScape Policy 2.0 achieves unprecedented long-horizon autonomous planning by integrating reasoning-augmented memory with multimodal instruction processing.
Current LLMs only achieve 27.3% accuracy in reasoning about scientific lineage, revealing a critical gap in their compositional capabilities.
Role-aware training can boost video diffusion models' physical consistency by up to 39.4% without sacrificing visual fidelity.
ACE-Brain-0.5 unifies spatial reasoning and action generation in embodied AI, achieving remarkable performance improvements across multiple benchmarks.
Large-scale structured academic visual data can transform image generation from mere aesthetic appeal to verifiable knowledge-grounded creation.
ACE revolutionizes context management for LLM agents, enabling them to adaptively retain critical information without loss, leading to superior decision-making performance.
Current video MLLMs struggle to grasp fleeting visual events, with top models barely surpassing 39% accuracy on critical momentary tasks.
Current AI agents struggle to reliably rediscover scientific knowledge, with top performers averaging only 21.5 out of a possible score, revealing critical gaps in their research capabilities.
Current omnimodal LLMs that ace offline benchmarks still fumble basic real-time interactions, highlighting a critical gap in their ability to handle streaming audio-visual data.
LLM-powered agents can now produce surprisingly strong photographs in complex 3D environments, suggesting a path towards embodied AI with aesthetic awareness.
Model-generated skills can actually hurt agent performance, and bigger models don't necessarily make for better skill extractors or consumers.
SkillOpt transforms agent skill development into a reproducible optimization process, achieving state-of-the-art results by treating skills as trainable parameters.
Visual degradations can cripple the spatial reasoning abilities of even state-of-the-art MLLMs, but targeted finetuning can restore—and even surpass—human-level performance.
LLMs can have their personalities surgically altered by tweaking just 0.5% of their neurons, preserving general capabilities while achieving competitive control.
Current image editing models stumble when domain-specific knowledge is required, as revealed by a new benchmark spanning disciplines from natural science to social science.
Hallucinations in RL-based image editing and generation are tamed with FIRM, a new framework that trains robust reward models on curated datasets to provide more accurate guidance.