Search papers, labs, and topics across Lattice.
Achieving a 25% reduction in prediction error, MOSH-WM redefines the landscape of object-centric video forecasting by grounding state representations in visual support.
Command-inconsistent memory can lead to a 15.6% increase in collision rates during autonomous driving, but MomADv2 effectively filters this noise for safer long-horizon planning.
Achieving a 99.8% success rate in reconnecting initially disconnected convex regions transforms the feasibility of GCS-based motion planning.
RuleMaze reveals that separating perception, execution, and rule verification can dramatically enhance MLLMs' ability to follow complex natural-language instructions in spatial planning tasks.
DECOWAM achieves a 21.7% reduction in action prediction error while maintaining robust task performance, showcasing the power of embodiment-aware factorization in mobile manipulation.
GS-Voxel revolutionizes 3D scene generation by allowing scalable, fitting-free structured latents that adapt to the complexity of the scene without the need for per-scene optimization.
The empirical analysis reveals that larger models may excel in generating narratives but often fail to maintain coherence and depth, exposing a critical trade-off in LLM storytelling capabilities.
Visual perturbations can significantly alter predictions in world models, but ACPC offers a quantifiable way to diagnose and mitigate these effects.
Bypassing RGB entirely, Latent-to-4D achieves significant improvements in 4D scene generation while maintaining a reusable framework across different video models.
Robots can now predict not just immediate actions but also the next stages of complex tasks, leading to more efficient manipulation strategies.
ChemWorld allows researchers to isolate the impact of hidden chemical laws on agent behavior, enabling unprecedented control and replayability in autonomous chemistry experiments.
JEPA-WAM achieves a remarkable 79.2% on LIBERO-Plus without large-scale pretraining, setting a new benchmark for efficient robot control.
Achieving a 14.6% improvement in planning accuracy while slashing communication costs by over half, DH-VLM redefines the potential for cooperative autonomous driving.
Simulation traces can transform LLMs into powerful tools for diagnosing and improving complex scheduling policies, achieving unprecedented performance gains.
FactorDrive redefines autonomous driving by seamlessly integrating spatial-physical evidence into adaptive reasoning, achieving state-of-the-art planning performance.
Achieving state-of-the-art performance in both 4D reconstruction and point tracking, Uni4R leverages continuous velocity fields to model dynamics at any timestamp, breaking free from traditional limitations.
Revisiting the same locations across time in Sekai2 enables the learning of persistent scene representations, a game-changer for interactive world modeling.
SLIM achieves state-of-the-art performance in robot manipulation with just 0.5B parameters, outperforming larger models while slashing GPU memory usage and inference latency.
EffectLearner achieves unprecedented video object removal quality by effectively reasoning about complex object-induced effects in dynamic scenes.
Early trajectory decoding from video diffusion models can cut planning latency by nearly 50% without sacrificing decision quality.
Ignoring future cross-variable dependencies can lead to significant forecasting errors, but a new structural regularizer, CvLoss, bridges this gap and boosts model accuracy.
TrajDebug uncovers the root causes of failures in long-horizon agent trajectories, enabling targeted improvements that could significantly boost agent performance.
GeniWorld achieves robust zero-shot generalization in robotic manipulation, outperforming traditional models even with minimal training data.
Trajectory scoring in aerial navigation can be revolutionized by focusing on unexplainable prediction discrepancies, leading to more robust and efficient UAV navigation.
Training LLM agents without expert supervision can lead to better performance and generalization across diverse environments.
Argus achieves a 78% success rate on long-horizon reasoning tasks while using 21% fewer tokens in mature workflows, showcasing a revolutionary approach to agentic autonomy.
Tactile supervision can boost robot manipulation success rates by over 37% compared to traditional visual-only models.
ODEWorld achieves high-quality long-horizon predictions and planning-oriented dynamics by seamlessly integrating continuous-time modeling with ODE-based representations.
Current evaluation practices for agentic AI in medicine are misaligned with clinical needs, risking the reliability of these systems in real-world applications.
Achieving up to 100% success on complex tasks without the need for search, INTACT redefines how we approach intent-to-action learning in dynamic environments.
Shifting the focus from photorealistic rendering to dynamic visual changes, DC-WAM enhances robot policy performance by 20% in challenging environments.
FCPAgent redefines how web agents validate their actions, achieving a 13.8% boost in success rates on complex tasks by integrating falsifiable commitments into planning.
Sketching at just 25% completion can yield better spatial rationality than fully specified baselines in 3D scene generation.
Lumera achieves state-of-the-art performance in 3D scene reconstruction, revealing critical gaps in light localization and cross-engine generalization.
Adaptive routing of perception priors allows PerceptDrive to generate optimal driving trajectories in real-time without complex post-processing.
The Cram茅r-geometric Bellman operator reveals a unique fixed point that could transform how we approach evaluation errors in distributional reinforcement learning.
DRIFT achieves near real-time trajectory planning with 89.6 PDMS and 90.4 EPDMS by efficiently aggregating multiple driving behavior proposals without requiring extensive quality labels.
Action-only decoding in GigaWorld-Policy-0.5 slashes inference latency to 85 ms, revolutionizing real-time robot control efficiency.
ABot-3DWorld 0 achieves state-of-the-art scene fidelity in 3D content creation, even outperforming established models like Marble with rich multimodal inputs.
PanoWorld achieves unprecedented performance in panoramic generation by effectively integrating long-range memory with rotation-equivariant representations.
Extracting interaction cues from a frozen video model enables robots to achieve up to 90.6% success in manipulation tasks without costly rollout processes.
HALO-WA boosts robotic manipulation success rates from 26.4% to 87.1% by effectively adapting to real-world errors in just over an hour of training.
State-of-the-art planners falter in long-tail scenarios, revealing critical gaps in autonomous driving safety and effectiveness.
Bridge-WA achieves superior task performance by predicting where and how the world will change, enabling robots to focus on relevant scene dynamics rather than irrelevant visual details.
WorldSample achieves a 28% boost in policy success rates while slashing training steps by nearly 60% through innovative real-synthetic data integration.
A single BDDL specification can drastically enhance the efficiency and effectiveness of embodied task planners, achieving a 25.9% performance improvement over existing baselines.
Current interactive world models fall short, with none passing the rigorous tests of WorldRoamBench designed to assess long-horizon stability across action, vision, physics, and memory.
PhysEditWorld reveals that explicit control over physical parameters can transform how game world models interact with their environments, leading to more realistic and manipulable simulations.
Teacher-forcing consistency models can accelerate autoregressive video generation by ten times, revolutionizing the training landscape for streaming applications.
ASSCG cuts inference latency by 60% while boosting performance scores in autonomous driving systems, redefining how LLMs can be efficiently integrated into fast-slow planning architectures.