Search papers, labs, and topics across Lattice.
A staggering 42.5% of correct answers in long-document VQA can't be derived from terminal working memory alone, revealing a crucial oversight in current evaluation methods.
Reducing visual token exposure by leveraging selective retrieval can dramatically enhance the performance of 3D medical image question answering systems.
Achieving a 25% reduction in prediction error, MOSH-WM redefines the landscape of object-centric video forecasting by grounding state representations in visual support.
Typed spatial readouts can elevate robotic execution success rates by over 50% while significantly reducing processing time.
Object-Uni transforms how we understand and generate spatial representations of objects, enabling precise manipulation of their poses in generated images.
"Prompt-form collapse" can drastically reduce success rates, but TOWN-VLA's controlled intervention boosts performance by ensuring only meaningful prompts are executed.
Emotion-aware art generation just got a major upgrade with ReART, achieving perfect alignment with target emotions through structured visual retrieval.
RuleMaze reveals that separating perception, execution, and rule verification can dramatically enhance MLLMs' ability to follow complex natural-language instructions in spatial planning tasks.
DECOWAM achieves a 21.7% reduction in action prediction error while maintaining robust task performance, showcasing the power of embodiment-aware factorization in mobile manipulation.
By combining images with text queries, ID-VTG significantly improves the accuracy of video grounding in scenarios with visually similar entities.
MedUAG sets a new standard in medical multimodal models, achieving strong performance across diverse understanding and generation tasks with the largest dataset yet.
Policies may succeed in tasks but still violate deformation tolerances, revealing a critical gap in current evaluation methods for deformable-object manipulation.
MSEditor achieves unprecedented consistency in multi-shot video editing, outperforming existing methods by effectively managing identity drift and temporal coherence.
GeoWeaver achieves unprecedented accuracy in long-sequence 3D reconstruction by correcting accumulated errors through a novel combination of geometric priors and adaptive refinement techniques.
Achieving the best macro-average results on dynamic scene reconstruction benchmarks, UniQuery4R redefines efficiency in 4D scene understanding.
Selectively routed stereo evidence boosts humanoid VLA control success rates, achieving 100% grasp success even under severe occlusion.
A non-invasive framework that enhances pretrained generative policies with force control, achieving remarkable improvements in contact stability and execution speed.
StreamOPD achieves near teacher-level performance in streaming video understanding without relying on memory or retrieval, reshaping the landscape of post-training techniques.
Domain adaptation can dramatically enhance sign language recognition, achieving superior results over conventional transfer learning methods.
Achieving 66.9 MOTA in 3D tracking without LiDAR reveals a new frontier in roadside infrastructure understanding.
SparkVLA redefines task execution in hierarchical VLA systems, achieving over 30% improvement in success rates by integrating stop and action length decisions into a single ranking process.
Compositional zero-shot singing-and-dancing generation is now possible, allowing for unprecedented flexibility in video creation from audio and visual prompts.
CineDub achieves unprecedented accuracy in multi-speaker dialogue dubbing from uncropped videos, setting a new standard for video dubbing technology.
Vision-based tactile sensors could revolutionize robotic interaction by providing high-resolution tactile data that enhances perception and manipulation capabilities.
Detecting partially forged videos is now feasible with a novel framework that leverages static images for enhanced supervision and accuracy.
AutoDesign outperforms existing design systems by aligning with human design principles and achieving superior poster generation quality through recursive self-improvement.
SCOUT achieves a remarkable 16.85% improvement in spatial reasoning benchmarks, setting a new standard for Vision-Language Models.
Leading MLLMs falter on the new VideoGAIA benchmark, scoring under 60% accuracy in complex, multi-turn video understanding tasks.
Pruning 50% of channels in RGB-infrared object detectors can actually boost performance by 0.6% mAP, challenging conventional wisdom about redundancy.
Bypassing RGB entirely, Latent-to-4D achieves significant improvements in 4D scene generation while maintaining a reusable framework across different video models.
Robots can now predict not just immediate actions but also the next stages of complex tasks, leading to more efficient manipulation strategies.
JEPA-WAM achieves a remarkable 79.2% on LIBERO-Plus without large-scale pretraining, setting a new benchmark for efficient robot control.
Achieving a 14.6% improvement in planning accuracy while slashing communication costs by over half, DH-VLM redefines the potential for cooperative autonomous driving.
Understanding how regularization influences model learning could revolutionize our approach to designing robust AI systems.
Cross-frame feedback boosts Transformer tracking performance by leveraging historical information, outperforming traditional same-frame methods by up to 3.2 AO points.
FactorDrive redefines autonomous driving by seamlessly integrating spatial-physical evidence into adaptive reasoning, achieving state-of-the-art planning performance.
Revisiting the same locations across time in Sekai2 enables the learning of persistent scene representations, a game-changer for interactive world modeling.
Unified audio generation just got a major upgrade鈥擲onicWeave's innovative routing mechanism boosts compositional quality and expert specialization, outperforming traditional models.
SLIM achieves state-of-the-art performance in robot manipulation with just 0.5B parameters, outperforming larger models while slashing GPU memory usage and inference latency.
GameAlpha-2.4K not only revolutionizes RGBA video generation for gaming but also achieves a significant efficiency boost by intelligently bypassing redundant computations.
CED reveals that VLMs can be trained to prioritize evidence-based reasoning over language shortcuts, leading to more reliable visual understanding.
A multi-agent forensic reasoning framework outperforms leading closed-source models in deepfake detection by leveraging diverse analytical perspectives on forgery cues.
PaDoc achieves a remarkable 94.24 F1 score among end-to-end parsers while being the fastest parser at five concurrency levels, revolutionizing document parsing efficiency.
SkillMemo transforms robotic manipulation by enabling models to leverage reusable skill structures, leading to unprecedented compositional generalization in complex tasks.
Smart-home agents struggle to differentiate between real commands and misleading ambient noise, with traditional detectors and MLLMs both failing in complementary ways.
TTA can boost accuracy but often at the cost of calibration, and ZAEC is the key to restoring reliable confidence without labeled data.
Grounded language comprehension, rather than free-form reasoning, is the key to unlocking superior performance in Vision-Language-Action models.
RTCF boosts the success of frozen VLA policies by leveraging past experiences without the need for retraining or extra GPU power.
Leveraging temporal differences can dramatically enhance video-to-audio generation quality, outperforming even dedicated multimodal representations.
ContextMaster achieves unprecedented consistency in multi-shot video creation, outperforming specialized models while processing at 16 FPS on a single GPU.