Search papers, labs, and topics across Lattice.
Object-Uni transforms how we understand and generate spatial representations of objects, enabling precise manipulation of their poses in generated images.
Emotion-aware art generation just got a major upgrade with ReART, achieving perfect alignment with target emotions through structured visual retrieval.
By combining images with text queries, ID-VTG significantly improves the accuracy of video grounding in scenarios with visually similar entities.
MSEditor achieves unprecedented consistency in multi-shot video editing, outperforming existing methods by effectively managing identity drift and temporal coherence.
Separating spatial and temporal modeling in action recognition leads to significant performance gains, challenging the effectiveness of existing implicit coupling methods.
GeoWeaver achieves unprecedented accuracy in long-sequence 3D reconstruction by correcting accumulated errors through a novel combination of geometric priors and adaptive refinement techniques.
Achieving the best macro-average results on dynamic scene reconstruction benchmarks, UniQuery4R redefines efficiency in 4D scene understanding.
GS-Voxel revolutionizes 3D scene generation by allowing scalable, fitting-free structured latents that adapt to the complexity of the scene without the need for per-scene optimization.
StreamOPD achieves near teacher-level performance in streaming video understanding without relying on memory or retrieval, reshaping the landscape of post-training techniques.
Achieving 66.9 MOTA in 3D tracking without LiDAR reveals a new frontier in roadside infrastructure understanding.
Compositional zero-shot singing-and-dancing generation is now possible, allowing for unprecedented flexibility in video creation from audio and visual prompts.
Vision-based tactile sensors could revolutionize robotic interaction by providing high-resolution tactile data that enhances perception and manipulation capabilities.
Detecting partially forged videos is now feasible with a novel framework that leverages static images for enhanced supervision and accuracy.
HumanScore reveals that traditional kinematic metrics overlook critical failures in humanoid motion tracking, such as unstable support and incorrect contacts.
Verifiable temporal grounding in video forensics can drastically improve the detection of AI-generated content, outperforming traditional methods reliant on coarse supervision.
Pruning 50% of channels in RGB-infrared object detectors can actually boost performance by 0.6% mAP, challenging conventional wisdom about redundancy.
Bypassing RGB entirely, Latent-to-4D achieves significant improvements in 4D scene generation while maintaining a reusable framework across different video models.
Cross-frame feedback boosts Transformer tracking performance by leveraging historical information, outperforming traditional same-frame methods by up to 3.2 AO points.
Achieving state-of-the-art performance in both 4D reconstruction and point tracking, Uni4R leverages continuous velocity fields to model dynamics at any timestamp, breaking free from traditional limitations.
Unpaired MRI translation can effectively preserve anatomical structures across varying magnetic field strengths, a significant leap in addressing domain shift challenges in medical imaging.
GameAlpha-2.4K not only revolutionizes RGBA video generation for gaming but also achieves a significant efficiency boost by intelligently bypassing redundant computations.
A multi-agent forensic reasoning framework outperforms leading closed-source models in deepfake detection by leveraging diverse analytical perspectives on forgery cues.
EffectLearner achieves unprecedented video object removal quality by effectively reasoning about complex object-induced effects in dynamic scenes.
Early trajectory decoding from video diffusion models can cut planning latency by nearly 50% without sacrificing decision quality.
TTA can boost accuracy but often at the cost of calibration, and ZAEC is the key to restoring reliable confidence without labeled data.
Trajectory scoring in aerial navigation can be revolutionized by focusing on unexplainable prediction discrepancies, leading to more robust and efficient UAV navigation.
Leveraging temporal differences can dramatically enhance video-to-audio generation quality, outperforming even dedicated multimodal representations.
ContextMaster achieves unprecedented consistency in multi-shot video creation, outperforming specialized models while processing at 16 FPS on a single GPU.
Image-level deepfake detectors can outperform traditional video-level detectors, with one achieving a remarkable 93.80% AUC when adapted for video analysis.
Fourier decomposition of motion enables unprecedented accuracy in rendering dynamic scenes, tackling the limitations of traditional polynomial models.
Tailoring deepfake detection to individual facial characteristics boosts accuracy and adaptability beyond traditional fixed architectures.
Synthesizing realistic 3D hand-object interactions from just a single image and text instruction could revolutionize AR/VR applications by enabling seamless integration of dynamic human behaviors.
CMIG-Net achieves up to a 0.619 dB gain in PSNR over existing methods by effectively leveraging conditional mutual information for low-light image enhancement.
TARS achieves robust camera and viewpoint control in video re-shooting without relying on 3D reconstruction or paired data, enabling the plausible synthesis of previously unseen regions.
Temporal reconstruction errors can be harnessed to significantly boost video quality, eliminating the need for expensive human annotations in preference optimization.
Achieving better video distillation quality isn't just about precision; it's about ensuring broad mode coverage during training.
LaP-Forensics reveals that leveraging reconstruction-based evidence can significantly enhance deepfake detection accuracy against state-of-the-art generative models.
BeyondFusion achieves high-quality infrared-visible fusion without the need for calibration, leveraging self-alignment to overcome sensor misalignment challenges.
CADER redefines long-video reasoning by enabling systems to adaptively allocate resources based on confidence, significantly improving efficiency and accuracy.
CameraAnything enables filmmakers to reshoot videos with arbitrary camera angles and focal lengths in a single generation, revolutionizing video editing workflows.
High-curvature regions in point clouds can be effectively represented by a novel hyperbolic rectification method that boosts discriminative power and captures fine geometric details.
Local interaction control in self-supervised denoising can dramatically enhance performance, revealing hidden regional dynamics that traditional methods overlook.
RODR effectively disentangles optimization objectives to preserve geometric fidelity in point cloud denoising, overcoming a critical challenge in manifold representation.
MoNO achieves unprecedented per-prompt diversity in diffusion sampling while eliminating the need for auxiliary quality-control objectives.
Achieving a 17.5x speedup in capacitance extraction while maintaining high accuracy could revolutionize electronic design automation workflows.
Lumera achieves state-of-the-art performance in 3D scene reconstruction, revealing critical gaps in light localization and cross-engine generalization.
Vera achieves unprecedented identity consistency in human-centric video generation, drastically reducing identity confusion in multi-person scenarios.
Achieving high-fidelity face reconstruction in under 3 seconds, UVFaceFusion redefines the balance between speed and accuracy in digital avatar creation.
Transitioning from 2D to 3D modeling reveals that fine details in monocular geometry can be captured with unprecedented fidelity.
A training-free anomaly detection framework that tracks object transformations in industrial videos outperforms existing methods while providing interpretable insights.