Search papers, labs, and topics across Lattice.
11
0
9
10
Realistic predictions in robotic manipulation now come with precise arm control, ensuring the right actions yield the expected outcomes.
Multi-dimensional Evaluation-Verification Reward transforms multi-reference image editing by providing a structured approach to evaluate and enhance visual consistency, yielding superior results over existing models.
Achieving high visual fidelity while ensuring physical consistency in video generation could redefine the standards for simulating realistic interactions in AI-generated content.
Aesthetic assessments can be dramatically improved by focusing on key moments and endings, as shown by Peak-End-Net's state-of-the-art results in video evaluation.
Memory overload in autoregressive video generation can be tackled by absorbing historical context into model weights, achieving up to 50% cache reduction with minimal quality loss.
HEE transforms static image understanding into a dynamic, query-guided exploration that significantly boosts accuracy in high-resolution perception tasks.
Optimal block size selection can lead to a 4.20x speedup in diffusion-based speculative decoding, revolutionizing inference efficiency.
CIPE-Dance is a game-changer, providing the largest dataset for dance video generation and enabling OmniDance to set new benchmarks in multimodal video synthesis.
Achieving an overall score of 84.76, DreamX-World 1.0 sets a new benchmark for interactive video generation, outperforming leading models in both camera control and visual fidelity.
Forget painstakingly curating 3D-consistent datasets: this RL approach uses a 3D foundation model to *verify* consistency, turning 3D scene editing into a tractable reinforcement learning problem.
ImagineAgent's clever combination of cognitive maps and generative tools lets it crush previous state-of-the-art on OV-HOI tasks while needing only 20% of the training data.