Search papers, labs, and topics across Lattice.
15
0
10
10
ContextMaster achieves unprecedented consistency in multi-shot video creation, outperforming specialized models while processing at 16 FPS on a single GPU.
Existing models mismanage tool use, but Beacon achieves a balance that enhances performance on complex tasks while preserving accuracy on simpler ones.
Temporal reconstruction errors can be harnessed to significantly boost video quality, eliminating the need for expensive human annotations in preference optimization.
AMFD not only outperforms traditional moment matching methods but also enables unprecedented gains in instruction-following for text-to-image generation.
StreamHOI achieves real-time HOI video generation with impressive interaction fidelity, hitting 17.6 FPS and just 0.75 seconds latency.
Captions generated by PercepCap are not only more accurate but also grounded in explicit spatio-temporal perception, revealing the underlying reasoning behind each description.
Embedding reference tokens at semantic positions allows for unprecedented precision in multi-reference video editing, setting a new benchmark for instruction quality.
Stop reinventing the wheel: OpenWorldLib offers a unified framework and codebase for advanced world models, finally bringing standardization to a fragmented field.
Achieve robust, real-time 3D multi-object tracking in panoramic views by representing object states on a sphere, sidestepping the limitations of image-plane trackers and redundant Euclidean formulations.
World models can now remember and realistically regenerate dynamic objects that temporarily disappear from view, thanks to a novel hybrid memory architecture.
Generate multi-shot videos at 16 FPS with a single GPU and interactively steer the narrative in real-time, thanks to a novel causal architecture that overcomes the limitations of bidirectional models.
Achieve lifelike character animation with 10x faster inference using Kling-MotionControl, a DiT-based framework that intelligently handles body, face, and hand motions.
Forget single-number video quality scores: UltraVQA and Analytic Score Optimization (ASO) unlock richer, multi-faceted evaluations that better align with human preferences.
Forget generic CoT: Embed-RL uses reinforcement learning to generate reasoning traces that are explicitly optimized for multimodal embedding tasks, leading to significant performance gains.
Forget separate image and video models: VINO's single diffusion backbone handles both, opening the door to truly unified visual creation and editing.