Search papers, labs, and topics across Lattice.
MotionCraft introduces a novel video super-resolution (VSR) framework that leverages motion-aware latent state prediction and adaptive sparse attention to enhance the quality of high-resolution video reconstructions from low-resolution inputs. This approach addresses the limitations of existing methods by balancing local detail fidelity with long-range spatio-temporal modeling, allowing for effective handling of complex motion and degradation scenarios. Empirical results demonstrate that MotionCraft not only achieves superior reconstruction and perceptual quality but also provides users with controllable trade-offs between temporal smoothness and fidelity, making it suitable for real-time applications.
MotionCraft achieves high-quality video super-resolution with predictable control over temporal smoothness and fidelity, outperforming traditional methods in complex motion scenarios.
Video super-resolution (VSR) aims to recover high-fidelity high-resolution videos from low-resolution inputs and is central to applications ranging from mobile capture to streaming and archival restoration. Existing approaches trade off among local-detail fidelity, long-range spatio-temporal modeling, perceptual realism, and efficiency: convolutional alignment techniques preserve local structure but suffer when motion is large or degradations are complex; transformer-based methods capture long-range dependencies yet require architectural or algorithmic adaptations to remain computationally feasible; and recent latent or diffusion-based generators synthesize rich texture but require specialized temporal constraints to maintain coherence. We present MotionCraft, a controllable VSR framework that formulates restoration as motion-aware latent state prediction inspired by world models and integrates adaptive sparse attention with an explicit user-accessible control interface. MotionCraft combines robust motion fusion, a Latent World Transformer that balances locality and targeted non-local interactions, and a compact conditional decoder to deliver temporally consistent, high-quality reconstructions under streaming constraints. Empirical evaluations show that MotionCraft achieves strong reconstruction and perceptual performance while enabling predictable trade-offs between temporal smoothness and reconstruction fidelity.