Feb 26, 2026arXiv:2602.23191

Uni-Animator: Towards Unified Visual Colorization

Xinyuan Chen, Xinyuan Chen, Yao Xu, Yao Xu, Shaowen Wang, Pengjie Song, Pengjie Song, Bowen Deng, Bowen Deng

AI Summary

The paper introduces Uni-Animator, a Diffusion Transformer (DiT) framework designed to unify image and video sketch colorization tasks. It addresses limitations in existing methods by incorporating visual reference enhancement via instance patch embedding for precise color transfer, physical detail reinforcement using physical features for preserving high-frequency textures, and sketch-based dynamic RoPE encoding to mitigate motion-induced temporal inconsistency. Experiments demonstrate that Uni-Animator achieves competitive performance on both image and video sketch colorization, matching task-specific methods while enabling unified cross-domain capabilities with high detail fidelity and robust temporal consistency.

Key Contribution

A single model now handles both image and video sketch colorization with high fidelity and temporal consistency, outperforming specialized methods.

Abstract

We propose Uni-Animator, a novel Diffusion Transformer (DiT)-based framework for unified image and video sketch colorization. Existing sketch colorization methods struggle to unify image and video tasks, suffering from imprecise color transfer with single or multiple references, inadequate preservation of high-frequency physical details, and compromised temporal coherence with motion artifacts in large-motion scenes. To tackle imprecise color transfer, we introduce visual reference enhancement via instance patch embedding, enabling precise alignment and fusion of reference color information. To resolve insufficient physical detail preservation, we design physical detail reinforcement using physical features that effectively capture and retain high-frequency textures. To mitigate motion-induced temporal inconsistency, we propose sketch-based dynamic RoPE encoding that adaptively models motion-aware spatial-temporal dependencies. Extensive experimental results demonstrate that Uni-Animator achieves competitive performance on both image and video sketch colorization, matching that of task-specific methods while unlocking unified cross-domain capabilities with high detail fidelity and robust temporal consistency.

Architecture Design (Transformers, SSMs, MoE)Computer Vision Multimodal Models

Citation Metrics

Citations0

Influential citations0

References60

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Uni-Animator: Towards Unified Visual Colorization

Related Papers