Search papers, labs, and topics across Lattice.
This paper introduces a novel approach to neural video compression that enhances temporal context quality by integrating deformable temporal alignment with difference-aware spatial selective fusion. By leveraging a Context-aware Temporal Alignment Module, the method generates complementary temporal context while effectively suppressing misalignment through a Difference-aware Spatial Selective Fusion module. Experimental results demonstrate a significant improvement in rate-distortion performance compared to existing methods, particularly in challenging scenarios with complex motion and occlusion.
Achieving superior video compression performance by addressing misalignment issues in complex motion scenarios could redefine standards in neural video coding.
In conditional coding-based neural video compression, the quality of temporal context directly affects compression per- formance. Existing methods mostly construct context from prop- agated reference features, but they are vulnerable to motion esti- mation and local alignment errors in regions with complex mo- tion, occlusion, and high-frequency textures, resulting in inaccu- rate temporal information. To address this issue, this paper pro- poses a method combining deformable temporal alignment and difference-aware spatial selective fusion. A Context-aware Tem- poral Alignment Module is used to generate complementary tem- poral context, while a Difference-aware Spatial Selective Fusion module adaptively selects reliable temporal information and sup- presses misalignment. Experiments show that the proposed method achieves certain rate-distortion performance improve- ment over DCVC-DC.