Search papers, labs, and topics across Lattice.
This paper introduces a novel cascaded framework, Proxy Avatar Meets Low-Rank Caching, for real-time one-shot emotion-controllable portrait animation that leverages a Gaussian-based emotion proxy avatar to generate expressive motion from audio and emotion labels. By decoupling motion generation from appearance, the method utilizes a large-scale retargeting model to adapt identity-independent motion to various target portraits, significantly enhancing emotional expressiveness while preserving identity. The approach also incorporates low-rank caching to optimize inference efficiency, resulting in a substantial reduction in computational costs and enabling real-time performance.
Emotion-controllable portrait animation can now be achieved in real-time with a single training instance, revolutionizing the efficiency and expressiveness of animated avatars.
Audio-driven portrait animation has advanced rapidly with diffusion-based generative models, yet real-time one-shot generation with expressive emotion control remains challenging. Existing methods often suffer from insufficient emotion-aware motion priors and expensive appearance computation during multi-step denoising. To address these issues, we propose Proxy Avatar Meets Low-Rank Caching, a cascaded framework for real-time one-shot emotion-controllable portrait animation. Instead of directly generating the target portrait from audio, our method uses a Gaussian-based emotion proxy avatar as a reusable motion generator, which is trained once on a single identity to produce expressive driving videos from audio and emotion labels. Since the proxy avatar only provides motion rather than target appearance or geometry, a large-scale one-shot retargeting model further extracts identity-independent motion from the proxy performance and adapts it to arbitrary target portraits. To improve inference efficiency, we introduce zero-shot appearance reuse with low-rank caching, which caches reference appearance features at the initial denoising step and models subsequent feature variations using lightweight low-rank adapters. Extensive experiments demonstrate that our method achieves stronger emotional expressiveness, better identity-preserving animation, and substantially reduced inference cost, enabling real-time one-shot portrait animation.