Search papers, labs, and topics across Lattice.
This paper introduces GroupVideo, a novel framework for multi-identity customized text-to-video generation that addresses identity confusion and unnatural expressions common in existing methods. By leveraging Video Diffusion Transformers and incorporating multimodal identity alignment, GroupVideo enhances the robustness of identity references and improves motion naturalness. Extensive experiments show that GroupVideo significantly outperforms prior approaches, producing videos with consistent identities and natural movements, supported by a newly curated dataset of 20,000 videos for multi-ID scenarios.
GroupVideo achieves unprecedented fidelity in multi-character video generation, resolving identity confusion and unnatural motions that plague existing methods.
Current identity customized video generation methodologies are predominantly limited to single-identity scenarios, as the lack of explicit identity separation mechanisms often leads to identity confusion in multi-identity settings. Existing multi-identity approaches, which directly extend single-identity frameworks by concatenating face images as input conditions, frequently result in unnatural facial expressions and motions, manifesting as the"copy-paste"phenomenon. To overcome these limitations, we introduce GroupVideo, a novel framework that leverages multiple individual photographs to generate identitycustomized video. Built upon Video Diffusion Transformers, GroupVideo incorporates multimodal identity alignment: visual alignment jointly encodes multiple face images to provide robust identity references, while semantic alignment introduces a semantic perceiver to enhance the naturalness of motions. An ID localization module with spatial guidance is introduced to address identity blending and enhance identity fidelity, along with bounding box constraints and mask regularization loss, to focus on facial regions and improve training efficiency. In response to the shortage of multi-ID video datasets, we have curated a comprehensive high-quality dataset of 20,000 videos, thereby establishing a crucial resource to advance future research in multi-ID video generation. Extensive experiments demonstrate that GroupVideo outperforms existing methods in generating multi-character videos with consistent identities and natural motions.