Search papers, labs, and topics across Lattice.
This paper introduces Vera, a novel framework for human subject-to-video generation that addresses the critical issue of identity consistency across frames, particularly in multi-person scenarios. By constructing a million-pair identity-aligned dataset and employing two innovative designs鈥擨dentity-Focal Masked Supervision (IFMS) and Reference-Aware Layer-wise Attention (RALA)鈥擵era enhances identity-aware learning and stabilizes identity anchors in generated videos. Experimental results show significant improvements in human identity consistency and motion naturalness, while minimizing identity confusion and excessive copying of reference attributes.
Vera achieves unprecedented identity consistency in human-centric video generation, drastically reducing identity confusion in multi-person scenarios.
Subject-to-video (S2V) generation has made substantial progress in preserving reference subjects across diverse categories, yet generic subject consistency remains insufficient for human-centric generation. A video may appear globally consistent while identity-critical human details still drift across frames, poses, and interactions. This issue becomes more severe in multi-person scenarios, where incorrect identity-role binding leads to subject confusion, attribute swapping, and excessive copying of reference-specific appearance cues. We propose Vera, a unified human-centric S2V framework for single- and multi-person generation. We first construct a million-pair identity-aligned human image-video dataset through person-level cross-clip retrieval, providing explicit identity correspondence and diverse references. Built on this dataset, Vera introduces two complementary designs. Identity-Focal Masked Supervision (IFMS) strengthens identity-aware learning with spatially focused supervision while reducing interference from irrelevant artifacts. Reference-Aware Layer-wise Attention (RALA) regulates how video tokens interact with reference identity cues in the DiT backbone, preserving stable identity anchors and enhancing layer-aware identity readout. Extensive experiments demonstrate that Vera improves human identity consistency, multi-person subject binding, and motion naturalness, while reducing identity confusion and excessive reference-image copying.