Search papers, labs, and topics across Lattice.
This paper introduces EmoWorld, a novel framework for controllable emotional video generation that decouples global atmosphere, semantic cues, and temporal progression within a video diffusion transformer. By employing a one-time preparation stage to extract affect directions and a reusable cue library, EmoWorld enhances emotional alignment and reduces temporal fluctuations during video generation. The framework achieves significant improvements in target-emotion alignment and cue detection across various emotion categories, demonstrating its versatility and effectiveness in both text-to-video and image-to-video settings.
EmoWorld achieves up to 37% improvement in emotional alignment while decoupling atmosphere, semantics, and temporal progression in video generation.
Emotion shapes how viewers interpret a scene, yet existing video generators entangle global atmosphere, affect-bearing semantic cues, and temporal progression within a single text condition. We present EmoWorld, a framework that decouples these factors within a frozen flow-matching video diffusion transformer (Video DiT). A one-time preparation stage extracts layer-specific affect directions and a reusable cue library from geometry-preserving neutral and emotion-edited panoramas. At inference, Visual Atmosphere Steering (VAS) injects atmosphere directions into hidden states, Semantic Affective Steering (SAS) isolates a separately scalable prompt residual for semantic cues, and Temporal Affective Steering (TAS) interpolates endpoint residual fields across denoising and video time. On Wan2.2, VAS improves target-emotion alignment by 19% while reducing a temporal-fluctuation proxy by 48%; SAS improves target-emotion alignment by 37% and increases detected affect-bearing cues by 36%; and TAS improves transition monotonicity by 15% over the strongest baseline. EmoWorld is evaluated across 27 emotion categories in text-to-video and image-to-video settings, demonstrates portability across multiple Video-DiT backbones, and supports camera-conditioned composition without updating generator parameters.