Search papers, labs, and topics across Lattice.
This paper introduces Sidecar, a semantic augmentation module designed to enhance character consistency in free-form visual storytelling without requiring additional training. By preserving and injecting essential identity-related semantics from the initial character description into subsequent prompts, Sidecar addresses the challenge of maintaining character identity across frames. Experiments demonstrate that Sidecar significantly improves prompt-image alignment and character consistency in multiple diffusion model baselines, while incurring minimal computational cost.
Sidecar boosts character consistency in visual storytelling by seamlessly infusing missing identity semantics into prompts, all without the need for retraining.
Visual storytelling requires generating images that follow a narrative while preserving consistent character identities across frames. In free-form story generation, a character is fully described only when first introduced and is later referred to by a type-level mention or pronoun. Although this setting better reflects natural storytelling, later prompts may omit important identity-related semantics, making character consistency more difficult to maintain. We propose \textbf{Sidecar}, a plug-and-play semantic augmentation module that preserves entity-level information from the initial description and injects the missing semantics into later prompt embeddings. Sidecar requires no additional training and does not modify the architecture of the base diffusion model. Experiments on FreeStoryBench show that Sidecar consistently improves prompt-image alignment and character consistency across multiple SDXL- and FLUX-based baselines, with negligible computational overhead.