Search papers, labs, and topics across Lattice.
This paper introduces BackgroundMellow, a novel framework for generating cohesive and cinematic soundscapes from long-form narratives, addressing the limitations of existing Text-to-Audio systems in achieving narrative alignment and emotional depth. By employing a master-specialist agent architecture, the framework decomposes text into multi-layered audio cues and utilizes a Tango2 latent diffusion model alongside a Cinematic BGM Retriever to synthesize and mix audio without requiring ground-truth data. Empirical evaluations demonstrate that BackgroundMellow significantly improves temporal synchronization, coverage, and spectral richness compared to traditional methods, marking a substantial advancement in multi-modal audio generation.
BackgroundMellow achieves unprecedented narrative cohesion in audio generation, transforming how we create immersive soundscapes for storytelling.
Generating immersive, synchronized and cinematic audio for long-form textual narratives remains a significant challenge in multi-modal AI. While current Text-to-Audio (TTA) frameworks successfully synthesize isolated sound effects, they struggle with narrative cohesion, temporal alignment, and cinematic emotional depth. We present BackgroundMellow, a framework that treats story-to-audio generation as a precise orchestration and signal processing problem. This framework is enabled without ground-truth through a master-specialist agent architecture that decomposes text into precise and multi-layered audio cues, generates each category of sounds with suitable specialist model, and superimposes the soundscapes to create a unified and aligned audio segment. Our pipeline is built over Tango2 latent diffusion model for environmental synthesis alongside a novel Cinematic BGM Retriever mined from professional soundtracks. To automate the sound mixing process, we use an NLP based module that predicts precise audio parameters, like start time, duration, and relative loudness, based on the narrative timeline. We further empirically evaluate and show the efficacy of the proposed framework leveraging nearest-neighbor retrieval against a curated dataset of YouTube cinematic trailers to measure temporal synchronization, coverage, and spectral richness.