Search papers, labs, and topics across Lattice.
This paper addresses the limitations of current text-to-audio models in following complex instructions involving multiple sound events and their temporal order. By leveraging audio-aware large language models (ALLMs) as fine-grained evaluators, the authors create a framework that enhances instruction-level correctness through preference optimization based on ALLM feedback. The results demonstrate significant improvements in event completeness, temporal accuracy, and overall instruction-following performance without compromising audio quality.
Instruction-level feedback from audio-aware LLMs can drastically enhance the accuracy of multi-event audio generation, bridging a critical gap in current models.
Recent text-to-audio models generate high-quality audio, but often fail to follow instructions involving multiple sound events and temporal order. This gap arises because existing evaluation and training signals mainly emphasize global similarity or perceptual quality, with limited supervision on instruction-level correctness. We propose an instruction-level framework that uses audio-aware large language models (ALLMs) as fine-grained judges to verify target event presence and temporal relations in generated audio. After validating ALLM judgments on benchmarks and through human verification, we use their feedback to construct preference pairs for direct preference optimization. We further introduce S3Bench, a narrative benchmark for evaluating multi-event temporal instruction following. Experiments show that our method improves event completeness, temporal ordering, and joint instruction-following accuracy across existing benchmarks and S3Bench, while maintaining audio quality.