Search papers, labs, and topics across Lattice.
This study evaluates the controllability of text-to-music models by employing a matched counterfactual evaluation method that distinguishes between naturally occurring outputs and those attributable to specific instructions. The findings reveal that while models like ACE-Step 1.5 and Stable Audio 3 Medium demonstrate significant control over musical key, their ability to follow beat grouping instructions is less reliable, with high agreement often stemming from pre-existing output distributions. This nuanced approach challenges the assumption that prompted attribute agreement is a definitive measure of model controllability, highlighting the need for more rigorous evaluation methods in assessing generative models.
Text-to-music models may appear controllable, but a closer look reveals that much of their output is simply a reflection of their training data rather than genuine instruction-following.
Prompted attribute agreement is widely used as evidence of text-to-music controllability, yet a requested attribute may occur simply because it is already common in the model's output distribution. We introduce a matched counterfactual evaluation that separates target occurrence from instruction-attributable control. Each family contains a neutral input that omits the scored attribute and two otherwise matched inputs that swap the requested target. All three are rendered through frozen native-interface adapters with a shared seed. Applied to global key and beat grouping in three open systems, this design changes the empirical conclusion. ACE-Step 1.5 and Stable Audio 3 Medium exhibit substantial key control, whereas LeVo2 does not. For beat grouping, the same models redirect toward the rare three-beat target, but high four-beat agreement is largely inherited from neutral outputs: Stable Audio 3 produces four-beat grouping in 0.97 of neutral cases but only 0.56 under its explicit four-beat treatment. Off-attribute placebos, external recognizer validation, blind expert annotation, and multi-seed sentinels support the attribution. When targets have unequal output priors, agreement describes what a model produced, while matched neutral and target-swap contrasts test whether the instruction changed it.