Search papers, labs, and topics across Lattice.
This study investigates the effects of synthetic chord corruption versus complete automatic chord recognition (ACR) replay on singing accompaniment generation (SAG) using a fixed MIDI-SAG generator. The authors found that localized chord corruption led to significantly larger changes in output compared to ACR replay, with a notable positive target gap in 29 out of 30 tracks. By employing joint matching and relation composition, they reduced mismatches in replay, revealing that local corruption tests underlying mechanisms while ACR replay assesses the propagation of chord conditions in deployed systems.
Localized chord corruption can produce over 2.88 times the output change compared to traditional ACR replay, challenging assumptions about chord recognition in music generation systems.
Synthetic chord corruption provides a controlled stress test for singing accompaniment generation (SAG), whereas complete automatic chord recognition (ACR) replay measures the condition delivered to a deployed system. We compare them by replaying CNN-CRF and DeepChroma+CRF predictions through one fixed MIDI-SAG generator, holding track, seed, context, and scoring window constant. Across 30 paired tracks and three seeds, a central four-second tritone produced larger changed-target and inside-output effects than CNN-CRF replay in STFT, CQT, and CENS; the CENS target gap was positive on 29/30 tracks (mean 0.462). Matching replay support and relation composition reduced this mismatch, with joint matching giving the lowest replay distance in the full-30 analysis. Relative-root substitutions produced 2.88-fold larger CENS full-window output change than same-root quality flips at near-equal dose. The matched surrogate was closer to replay for both recognizer paths. We conclude that localized corruption tests mechanisms, whereas complete replay evaluates deployed chord-condition propagation.