Mar 1, 2026arXiv:2603.01006

AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching

AI Summary

This paper introduces Attribution-Guided REPresentation Alignment (AG-REPA), a novel method for selecting which layers to align in token-conditioned audio Flow Matching models. AG-REPA addresses the "Store-Contribute Dissociation" (SCD) phenomenon, where layers with high teacher-space similarity don't necessarily contribute most to the velocity field. By using a forward-only gate ablation (FoG-A) to quantify each layer's causal contribution to the velocity field, AG-REPA enables sparse layer selection and adaptive weighting for representation alignment, leading to improved performance on unified speech and general-audio generation tasks.

Key Contribution

Forget blindly aligning layers in audio Flow Matching: AG-REPA reveals that targeting the *causally dominant* layers driving the velocity field, not just representationally rich ones, unlocks superior generation.

Abstract

REPresentation Alignment (REPA) improves the training of generative flow models by aligning intermediate hidden states with pretrained teacher features, but its effectiveness in token-conditioned audio Flow Matching critically depends on the choice of supervised layers, which is typically made heuristically based on the depth. In this work, we introduce Attribution-Guided REPresentation Alignment (AG-REPA), a novel causal layer selection strategy for representation alignment in audio Flow Matching. Firstly, we find that layers that best store semantic/acoustic information (high teacher-space similarity) are not necessarily the layers that contribute most to the velocity field that drives generation, and we call it Store-Contribute Dissociation (SCD). To turn this insight into an actionable training guidance, we propose a forward-only gate ablation (FoG-A) that quantifies each layer's causal contribution via the induced change in the predicted velocity field, enabling sparse layer selection and adaptive weighting for alignment. Across unified speech and general-audio training (LibriSpeech + AudioSet) under different token-conditioning topologies, AG-REPA consistently outperforms REPA baselines. Overall, our results show that alignment is most effective when applied to the causally dominant layers that drive the velocity field, rather than to layers that are representationally rich but functionally passive.

Architecture Design (Transformers, SSMs, MoE)Speech & Audio Training Efficiency & Optimization

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching

Related Papers