Search papers, labs, and topics across Lattice.
This paper introduces DINO-A, a novel adaptation of the self-distillation method from vision to general audio representation learning, specifically utilizing log-mel spectrograms and a modified augmentation block. The authors pretrain multiple architectures, including Vision Transformers and convolutional encoders, on the FSD50K dataset and evaluate their performance across various audio classification tasks. Key findings reveal that smaller patch sizes in Vision Transformers consistently enhance representation quality, while the choice of backbone architecture significantly influences task performance, with DINO-A outperforming BYOL-A v2 by nearly 12 percentage points on average due to the interplay of projection space and augmentation strategies.
DINO-A reveals that smaller patch sizes in Vision Transformers consistently yield better audio representation quality, challenging assumptions about model architecture in audio tasks.
We present DINO-A, an adaptation of self-distillation from vision to general audio representation learning. While DINO has become a canonical method in self-supervised vision and prior audio work has explored latent prediction (BYOL-A) and masked modeling (Audio-MAE, BEATs), no prior work has brought canonical DINO to general audio classification in the way BYOL-A brought BYOL. DINO-A retains DINO's multi-crop, EMA teacher, and high-dimensional projection, replacing only the input modality and augmentations with log-mel spectrograms and the BYOL-A v2 augmentation block. We pretrain three backbones, two Vision Transformers with 8x8 and 16x16 patches and a convolutional encoder, on FSD50K and evaluate them with linear probing on ESC-50, Speech Commands v2, UrbanSound8K, and GTZAN. Three findings characterize the resulting representations. Patch resolution within the Vision Transformer family has consistent effect on representation quality, with smaller patches winning across all four tasks. The choice between Vision Transformer and convolutional backbone interacts with task type: convolutional networks lead on speech while Vision Transformers lead on environmental sounds and music. Under identical pretraining and evaluation conditions, DINO-A and BYOL-A v2 differ by 11.96 percentage points on average, and we trace this difference to two mechanisms: the interaction between DINO's high-dimensional projection space and FSD50K's limited scale, and the additional cost of multi-crop augmentation, which DINO uses but BYOL-A v2 does not. The high-dimensional projection space, central to DINO's success in vision, becomes a liability at FSD50K scale.