Search papers, labs, and topics across Lattice.
This paper introduces AudioLens, a novel approach to multi-perspective audio clustering that leverages reasoning audio-language models to organize speech collections based on user-defined linguistic and paralinguistic cues. By developing the AudioLens-Bench benchmark, the authors evaluate both in-perspective and cross-perspective generalization, showcasing the model's ability to infer the number of clusters and their assignments directly from natural language perspectives. Experimental results indicate that AudioLens-R1 significantly outperforms existing methods, achieving a 12.99-point improvement in Adjusted Rand Index (ARI) and an 11.62-point increase in V-measure, highlighting the effectiveness of audio-language models in flexible structure discovery.
AudioLens-R1 redefines audio clustering by allowing models to adaptively organize speech based on user-specified perspectives, achieving unprecedented accuracy improvements.
Audio clustering is a fundamental task for organizing rapidly growing speech collections, supporting applications such as conversational analysis and speech-driven discovery. However, existing methods rely on fixed acoustic similarity metrics or ASR-based text pipelines, limiting their ability to reorganize the same audio collection under different user-specified perspectives, especially when clustering depends on both linguistic and paralinguistic cues. We introduce audio multi-perspective clustering, where a model directly partitions speech recordings according to a natural-language perspective while inferring both the number of clusters and their assignments. To study this setting, we construct AudioLens-Bench, a benchmark spanning multiple application domains and evaluating both in-perspective and cross-perspective generalization. We further propose AudioLens-R1, an end-to-end large audio-language model trained with reasoning distillation and preference optimization. Experiments show that AudioLens-R1 consistently outperforms all baselines, improving overall ARI by 12.99 points and V-measure by 11.62 points. These results demonstrate the promise of native audio-language models for flexible, perspective-conditioned structure discovery over speech collections.