Search papers, labs, and topics across Lattice.
This paper introduces INSPIRE, a novel benchmark designed for instruction-aware speech retrieval that allows users to specify diverse relevance criteria through natural language instructions. The evaluation of four different retrieval paradigms reveals that existing methods do not adequately address all retrieval intents, with text-based models excelling in semantic retrieval but lacking in paralinguistic attributes, while speech-based models perform better in capturing acoustic properties but struggle with instruction adherence. These insights underscore the necessity for unified architectures that can effectively integrate instruction awareness into speech retrieval systems.
No existing speech retrieval method can robustly handle the full spectrum of user intents, revealing a critical gap in current technologies.
Existing speech retrieval systems rely on fixed similarity matching and cannot adapt to diverse user intents. We introduce INSPIRE, the first benchmark for instruction-aware speech retrieval, in which natural-language instructions dynamically specify relevance criteria, including semantic content, speaker identity, speaking style, environmental sounds, and their combinations. We evaluate four retrieval paradigms: large audio-language models, cascaded pipelines, self-supervised speech models, and contrastive audio-language models. Our results reveal that no current method robustly handles all retrieval intents. Text-based approaches perform relatively better at semantic retrieval but struggle with paralinguistic attributes, while speech-based models are moderately better at capturing acoustic properties but falter at following instructions. These findings highlight the need for unified architectures capable of instruction-aware speech retrieval.