Search papers, labs, and topics across Lattice.
The VAANI Noise Event Timestamp Dataset introduces a novel layer of annotations for spontaneous speech recordings, capturing real-world audio across 165 Indian districts in 105 languages. This dataset provides fine-grained, timestamped labels for overlapping noise events, categorized into seven classes, addressing a critical gap in existing sound-event corpora that typically focus on either clean speech or general audio tagging. By enabling tasks such as noise-robust Automatic Speech Recognition (ASR) and sound event detection (SED), VAANI enhances the ability to develop systems that perform well in complex auditory environments.
Real-world speech recordings with overlapping noise events are now timestamped and categorized, filling a crucial gap in audio datasets.
Most public sound-event corpora are optimized either for general audio tagging or for clean speech separation, and comparatively few provide strong timestamped noise annotations layered directly on top of spontaneous, real-world speech. We present the VAANI Noise Event Timestamp Dataset, a derived annotation layer built on Project VAANI field recordings of spontaneous speech collected across 165 Indian districts in 105 languages. Unlike synthetically mixed corpora, VAANI captures speech and ambient noise in situ and simultaneously, and annotates each recording with fine-grained start/end timestamps for overlapping background noise events organized into a compact seven-class semantic taxonomy: animal, traffic, baby/child, music, signal/alarm, appliance, and non-speech human. This combination of spontaneous multilingual Indic speech, authentic regional soundscapes, and span-level noise tags that may overlap with speech targets tasks that existing datasets address only partially: noise-robust Automatic Speech Recognition (ASR), sound event detection (SED), and speech enhancement. We position VAANI against nine widely used corpora and benchmarks, including WHAM!, AVA-Speech, MUSAN, FSD50K, CHiME-6, AudioSet, DESED, the India-specific iNoise noise database, and the Kathbath-Noisy noisy-ASR benchmarks, and describe the annotation protocol and quality-control procedure used to produce the timestamped tags.