Search papers, labs, and topics across Lattice.
This paper introduces a family-conditioned, multi-tier audio tagging system that leverages a LoRA-finetuned Whisper encoder and a lightweight Transformer to enhance infant-centered audio understanding in challenging naturalistic settings. By implementing a sequence-level smoothing loss and a factorized speaker-token design, the model effectively reduces family bias and improves temporal coherence in long-context audio recordings. The result is a robust system capable of accurately tagging daylong audio recordings from diverse households, addressing the limitations posed by low signal-to-noise ratios and limited labeled data.
Family-specific audio tagging can now achieve unprecedented accuracy in noisy, naturalistic environments, bridging the gap in infant-centered audio understanding.
Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challenging due to limited labeled data, low signal-to-noise ratio, and cross-family domain shifts. We present a family-conditioned, multi-tier audio tagger that combines a LoRA-finetuned Whisper encoder with a lightweight, target-speaker-aware Transformer for long-context inference and framewise prediction across tiers. To improve temporal coherence, we incorporate a simple sequence-level smoothing loss, and to enhance robustness across households, we introduce a factorized speaker-token design with a shared tier token and a learned family-specific offset, reducing family bias and promoting generalizable representations. Together, these choices enable efficient and effective infant-centered audio tagging of daylong audio recordings in home environments.