Search papers, labs, and topics across Lattice.
This paper tackles the challenge of first-shot anomalous sound detection by introducing a training-free post-hoc layer that calibrates scores based on domain-specific characteristics, effectively addressing the negative correlation between source and target domain AUC. By employing a per-domain quantile calibration and a label-free cross-validated domain-balance criterion, the authors achieve a significant improvement in evaluation scores on the DCASE 2025 challenge, demonstrating a robust predictive capability across multiple configurations. The proposed method not only enhances performance but also reveals that traditional development-set scores are uninformative, emphasizing the need for alternative selection criteria in this context.
A novel training-free approach boosts anomalous sound detection scores by over 3.5 points while revealing that conventional development-set metrics can mislead model selection.
First-shot anomalous sound detection in DCASE Challenge Task 2 must flag anomalies of unseen machine types with a single threshold, without knowing whether a test clip comes from the data-rich source domain (990 normal training clips) or the data-scarce target domain (10). Two organizer-reported problems remain open: source- and target-domain AUC are negatively correlated across systems, and development-set performance does not predict evaluation-set performance. We address both with a training-free post-hoc layer over frozen audio embeddings: (i) per-domain quantile calibration shrunk toward a pooled map by a prior strength m, tracing a source/target balance frontier, and (ii) a label-free cross-validated domain-balance criterion that ranks candidate configurations from training normals only, paired with a coarse development-labeled viability veto. On DCASE 2025, the criterion rank-predicts the official evaluation score across a 45-configuration grid (Spearman rho = +0.91; family-block bootstrap 95% CI [+0.83, +0.95]) while development score is uninformative (+0.06). Criterion-based selection raises the evaluation score from 55.83 to 59.34 (jackknife CI [2.2, 4.8]) and, on an extended grid, to 61.05 -- retrospectively fourth of 35 teams. Replicating on DCASE 2023 and 2024 bounds the claim: development score is uninformative in all three years and degenerate configurations recur (vetoed every time), but under family-clustered uncertainty the criterion's predictive evidence survives only in 2025; in both replication years a fixed full-equalization default matches or beats criterion-based selection. A DCASE 2026 forward test is frozen before the 2026 evaluation ground truth is released; all headline numbers are reproduced by the official evaluator.