Search papers, labs, and topics across Lattice.
This paper introduces MADS, a 19-dimensional descriptor set that captures the physical dynamics of sound-generating events by integrating spectral, temporal, mechanical, and stochastic features. Evaluated against standard audio classification benchmarks (ESC-10, ESC-50, and MSoS), MADS outperforms traditional handcrafted baselines while utilizing significantly fewer dimensions. The results indicate that MADS not only excels in classification accuracy but also serves as a foundational layer for future audio modeling approaches grounded in acoustic principles.
MADS achieves superior audio classification performance while halving the dimensionality compared to conventional spectral summaries, redefining efficiency in audio representation.
Dominant audio classification pipelines rely either on compact handcrafted summaries or on fixed time-frequency frontends such as log-mel representations prior to deep modeling. While highly successful, these representations do not explicitly expose the physical dynamics of the underlying sound-generating event. We introduce MADS (Multi-view Acoustic Descriptor Set), a compact 19-dimensional physics-informed descriptor set de- signed to capture complementary spectral, temporal, mechanical, and stochastic structure in audio signals. Rather than treating sound only as a spectral pattern, MADS encodes properties related to excitation, damping, periodicity, impulsiveness, and structural consistency within a unified multi-view representation. We evaluate MADS using standard classical machine learning models on ESC-10, ESC-50, and MSoS, and compare it against two conventional handcrafted baselines: a compact 26D MFCC- based baseline and an expanded 38D spectral-summary baseline. Across ESC-10 and ESC-50, MADS achieves the strongest peak results overall, reaching 81.00% and 52.78%, respectively, while using roughly half the dimensionality of the 38D baseline. On MSoS, MADS again delivers the strongest top-end performance, reaching 67.48%. These results establish MADS not merely as a competitive standalone descriptor set, but as the foundational descriptor layer of a broader acoustically grounded representation program for future frame-level and deep-learning-compatible audio modeling.