Search papers, labs, and topics across Lattice.
This paper introduces a multi-axis evaluation framework for structured audio captions, addressing the limitations of existing metrics that focus solely on flat textual outputs. By leveraging Large Language Model judges alongside deterministic computational metrics, the framework evaluates audio descriptions across five distinct axes, including logical reasoning and spectral profiles. The controlled perturbation testing protocol validates the framework's reliability, showing its effectiveness in differentiating between meaning-preserving paraphrases and actual semantic or acoustic distortions.
This new evaluation framework can reliably distinguish between meaningful audio caption variations and genuine corruptions, transforming how we assess automated audio captioning systems.
Recent advancements in automated audio captioning (AAC) have shifted from monolithic sentence generation toward structured formats that explicitly disentangle distinct acoustic and semantic properties. However, evaluating this heterogeneous data remains a significant challenge. Existing caption metrics focus on flat textual outputs and fail to reliably assess multimodal attributes. To bridge this gap, we propose a multi-axis evaluation framework tailored for structured audio descriptions. Building on the AudioCards dataset, we evaluate outputs across five orthogonal axes: tag-sets, descriptions, logical reasoning, numeric measurements, and spectral profiles. Our approach combines Large Language Model (LLM) judges to capture semantic nuance with deterministic computational metrics to precisely measure acoustic deviations. To rigorously validate the reliability of this framework, we introduce a controlled perturbation testing protocol that injects typed, graded errors into groundtruth annotations. Our results demonstrate that this framework successfully distinguishes meaning-preserving paraphrases from genuine semantic and acoustic corruptions.