Search papers, labs, and topics across Lattice.
This study critically evaluates the effectiveness of machine-learnt models in predicting room acoustic parameters, revealing that reported high accuracy metrics are heavily influenced by the evaluation protocols employed. By conducting a multi-condition measurement campaign in two different hall settings, the authors demonstrate that validation strategies and input feature selection can drastically alter the perceived performance of various model families. Notably, while hybrid CNNs show potential, their reliance on specific input configurations limits their generalizability, underscoring the importance of protocol consistency in model evaluation.
Evaluation protocols can inflate model performance metrics by up to an order of magnitude, challenging the reliability of current acoustic prediction models.
Machine-learnt models are increasingly used to predict ISO 3382-1 room acoustic parameters from sparse measurements, with reported coefficients of determination frequently above 0.85. This paper shows that such figures are often determined by the evaluation protocol rather than by the model. Using a multi-condition measurement campaign in a 264-seat conference hall and a 180-seat concert hall, three model families were evaluated under a factorial protocol ablation: validation splits either row-based or grouped by receiver position, and input features either including measured-at-test quantities or restricted to source-receiver geometry and environmental state. Row-based splits with measured-at-test inputs reproduce the high reported accuracies (mean $R^2$ 0.81 for the core parameters); grouping the splits by position and restricting inputs to information available at an unmeasured position reduces these to 0.09-0.57, reordering the apparent difficulty of parameter classes. A hybrid CNN evaluated with the target's own impulse response as input is shown to exploit it as a position fingerprint rather than as transferable acoustic information; training-only signal access yields no gain for any parameter tested, including reverberation time. Under the deployment-consistent protocol, the spread between Random Forest, the hybrid CNN, and inverse-distance weighting is an order of magnitude smaller than the spread between protocols for a fixed model; the learnt models retain a genuine advantage for sound strength and reverberation time, and the high accuracy of the original pipelines re-emerges as condition interpolation at measured positions (band means 0.80-0.88), a distinct and operationally useful task.