NTU TaiwanMar 19, 2026arXiv:2603.19195

How Auditory Knowledge in LLM Backbones Shapes Audio Language Models: A Holistic Evaluation

Ke-Han Lu, Szu-Wei Fu, Chao-Han Huck Yang, Zhehuai Chen, Sung-Feng Huang, Chih-Kai Yang, Yi-Cheng Lin, Chi-Yuan Hsiao, Wenze Ren, En-Pei Hu, Yu-Han Huang, Anqi Cheng, An-Yu Cheng, Cheng-Han Chiang, Yu Tsao, Yu-Chiang Frank Wang, Y. Wang, Hung-yi Lee

AI Summary

This paper investigates the extent to which LLMs, pre-trained on text alone, possess auditory knowledge and how this impacts their performance when adapted into Large Audio Language Models (LALMs). They evaluate LLMs using direct probing with the AKB-2000 auditory knowledge benchmark, cascade evaluation using audio captioners, and fine-tuning into LALMs. The study reveals significant variance in auditory knowledge across different LLM families and a strong correlation between text-only auditory knowledge and downstream audio task performance.

Key Contribution

LLMs' text-only pre-training secretly encodes surprisingly different levels of auditory knowledge, directly impacting their effectiveness as backbones for audio language models.

Abstract

Large language models (LLMs) have been widely used as knowledge backbones of Large Audio Language Models (LALMs), yet how much auditory knowledge they encode through text-only pre-training and how this affects downstream performance remains unclear. We study this gap by comparing different LLMs under two text-only and one audio-grounded setting: (1) direct probing on AKB-2000, a curated benchmark testing the breadth and depth of auditory knowledge; (2) cascade evaluation, where LLMs reason over text descriptions from an audio captioner; and (3) audio-grounded evaluation, where each LLM is fine-tuned into a Large Audio Language Model (LALM) with an audio encoder. Our findings reveal that auditory knowledge varies substantially across families, and text-only results are strongly correlated with audio performance. Our work provides empirical grounding for a comprehensive understanding of LLMs in audio research.

Eval Frameworks & Benchmarks Multimodal Models Speech & Audio

Citation Metrics

Citations0

Influential citations0

References65

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

How Auditory Knowledge in LLM Backbones Shapes Audio Language Models: A Holistic Evaluation

Related Papers