Search papers, labs, and topics across Lattice.
This study investigates the impact of language-model fine-tuning on the internal roles of audio tokens in large audio-language models (LALMs) during temporal audio grounding. Through a series of analyses, the authors reveal that fine-tuning enhances the accessibility of event-related information in audio tokens, particularly in the early and middle layers of the decoder, while also demonstrating that significant latent evidence for queried events exists prior to fine-tuning. The findings indicate that grounding fine-tuning primarily improves the decoder's ability to read and align existing event evidence with temporal outputs, thus advancing our understanding of audio token semantics in LALMs.
Fine-tuning not only enhances the decoder's access to existing audio token evidence but also reveals that crucial event information is already present before any adjustments are made.
Large audio-language models (LALMs) convey acoustic evidence to language decoders through native audio tokens, yet the internal roles of these tokens remain poorly understood. Using temporal audio grounding as a diagnostic setting, we examine how language-model fine-tuning affects the layerwise semantics, decoder accessibility, and temporal output alignment of native audio-token states through four complementary analyses: query-conditioned token semantics, calibrated token readout, temporal-window probes, and residual-delta erasure during generation. Alongside substantial improvements in temporal localization, semantic analysis of Qwen2.5-Omni shows that latent evidence for queried events is already present before fine-tuning and that the audio tokens most strongly aligned with the queried event appear at similar temporal positions before and after fine-tuning. After fine-tuning, event-related information in audio tokens becomes more accessible to the decoder, especially in early and middle layers, and a cross-checkpoint control shows that this improvement arises primarily from decoder adaptation. Temporal probes show that the base checkpoint already contains recoverable information about annotated windows and that fine-tuning mainly improves alignment with each checkpoint's own predicted temporal support. Residual-delta erasure further shows that removing audio-token updates within predicted windows harms timestamp generation more than removing the same number of randomly selected updates. The same broad improvements in decoder readability and prediction alignment also appear in Qwen2-Audio. Together, these results support a semantics-to-readout account in which grounding fine-tuning helps the decoder read existing event evidence and connect it more reliably to temporal outputs.