Search papers, labs, and topics across Lattice.
This paper introduces GigaChat Audio, a time-aware large audio language model designed to tackle the challenges of temporal grounding in lengthy audio recordings. By interleaving periodic time markers with continuous audio tokens and employing large-scale synthetic supervision, the model effectively answers questions with explicit timestamps over recordings of up to 120 minutes. The results demonstrate significant improvements in temporal-grounding accuracy across various benchmarks, along with detailed analyses of factors influencing performance and computational efficiency.
GigaChat Audio achieves remarkable temporal grounding, accurately answering questions with timestamps even in lengthy audio inputs, setting a new standard for audio language models.
Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input. Our approach interleaves periodic time markers with continuous audio tokens using large-scale synthetic supervision from a cascaded pipeline. Our model achieves strong temporal-grounding accuracy on short and long benchmarks and supports time-anchored fragment descriptions and summaries. Extensive ablations examine how time representation, marker frequency, tokenization, and duration-mixture design affect accuracy and computational cost. We release model weights and datasets to support further research on time-aware audio understanding, available at https://huggingface.co/ai-sage/GigaChat3.1-Audio-10B-A1.8B.