Search papers, labs, and topics across Lattice.
This paper introduces the LongAudioQA dataset and the GRGA model, which enhances long-form audio meeting understanding by integrating heterogeneous audio features into a multi-dimensional graph. By employing agent planning for retrieval and answer generation, the model effectively mitigates issues related to acoustic information loss and long-term context memory in existing speech QA systems. The results demonstrate significant improvements in understanding and responding to audio meeting content compared to current state-of-the-art methods.
Transforming long-form audio meeting comprehension, the GRGA model leverages graph-based planning to overcome acoustic loss and memory challenges, achieving superior QA performance.
While long-form audio meeting understanding (LAMU) is garnering growing attention, task-specific question answering (QA) datasets remain scarce. Existing speech QA paradigms and state-of-the-art Speech LLMs suffer from acoustic information loss and poor long-term context memory. To address these issues, we construct the LongAudioQA dataset and propose the GRGA model, which models heterogeneous audio features into a multi-dimensional graph and leverages agent planning for retrieval and answer generation.