Search papers, labs, and topics across Lattice.
This paper addresses the challenge of summarizing medical dialogues by developing a scalable data augmentation pipeline that integrates diverse medical dialogue datasets through synthetic speech generation and automated SOAP note supervision. The approach aims to streamline the clinical documentation process by enabling direct generation of clinical notes from speech, thereby minimizing the reliance on intermediate transcripts and preserving critical paralinguistic information. The key result is a robust adaptation of a speech foundation model for end-to-end speech-to-SOAP generation, which significantly reduces the documentation burden on healthcare workers.
Automating clinical note generation from speech could drastically reduce healthcare workers' documentation time while preserving essential patient information.
With the advent of Large Language Models and its instruction following capabilities a promising application is the task of summarization. Within this domain of task the extractive sub-task of clinical protocolling has emerged as a topic of particular interest as it can significantly reduce the downtime and protocolling burden of health-care workers thus enabling them to focus on their core work helping humans. A further step towards automation is the direct generation of clinical notes from speech without intermediate transcripts, reducing processing time while preserving information such as coughing or other paralinguistic cues that may be lost in transcript-based systems. To this end, we present KIT's submission to this years BeTraC challenge in the lightweight track. Our main contribution is a scalable data augmentation pipeline that unifies heterogeneous medical dialogue datasets through synthetic speech generation and automatically generated SOAP supervision, enabling robust adaptation of a speech foundation model for end-to-end speech-to-SOAP generation.