Search papers, labs, and topics across Lattice.
This paper introduces DocTalkBN, a comprehensive dataset of expert telemedicine conversations in Bengali, comprising 557.63 hours of audio and text from real patient interactions. The dataset is significant as it captures the spontaneity and contextual richness of authentic medical dialogues, addressing the scarcity of such resources in low-resource languages. Benchmarking on three downstream tasks reveals that DocTalkBN enhances the performance of large language models in clinically relevant reasoning tasks, underscoring its utility for advancing medical NLP applications.
Authentic telemedicine conversations in Bengali reveal that existing models can significantly improve their performance on clinical reasoning tasks with the right data.
Reliable medical conversational AI requires authentic expert--patient interaction data, yet such datasets remain scarce, especially for low-resource languages such as Bengali. We present DocTalkBN, a large-scale multimodal dataset of real-world expert telemedicine conversations in Bengali, collected from nationally broadcast telemedicine programs featuring board-certified physicians. DocTalkBN contains 557.63 hours of paired audio and text, 1,515 multi-turn patient calls, 10,274 host--doctor question--answer exchanges, totaling 1.7M tokens, spanning 26 medical specialties. Unlike prior resources derived from medical forums, written health content, or synthetic data, our dataset preserves the spontaneity, contextual richness, and spoken characteristics of authentic medical interactions in a low-resource setting. To support benchmark-driven research, we further construct three downstream tasks from the corpus, medical triage classification, advice safety evaluation, and medical named entity recognition, and benchmark a diverse set of large language models and encoder-based baselines. Our results show that DocTalkBN is a practically useful resource, particularly for clinically grounded reasoning tasks. We release this resource to facilitate future research on reliable medical NLP and safer, more culturally grounded healthcare systems for low-resource languages. Our source codes and dataset are publicly available at https://anonymous.4open.science/r/doctalk.