Search papers, labs, and topics across Lattice.
This paper addresses the limitations of existing Multimodal Large Language Models (MLLMs) in the medical domain by introducing MedUAG, a unified understanding and generation framework. The authors construct the MedUAGCorpus, the largest dataset for medical UAG with over 6 million instances across 14 imaging modalities, and develop MedUAGBench, a benchmark for evaluating medical generation across 12 tasks. Experimental results show that MedUAG sets a competitive baseline, significantly advancing the state of medical multimodal systems.
MedUAG sets a new standard in medical multimodal models, achieving strong performance across diverse understanding and generation tasks with the largest dataset yet.
Recent Multimodal Large Language Models (MLLMs) are rapidly evolving into unified understanding and generation (UAG) frameworks. However, extending these unified paradigms to the medical domain is hindered by: the absence of comprehensive training and evaluation benchmarks, and the lack of broadly validated unified medical model. To address these gaps, we present a comprehensive foundation for medical UAG. First, we construct MedUAGCorpus, the largest unified medical understanding and generation dataset to date, comprising over 6 million instances across 14 imaging modalities. Second, we introduce MedUAGBench, a systematic benchmark that expands medical generation evaluation to 12 diverse tasks under standardized protocols. Finally, leveraging these resources, we develop MedUAG, an end-to-end trained unified medical model. Extensive experiments demonstrate that MedUAG achieves strong performance across a wide array of understanding and generation tasks, establishing a competitive baseline and paving the way for next-generation medical multimodal systems.