Search papers, labs, and topics across Lattice.
This paper introduces Mr3D-VL, a vision-language foundation model specifically designed for multi-parametric 3D magnetic resonance imaging (mpMRI), addressing the limitations of existing AI models that struggle with natural language interaction and interpretability in clinical settings. By utilizing a 4 billion parameter architecture with an unsupervised pre-trained shared 3D encoder and 4D rotational positional embedding, Mr3D-VL effectively integrates spatial information across multiple imaging modalities. The model demonstrates substantial performance improvements in text generation, achieving a BERTScore of 0.856 and high accuracy in question-answering and multiple-choice tasks, showcasing its potential for enhancing clinical decision-making in brain tumor diagnosis and treatment.
Mr3D-VL outperforms existing models in cross-modal reasoning for brain tumor imaging, achieving a BERTScore of 0.856 and setting a new standard for interpretability in mpMRI applications.
Multi-parametric magnetic resonance imaging (mpMRI) is a cornerstone for brain tumor diagnosis and treatment, yet current AI models face critical limitations: their lack of natural language interaction and interpretability impedes spatial information integration and cross-modal reasoning required clinically. Key challenges arise from significant physical meaning differences across modalities, spatial misalignment due to scan intervals, and the need for complex multi-feature interpretation in tasks like glioma grading. While visual-language models (VLMs) show promise in cross-modal understanding, existing methods focus mainly on 2D image modeling, neglecting direct perception of 3D volumetric space. Although 3D VLMs have been proposed for report generation and feature alignment in 3D CT imaging, mpMRI applications demand collaborative inference across multiple imaging modalities-a requirement unmet by current solutions. To address this, we introduce Mr3D-VL, a dedicated visual-language foundation model for multi-parametric 3D MRI. With 4 billion parameters, it employs an unsupervised pre-trained shared 3D encoder and 4D rotational positional embedding for dual modality-spatial integration. Its cross-modal projection layer uses a multi-resolution feature implantation strategy to enhance feature perception across resolutions. Experimental results show significant improvements over existing 4B/7B/30B domain-specific and general-purpose models in text generation tasks, achieving a BERTScore of 0.856 for report generation, with question-answering accuracy at 0.713 and multiple-choice accuracy at 0.912.