Tsinghua AIECNUGIST GuangdongFeb 26, 2026arXiv:2602.22955

MM-NeuroOnco: A Multimodal Benchmark and Instruction Dataset for MRI-Based Brain Tumor Diagnosis

Feng Guo, Feng Guo, Jiaxiang Liu, Jiaxiang Liu, Yang Li, Yang Li, Qianqian Shi, Mingkun Xu, Mingkun Xu

AI Summary

The authors introduce MM-NeuroOnco, a large-scale multimodal dataset for brain tumor MRI understanding, comprising 24,726 MRI slices and approximately 200,000 semantically enriched multimodal instructions. They employ a multi-model collaborative pipeline for automated medical information completion and quality control to generate diagnosis-related semantics beyond mask annotations. Experiments on MM-NeuroOnco-Bench, a manually annotated evaluation benchmark, reveal that even strong baselines like Gemini 3 Flash struggle with diagnosis-related questions, while fine-tuning NeuroOnco-GPT on MM-NeuroOnco yields a significant 27% accuracy improvement.

Key Contribution

Even the best vision-language models struggle to diagnose brain tumors from MRI scans, but a new dataset and benchmark reveals a path to significant accuracy gains through instruction tuning.

Abstract

Accurate brain tumor diagnosis requires models to not only detect lesions but also generate clinically interpretable reasoning grounded in imaging manifestations, yet existing public datasets remain limited in annotation richness and diagnostic semantics. To bridge this gap, we introduce MM-NeuroOnco, a large-scale multimodal benchmark and instruction-tuning dataset for brain tumor MRI understanding, consisting of 24,726 MRI slices from 20 data sources paired with approximately 200,000 semantically enriched multimodal instructions spanning diverse tumor subtypes and imaging modalities. To mitigate the scarcity and high cost of diagnostic semantic annotations, we develop a multi-model collaborative pipeline for automated medical information completion and quality control, enabling the generation of diagnosis-related semantics beyond mask-only annotations. Building upon this dataset, we further construct MM-NeuroOnco-Bench, a manually annotated evaluation benchmark with a rejection-aware setting to reduce biases inherent in closed-ended question formats. Evaluation across ten representative models shows that even the strongest baseline, Gemini 3 Flash, achieves only 41.88% accuracy on diagnosis-related questions, highlighting the substantial challenges of multimodal brain tumor diagnostic understanding. Leveraging MM-NeuroOnco, we further propose NeuroOnco-GPT, which achieves a 27% absolute accuracy improvement on diagnostic questions following fine-tuning. This result demonstrates the effectiveness of our dataset and benchmark in advancing clinically grounded multimodal diagnostic reasoning. Code and dataset are publicly available at: https://github.com/gfnnnb/MM-NeuroOnco

Computer Vision Eval Frameworks & Benchmarks Multimodal Models

Citation Metrics

Citations0

Influential citations0

References59

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

MM-NeuroOnco: A Multimodal Benchmark and Instruction Dataset for MRI-Based Brain Tumor Diagnosis

Related Papers