Search papers, labs, and topics across Lattice.
This paper introduces MVC-Bench, a novel benchmark designed specifically to evaluate the calibration of medical vision-language models (Medical-VLMs) under realistic clinical conditions. The benchmark assesses calibration across multiple dimensions, including robustness to various shifts and the effectiveness of different calibration strategies, using over 1638 controlled experiments. Key findings reveal that the proposed Multi-Class Margin (MCM) regularization method significantly reduces Expected Calibration Error (ECE) across most settings, highlighting its potential for enhancing model reliability in safety-critical medical applications.
Calibration in medical vision-language models is crucial, and MVC-Bench reveals that a simple train-time calibration method can outperform existing approaches in most scenarios.
Reliable evaluation of vision-language models (VLMs) and medical vision-language models (Medical-VLMs) requires calibrated confidence, particularly under realistic clinical conditions. However, existing efforts mainly focused on improving accuracy, leaving calibration in the medical domain underexplored. To this end, we propose MVC-Bench, a calibration-centric benchmark for medical image classification with VLMs and Medical-VLMs. MVC-Bench assesses the calibration across three axes: (i) robustness to modality, backbone, and domain shift (ii) effectiveness of calibration strategies and prompt-tuning methods (iii) stability under prompt-template and random-seed variations. The benchmark covers eight different backbones, three medical modalities, including fundus imaging, histopathology, and chest X-ray under in-domain and domain shift settings. It compares post-hoc calibration, train-time calibration, and zero-shot inference methods, together with six prompt-tuning methods. Across more than 1638 controlled experiments, we report accuracy and Expected Calibration Error (ECE) as primary metrics, and further report results with complementary calibration measures, including Maximum Calibration Error (MCE) and Adaptive Calibration Error (ACE). We further investigate the underlying causes of miscalibration in VLMs and Medical-VLMs and propose a simple train-time calibration method, Multi-Class Margin (MCM) regularization, which achieves lowest ECE on 10 out of 12 settings in in-domain and remains competitive under domain shifts. Collectively, MVC-Bench provides a structured evaluation framework and actionable guidance for improving calibration in safety-critical medical workflows.