Search papers, labs, and topics across Lattice.
This paper introduces PathVU, a novel benchmark designed to assess the fine-grained and multiscale visual understanding of multimodal large language models (MLLMs) in the context of pathology images. By leveraging 23 public datasets and converting raw annotations into deterministic task targets, PathVU facilitates a comprehensive evaluation of MLLM performance across various tasks, including region localization and spatial reasoning. The evaluation of 18 MLLMs reveals significant shortcomings in their ability to perform fine-grained visual tasks, underscoring the need for improved models in computational pathology.
Despite the rise of multimodal models, even the most advanced struggle with fine-grained visual tasks in pathology, revealing critical gaps in their understanding.
Multimodal large language models (MLLMs) are increasingly used to analyze pathology images. However, dominant multimodal benchmarks in pathology mainly score final diagnostic answers, captions, or reports. These evaluations provide limited insight into whether a model understands the multiscale visual content needed for pathology reasoning and decision-making. We introduce PathVU, a vision-anchored benchmark for fine-grained and multiscale visual understanding in computational pathology. Built from 23 public pathology imaging datasets with human-supervised labels and spatial annotations, PathVU evaluates MLLM understanding in two fields of view: Region FOV for high-resolution local regions and Slide FOV for macro whole-slide views. By converting raw annotations into deterministic task targets, PathVU enables programmatic scoring of region localization, visual recognition, quantity estimation, spatial reasoning, and insufficient-context judgment. The benchmark contains 14 VQA-style tasks, 61,673 images, and 308,070 samples across 28 organs and 7,253,526 annotations. Evaluating 18 representative general-purpose, medical-domain, and pathology-oriented MLLMs, we observe substantial limitations even in advanced models on fine-grained visual tasks across multiscale pathology images. PathVU provides a reproducible basis for developing and evaluating pathology MLLMs with explicit multiscale visual understanding.