Search papers, labs, and topics across Lattice.
This paper introduces MV-dVRK, a pioneering ex-vivo surgical dataset that integrates multiple exposure-synchronized stereo viewpoints with precise surface geometry and camera poses, specifically designed for evaluating 3D reconstruction methods in surgical contexts. The dataset facilitates a comprehensive comparison of various reconstruction approaches, revealing that optimization-based multi-view methods outperform feed-forward models, achieving 67% coverage of ground-truth surface points within a 1 mm tolerance when three viewpoints are utilized. Additionally, MV-dVRK includes dynamic sequences that reflect real surgical tasks, laying the groundwork for advancements in multi-viewpoint surgical perception research.
Multi-stereo reconstruction outperforms feed-forward models by achieving 67% coverage of ground-truth surgical surfaces with just three viewpoints, highlighting the critical role of viewpoint diversity in surgical perception.
Large-scale training and refined optimization techniques have greatly improved sparse multi-view 3D reconstruction. Despite their relevance to surgery, such methods have never before been rigorously evaluated on real endoscopic images. Current clinical telerobots deploy a single stereo camera inside the patient, making multi-viewpoint data extremely rare. This paper presents MV-dVRK, the first ex-vivo surgical dataset to combine multiple exposure-synchronized stereo viewpoints with accurate surface geometry and camera poses. The static subset of the benchmark provides dense SfM reference geometry, validated against an industrial 3D scanner, together with ground-truth camera poses and sparse-view test sets. We use MV-dVRK to systematically compare zero-shot monocular, stereo, multi-stereo, and multi-view 3D reconstruction methods as the number of viewpoints increases. With two endoscopes, multi-stereo reconstruction achieves the highest coverage. With a third viewpoint, optimization-based multi-view methods perform best, covering 67% of ground-truth surface points within a 1 mm tolerance and recovering highly accurate relative camera poses. By contrast, feed-forward foundation models cover only 43% of the ground-truth surface in the same setting. MV-dVRK also includes ten dynamic sequences spanning multiple surgical tasks, with increasing kinematic complexity and tissue deformation, providing a basis for future research in multi-viewpoint surgical perception. The project is available at: https://mv-dvrk.is.mpg.de.