Search papers, labs, and topics across Lattice.
This paper introduces a global-to-local pose estimation framework for freehand 3-D ultrasound imaging that combines external camera observations with B-mode ultrasound images to enhance probe pose accuracy. By utilizing a dual-camera branch for global localization and a B-mode branch for local anatomical refinement, the method effectively reduces accumulated pose errors during extended scanning. Validation on phantom and in vivo datasets shows significant improvements in trajectory drift, achieving reductions of up to 27.12% compared to existing methods, and demonstrating superior performance in clinical scenarios.
Achieving a remarkable 27.12% reduction in trajectory drift, this framework redefines pose estimation accuracy in freehand 3-D ultrasound imaging.
Freehand 3-D ultrasound (US) imaging has attracted increasing attention owing to its intuitive volumetric visualization, ease of use, and low cost. However, accurate 3-D reconstruction critically depends on stable probe pose estimation, yet existing trackerless methods remain susceptible to accumulated pose errors, particularly over long scanning trajectories. To address this limitation, we propose a global-to-local pose estimation framework that exploits external camera observations for globally stable localization and B-mode US images for anatomy-aware local refinement. Specifically, the framework comprises a dual-camera branch that performs contextual feature aggregation across camera views and temporal observations to estimate a globally consistent probe trajectory, and a B-mode branch that performs anatomical feature aggregation from sequential US images to capture tissue-dependent local motion cues. A cross-modal fusion module subsequently integrates the contextual camera features and anatomical US features to predict pose residuals and refine the camera-derived estimates in the transformation space. Furthermore, a multi-scale pose loss constrains relative motion over multiple temporal horizons to suppress accumulated drift during extended scans. The proposed framework is validated on phantom and in vivo datasets. On two in-house datasets (FUSION-J and FUSION-L) collected using different machines, the proposed US + Dual-Cam model reduces average trajectory drift to 1.67 mm and 1.29 mm, representing improvement of 16.50% and 27.12%, respectively, over a strong dual-camera baseline, while substantially outperforming US-only pose estimation (>13 mm drift). In in vivo forearm arteries reconstruction, it achieves Hausdorff distances of 1.58 mm, demonstrating the effectiveness of the proposed method on real clinical scenarios.