Search papers, labs, and topics across Lattice.
This paper introduces StereoFlow, a novel framework for stereo matching that combines deterministic regression with generative distribution modeling to address the limitations of traditional approaches, particularly in ambiguous regions. By employing a two-stage progressive cascade matching network, a frequency-decoupled pixel diffusion transformer (StereoDiT), and a few-step flow matching objective, StereoFlow significantly improves geometric consistency and detail in challenging scenarios. Extensive evaluations reveal that StereoFlow sets new state-of-the-art performance across multiple benchmarks, including Scene Flow and KITTI, demonstrating its effectiveness in zero-shot generalization.
StereoFlow shatters the regression-to-mean bias in stereo matching, achieving state-of-the-art results even in the most ambiguous scenarios.
Stereo matching is a fundamental task in 3D reconstruction. Despite remarkable advances, the prevailing paradigms formulate stereo matching as a deterministic regression problem, collapsing the multimodal distribution modeling into a single-point estimation. This formulation suffers from a regression-to-mean bias, frequently struggling with ambiguous regions. In contrast, we introduce a prior-guided generative framework that integrates deterministic matching regression and generative distribution modeling within a complementary formulation. Built upon this formulation, we introduce StereoFlow through three key components: (i) a two-stage progressive cascade matching network that progressively produces multi-resolution stereo conditions with complementary matching cues; (ii) a pixel diffusion transformer (termed StereoDiT) with a frequency-decoupled architecture for modeling correspondence ambiguity; (iii) a few-step flow matching objective (termed Transition Flow Matching) for efficient optimization. In summary, \textsc{\textbf{StereoFlow}} achieves strong geometric consistency and rich fine-grained details in ill-posed, discontinuous regions and under zero-shot generalization. Extensive experiments demonstrate that the proposed StereoFlow establishes multiple state-of-the-art results across benchmarks, including Scene Flow, KITTI, ETH3D, and Middlebury.