Search papers, labs, and topics across Lattice.
Marigold V2 repurposes pretrained flow-matching Diffusion Transformers (DiTs) into efficient, single-step monocular depth estimators by mitigating the structural artifacts typical of naive fine-tuning. By coupling internal representation alignment to semantic features with a two-stage fine-tuning schedule, the framework resolves fine geometric details like fur and foliage that typically elude regression models. Across KITTI and ETH3D, the model reduces AbsRel error by 16–26% over prior state-of-the-art baselines while generalizing zero-shot to surface normal estimation and intrinsic image decomposition.
Single-step flow-matching DiTs can surpass multi-step generative models in geometric accuracy, slashing depth error by up to 26% while recovering fine-grained details down to individual strands of hair.
Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computational photography, and robotics, among others. Despite the field's maturity, recent models still struggle to generalize to out-of-distribution inputs and to produce sharp and detailed depth maps. In this paper, we revisit Marigold, a set of techniques for repurposing modern image generation and editing models, powered by the diffusion transformer (DiT) architecture, into state-of-the-art monocular depth estimators. Our recipes target single-step inference from pretrained multi-step flow-matching models, with quantization where needed, preserving model capacity while remaining cheap to run. We analyze the artifacts of naive training and identify two effective remedies: aligning the model's internal representations with semantic features extracted from ground-truth, and adopting a 2-stage fine-tuning protocol built around a novel Sinkhorn-based loss. The results are crisper, cleaner depth maps that generalize well out-of-distribution, with 16-26% improvement in AbsRel over the previous best on KITTI and ETH3D. Qualitatively, our model resolves fur, foliage, and hair-thin edges that have eluded prior models. Furthermore, Marigold V2 achieves state-of-the-art results when applied to other dense regression tasks, such as surface normals estimation and intrinsic image decomposition. Project website: https://hf.co/spaces/huawei-bayerlab/marigold-v2-web