Search papers, labs, and topics across Lattice.
The Depth Anything V4 (DAV4) framework advances dynamic 4D scene reconstruction from monocular video by integrating Riemannian Flow Matching (RFM) with 4D Gaussian Splatting parameters, enabling valid intermediate states on non-Euclidean manifolds. Controlled experiments reveal that RFM significantly enhances performance, achieving an F-score of 0.806 compared to a deterministic MLP baseline's 0.762, isolating a +0.044 gain attributable solely to RFM. Additionally, DAV4 demonstrates superior performance in dynamic reconstruction and novel-view synthesis without relying on human-annotated depth labels for training losses, while providing a comprehensive computational cost analysis for large-scale deployment.
Riemannian Flow Matching boosts dynamic 4D scene reconstruction performance by over 5% without requiring human-annotated depth labels.
We present Depth Anything V4 (DAV4), a framework for dynamic 4D scene reconstruction from monocular video. Our key contribution is the application of Riemannian Flow Matching (RFM) to 4D Gaussian Splatting parameters, defining probability paths directly on non-Euclidean manifolds (scale, rotation, opacity), ensuring all intermediate states are valid. Through controlled experiments, we isolate RFM's contribution from test-time optimization (TTO) and pre-training. A deterministic MLP baseline with the same data, architecture, and TTO achieves F-score 0.762; RFM achieves 0.806 - the +0.044 gain is RFM's isolated contribution. We provide corrected computational cost analysis: pre-training is 360 GPU-hours, amortizing for large-scale deployment (over 10,000 scenes). Uncertainty is quantified via Negative Gaussian Log-Likelihood and Expected Calibration Error. DAV4 outperforms prior Depth Anything models and per-scene 4D-GS on dynamic reconstruction and novel-view synthesis, while using no human-annotated depth labels as training losses.