Search papers, labs, and topics across Lattice.
This paper introduces CausalSplat, a novel framework that enhances 3D Gaussian Splatting (3DGS) by integrating vision-language models with 3D scene graphs to improve hierarchical reasoning capabilities. The authors identify significant limitations in existing methods, which struggle with implicit intents and complex reasoning, and construct two benchmarks鈥擟ausal-LERF and Causal-ScanNet鈥攖o systematically evaluate these reasoning challenges. Experimental results show that CausalSplat not only outperforms state-of-the-art methods on the proposed benchmarks but also demonstrates strong generalizability across standard 3D segmentation tasks.
Current state-of-the-art methods falter in commonsense reasoning for 3D scene understanding, but CausalSplat redefines the landscape by achieving superior performance on complex reasoning tasks.
While 3D Gaussian Splatting (3DGS) has advanced open vocabulary scene understanding, existing methods remain confined to explicit queries. They struggle to interpret implicit intents, complex spatial constraints, and commonsense reasoning required for practical embodied interactions. To address this gap, we introduce the task of reasoning 3D Gaussian segmentation and construct two benchmarks, Causal-LERF and Causal-ScanNet. These benchmarks systematically evaluate commonsense, spatial, affordance, and counterfactual reasoning. Evaluations reveal that current state of the art methods perform poorly on these reasoning challenges. Therefore, we propose CausalSplat, a framework that integrates vision-language models with 3D scene graphs to disentangle explicit structural perception from implicit logical inference. Extensive experiments demonstrate that CausalSplat achieves state of the art performance on our reasoning benchmarks while showing strong generalizability on standard referring and open vocabulary 3D segmentation tasks. Project Page: https://jiayuding031020.github.io/CausalSplat