Search papers, labs, and topics across Lattice.
This paper addresses the limitations of existing 3D language models (3D-LLMs) that primarily focus on single-object descriptions by introducing a framework for detailed multi-object reasoning. The authors develop MO3D, a dataset designed for fine-grained comparisons among multiple objects, and the Multi-3DLLM, a Patch-Interaction Transformer that effectively models inter- and intra-object relationships while maintaining local geometry. The results show that Multi-3DLLM significantly outperforms current 3D-LLMs and 2D vision-language models on the proposed benchmarks, demonstrating enhanced geometric reasoning capabilities and positive transfer effects to single-object tasks.
Multi-3DLLM not only excels at multi-object reasoning but also boosts performance on single-object classification tasks, revealing a surprising synergy in geometric understanding.
We address a fundamental gap in 3D-LLMs: existing models focus on single-object/scene description, struggling with detailed, inter-object comparison. We propose a framework for detailed object-level reasoning across multiple objects with three components: (1) MO3D (Multi-Object in 3D), an instruction dataset requiring fine-grained multi-object comparison; (2) Multi-3DLLM, using a minimal Patch-Interaction Transformer (PIT) that models inter-/intra-object relationships while preserving local geometry; (3) Mini-apps, two application-driven benchmarks (Shape Mating, Change Captioning) that probe geometric understanding for practical use. Recent 3D-LLMs and 2D-VLMs perform poorly on these tasks, lacking both comparison-centric design and geometric awareness. In contrast, Multi-3DLLM trained on our mixture data learns geometric reasoning, surpasses all baselines on MO3D, and provides positive transfer to single-object classification.