Search papers, labs, and topics across Lattice.
This paper introduces ReFine3D, a regularized fine-tuning framework aimed at enhancing the domain generalization of 3D vision-language models. By employing selective layer tuning and targeted regularization strategies鈥攕uch as multi-view consistency and synonym-based prompts鈥擱eFine3D effectively mitigates overfitting and catastrophic forgetting in multimodal models. Extensive evaluations demonstrate that ReFine3D significantly outperforms existing methods, achieving notable improvements in generalization, transfer, and robustness metrics with minimal computational cost.
ReFine3D boosts 3D vision-language model generalization by over 3% while keeping computational demands low.
Domain adaptation remains a central challenge in 3D vision, especially for multimodal foundation models that align 3D point clouds with visual and textual data. While these models demonstrate strong general capabilities, adapting them to downstream domains with limited data often leads to overfitting and catastrophic forgetting. To address this, we introduce ReFine3D, a regularized fine-tuning framework designed for domain-generalizable tuning of 3D large multimodal models (LMMs). ReFine3D combines selective layer tuning with two targeted regularization strategies: multi-view consistency across augmented point clouds and text diversity through synonym-based prompts generated by large language models. Additionally, we incorporate point-rendered vision supervision and a test-time augmentation mechanism with confidence-based aggregation to further enhance robustness. Extensive experiments across different 3D domain generalization benchmarks show that ReFine3D improves base-to-novel class generalization by 1.36%, cross-dataset transfer by 2.43%, robustness to corruption by 1.80%, and few-shot accuracy by up to 3.11%, outperforming prior state-of-the-art methods with minimal added computational overhead.