Search papers, labs, and topics across Lattice.
This paper investigates the effectiveness of query rewriting in enhancing complex object segmentation within 4D Gaussian representations, addressing the challenges posed by verbose and noisy queries. By employing a training-free reinterpretation strategy that condenses lengthy descriptive queries into concise, keyword-focused forms, the authors demonstrate significant improvements in both temporal localization and spatial segmentation performance. Experimental results on HyperNeRF and Neu3D reveal that this approach boosts average temporal accuracy from 60.92% to 92.21% and average vIoU from 20.08% to 76.94%, highlighting the method's robustness without requiring additional fine-tuning.
Transforming verbose queries into concise keywords can elevate segmentation accuracy from 20% to nearly 77% in complex dynamic scenes.
Recent 4D Gaussian representation frameworks have demonstrated strong performance in language-guided dynamic scene understanding. However, these methods remain highly sensitive to verbose and narrative-style queries that contain noisy contextual information. In this paper, we investigate the impact of query rewriting for complex object segmentation in 4D Gaussian representations. Inspired by recent findings in retrieval-augmented language models and keyword-guided query reformulation, we propose a training-free reinterpretation strategy that transforms long descriptive queries into concise keyword-grounded forms. Our approach progressively reduces linguistic noise while preserving semantic anchors relevant to object-centric representations. Experiments on HyperNeRF and Neu3D demonstrate that concise rewritten queries significantly improve both temporal localization and spatial segmentation performance. In particular, our method improves average temporal accuracy from 60.92% to 92.21% and average vIoU from 20.08% to 76.94% without any additional fine-tuning. Extensive ablation studies further reveal that shorter, keyword-focused queries consistently yield stable video-feature similarity distributions and better alignment with object-centric Gaussian representations