Search papers, labs, and topics across Lattice.
This paper introduces SONG, a novel simulation platform that utilizes 3D Gaussian splatting to create photorealistic environments for benchmarking social navigation in robots. By integrating semantically grounded trajectories generated by a large language model and a trajectory-conditioned generator for natural motion synthesis, SONG addresses the limitations of existing platforms in visual fidelity and pedestrian behavior. The evaluation of navigation baselines reveals significant gaps in current methodologies, emphasizing the importance of real-world data and the need for improved safety and social compliance in robotic navigation systems.
Vision-based social navigation is still far from solved, with critical safety deficits overshadowing social etiquette in robotic interactions.
Social navigation has progressed from simplified 2D environments toward a more general vision-based setting, in which a robot needs to achieve socially compliant behavior purely from onboard visual observations. Yet supporting simulation platforms have not kept pace: existing options either lack visual observations, lack moving human avatars, or fall short of real-world fidelity in appearance and pedestrian behavior, offering limited support for advancing vision-based social navigation. We introduce SONG, a SOcial Navigation platform powered by 3D Gaussian splatting (3DGS). It leverages 3DGS for both scene and avatar representations, drives pedestrians using semantically grounded trajectories generated by a large language model, and synthesizes their full-body motion with a trajectory-conditioned generator to produce continuous, natural movement. On top of the platform, we curate SONG-Bench, a set of evaluation episodes stratified by difficulty, and propose a multi-dimensional metric suite covering effectiveness, safety, and social compliance. A systematic evaluation of representative navigation baselines reveals three findings: (a) vision-based social navigation is far from solved; (b) a critical safety deficit precedes social etiquette; (c) real-world data matters more than model scale. Crucially, we demonstrate that fine-tuning on our curated data effectively improves the success rate in real-world environments. We hope our platform provides a faithful and rigorous testbed for the next generation of vision-based social navigation research.