Search papers, labs, and topics across Lattice.
This paper introduces Object-Conditioned Social Diffusion (OCSD), a novel conditional diffusion model designed for accurately forecasting multi-person human motion in complex scenes by integrating motion history, social interactions, and object information. The model employs an object-conditioning mechanism that enhances denoising at each timestep, facilitating detailed human-object reasoning, while a social encoder captures interactions among individuals. Experimental results demonstrate that OCSD outperforms previous methods, achieving significant reductions in path error on the Humans in Kitchens and HOI-M3 benchmarks, while generating more realistic long-term forecasts.
OCSD reduces path error by over 30% and generates more realistic long-term human motion forecasts by effectively integrating object cues and social interactions.
Accurately forecasting the movement of people in complex scenes requires reasoning over the past and present state of the entire environment. In this context, effectively incorporating object information and social interactions into a unified framework remains particularly challenging. To address this, we propose Object-Conditioned Social Diffusion (OCSD), a conditional diffusion model that integrates motion history, multi-person interactions, and object cues into a single framework. OCSD uses an object-conditioning mechanism that modulates denoising at every timestep, enabling fine-grained human-object reasoning, and a social encoder that models the interactions between all humans in the scene. As a result, our model naturally handles varying group sizes, complex social interactions, and supports sampling multiple plausible futures. Extensive experiments show that OCSD achieves state-of-the-art results on the Humans in Kitchens (HiK) and HOI-M3 benchmarks. It reduces the two-second path error by 121.5 mm (31.3%) on HiK and 130.5 mm (33.2%) on HOI-M3 compared to prior work, and produces more realistic long-term forecasts.