Search papers, labs, and topics across Lattice.
This paper introduces surface keypoint trajectories as a novel representation for modeling object motion in multi-object and articulated human-object interactions. By tracking a small set of non-collinear surface points over time, the method effectively captures the dynamics of both rigid and articulated objects without the need for explicit joint-type specifications. The proposed approach outperforms or matches existing techniques in generating realistic interactions across various datasets, highlighting its versatility in complex scenarios.
Surface keypoint trajectories enable seamless coordination between human motion and multiple articulated objects, outperforming traditional methods in complex interaction scenarios.
Daily activities require humans to coordinate whole-body motion with the motion of surrounding objects. Despite recent progress in human-object interaction (HOI) generation, most existing methods assume interactions with a single rigid object and do not extend well to scenarios involving a variable number of objects or articulated objects with diverse joint mechanisms. We propose surface keypoint trajectories as an object motion representation: for each rigid component, whether a standalone object or one part of an articulated assembly, we track a small set of non-collinear surface points over time. This representation handles multi-object coordination and diverse articulation mechanisms directly from point dynamics without requiring explicit joint-type specification. To model when and where each body region contacts each object, we introduce a spatio-temporal contact distance field that extends distance-based contact modeling to whole-body, multi-object, and articulated settings. We factorize HOI generation into three stages: generating object motions from text or waypoints, predicting the contact distance field, and synthesizing whole-body motion with contact-guided optimization. Experiments on ParaHome, HIMO, ARCTIC, and OMOMO demonstrate better or comparable performance to existing methods across single-object, multi-object, and articulated interaction settings.