Search papers, labs, and topics across Lattice.
This paper introduces PhotoHOI, a method for synthesizing 3D hand-object interaction sequences from a single RGB photograph and open-vocabulary language instructions. By leveraging a vision-language model to parse input images and instructions, PhotoHOI generates structured task specifications and plans collision-aware object trajectories, while learning transferable priors from extensive affordance and hand-object interaction data. Experiments show that PhotoHOI significantly improves contact quality and reduces penetration compared to existing methods, achieving higher task success rates and scene consistency in real-world applications.
Synthesizing realistic 3D hand-object interactions from just a single image and text instruction could revolutionize AR/VR applications by enabling seamless integration of dynamic human behaviors.
Hand-object interaction (HOI) is a fundamental human behavior with broad applications in AR/VR, digital humans, and embodied interaction. Existing methods typically require predefined object geometry, object trajectories, or task-specific conditions, limiting their use with natural real-world inputs. To address this, we study a more practical problem of synthesizing 3D hand-object interaction sequences from a single RGB photograph and an open-vocabulary language instruction, and introduce PhotoHOI. PhotoHOI first uses a vision-language model to parse the input image and instruction into a structured task specification, including the interaction object, target region, and spatial relation. It then recovers a compact task-relevant 3D scene and plans a smooth collision-aware object trajectory based on the recovered object states, support relations, and surrounding scene geometry. To synthesize hand motion that generalizes to real-world photographs and unseen objects, it learns transferable task-conditioned contact and contact-conditioned grasp priors from large-scale affordance and HOI data. The grasp is further refined in a learned latent space, constraining the optimization to a plausible hand-pose manifold. Experiments on GRAB and H2O demonstrate improved contact quality and reduced penetration over representative baselines. Results on real-world photographs further demonstrate higher task success and scene consistency, together with generalization to unseen objects and open-vocabulary instructions.