Search papers, labs, and topics across Lattice.
This paper introduces EgoPHI, a novel method that jointly estimates dense contact maps and 3D force distributions from a single monocular RGB image and object geometry, addressing a critical gap in understanding hand-object interactions. By leveraging a physics-based simulation pipeline to create augmented datasets with dense per-vertex force supervision, EgoPHI significantly enhances force estimation capabilities compared to existing methods. The evaluation reveals that EgoPHI not only improves performance on both in-distribution and out-of-distribution benchmarks but also successfully transfers to real-world scenarios, marking a significant advancement in physically grounded interaction reasoning.
EgoPHI transforms egocentric vision by enabling precise 3D force estimation from a single image, bridging the gap between contact localization and physical interaction reasoning.
Understanding hand-object interaction from egocentric vision is essential for modeling how people physically engage with the surrounding world. Yet reasoning about physically grounded interaction requires estimating the forces acting on hands and objects, beyond localizing contact. We present EgoPHI, the first method that jointly estimates dense contact maps and 3D force distributions on hand and object meshes from a single monocular RGB image and object geometry. To address the lack of scalable ground-truth force annotations, we introduce a physics-based simulation pipeline that augments existing hand-object datasets with dense per-vertex force supervision. EgoPHI then learns dense 3D contact and force on interacting hand and articulated object meshes, extending vision-based force estimation beyond image-space or planar settings. Our evaluation on in-distribution and out-of-distribution benchmarks shows that EgoPHI improves force estimation over existing approaches while generalizing to unseen datasets. To evaluate sim-to-real transfer, we constructed two physical objects that capture dense object contact and force magnitude and used them to record a dataset of interactions from eight participants across diverse touch and grasp types. Our results demonstrate that EgoPHI recovers meaningful 3D contact and force distributions in simulated, out-of-distribution, and real-world settings, advancing egocentric hand-object understanding from contact localization toward physically grounded interaction reasoning.