Search papers, labs, and topics across Lattice.
This survey systematically reviews the integration of foundation model priors into hand-object interaction (HOI) modeling, addressing the challenges of joint reasoning under visual uncertainty. By categorizing existing literature into six tasks and establishing a taxonomy of eight sub-priors鈥攇eometric, semantic, and visual鈥攖he authors clarify how these models enhance HOI pipelines. The findings highlight the diverse applications of HOI-derived knowledge in robot learning, paving the way for more generalizable systems and improved methodologies in the field.
Foundation models can revolutionize hand-object interaction by systematically categorizing and leveraging diverse priors to enhance robot learning and task performance.
Hand-object interaction (HOI) modeling remains challenging because it requires joint reasoning about hand articulation, object geometry, contact, semantics, and dynamics under severe visual uncertainty. Foundation models introduce transferable prior knowledge learned from large-scale cross-domain data, offering new ways to address these challenges beyond task-specific data and models. However, the rapidly growing literature remains fragmented, and existing studies typically describe these methods simply as ``using large models''without systematically characterizing what knowledge is introduced, where it enters the HOI pipeline, or which HOI uncertainty it helps reduce. This survey presents the first systematic review of foundation-model priors for HOI. We organize the literature into six HOI tasks spanning reconstruction and generation. More importantly, we establish a taxonomy of eight foundation-model sub-priors grouped into geometric, semantic, and visual families. Geometric priors encompass shape retrieval, shape reconstruction, and spatial reconstruction; semantic priors include semantic grounding and language reasoning; and visual priors cover visual representation, image generation, and video generation. Based on this taxonomy, we systematically analyze how different priors are represented, injected, and adapted across HOI pipelines and tasks. Beyond how foundation models empower HOI, we further examine how HOI-derived knowledge is used in robot learning, including human-data pretraining, human-to-robot skill transfer, and HOI-to-robot data generation. Finally, we summarize datasets and evaluation protocols, and discuss limitations and future directions toward more generalizable HOI systems. To support long-term progress, we curate a live repository that continuously aggregates emerging methods and benchmarks.