Search papers, labs, and topics across Lattice.
To predict 3D displacement vectors from 21 hand joints to manipulated object surfaces from egocentric stereo video, the authors developed GOLF, the winning solution to the ECCV 2026 SHOW3D Interaction Field Estimation Challenge. The framework integrates dense visual representations with localized hand-object sampling and explicit common-frame Pl眉cker-ray geometry using a LoRA-adapted DINOv3 ViT-H+/16 backbone. Combining the parameter-efficient model with a fully fine-tuned variant in an equal-weight ensemble achieved a winning average displacement error (ADE) of 27.82 mm on the hidden benchmark.
Injecting explicit Pl眉cker-ray geometry and local feature sampling into foundation vision backbones drives egocentric 3D hand-object interaction error down to sub-3cm precision.
We present GOLF, the first-place solution to the SHOW3D Interaction Field Estimation Challenge at HANDS@ECCV 2026. Given synchronized egocentric stereo views, the task is to predict a 3D vector from each of 21 hand joints to the closest point on the manipulated object. GOLF combines dense global context, locally sampled hand/object evidence, and common-frame Pl眉cker-ray geometry. We adapt DINOv3 ViT-H+/16 with LoRA and trainable LayerNorm parameters, then jointly decode both interaction fields. Our primary model achieves an official score of 27.61 and a mean ADE of 27.96 mm on the hidden test set. An equal-weight ensemble with a complementary directly fine-tuned variant improves these results to an official score of 27.47 and a mean ADE of 27.82 mm, securing first place.