Search papers, labs, and topics across Lattice.
This paper introduces PoseOFF, a novel pose-anchored optical flow representation designed to enhance early human action anticipation in human-robot interaction. By focusing on localized motion dynamics around human joints, PoseOFF improves action recognition accuracy while significantly reducing the computational load associated with full-frame optical flow processing. The results demonstrate that PoseOFF allows for effective early predictions, achieving comparable or improved performance with less observed action sequences, making it suitable for real-time applications in resource-constrained environments.
PoseOFF captures critical motion cues around human joints, enabling robots to anticipate actions with less data and faster response times.
Human-robot interaction (HRI) requires robots to interpret human actions early in their execution in order to respond safely, efficiently, and naturally. However, many existing approaches to human action recognition rely either on sparse skeletal representations, which lack fine-grained motion cues, or dense optical flow, which can be computationally expensive for low-latency perception pipelines. In this paper, we propose PoseOFF, a pose-anchored optical flow representation that captures local motion information around human joints to support earlier human intent understanding. By conditioning motion feature extraction on human pose, PoseOFF encodes localised motion dynamics at semantically meaningful body locations, forming a structured motion representation that is explicitly aligned with human kinematics. We evaluate PoseOFF across multiple benchmark datasets and backbone architectures for action anticipation, demonstrating consistent improvements in recognition accuracy, particularly at early observation ratios. Our results show that PoseOFF enables models to achieve comparable or improved performance while observing less of the action sequence, highlighting its effectiveness for early prediction. Importantly, these gains are achieved without requiring full-frame motion processing, making the approach practical for real-time and resource-constrained settings. These findings suggest that pose-centred motion representations such as PoseOFF can enhance the ability of interactive robot systems to infer human actions earlier, supporting more responsive and anticipatory behaviour in human-robot interaction scenarios.