BITCASGigaAIApr 20, 2026arXiv:2604.17887

StableIDM: Stabilizing Inverse Dynamics Model against Manipulator Truncation via Spatio-Temporal Refinement

Kerui Li, Zhe Jing, Xiaofeng Wang, Zheng Zhu, Yukun Zhou, Guan Huang, Dongze Li, Qingkai Yang, Huaibo Huang

AI Summary

This paper introduces StableIDM, a spatio-temporal framework designed to enhance the robustness of Inverse Dynamics Models (IDMs) against manipulator truncation by refining features from visual inputs. StableIDM uses robot-centric masking, Directional Feature Aggregation (DFA) for geometry-aware spatial reasoning, and Temporal Dynamics Refinement (TDR) to smooth action predictions. Experiments on the AgiBot benchmark and real-robot replay demonstrate that StableIDM significantly improves action accuracy, task success, and end-to-end grasp success under severe truncation.

Key Contribution

Even with partial observability from manipulator truncation, StableIDM recovers stable action predictions, boosting downstream VLA real-robot success by 17.6%.

Abstract

Inverse Dynamics Models (IDMs) map visual observations to low-level action commands, serving as central components for data labeling and policy execution in embodied AI. However, their performance degrades severely under manipulator truncation, a common failure mode that makes state recovery ill-posed and leads to unstable control. We present StableIDM, a spatio-temporal framework that refines features from visual inputs to stabilize action predictions under such partial observability. StableIDM integrates three complementary components: (1) auxiliary robot-centric masking to suppress background clutter, (2) Directional Feature Aggregation (DFA) for geometry-aware spatial reasoning, which extracts anisotropic features along directions inferred from the visible arm and (3) Temporal Dynamics Refinement (TDR) to smooth and correct predictions via motion continuity. Extensive evaluations validate our approach: StableIDM improves strict action accuracy by 12.1% under severe truncation on the AgiBot benchmark, and increases average task success by 9.7% in real-robot replay. Moreover, it boosts end-to-end grasp success by 11.5% when decoding video-generated plans, and improves downstream VLA real-robot success by 17.6% when functioning as an automatic annotator. These results demonstrate that StableIDM provides a robust and scalable backbone for both policy execution and data generation in embodied artificial intelligence.

Computer Vision Multimodal Models Robotics & Embodied AI World Models & Planning

Citation Metrics

Citations0

Influential citations0

References55

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

StableIDM: Stabilizing Inverse Dynamics Model against Manipulator Truncation via Spatio-Temporal Refinement

Related Papers