Search papers, labs, and topics across Lattice.
This paper introduces DyG$^2$T, a novel framework for modeling object dynamics that enhances motion trajectory prediction by addressing limitations in existing methods that oversimplify interactions and discard crucial local details. By spatially completing Key Point representations and employing a Temporal Disentangling Network to amplify inter-frame differences, DyG$^2$T captures fine-grained local details and long-range dependencies effectively. Experimental results show that DyG$^2$T outperforms traditional approaches in both synthetic and real-world scenarios, demonstrating superior accuracy in dynamics modeling and generalization across objects.
Achieving accurate trajectory predictions by recovering fine-grained details and capturing long-range dependencies could redefine how we model object dynamics in AI systems.
Modeling object dynamics from limited visual observations is a fundamental problem for enabling accurate motion trajectory prediction in embodied interaction scenarios. Existing dynamics modeling methods first compress reconstructed particle representations into sparse Key Points and model their evolution using locally constrained interactions, thereby discarding fine-grained local details and obscuring discriminative interaction modeling across spatial and temporal scales, leading to drifting trajectories and inaccurate appearance prediction. To tackle these issues, we propose DyG$^2$T, a dynamics modeling framework that infers object motion trajectories by spatially completing and temporally discriminating Key Point representations and modeling multi-scale interaction over particle graphs. Spatially, DyG$^2$T enriches each Key Point by aggregating neighboring raw particle positions to recover fine-grained local details, while explicitly encoding relative offsets among Key Points to enhance geometric structure perception. Temporally, we introduce a Temporal Disentangling Network (TDN) to identify dominant cross-frame variations in latent space and amplify inter-frame differences, yielding temporally discriminative representations that are subsequently aggregated via Temporal Attention to capture frame-wise temporal evolution cues. For comprehensive interaction modeling, a Particle Graph Transformer leverages global attention to preserve discriminative long-range dependencies among Key Points, mitigating representation homogenization induced by locality-constrained modeling and providing a robust basis for accurate trajectory prediction. Experiments on both synthetic and real-world datasets demonstrate that DyG$^2$T achieves accurate dynamics modeling and reasoning, and exhibits strong cross-object and real-world generalization.