Search papers, labs, and topics across Lattice.
This paper introduces GARFIELD, a novel probabilistic model that captures the distribution of potential scene kinematics from partial observations, allowing for more nuanced predictions of future motion. Unlike existing methods that either focus on appearance or sample limited trajectories, GARFIELD utilizes a structured spatio-temporal latent representation to jointly sample all possible trajectories while providing direct access to motion distributions. Experimental results show that GARFIELD achieves competitive motion planning performance with significantly faster trajectory sampling and motion density estimation, thus enhancing interactive exploration and uncertainty-aware planning capabilities.
GARFIELD enables real-time, uncertainty-aware motion planning by efficiently modeling the distribution of possible scene futures, outperforming traditional methods by a factor of 97 in trajectory sampling speed.
Predicting how a scene may evolve from partial observations requires reasoning about multiple possible futures rather than committing to a single trajectory. Existing approaches either generate appearance-dominated video predictions or sample a small number of trajectories without explicitly modeling the distribution of possible motion. We introduce Goal-Aware Representations of Future kInEmatic Latent Distributions (GARFIELD), a probabilistic model of scene kinematics that learns a structured spatio-temporal latent representation of the distribution over possible futures given an image and optional spatio-temporally sparse constraints. The same latent representation enables both joint sampling of all trajectories and direct access to the underlying motion distribution through an efficient deterministic density decoder. As a result, uncertainty about future motion can be localized to specific scene elements and timesteps and progressively refined through additional constraints. Experiments demonstrate strong motion planning performance competitive with large video generation models while sampling trajectories $97\times$ faster. Our method further estimates motion densities two orders of magnitude faster than Monte-Carlo sampling from motion generation models, enabling interactive exploration and uncertainty-aware planning.