Search papers, labs, and topics across Lattice.
This paper investigates the effectiveness of test-time augmentation strategies for frozen vision-language-action (VLA) policies by introducing the concepts of recoverable headroom and retrieval complementarity. The authors demonstrate that a selector can significantly enhance policy performance by leveraging existing latent capabilities, achieving up to 21.0 success-rate points on LIBERO, while also showing the importance of external action priors when gaps exist. Their findings provide a framework for determining when to reuse internal policy behaviors versus when to retrieve external demonstrations, thus optimizing deployment strategies for embodied multimodal policies.
Unlocking up to 21% more success in embodied AI tasks by strategically deciding when to leverage existing policy behaviors versus retrieving external demonstrations.
Frozen vision-language-action (VLA) policies are increasingly improved at test time by sampling additional policy behaviors or introducing external demonstrations. Yet there is little guidance for deciding which intervention a deployed policy actually needs. Additional sampling is useful only when better behavior already exists within the policy's stochastic rollouts and can be identified, whereas retrieval is most useful when the relevant action prior is not reliably represented by the policy. We study this decision through two measurable factors, recoverable headroom and retrieval complementarity, which characterize how much useful behavior is already available to recover and whether an external action prior fills a measurable gap. We evaluate an episode-level retry selector under retryable or parallel execution, together with retrieval across multiple frozen VLA policies and environments. The selector consistently recovers substantial latent capability across all tested VLA backbones on LIBERO, with gains of up to 21.0 success-rate points that closely track recoverable headroom. It also transfers to a different robot and simulator and remains effective under degraded observations, while experiments with autoregressive OpenVLA illustrate the distinction between available headroom and the ability to rank candidate rollouts. Retrieval behaves differently, improving the policy with the largest measured action-prior gap and providing further gains when combined with selection. Together, these results provide an empirical basis for characterizing test-time augmentation opportunities by separating capability that can be recovered from the frozen policy from behavioral priors that may need to be introduced externally.