Search papers, labs, and topics across Lattice.
SAFECAST enhances the robustness of vision-language-action (VLA) policies against deployment-time distribution shifts by employing contrast set perturbations for improved hidden-state probe training and calibration. This method addresses the critical challenge of failure detection in varying conditions, demonstrating statistically significant improvements in ROC-AUC scores compared to existing baselines in both real-world and simulated environments. The findings reveal that utilizing both visual and language contrast sets maximizes the effectiveness of the probes, outperforming traditional calibration methods reliant solely on real rollout data.
Contrast set perturbations can dramatically boost failure detection in vision-language-action systems, outperforming conventional calibration techniques.
Vision-language-action policies often fail under deployment-time distribution shifts such as clutter, distractor objects, lighting changes, novel objects, altered initial states, and reworded instructions. Hidden-state-based risk probes combined with functional conformal prediction can detect rollout failures, but their reliability depends on calibration data matching deployment conditions. We introduce SAFECAST, which leverages contrast set perturbations to improve hidden-state probe training and calibration for deployment time shift. SAFECAST statistically significantly improves failure detection ROC-AUC scores over a state of the art baseline in both real-world DROID and LIBERO simulation experiments across multiple VLM backbones. We further find that SAFECAST benefits most when both visual and language contrast set perturbations are used to augment data, and that with contrast set perturbations, sim-to-real calibration leads to better probes than using real rollout data only.