Search papers, labs, and topics across Lattice.
This paper introduces Gated VLA-Cache, an enhancement to existing KV-cache reuse methods in Vision-Language-Action models that incorporates neural introspection to improve decision-making during real-time control. By monitoring the logit margin between the top two predicted action tokens, the method dynamically invalidates the cache when uncertainty is detected, ensuring that accuracy is preserved without incurring significant computational costs. Evaluations on the LIBERO benchmark suites demonstrate that Gated VLA-Cache recovers over 100% of lost accuracy while maintaining 80% of the computational savings compared to traditional caching methods.
Gated VLA-Cache recovers lost accuracy in real-time control while slashing compute costs by leveraging model uncertainty.
Vision-Language-Action(VLA) models map camera images and language instructions directly to motor commands through a single autoregressive transformer. In real-time control, they still spend substantial compute recomputing key-value(KV) representations for visual tokens that barely change across neighboring frames. Recent work such as VLA-Cache reduces that cost by reusing KV states for visually static patches, but its policy relies only on observation-space heuristics and does not account for the model's own uncertainty. We propose Gated VLA-Cache, a lightweight, training-free extension that augments visual-similarity caching with neural introspection. The method monitors the logit margin between the top two predicted action tokens, a zero-cost confidence signal available during decoding. When the margin drops below a threshold, the cache is invalidated and a full recompute is triggered. Evaluated on four LIBERO benchmark suites with both OpenVLA and OpenVLA-OFT, Gated VLA-Cache improves reliability when blind caching hurts. On LIBERO-Goal and LIBERO-Long, it recovers over 100% of the lost accuracy while retaining 80% of the compute savings.