Search papers, labs, and topics across Lattice.
This paper introduces LAVLA, a framework for latent cluster analysis specifically designed for Vision-Language-Action (VLA) models, focusing on the GR00T N1.5 model's action decoder. By employing a cross-attention-based embedding-weighting method, the authors enhance the identification of relevant features in the latent space, leading to superior clustering performance compared to baseline methods. The study reveals that latent clusters effectively disentangle spatiotemporal and kinematic features, improving the interpretability of language-driven robotic systems as representations evolve through the model layers.
Weighted clustering in VLA models reveals that latent representations become increasingly refined, enhancing our understanding of how robots interpret language-driven actions.
Vision-Language-Action (VLA) Models are increasingly used in robotics for their ability to ground language and perception into action, yet the internal representations driving their behaviour remain poorly understood. We propose LAVLA, a framework for latent cluster analysis of VLA models, and conduct a layer-wise study of the state-of-the-art GR00T N1.5 model, with particular focus on its action decoder. To better characterise the latent space during action diffusion, we introduce a cross-attention-based embedding-weighting method that amplifies relevant features while suppressing less informative ones. Quantitative evaluation shows that weighted clustering consistently outperforms the baseline. To improve interpretability, we extract human-interpretable concepts for each cluster, linking latent representations to semantic descriptions. Our analysis shows that latent clusters progressively disentangle spatiotemporal and kinematic features, with representations becoming more refined in the middle layers and stabilising toward the output. As such, LAVLA advances the interpretability of language-driven robotic systems.