Mar 16, 2026arXiv:2603.15248

Mechanistic Foundations of Goal-Directed Control

AI Summary

This paper extends mechanistic interpretability techniques to embodied control systems, using infant motor learning as a model. It demonstrates how inductive biases lead to causal control circuits with learned gating mechanisms that converge towards uncertainty thresholds. The study identifies the context window size as a critical parameter governing circuit formation and reveals a phase transition in the arbitration gate's commitment behavior, described by a closed-form exponential moving-average surrogate.

Key Contribution

Infant motor learning reveals a sharp phase transition in control strategy arbitration, governed by context window size and predictable via a closed-form exponential moving average.

Abstract

Mechanistic interpretability has transformed the analysis of transformer circuits by decomposing model behavior into competing algorithms, identifying phase transitions during training, and deriving closed-form predictions for when and why strategies shift. However, this program has remained largely confined to sequence-prediction architectures, leaving embodied control systems without comparable mechanistic accounts. Here we extend this framework to sensorimotor-cognitive development, using infant motor learning as a model system. We show that foundational inductive biases give rise to causal control circuits, with learned gating mechanisms converging toward theoretically motivated uncertainty thresholds. The resulting dynamics reveal a clean phase transition in the arbitration gate whose commitment behavior is well described by a closed-form exponential moving-average surrogate. We identify context window k as the critical parameter governing circuit formation: below a minimum threshold (k$\leq$4) the arbitration mechanism cannot form; above it (k$\geq$8), gate confidence scales asymptotically as log k. A two-dimensional phase diagram further reveals task-demand-dependent route arbitration consistent with the prediction that prospective execution becomes advantageous only when prediction error remains within the task tolerance window. Together, these results provide a mechanistic account of how reactive and prospective control strategies emerge and compete during learning. More broadly, this work sharpens mechanistic accounts of cognitive development and provides principled guidance for the design of interpretable embodied agents.

Interpretability & Mechanistic Interp Robotics & Embodied AI World Models & Planning

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Mechanistic Foundations of Goal-Directed Control

Related Papers