Apr 6, 2026arXiv:2604.04394

Finite-Time Analysis of Q-Value Iteration for General-Sum Stackelberg Games

AI Summary

This paper analyzes the convergence of Stackelberg Q-value iteration in two-player general-sum Markov games, framing the learning dynamics as a switching system. They introduce a relaxed policy condition and use comparison systems to derive finite-time error bounds for the Q-functions. The key result is the first finite-time convergence guarantee for Q-value iteration in general-sum Markov games under Stackelberg interactions.

Key Contribution

Stackelberg Q-learning finally gets finite-time guarantees, opening the door to more reliable multi-agent RL in complex, general-sum games.

Abstract

Reinforcement learning has been successful both empirically and theoretically in single-agent settings, but extending these results to multi-agent reinforcement learning in general-sum Markov games remains challenging. This paper studies the convergence of Stackelberg Q-value iteration in two-player general-sum Markov games from a control-theoretic perspective. We introduce a relaxed policy condition tailored to the Stackelberg setting and model the learning dynamics as a switching system. By constructing upper and lower comparison systems, we establish finite-time error bounds for the Q-functions and characterize their convergence properties. Our results provide a novel control-theoretic perspective on Stackelberg learning. Moreover, to the best of the authors' knowledge, this paper offers the first finite-time convergence guarantees for Q-value iteration in general-sum Markov games under Stackelberg interactions.

Robotics & Embodied AI Training Efficiency & Optimization

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Finite-Time Analysis of Q-Value Iteration for General-Sum Stackelberg Games

Related Papers