Search papers, labs, and topics across Lattice.
This paper develops a distributed bandit online feedback optimization algorithm tailored for multi-agent dynamical systems facing constrained inputs and time-varying cost functions. By utilizing a smoothing zeroth-order one-point estimator for local gradient approximations and integrating a projection-free conditional gradient update, the method circumvents the need for precise system models, enhancing its practical applicability. The authors establish a sublinear dynamic regret bound that reflects the system's non-stationarity, supported by numerical simulations that validate the algorithm's effectiveness in large-scale settings.
A novel optimization algorithm achieves sublinear dynamic regret in multi-agent systems without requiring accurate system models, leveraging real-time data instead.
This paper investigates distributed online optimization for multi-agent dynamical systems with constrained inputs and time-varying cost functions. While online convex optimization offers a principal framework for sequential decision-making, existing online learning and optimization algorithms typically require accurate system models, limiting their applicability in practical settings. To overcome this challenge, we propose a distributed bandit online feedback optimization algorithm that relies solely on real-time input-output data. The algorithm employs a smoothing zeroth-order one-point estimator to construct local gradient approximations directly from cost evaluations. Additionally, to enforce input constraints effectively, we integrate a projection-free conditional gradient update, making the algorithm well-suited for online and large-scale settings. Furthermore, we establish a sublinear dynamic regret bound that depends on a temporal variation measure of system non-stationarity. Finally, numerical simulations demonstrate the effectiveness of the proposed algorithm.