Search papers, labs, and topics across Lattice.
This paper introduces a novel approach to constructing scenario trees for multistage stochastic model predictive control (MPC) by directly optimizing the tree based on its impact on downstream control decisions. Unlike traditional methods that focus on matching probability distributions, the proposed method employs reinforcement learning to parameterize scenario assignments, leading to improved control performance and robustness in decision-making. Evaluated on a battery arbitrage problem, the approach consistently outperforms classical methods, demonstrating superior profit and tail-risk characteristics through the construction of compact, decision-supportive scenario trees.
Control-oriented scenario trees can significantly enhance decision-making in uncertain environments, achieving higher profits and better risk management than traditional methods.
Multistage stochastic model predictive control (MPC) handles uncertainty by optimizing over a scenario tree, a finite branching approximation of future outcomes constructed from sampled forecasts. To build such a tree, conventional methods focus on matching the underlying probability distribution---e.g., via Wasserstein-based scenario reduction---but improved distributional accuracy does not necessarily yield better control performance. We propose a control-oriented approach that learns scenario tree construction directly from its impact on downstream decisions. Fixing the tree topology, we formulate tree construction as a sequential assignment of sampled scenarios to leaves. This assignment is parameterized by an attention-based policy over the scenario set and trained using reinforcement learning, with closed-loop control profit as the objective. Training is stabilized by an asymmetric critic that leverages realized future trajectories. We evaluate the method on a risk-averse battery arbitrage problem. Across a range of forecast set sizes, the learned construction consistently achieves the highest profit, outperforming classical forward and backward reduction methods and certainty-equivalent (single-trajectory forecast) control. The learned policy also exhibits greater robustness on challenging instances, consistently demonstrating better tail-risk characteristics. Analysis of the resulting trees indicates that our method constructs compact, selectively branching structures that capture high-impact events while keeping most trajectories nearly deterministic. These findings highlight that the value of a scenario tree depends critically on the decisions it supports, and provide an effective framework to train scenario tree constructors merely based on the closed-loop control optimization signal.