Search papers, labs, and topics across Lattice.
This study integrates distributed reinforcement learning (RL) agents with the Met Office Unified Model (UM) to enhance numerical weather prediction through adaptive machine-learnt corrections. By employing a Deep Deterministic Policy Gradient (DDPG) approach with rank-local tensors, the model achieves significant reductions in forecast errors, including a 45.8% decrease in Z$_{500}$ mean absolute error in the tropics. The findings indicate that this coupled workflow not only maintains numerical stability but also demonstrates the feasibility of implementing RL for bias correction in operational forecasting systems.
Reinforcement learning can cut weather forecast errors by up to 45% while maintaining numerical stability in operational models.
Machine-learnt corrections can complement numerical weather prediction only if they adapt to the evolving model state while preserving dynamical consistency and numerical stability. To test this within a global forecasting model, we couple the Met Office (UKMO) Unified Model (UM) with distributed RL agents through rank-local tensors. A DDPG actor shares weights across the 70 vertical model levels of each atmospheric column and applies bounded potential-temperature corrections to the model tendencies. Across ten nudged training forecasts, nudging calculations towards the UKMO operational analysis provides an immediate counterfactual target. The frozen policy is then evaluated in a non-nudged forecast for inference. The coupled workflow successfully completes training and remains numerically stable in the evaluated case. Relative to a matched native UM forecast at +6 h, the learnt policy reduces Z$_{500}$ MAE in four of six latitude bands, including reductions of 45.8% and 40.8% in the northern and southern tropics. MSLP error too decreases in three bands, with a maximum reduction of 27.3% at 0-30掳N. This single-case experiment demonstrates significant promise and feasibility of distributed online learning followed by non-nudged inference, laying the groundwork for RL-based bias correction and parametrisations within operational systems.