Search papers, labs, and topics across Lattice.
This study critically evaluates the generalizability of power outage prediction models by identifying three methodological choices that significantly impact their performance across varying conditions. Using data from the U.S. East Coast, the authors demonstrate that traditional evaluation methods, particularly random train-test splits, inflate model performance due to spatial and temporal autocorrelation, leading to misleadingly optimistic results. Ultimately, the research reveals that models often fail to outperform a simple null baseline under more realistic testing conditions, highlighting the urgent need for improved data coverage and evaluation protocols in this domain.
Current power outage prediction models may be overhyped, as they often fail to generalize beyond inflated performance metrics derived from flawed evaluation methods.
Power outage prediction models are increasingly used in assessments of climate-driven infrastructure risk, yet current evaluation practices obscure whether these models generalize to the novel conditions such applications require. We identify three common methodological choices in power outage prediction models that influence their ability to generalize across spatial, temporal, and event-based settings. We compare the predictive performance impacts of different methodological decisions using publicly available data for the U.S. East Coast from 2018 to 2023 and feature sets derived from weather reanalysis and land-cover data, and embeddings from a GeoAI foundation model (Prithvi WxC). Specifically, we assess model performance under multiple test selection strategies, including unfiltered random splits, leave-one-state-out, and leave-one-event-out designs, which increasingly approximate real-world deployment conditions. While random train-test splits yield strong performance, we show that these results are inflated by spatial and temporal autocorrelation. Under spatial and temporal holdout experiments, predictive accuracy degrades substantially, with models often failing to outperform a simple null baseline. Incorporating GeoAI foundation model embeddings yields limited and inconsistent improvements, primarily for spatial generalization, and does not resolve poor event-level transferability. These findings suggest that, given current data availability and evaluation practices, publicly trained outage prediction models offer limited and uncertain operational value. Progress will likely require improved data coverage, more realistic evaluation protocols, and a shift in focus from marginal modeling advances toward addressing structural data constraints.