Search papers, labs, and topics across Lattice.
This study investigates the biases and failure modes of time series foundation models by employing a causal analysis framework that assesses how well these models preserve time series patterns. By intervening on parameterized synthetic time series generators, the authors measure changes in model outputs under controlled conditions, revealing significant insights into the models' behavior. Key findings include safe configurations for certain patterns and notable biases, such as an overestimation of persistence and failures in specific scenarios, which are linked to the pretraining data used for the models.
Time series foundation models could expose multiple applications to the same biases, but our causal analysis reveals critical failure modes that must be addressed before deployment.
Transitioning from bespoke time series models towards time series foundation models changes the relationship of model and application from one-to-one to one-to-many. This shift introduces concentration risk as many, potentially high-risk, forecasting applications are exposed to the same biases and failure modes of a single time series foundation model. At the same time, this centralization allows for economies of scale in model development and validation. In this study we investigate how biases and failure modes of time series foundation models can be identified before deployment. We propose a causal analysis framework to investigate the ability of a time series foundation model to preserve time series patterns. To achieve this, we intervene on parameterized synthetic time series generators and measure the corresponding change in model output under ceteris paribus conditions. We apply our causal analysis framework to Chronos-2 and TimesFM-2.5 and test them across six distinct time series patterns. We find safe configurations for trend and harmonic oscillation patterns. The results also indicate a bias in both models towards overestimating persistence, sudden failures for both models against the regime switch pattern and failure for TimesFM-2.5 against the energy-release pattern. Our review of the original works for both models indicates that the findings might be explained by the data used for pretraining. We conclude our study with suggestions for further model development, recommendations for application-specific model selection, and a discussion of limitations and further research directions.