Search papers, labs, and topics across Lattice.
This study critically evaluates generative time-series models on datasets characterized by point masses, revealing significant discrepancies in model performance due to the standard rolling-origin evaluation protocol. The authors demonstrate that this protocol can yield misleading results, as evidenced by a case where the strongest model appeared weak when evaluated under mismatched conditions. Through a controlled analysis, they find that an autoregressive hurdle model consistently outperforms a conditional flow model across multiple datasets, highlighting the instability of flow model statistics across training seeds.
The standard evaluation protocol for generative time-series models can drastically misrepresent model performance, turning top performers into cautionary tales.
Many of the series that generative time-series models are benchmarked on place a large probability mass on a single value --- it does not rain, no ride is requested, no part is ordered. We report what happens when such data is evaluated carefully. First, the standard rolling-origin protocol can score a model on a window whose atom structure bears no resemblance to the dataset: on one benchmark the dataset is $42\%$ zeros and the evaluation windows are $13\%$, on another $47\%$ against $5\%$. This is not a cosmetic problem --- it reversed one of our own conclusions, turning the strongest occurrence model in our study into what looked like a cautionary tale. Second, we give a control in which CRPS is invariant \emph{by construction} while the temporal coupling is destroyed, which measures exactly how much that coupling contributes to a chosen statistic. Third, benchmarking seven models on a matched protocol over five seeds, an autoregressive hurdle beats a conditional flow on five of six datasets, by up to a factor of $153$, while the flow's own occurrence statistics vary by up to $62\%$ across training seeds and every baseline is deterministic. Finally, the model ordering is not the same under five different occurrence statistics, and the two that do not share a construction agree with each other least.