Time-Series Forecasting Evaluation Must Adopt Taxonomy-Specific Benchmarks
A new study on arXiv (2603.15506) takes a critical look at how we assess AI and machine learning models used for time-series forecasting. It points out that current evaluation methods rely heavily on benchmarks with strong seasonal patterns. The authors found that many of these datasets have significant autocorrelation, which simpler linear models can handle just as well as more complex deep learning approaches. They suggest that the field should either phase out or upgrade existing benchmarks to include datasets with various non-stationary features, like structural breaks and changing volatility. The researchers highlight that without evaluations tailored to these specific characteristics, the claimed advancements could be misleading, and the costs of complex models might not be justified.
Key facts
- Paper arXiv:2603.15506
- Published on arXiv
- Focus on AI/ML time-series forecasting
- Critiques current benchmarks for strong periodicities
- Shows classical models can match deep learning on such data
- Calls for new benchmarks with non-stationarities
- Mentions structural breaks and time-varying volatility
- Highlights computational overhead concerns
Entities
Institutions
- arXiv