A survey of deep time series forecasting backbone architectures: progress, pitfalls, and a systematic comparison

Xiang LI , Yanping ZHENG , Zhewei WEI

Front. Comput. Sci. ›› 2026, Vol. 20 ›› Issue (8) : 2008371

PDF (13248KB)
Front. Comput. Sci. ›› 2026, Vol. 20 ›› Issue (8) :2008371 DOI: 10.1007/s11704-026-50462-z
Excellent Young Computer Scientists Forum
REVIEW ARTICLE
A survey of deep time series forecasting backbone architectures: progress, pitfalls, and a systematic comparison
Author information +
History +
PDF (13248KB)

Abstract

Deep learning-based time series forecasting has largely converged to standardized training protocols and a narrow set of public benchmarks. While such standardization improves comparability, evaluations based on averaged metrics over fixed windows obscure variable-level differences, mask long-horizon degradation, and diverge from real-world rolling-forecasting scenarios. This study revisits these limitations by surveying seven major backbone architectures and systematically evaluating 21 representative models across diverse datasets. A fine-grained variable-level analysis shows that, in several settings, extending the input window contributes more to forecasting accuracy than architectural innovations, yet such gains diminish rapidly and may even become detrimental as the window grows excessively. Furthermore, models with strong average performance often behave inconsistently across variables and differ markedly in their ability to capture short- and long-term temporal dynamics. These findings highlight inherent constraints on predictability and call for new research directions, including domain-specific end-to-end forecasting pipelines, forecasting with exogenous drivers, and the development of large time series models.

Graphical abstract

Keywords

time series forecasting / backbone architectures / variable-level analysis

Cite this article

Download citation ▾
Xiang LI, Yanping ZHENG, Zhewei WEI. A survey of deep time series forecasting backbone architectures: progress, pitfalls, and a systematic comparison. Front. Comput. Sci., 2026, 20 (8) : 2008371 DOI:10.1007/s11704-026-50462-z

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Shao Z, Wang F, Xu Y, Wei W, Yu C, Zhang Z, Yao D, Sun T, Jin G, Cao X, Cong G, Jensen C S, Cheng X . Exploring progress in multivariate time series forecasting: comprehensive benchmarking and heterogeneity analysis. IEEE Transactions on Knowledge and Data Engineering, 2025, 37( 1): 291–305

[2]

Qiu X, Hu J, Zhou L, Wu X, Du J, Zhang B, Guo C, Zhou A, Jensen C S, Sheng Z, Yang B . TFB: towards comprehensive and fair benchmarking of time series forecasting methods. Proceedings of the VLDB Endowment, 2024, 17( 9): 2363–2377

[3]

Wang Y, Wu H, Dong J, Liu Y, Wang C, Long M, Wang J. Deep time series models: a comprehensive survey and benchmark. 2024, arXiv preprint arXiv: 2407.13278

[4]

Wen Q, Zhou T, Zhang C, Chen W, Ma Z, Yan J, Sun L. Transformers in time series: a survey. In: Proceedings of the 32nd International Joint Conference on Artificial Intelligence. 2023

[5]

Jiang J, Han C, Jiang W, Zhao W X, Wang J. LibCity: a unified library towards efficient and comprehensive urban spatial-temporal prediction. 2024, arXiv preprint arXiv: 2304.14343

[6]

Zhang J, Wen X, Zhang Z, Zheng S, Li J, Bian J. ProbTS: benchmarking point and distributional forecasting across diverse prediction horizons. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. 2024

[7]

Li Z, Qiu X, Chen P, Wang Y, Cheng H, Shu Y, Hu J, Guo C, Zhou A, Wen Q, Jensen C S, Yang B. FoundTS: comprehensive and unified benchmarking of foundation models for time series forecasting. 2024, arXiv preprint arXiv: 2410.11802v1

[8]

Roque L, Cerqueira V, Soares C, Torgo L. Cherry-picking in time series forecasting: how to select datasets to make your model shine. In: Proceedings of the 39th AAAI Conference on Artificial Intelligence. 2025

[9]

Brigato L, Morand R, Strømmen K, Panagiotou M, Schmidt M, Mougiakakou S. Position: there are no champions in long-term time series forecasting. 2025, arXiv preprint arXiv: 2502.14045v1

[10]

Abdelmalak I, Madhusudhanan K, Kloetergens C, Yalavarit V K, Stubbemann M, Schmidt-Thieme L. Channel dependence, limited lookback windows, and the simplicity of datasets: how biased is time series forecasting? 2025, arXiv preprint arXiv: 2502.09683

[11]

Kudrat D, Xie Z, Sun Y, Jia T, Hu Q. Patch-wise structural loss for time series forecasting. In: Proceedings of the 42nd International Conference on Machine Learning. 2025

[12]

Liu Z, Cheng M, Zhao G, Yang J, Liu Q, Chen E. Improving time series forecasting via instance-aware post-hoc revision. 2025, arXiv preprint arXiv: 2505.23583

[13]

Fu Y, Shao Z, Yu C, Li Y, An Z, Wang C, Xu Y, Wang F. Selective learning for deep time series forecasting. In: Proceedings of the 39th Conference on Neural Information Processing Systems. 2025

[14]

Oreshkin B N, Carpov D, Chapados N, Bengio Y. N-BEATS: neural basis expansion analysis for interpretable time series forecasting. In: Proceedings of the 8th International Conference on Learning Representations. 2020

[15]

Olivares K G, Challu C, Marcjasz G, Weron R, Dubrawski A . Neural basis expansion analysis with exogenous variables: forecasting electricity prices with NBEATSx. International Journal of Forecasting, 2023, 39( 2): 884–900

[16]

Challu C, Olivares K G, Oreshkin B N, Ramirez F G, Canseco M M, Dubrawski A. NHITS: neural hierarchical interpolation for time series forecasting. In: Proceedings of the 37th AAAI Conference on Artificial Intelligence. 2023, 6989−6997

[17]

Fan W, Zheng S, Yi X, Cao W, Fu Y, Bian J, Liu T Y. DEPTS: deep expansion learning for periodic time series forecasting. In: Proceedings of the 10th International Conference on Learning Representations. 2022

[18]

Lin S, Lin W, Hu X, Wu W, Mo R, Zhong H. CycleNet: enhancing time series forecasting through modeling periodic patterns. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. 2024, 3373

[19]

Xu Z, Zeng A, Xu Q. FITS: modeling time series with 10k parameters. 2023, arXiv preprint arXiv: 2307.03756

[20]

Wang S, Wu H, Shi X, Hu T, Luo H, Ma L, Zhang J Y, Zhou J. TimeMixer: decomposable multiscale mixing for time series forecasting. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[21]

Ekambaram V, Jati A, Nguyen N, Sinthong P, Kalagnanam J. TSMixer: lightweight MLP-mixer model for multivariate time series forecasting. In: Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 2023, 459−469

[22]

Lu H, Chen X Y, Ye H J, Zhan D C. SOFTS: efficient multivariate time series forecasting with series-core fusion. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. 2024, 2046

[23]

Zeng A, Chen M, Zhang L, Xu Q. Are transformers effective for time series forecasting? In: Proceedings of the 37th AAAI Conference on Artificial Intelligence. 2023, 11121−11128

[24]

Li Z, Qi S, Li Y, Xu Z. Revisiting long-term time series forecasting: an investigation on linear mapping. 2023, arXiv preprint arXiv: 2305.10721

[25]

Lin S, Lin W, Wu W, Chen H, Yang J. SparseTSF: modeling long-term time series forecasting with 1k parameters. In: Proceedings of the 41st International Conference on Machine Learning. 2024

[26]

Yi K, Fei J, Zhang Q, He H, Hao S, Lian D, Fan W. FilterNet: harnessing frequency filters for time series forecasting. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. 2024

[27]

Li Z, Li H, Wang H, Fang J, Tan Y, Qin X C Y. FSMLP: modelling channel dependencies with simplex theory based multi-layer perceptions in frequency domain. 2024, arXiv preprint arXiv: 2412.01654

[28]

Fei J, Yi K, Fan W, Zhang Q, Niu Z. Amplifier: bringing attention to neglected low-energy components in time series forecasting. In: Proceedings of the 39th AAAI Conference on Artificial Intelligence. 2025

[29]

Bai S, Kolter J Z, Koltun V. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. 2018, arXiv preprint arXiv: 1803.01271

[30]

Lai G, Chang W C, Yang Y, Liu H. Modeling long-and short-term temporal patterns with deep neural networks. In: Proceedings of the 41st International ACM SIGIR Conference on Research & Development in Information Retrieval. 2018, 95−104

[31]

Wang H, Peng J, Huang F, Wang J, Chen J, Xiao Y. MICN: multi-scale local and global context modeling for long-term series forecasting. In: Proceedings of the 11th International Conference on Learning Representations. 2023

[32]

Liu M, Zeng A, Chen M, Xu Z, Lai Q, Ma L, Xu Q. SCINet: time series modeling and forecasting with sample convolution and interaction. In: Proceedings of the 36th Conference on Neural Information Processing Systems. 2022

[33]

Wu H, Hu T, Liu Y, Zhou H, Wang J, Long M. TimesNet: temporal 2D-variation modeling for general time series analysis. In: Proceedings of the 11th International Conference on Learning Representations. 2023

[34]

Luo D, Wang X. ModernTCN: a modern pure convolution structure for general time series analysis. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[35]

Gong Z, Tang Y, Liang J. PatchMixer: a patch-mixing architecture for long-term time series forecasting. 2024, arXiv preprint arXiv: 2310.00655

[36]

Cheng M, Yang J, Pan T, Liu Q, Li Z, Wang S. ConvTimenet: a deep hierarchical fully convolutional model for multivariate time series analysis. In: Proceedings of the ACM on Web Conference 2025. 2025

[37]

Chen M, Shen L, Li Z, Wang X J, Sun J, Liu C. VisionTS: visual masked autoencoders are free-lunch zero-shot time series forecasters. In: Proceedings of the 42nd International Conference on Machine Learning. 2025

[38]

Hewamalage H, Bergmeir C, Bandara K . Recurrent Neural Networks for Time Series Forecasting: current status and future directions. International Journal of Forecasting, 2021, 37( 1): 388–427

[39]

Hochreiter S, Schmidhuber J . Long short-term memory. Neural Computation, 1997, 9( 8): 1735–1780

[40]

Chung J, Gulcehre C, Cho K, Bengio Y. Empirical evaluation of gated recurrent neural networks on sequence modeling. 2014, arXiv preprint arXiv: 1412.3555

[41]

Koutník J, Greff K, Gomez F, Schmidhuber J. A clockwork RNN. In: Proceedings of the 31st International Conference on Machine Learning. 2014, 1863−1871

[42]

Neil D, Pfeiffer M, Liu S C. Phased LSTM: accelerating recurrent network training for long or event-based sequences. In: Proceedings of the 30th International Conference on Neural Information Processing Systems. 2016

[43]

Li S, Li W, Cook C, Zhu C, Gao Y. Independently recurrent neural network (IndRNN): building a longer and deeper RNN. In: Proceedings of 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2018, 5457−5466

[44]

Jia Y, Lin Y, Hao X, Lin Y, Guo S, Wan H. WITRAN: water-wave information transmission and recurrent acceleration network for long-range time series forecasting. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. 2023, 544

[45]

Jia Y, Lin Y, Yu J, Wang S, Liu T, Wan H. PGN: the RNN’s new successor is effective for long-range time series forecasting. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. 2024, 2674

[46]

Gu A, Dao T. Mamba: linear-time sequence modeling with selective state spaces. 2023, arXiv preprint arXiv: 2312.00752

[47]

Nie X, Zhou X, Li Z, Wang L, Lin X, Tong T. LogTrans: providing efficient local-global fusion with transformer and CNN parallel network for biomedical image segmentation. In: Proceedings of the 24th IEEE International Conference on High Performance Computing & Communications; 8th International Conference on Data Science & Systems; 20th International Conference on Smart City; 8th International Conference on Dependability in Sensor, Cloud & Big Data Systems & Application. 2022, 769−776

[48]

Zhou H, Zhang S, Peng J, Zhang S, Li J, Xiong H, Zhang W. Informer: beyond efficient transformer for long sequence time-series forecasting. In: Proceedings of the 35th AAAI Conference on Artificial Intelligence. 2021

[49]

Wu H, Xu J, Wang J, Long M. Autoformer: decomposition transformers with auto-correlation for long-term series forecasting. In: Proceedings of the 35th International Conference on Neural Information Processing Systems. 2021, 1717

[50]

Zhou T, Ma Z, Wen Q, Wang X, Sun L, Jin R. FEDformer: frequency enhanced decomposed transformer for long-term series forecasting. In: Proceedings of the 39th International Conference on Machine Learning. 2022, 27268−27286

[51]

Liu S, Yu H, Liao C, Li J, Lin W, Liu A X, Dustdar S. Pyraformer: low-complexity pyramidal attention for long-range time series modeling and forecasting. In: Proceedings of the 10th International Conference on Learning Representations. 2022

[52]

Kim T, Kim J, Tae Y, Park C, Choi J H, Choo J. Reversible instance normalization for accurate time-series forecasting against distribution shift. In: Proceedings of the 10th International Conference on Learning Representations. 2022

[53]

Zhao Y, Ma Z, Zhou T, Ye M, Sun L, Qian Y. GCformer: an efficient solution for accurate and scalable long-term multivariate time series forecasting. In: Proceedings of the 32nd ACM International Conference on Information and Knowledge Management. 2023

[54]

Yu C, Wang F, Shao Z, Sun T, Wu L, Xu Y. DSformer: a double sampling transformer for multivariate time series long-term prediction. In: Proceedings of the 32nd ACM International Conference on Information and Knowledge Management. 2023, 3062−3072

[55]

Ilbert R, Odonnat A, Feofanov V, Virmaux A, Paolo G, Palpanas T, Redko I. SAMformer: unlocking the potential of transformers in time series forecasting with sharpness-aware minimization and channel-wise attention. In: Proceedings of the 41st International Conference on Machine Learning. 2024

[56]

Liu Y, Wu H, Wang J, Long M. Non-stationary transformers: exploring the stationarity in time series forecasting. In: Proceedings of the 36th Conference on Neural Information Processing Systems. 2022

[57]

Nie Y, Nguyen N H, Sinthong P, Kalagnanam J. A time series is worth 64 words: long-term forecasting with transformers. In: Proceedings of the 11th International Conference on Learning Representations. 2023

[58]

Liu Y, Hu T, Zhang H, Wu H, Wang S, Ma L, Long M. iTransformer: inverted transformers are effective for time series forecasting. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[59]

Zhang Y, Yan J. Crossformer: transformer utilizing cross-dimension dependency for multivariate time series forecasting. In: Proceedings of the 11th International Conference on Learning Representations. 2023

[60]

Wang X, Zhou T, Wen Q, Gao J, Ding B, Jin R. CARD: channel aligned robust blend transformer for time series forecasting. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[61]

Piao X, Chen Z, Murayama T, Matsubara Y, Sakurai Y. Fredformer: frequency debiased transformer for time series forecasting. In: Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 2024, 2400−2410

[62]

Kim D, Park J, Lee J, Kim H. Are self-attentions effective for time series forecasting? In: Proceedings of the 38th International Conference on Neural Information Processing Systems. 2024, 3627

[63]

Chen P, Zhang Y, Cheng Y, Shu Y, Wang Y, Wen Q, Yang B, Guo C. Pathformer: multi-scale transformers with adaptive pathways for time series forecasting. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[64]

Gu A, Goel K, C. Efficiently modeling long sequences with structured state spaces. In: Proceedings of the 10th International Conference on Learning Representations. 2022

[65]

Montgomery D C, Jennings C L, Kulahci M. Introduction to Time Series Analysis and Forecasting. 2nd ed. Hoboken: John Wiley & Sons, 2015

[66]

Rangapuram S S, Seeger M, Gasthaus J, Stella L, Wang Y, Januschowski T. Deep state space models for time series forecasting. In: Proceedings of the 32nd International Conference on Neural Information Processing Systems. 2018

[67]

Gu A, Dao T, Ermon S, Rudra A, C. HiPPO: recurrent memory with optimal polynomial projections. In: Proceedings of the 34th International Conference on Neural Information Processing Systems. 2020, 125

[68]

Gu A, Johnson I, Goel K, Saab K, Dao T, Rudra A, C. Combining recurrent, convolutional, and continuous-time models with linear state-space layers. In: Proceedings of the 35th International Conference on Neural Information Processing Systems. 2021, 44

[69]

Gu A, Gupta A, Goel K, C. On the parameterization and initialization of diagonal state space models. In: Proceedings of the 36th International Conference on Neural Information Processing Systems. 2022, 2607

[70]

Gu A, Johnson I, Timalsina A, Rudra A, C. How to train your HIPPO: state space models with generalized orthogonal basis projections. In: Proceedings of the 11th International Conference on Learning Representations. 2023

[71]

Gupta A, Gu A, Berant J. Diagonal state spaces are as effective as structured state spaces. In: Proceedings of the 36th International Conference on Neural Information Processing Systems. 2022, 1670

[72]

Wang Z, Kong F, Feng S, Wang M, Yang X, Zhao H, Wang D, Zhang Y. Is Mamba effective for time series forecasting? Neurocomputing, 2025, 619: 129178

[73]

Liang A, Jiang X, Sun Y, Shi X, Li K. Bi-mamba+: bidirectional mamba for time series forecasting. 2024, arXiv preprint arXiv: 2404.15772

[74]

Kipf T N, Welling M. Semi-supervised classification with graph convolutional networks. In: Proceedings of the 5th International Conference on Learning Representations. 2017

[75]

Veličković P, Cucurull G, Casanova A, Romero A, Liò P, Bengio Y. Graph attention networks. In: Proceedings of the 6th International Conference on Learning Representations. 2018

[76]

Yu B, Yin H, Zhu Z. Spatio-temporal graph convolutional networks: a deep learning framework for traffic forecasting. In: Proceedings of the 27th International Joint Conference on Artificial Intelligence. 2018

[77]

Li Y, Yu R, Shahabi C, Liu Y. Diffusion convolutional recurrent neural network: data-driven traffic forecasting. In: Proceedings of the 6th International Conference on Learning Representations. 2018, 1−16

[78]

Wu Z, Pan S, Long G, Jiang J, Zhang C. Graph wavenet for deep spatial-temporal graph modeling. In: Proceedings of the 28th International Joint Conference on Artificial Intelligence. 2019, 1907−1913

[79]

Wu Z, Pan S, Long G, Jiang J, Chang X, Zhang C. Connecting the dots: multivariate time series forecasting with graph neural networks. In: Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 2020

[80]

Cai W, Liang Y, Liu X, Feng J, Wu Y. MSGNet: learning multi-scale inter-series correlations for multivariate time series forecasting. In: Proceedings of the 38th AAAI Conference on Artificial Intelligence. 2024, 11141−11149

[81]

Cai W, Wang K, Wu H, Chen X, Wu Y. ForecastGrapher: redefining multivariate time series forecasting with graph neural networks. 2024, arXiv preprint arXiv: 2405.18036

[82]

Coskunuzer B, Segovia-Dominguez I, Chen Y, Gel Y R. Time-aware knowledge representations of dynamic objects with multidimensional persistence. In: Proceedings of the 38th AAAI Conference on Artificial Intelligence. 2024, 11678−11686

[83]

Liu Y, Liu Q, Zhang J W, Feng H, Wang Z, Zhou Z, Chen W. Multivariate time-series forecasting with temporal polynomial graph neural networks. In: Proceedings of the 36th International Conference on Neural Information Processing Systems. 2022, 1411

[84]

Chen X, Li X, Chen X, Li Z. Structured matrix basis for multivariate time series forecasting with interpretable dynamics. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. 2024, 766

[85]

Yang L, Zhang Z, Song Y, Hong S, Xu R, Zhao Y, Zhang W, Cui B, Yang M H . Diffusion models: a comprehensive survey of methods and applications. ACM Computing Surveys, 2024, 56( 4): 105

[86]

Lin L, Li Z, Li R, Li X, Gao J . Diffusion models for time-series applications: a survey. Frontiers of Information Technology & Electronic Engineering, 2024, 25( 1): 19–41

[87]

Rasul K, Seward C, Schuster I, Vollgraf R. Autoregressive denoising diffusion models for multivariate probabilistic time series forecasting. In: Proceedings of the 38th International Conference on Machine Learning. 2021, 8857−8868

[88]

Shen L, Kwok J T. Non-autoregressive conditional diffusion models for time series prediction. In: Proceedings of the 40th International Conference on Machine Learning. 2023, 1284

[89]

Wen H, Lin Y, Xia Y, Wan H, Wen Q, Zimmermann R, Liang Y. DiffSTG: probabilistic spatio-temporal graph forecasting with denoising diffusion models. In: Proceedings of the 31st ACM International Conference on Advances in Geographic Information Systems. 2023, 60

[90]

Huang H, Chen M, Qiao X. Generative learning for financial time series with irregular and scale-invariant patterns. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[91]

Li Y, Lu X, Wang Y, Dou D. Generative time series forecasting with diffusion, denoise, and disentanglement. In: Proceedings of the 36th International Conference on Neural Information Processing Systems. 2022, 23009−23022

[92]

Yuan X, Qiao Y. Diffusion-TS: interpretable diffusion for general time series generation. 2024, arXiv preprint arXiv: 2403.01742

[93]

Li Y, Chen W, Hu X, Chen B, Sun B, Zhou M. Transformer-modulated diffusion models for probabilistic multivariate time series forecasting. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[94]

Fan X, Wu Y, Xu C, Huang Y, Liu W, Bian J. MG-TSD: multi-granularity time series diffusion models with guided learning process. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[95]

Guo S, Lin Y, Wan H, Li X, Cong G . Learning dynamics and heterogeneity of spatial-temporal graph data for traffic forecasting. IEEE Transactions on Knowledge and Data Engineering, 2022, 34( 11): 5415–5428

[96]

Liu X, Xia Y, Liang Y, Hu J, Wang Y, Bai L, Huang C, Liu Z, Hooi B, Zimmermann R. LargeST: a benchmark dataset for large-scale traffic forecasting. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. 2023, 3293

[97]

IQAir . China national environmental monitoring centre. See iqair.cn/cn/profile/china-national-environmental-monitoring-centre website, 2025

[98]

Yi K, Zhang Q, Fan W, Wang S, Wang P, He H, Lian D, An N, Cao L, Niu Z. Frequency-domain MLPs are more effective learners in time series forecasting. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. 2023, 3349

[99]

Liu Y, Li C, Wang J, Long M. Koopa: learning non-stationary time series dynamics with koopman predictors. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. 2023, 538

[100]

Zhou T, Ma Z, Wang X, Wen Q, Sun L, Yao T, Yin W, Jin R. FiLM: frequency improved legendre memory model for long-term time series forecasting. In: Proceedings of the 36th International Conference on Neural Information Processing Systems. 2022, 921

[101]

Dai T, Wu B, Liu P, Li N, Bao J, Jiang Y, Xia S T. Periodicity decoupling framework for long-term series forecasting. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[102]

Wang Y, Wu H, Dong J, Qin G, Zhang H, Liu Y, Qiu Y, Wang J, Long M. TimeXer: empowering transformers for time series forecasting with exogenous variables. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. 2024

[103]

Shi J, Ma Q, Ma H, Li L. Scaling law for time series forecasting. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. 2024

[104]

Wang Y, Wu H, Ma Y, Fang Y, Zhang Z, Liu Y, Wang S, Ye Z, Xiang Y, Wang J, Long M. Accuracy law for the future of deep time series forecasting. 2025, arXiv preprint arXiv: 2510.02729

[105]

Zhou P, Liu Y, Liang J, Song Q, Li X. CrossLinear: plug-and-play cross-correlation embedding for time series forecasting with exogenous variables. In: Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 2025, 4120−4131

[106]

Chen W, Wu Y, Zhu Y, Hao X, Wang S, Liang Y. Select, then balance: a plug-and-play framework for exogenous-aware spatio-temporal forecasting. 2025, arXiv preprint arXiv: 2509.05779v1

[107]

Boussif O, Boukachab G, Assouline D, Massaroli S, Yuan T, Benabbou L, Bengio Y. Improving day-ahead solar irradiance time series forecasting by leveraging spatio-temporal context. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. 2023, 109

[108]

Gerard S, Zhao Y, Sullivan J. WildfireSpreadTS: a dataset of multi-modal time series for wildfire spread prediction. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. 2023, 3258

[109]

Xia H, Chen X, Wang Z, Chen X, Dong F . A multi-modal deep-learning air quality prediction method based on multi-station time-series data and remote-sensing images: case study of Beijing and Tianjin. Entropy, 2024, 26( 1): 91

[110]

Liang Y, Xia Y, Ke S, Wang Y, Wen Q, Zhang J, Zheng Y, Zimmermann R. AirFormer: predicting nationwide air quality in China with transformers. In: Proceedings of the 37th AAAI Conference on Artificial Intelligence. 2023, 14329−14337

[111]

Wang S, Li Y, Zhang J, Meng Q, Meng L, Gao F. PM2.5-GNN: a domain knowledge enhanced graph neural network for PM2.5 forecasting. In: Proceedings of the 28th International Conference on Advances in Geographic Information Systems. 2020, 163−166

[112]

Gruver N, Finzi M, Qiu S, Wilson A G. Large language models are zero-shot time series forecasters. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. 2023, 861

[113]

Zhou T, Niu P, Wang X, Sun L, Jin R. One fits all: power general time series analysis by pretrained LM. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. 2023, 1877

[114]

Pan Z, Jiang Y, Garg S, Schneider A, Nevmyvaka Y, Song D. S2IP-LLM: semantic space informed prompt learning with LLM for time series forecasting. In: Proceedings of the 41st International Conference on Machine Learning. 2024, 1588

[115]

Chang C, Wang W Y, Peng W C, Chen T F . LLM4TS: aligning pre-trained LLMs as data-efficient time-series forecasters. ACM Transactions on Intelligent Systems and Technology, 2025, 16( 3): 60

[116]

Sun C, Li H, Li Y, Hong S. TEST: text prototype aligned embedding to activate LLM’s ability for time series. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[117]

Liu X, Hu J, Li Y, Diao S, Liang Y, Hooi B, Zimmermann R. UniTime: a language-empowered unified model for cross-domain time series forecasting. In: Proceedings of the ACM Web Conference 2024. 2024, 4095−4106

[118]

Jin M, Wang S, Ma L, Chu Z, Zhang J Y, Shi X, Chen P Y, Liang Y, Li Y F, Pan S, Wen Q. Time-LLM: time series forecasting by reprogramming large language models. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[119]

Xue H, Salim F D . PromptCast: a new prompt-based learning paradigm for time series forecasting. IEEE Transactions on Knowledge and Data Engineering, 2024, 36( 11): 6851–6864

[120]

Liu Y, Qin G, Huang X, Wang J, Long M. AutoTimes: autoregressive time series forecasters via large language models. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. 2024

[121]

Cao D, Jia F, Arik S Ö, Pfister T, Zheng Y, Ye W, Liu Y. TEMPO: prompt-based generative pre-trained transformer for time series forecasting. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[122]

Jin M, Zhang Y, Chen W, Zhang K, Liang Y, Yang B, Wang J, Pan S, Wen Q. Position: what can large language models tell us about time series analysis. In: Proceedings of the 41st International Conference on Machine Learning. 2024

[123]

Tan M, Merrill M A, Gupta V, Althoff T, Hartvigsen T. Are language models actually useful for time series forecasting? In: Proceedings of the 38th International Conference on Neural Information Processing Systems. 2024, 1922

[124]

Tang H, Zhang C, Jin M, Yu Q, Wang Z, Jin X, Zhang Y, Du M . Time series forecasting with LLMs: understanding and enhancing model capabilities. ACM SIGKDD Explorations Newsletter, 2025, 26( 2): 109–118

[125]

Liang Y, Wen H, Nie Y, Jiang Y, Jin M, Song D, Pan S, Wen Q. Foundation models for time series analysis: a tutorial and survey. In: Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 2024, 6555−6565

[126]

Garza A, Challu C, Mergenthaler-Canseco M. TimeGPT-1. 2024, arXiv preprint arXiv: 2310.03589

[127]

Rasul K, Ashok A, Williams A R, Ghonia H, Bhagwatkar R, Khorasani A, Bayazi M J D, Adamopoulos G, Riachi R, Hassen N, Biloš M, Garg S, Schneider A, Chapados N, Drouin A, Zantedeschi V, Nevmyvaka Y, Rish I. Lag-llama: towards foundation models for probabilistic time series forecasting. 2023, arXiv preprint arXiv: 2310.08278

[128]

Liu Y, Zhang H, Li C, Huang X, Wang J, Long M. Timer: generative pre-trained transformers are large time series models. In: Proceedings of the 41st International Conference on Machine Learning. 2024

[129]

Goswami M, Szafer K, Choudhry A, Cai Y, Li S, Dubrawski A. MOMENT: a family of open time-series foundation models. In: Proceedings of the 41st International Conference on Machine Learning. 2024

[130]

Liu Y, Qin G, Shi Z, Chen Z, Yang C, Huang X, Wang J, Long M. Sundial: a family of highly capable time series foundation models. 2025, arXiv preprint arXiv: 2502.00816

[131]

Ansari A F, Stella L, Turkmen C, Zhang X, Mercado P, Shen H, Shchur O, Rangapuram S S, Arango S P, Kapoor S, Zschiegner J, Maddix D C, Wang H, Mahoney M W, Torkkola K, Wilson A G, Bohlke-Schneider M, Wang Y. Chronos: learning the language of time series. 2024, arXiv preprint arXiv: 2403.07815

[132]

Shi X, Wang S, Nie Y, Li D, Ye Z, Wen Q, Jin M. Time-MoE: billion-scale time series foundation models with mixture of experts. In: Proceedings of the 13th International Conference on Learning Representations. 2025

[133]

Cao D, Ye W, Zhang Y, Liu Y. TimeDiT: general-purpose diffusion transformers for time series foundation model. 2025, arXiv preprint arXiv: 2409.02322

[134]

Xiao C, Zhou J, Xiao Y, Lu X, Zhang L, Xiong H. TimeFound: a foundation model for time series forecasting. 2025, arXiv preprint arXiv: 2503.04118

[135]

Das A, Kong W, Sen R, Zhou Y. A decoder-only foundation model for time-series forecasting. In: Proceedings of the 41st International Conference on Machine Learning. 2024, 404

[136]

Woo G, Liu C, Kumar A, Xiong C, Savarese S, Sahoo D. Unified training of universal time series forecasting transformers. In: Proceedings of the 41st International Conference on Machine Learning. 2024, 2178

[137]

Ni J, Zhao Z, Shen C A, Tong H, Song D, Cheng W, Luo D, Chen H. Harnessing vision models for time series analysis: a survey. In: Proceedings of the 34th International Joint Conference on Artificial Intelligence. 2025, 1178

[138]

Yang L, Wang Y, Fan X, Cohen I, Zhao Y, Zhang Z. ViTime: a visual intelligence-based foundation model for time series forecasting. 2025, arXiv preprint arXiv: 2407.07311v1

Rights & permissions

Higher Education Press

PDF (13248KB)

Supplementary files

Highlights

820

Accesses

0

Citation

Detail

Sections
Recommended

/