A comprehensive and practical benchmark for financial time series forecasting

Yifan HU , Yuante LI , Peiyuan LIU , Yuxia ZHU , Naiqi LI , Tao DAI , Shutao XIA , Dawei CHENG , Changjun JIANG

Front. Comput. Sci. ›› 2026, Vol. 20 ›› Issue (10) : 2010629

PDF (2386KB)
Front. Comput. Sci. ›› 2026, Vol. 20 ›› Issue (10) :2010629 DOI: 10.1007/s11704-026-51064-5
Information Systems
RESEARCH ARTICLE
A comprehensive and practical benchmark for financial time series forecasting
Author information +
History +
PDF (2386KB)

Abstract

Financial time series (FinTS) record the behavior of human-brain-augmented decision-making, capturing valuable historical information that can be leveraged for profitable investment strategies. Not surprisingly, this area has attracted considerable attention from researchers, who have proposed a wide range of methods based on various backbones. However, the evaluation of the area often exhibits three systemic limitations: 1) Failure to account for the full spectrum of stock movement patterns observed in dynamic financial markets. (Diversity Gap); 2) The absence of unified assessment protocols undermines the validity of cross-study performance comparisons. (Standardization Deficit); 3) Neglect of critical market structure factors, resulting in inflated performance metrics that lack practical applicability. (Real-World Mismatch). Addressing these limitations, we propose Financial Time Series Benchmark (FinTSB), a comprehensive and practical benchmark for financial time series forecasting (FinTSF). To increase the variety, we categorize movement patterns into four specific parts, tokenize and pre-process the data, and assess the data quality based on some sequence characteristics. To eliminate biases due to different evaluation settings, we standardize the metrics across three dimensions and build a user-friendly, lightweight pipeline incorporating methods from various backbones. To accurately simulate real-world trading scenarios and facilitate practical implementation, we extensively model various regulatory constraints, including transaction fees, among others. Finally, we conduct extensive experiments on FinTSB, highlighting key insights to guide model selection under varying market conditions. Overall, FinTSB provides researchers with a novel and comprehensive platform for improving and evaluating FinTSF methods. The code is available at the website of github.com/TongjiFinLab/FinTSB.

Graphical abstract

Keywords

financial time series / computational finance / quantitative trading / benchmark / data mining

Cite this article

Download citation ▾
Yifan HU, Yuante LI, Peiyuan LIU, Yuxia ZHU, Naiqi LI, Tao DAI, Shutao XIA, Dawei CHENG, Changjun JIANG. A comprehensive and practical benchmark for financial time series forecasting. Front. Comput. Sci., 2026, 20 (10) : 2010629 DOI:10.1007/s11704-026-51064-5

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Zou J, Zhao Q, Jiao Y, Cao H, Liu Y, Yan Q, Abbasnejad E, Liu L, Shi J Q. Stock market prediction via deep learning techniques: a survey. 2023, arXiv preprint arXiv: 2212.12717v2

[2]

Xia H, Ao H, Li L, Liu Y, Liu S, Ye G, Chai H. CI-STHPAN: pre-trained attention network for stock selection with channel-independent spatio-temporal hypergraph. In: Proceedings of the 38th AAAI Conference on Artificial Intelligence. 2024, 1022

[3]

Cheng D, Liu Y, Niu Z, Zhang L . Modeling similarities among multi-dimensional financial time series. IEEE Access, 2018, 6: 43404–43413

[4]

Cheng D, Yang F, Xiang S, Liu J . Financial time series forecasting with multi-modality graph neural network. Pattern Recognition, 2022, 121: 108218

[5]

Zhu P, Li Y, Hu Y, Xiang S, Liu Q, Cheng D, Liang Y . MCI-GRU: stock prediction model based on multi-head cross-attention and improved GRU. Neurocomputing, 2025, 638: 130168

[6]

Zhu P, Li Y, Hu Y, Liu Q, Cheng D, Liang Y. LSR-IGRU: Stock trend prediction based on long short-term relationships and improved GRU. In: Proceedings of the 33rd ACM International Conference on Information and Knowledge Management. 2024, 5135−5142

[7]

Hu Y, Liu P, Zhu P, Cheng D, Dai T. Adaptive multi-scale decomposition framework for time series forecasting. In: Proceedings of the 39th AAAI Conference on Artificial Intelligence. 2025, 1935

[8]

Liu P, Wu B, Hu Y, Li N, Dai T, Bao J, Xia S T. TimeBridge: Non-stationarity matters for long-term time series forecasting. In: Proceedings of the 42nd International Conference on Machine Learning. 2025, 39815−39840

[9]

Dai T, Wu B, Liu P, Li N, Bao J, Jiang Y, Xia S T. Periodicity decoupling framework for long-term series forecasting. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[10]

Hu Y, Zhang G, Liu P, Lan D, Li N, Cheng D, Dai T, Xia S T, Pan S. TimeFilter: patch-specific spatial-temporal graph filtration for time series forecasting. In: Proceedings of the 42nd International Conference on Machine Learning. 2025, 24893–24911

[11]

Yu Y, Yao Z, Li H, Deng Z, Jiang Y, Cao Y, Chen Z, Suchow J W, Cui Z, Liu R, Xu Z, Zhang D, Subbalakshmi K, Xiong G, He Y, Huang J, Li D, Xie Q. FINCON: A synthesized LLM multi-agent system with conceptual verbal reinforcement for enhanced financial decision making. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. 2024, 4354

[12]

Li Y, Yu Y, Li H, Chen Z, Khashanah K. TradingGPT: multi-agent system with layered memory and distinct characters for enhanced financial trading performance. 2023, arXiv preprint arXiv: 2309.03736

[13]

Li Y, Yang X, Yang X, Xu M, Wang X, Liu W, Bian J. R&D-Agent-Quant: a multi-agent framework for data-centric factors and model joint optimization. 2025, arXiv preprint arXiv: 2505.15155

[14]

Shi X, Wang S, Nie Y, Li D, Ye Z, Wen Q, Jin M. Time-MoE: billion-scale time series foundation models with mixture of experts. In: Proceedings of the 13th International Conference on Learning Representations. 2025

[15]

Li S, Sun Y, Lin Y, Gao X, Shang S, Yan R. CausalStock: deep end-to-end causal discovery for news-driven stock movement prediction. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. 2024, 1504

[16]

Zhang W, Zhao L, Xia H, Sun S, Sun J, Qin M, Li X, Zhao Y, Zhao Y, Cai X, Zheng L, Wang X, An B. A multimodal foundation agent for financial trading: Tool-augmented, diversified, and generalist. In: Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 2024, 4314–4325

[17]

Duan Y, Wang L, Zhang Q, Li J. FactorVAE: a probabilistic dynamic factor model based on variational autoencoder for predicting cross-sectional stock returns. In: Proceedings of the 36th AAAI Conference on Artificial Intelligence. 2022, 4468–4476

[18]

Daiya D, Yadav M, Rao H S. Diffstock: Probabilistic relational stock market predictions using diffusion models. In: Proceedings of ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2024, 7335–7339

[19]

Gao Y, Chen H, Wang X, Wang Z, Wang X, Gao J, Ding B. DiffsFormer: a diffusion transformer on stock factor augmentation. 2024, arXiv preprint arXiv: 2402.06656

[20]

Xia H, Sun S, Wang X, An B. Market-GAN: adding control to financial market data generation with semantic context. In: Proceedings of the 38th AAAI Conference on Artificial Intelligence. 2024, 1783

[21]

Wang J, Zhang Y, Tang K, Wu J, Xiong Z. AlphaStock: a buying-winners-and-selling-losers investment strategy using interpretable deep reinforcement attention networks. In: Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 2019, 1900–1908

[22]

Liu X Y, Yang H, Gao J, Wang C D. FinRL: deep reinforcement learning framework to automate trading in quantitative finance. In: Proceedings of the 2nd ACM International Conference on AI in Finance. 2021, 1

[23]

Wang Z, Huang B, Tu S, Zhang K, Xu L. DeepTrader: A deep reinforcement learning approach for risk-return balanced portfolio management with market conditions embedding. In: Proceedings of the 35th AAAI Conference on Artificial Intelligence. 2021, 643–650

[24]

Jeon J, Park J, Park C, Kang U. FreQuant: a reinforcement-learning based adaptive portfolio optimization with multi-frequency decomposition. In: Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 2024, 1211–1221

[25]

Niu H, Li S, Zheng J, Lin Z, An B, Li J, Guo J. IMM: an imitative reinforcement learning approach with predictive representation learning for automatic market making. In: Proceedings of the 33rd International Joint Conference on Artificial Intelligence. 2024, 663

[26]

Li T, Liu Z, Shen Y, Wang X, Chen H, Huang S. MASTER: Market-guided stock transformer for stock price forecasting. In: Proceedings of the 38th AAAI Conference on Artificial Intelligence. 2024, 162–170

[27]

Xing R, Cheng R, Huang J, Li Q, Zhao J . Learning to understand the vague graph for stock prediction with momentum spillovers. IEEE Transactions on Knowledge and Data Engineering, 2024, 36( 4): 1698–1712

[28]

Chen T, Guestrin C. XGBoost: a scalable tree boosting system. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 2016, 785−794

[29]

Lin Y, Guo H, Hu J. An SVM-based approach for stock market trend prediction. In: Proceedings of 2013 International Joint Conference on Neural Networks (IJCNN). 2013

[30]

Jegadeesh N, Titman S . Returns to buying winners and selling losers: Implications for stock market efficiency. The Journal of Finance, 1993, 48( 1): 65–91

[31]

Poterba J M, Summers L H . Mean reversion in stock prices: evidence and implications. Journal of Financial Economics, 1988, 22( 1): 27–59

[32]

Deng S, Huang X, Zhu Y, Su Z, Fu Z, Shimada T . Stock index direction forecasting using an explainable extreme gradient boosting and investor sentiments. The North American Journal of Economics and Finance, 2023, 64: 101848

[33]

Illa P K, Parvathala B, Sharma A K . Stock price prediction methodology using random forest algorithm and support vector machine. Materials Today: Proceedings, 2022, 56: 1776–1782

[34]

Box G E P, Jenkins G M, MacGregor J F . Some recent advances in forecasting and control. Journal of the Royal Statistical Society Series C: Applied Statistics, 1974, 23( 2): 158–179

[35]

Ke G, Meng Q, Finley T, Wang T, Chen W, Ma W, Ye Q, Liu T Y. LightGBM: a highly efficient gradient boosting decision tree. In: Proceedings of the 31st International Conference on Neural Information Processing Systems. 2017, 3149–3157

[36]

Zheng J, Xin D, Cheng Q, Tian M, Yang L. The Random Forest Model for analyzing and Forecasting the US Stock Market under the background of smart finance. In: Proceedings of the 3rd International Academic Conference on Blockchain, Information Technology and Smart Finance (ICBIS 2024). 2024, 82–90

[37]

Xiang S, Cheng D, Shang C, Zhang Y, Liang Y. Temporal and heterogeneous graph neural network for financial time series prediction. In: Proceedings of the 31st ACM International Conference on Information & Knowledge Management. 2022, 3584–3593

[38]

Zhu P, Li Y, Liu Q, Cheng D, Jiang C . Financial time series prediction with multi-granularity graph augmented learning. IEEE Transactions on Knowledge and Data Engineering, 2025, 37( 11): 6436–6449

[39]

Ying Z, Cheng D, Chen C, Li X, Zhu P, Luo Y, Liang Y . Predicting stock market trends with self-supervised learning. Neurocomputing, 2024, 568: 127033

[40]

Schulman J, Wolski F, Dhariwal P, Radford A, Klimov O. Proximal policy optimization algorithms. 2017, arXiv preprint arXiv: 1707.06347

[41]

Soleymani F, Paquet E . Deep graph convolutional reinforcement learning for financial portfolio management −deeppocket. Expert Systems with Applications, 2021, 182: 115127

[42]

Pan K, Hu Y, Han L, Sun H, Cheng D, Liang Y. Cross-contextual sequential optimization via deep reinforcement learning for algorithmic trading. In: Proceedings of the 33rd ACM International Conference on Information and Knowledge Management. 2024, 4811−4818

[43]

Hirshleifer D, Peng L, Wang Q . News diffusion in social networks and stock market reactions. The Review of Financial Studies, 2025, 38( 3): 883–937

[44]

Huang Y H, Xu C, Liu Y, Liu W, Li W J, Bian J. Controllable financial market generation with diffusion guided meta agent. In: Proceedings of the 40th AAAI Conference on Artificial Intelligence. 2026, 462−470

[45]

Li J, Lei Y, Bian Y, Cheng D, Ding Z, Jiang C . RA-CFGPT: Chinese financial assistant with retrieval-augmented large language model. Frontiers of Computer Science, 2024, 18( 5): 185350

[46]

Liu P, Guo H, Dai T, Li N, Bao J, Ren X, Jiang Y, Xia S T. CALF: aligning LLMs for time series forecasting via cross-modal fine-tuning. In: Proceedings of the 39th AAAI Conference on Artificial Intelligence. 2025, 2109

[47]

Yang X, Liu W, Zhou D, Bian J, Liu T Y. Qlib: an AI-oriented quantitative investment platform. 2020, arXiv preprint arXiv: 2009.11189

[48]

Wang Y, Wu H, Dong J, Liu Y, Long M, Wang J. Deep time series models: a comprehensive survey and benchmark. 2024, arXiv preprint arXiv: 2407.13278

[49]

Qiu X, Hu J, Zhou L, Wu X, Du J, Zhang B, Guo C, Zhou A, Jensen C S, Sheng Z, Yang B . TFB: towards comprehensive and fair benchmarking of time series forecasting methods. Proceedings of the VLDB Endowment, 2024, 17( 9): 2363–2377

[50]

Deng J, Dong W, Socher R, Li L J, Li K, Li F F. ImageNet: a large-scale hierarchical image database. In: Proceedings of 2009 IEEE Conference on Computer Vision and Pattern Recognition. 2009, 248−255

[51]

Lin T Y, Maire M, Belongie S, Hays J, Perona P, Ramanan D, Dollár P, Zitnick C L. Microsoft COCO: common objects in context. In: Proceedings of the 13th European Conference on Computer Vision -- ECCV 2014. 2014, 740−755

[52]

Wang H, Wang T, Li S, Zheng J, Guan S, Chen W. Adaptive long-short pattern transformer for stock investment selection. In: Proceedings of the 31st International Joint Conference on Artificial Intelligence. 2022, 3970−3977

[53]

Chen W, Li S, Yu X, Wang H, Chen W, Wang T. Automatic de-biased temporal-relational modeling for stock investment recommendation. In: Proceedings of the 33rd International Joint Conference on Artificial Intelligence. 2024, 221

[54]

Zeng L, Wang L, Niu H, Zhang R, Wang L, Li J. Trade when opportunity comes: price movement forecasting via locality-aware attention and iterative refinement labeling. In: Proceedings of the 33rd International Joint Conference on Artificial Intelligence. 2024, 6134−6142

[55]

Xiang Q, Chen Z, Sun Q, Jiang R. RSAP-DFM: regime-shifting adaptive posterior dynamic factor model for stock returns prediction. In: Proceedings of the 33rd International Joint Conference on Artificial Intelligence. 2024, 6116–6124

[56]

Hu Y, Liu P, Li Y, Cheng D, Li N, Dai T, Bao J, Shu-Tao X. FinMamba: market-aware graph enhanced multi-level mamba for stock movement prediction. 2025, arXiv preprint arXiv: 2502.06707

[57]

Abdi H, Williams L J . Principal component analysis. WIREs Computational Statistics, 2010, 2( 4): 433–459

[58]

Hendrycks D, Burns C, Kadavath S, Arora A, Basart S, Tang E, Song D, Steinhardt J. Measuring mathematical problem solving with the math dataset. 2021, arXiv preprint arXiv: 2103.03874

[59]

Cobbe K, Kosaraju V, Bavarian M, Chen M, Jun H, Kaiser L, Plappert M, Tworek J, Hilton J, Nakano R, Hesse C, Schulman J. Training verifiers to solve math word problems. 2021, arXiv preprint arXiv: 2110.14168

[60]

Mialon G, Fourrier C, Wolf T, LeCun Y, Scialom T. GAIA: a benchmark for general AI assistants. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[61]

Yoran O, Amouyal S J, Malaviya C, Bogin B, Press O, Berant J. AssistantBench: can web agents solve realistic and time-consuming tasks?. In: Proceedings of 2024 Conference on Empirical Methods in Natural Language Processing. 2024, 8938–8968

[62]

Hendrycks D, Burns C, Basart S, Zou A, Mazeika M, Song D, Steinhardt J. Measuring massive multitask language understanding. In: Proceedings of the 9th International Conference on Learning Representations. 2021

[63]

Madaan A, Tandon N, Gupta P, Hallinan S, Gao L, Wiegreffe S, Alon U, Dziri N, Prabhumoye S, Yang Y, Welleck S, Majumder B P, Hermann K, Welleck S, Yazdanbakhsh A, Clark P. SELF-REFINE: iterative refinement with self-feedback. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. 2023, 2023

[64]

Cui G, Yuan L, Ding N, Yao G, He B, Zhu W, Ni Y, Xie G, Xie R, Lin Y, Liu Z, Sun M. ULTRAFEEDBACK: boosting language models with scaled AI feedback. In: Proceedings of the 41st International Conference on Machine Learning. 2024, 9722–9744

[65]

Austin J, Odena A, Nye M, Bosma M, Michalewski H, Dohan D, Jiang E, Cai C, Terry M, Le Q, Sutton C. Program synthesis with large language models. 2021, arXiv preprint arXiv: 2108.07732

[66]

Chen M, Tworek J, Jun H, Yuan Q, De Oliveira Pinto H P, et al. Evaluating large language models trained on code. 2021, arXiv preprint arXiv: 2107.03374

[67]

Liu J, Xia C S, Wang Y, Zhang L. Is your code generated by ChatGPT really correct? Rigorous evaluation of large language models for code generation. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. 2023, 943

[68]

Mushtaq R. Augmented dickey fuller test. 2011

[69]

Madsen H. Time Series Analysis. New York: Chapman and Hall/CRC, 2007

[70]

Goerg G M. Forecastable component analysis. In: Proceedings of the 30th International Conference on Machine Learning. 2013, 64–72

[71]

Zhang C, Li Y, Chen X, Jin Y, Tang P, Li J. DoubleEnsemble: a new ensemble method based on sample reweighting and feature selection for financial data analysis. In: Proceedings of 2020 IEEE International Conference on Data Mining (ICDM). 2020

[72]

Hochreiter S, Schmidhuber J . Long short-term memory. Neural Computation, 1997, 9( 8): 1735–1780

[73]

Qin Y, Song D, Chen H, Cheng W, Jiang G, Cottrell G W. A dual-stage attention-based recurrent neural network for time series prediction. In: Proceedings of the 26th International Joint Conference on Artificial Intelligence. 2017, 2627−2633

[74]

Chung J, Gulcehre C, Cho K, Bengio Y. Empirical evaluation of gated recurrent neural networks on sequence modeling. 2014, arXiv preprint arXiv: 1412.3555

[75]

Kipf T N, Welling M. Semi-supervised classification with graph convolutional networks. In: Proceedings of the 5th International Conference on Learning Representations. 2017

[76]

Velickovic P, Cucurull G, Casanova A, Romero A, Liò P, Bengio Y. Graph attention networks. In: Proceedings of the 6th International Conference on Learning Representations. 2018

[77]

Bai S, Kolter J Z, Koltun V. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. 2018, arXiv preprint arXiv: 1803.01271

[78]

Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez A N, Kaiser L, Polosukhin I. Attention is all you need. In: Proceedings of the 31st International Conference on Neural Information Processing Systems. 2017, 6000–6010

[79]

Gu A, Dao T. Mamba: Linear-time sequence modeling with selective state spaces. In: Proceedings of Conference Paper at COLM 2024. 2024

[80]

Nie Y, Nguyen N H, Sinthong P, Kalagnanam J. A time series is worth 64 words: long-term forecasting with transformers. In: Proceedings of the 11th International Conference on Learning Representations. 2023

[81]

Zhang Y, Yan J. Crossformer: transformer utilizing cross-dimension dependency for multivariate time series forecasting. In: Proceedings of the 11th International Conference on Learning Representations. 2023

[82]

Liu Y, Hu T, Zhang H, Wu H, Wang S, Ma L, Long M. iTransformer: inverted transformers are effective for time series forecasting. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[83]

Zhao Q, Wang Y, Zhou Z, Miao D, Wang L, Qiao Y, Zhao C. Rethinking the zigzag flattening for image reading. 2024, arXiv preprint arXiv: 2202.10240v8

[84]

Lillicrap T P, Hunt J J, Pritzel A, Heess N, Erez T, Tassa Y, Silver D, Wierstra D. Continuous control with deep reinforcement learning. 2019, arXiv preprint arXiv: 1509.02971v6

[85]

Haarnoja T, Zhou A, Hartikainen K, Tucker G, Ha S, Tan J, Kumar V, Zhu H, Gupta A, Abbeel P, Levine S. Soft actor-critic algorithms and applications. 2019, arXiv preprint arXiv: 1812.05905v2

[86]

Carta S, Ferreira A, Podda A S, Reforgiato Recupero D, Sanna A . Multi-DQN: an ensemble of deep Q-learning agents for stock market forecasting. Expert Systems with Applications, 2021, 164: 113820

[87]

Ho J, Jain A, Abbeel P. Denoising diffusion probabilistic models. In: Proceedings of the 34th International Conference on Neural Information Processing Systems. 2020, 574

[88]

Song J, Meng C, Ermon S. Denoising diffusion implicit models. In: Proceedings of the 9th International Conference on Learning Representations. 2021

[89]

Liu Y, Zhang H, Li C, Huang X, Wang J, Long M. Timer: Generative pre-trained transformers are large time series models. In: Proceedings of the 41st International Conference on Machine Learning. 2024, 1313

[90]

Ansari A F, Stella L, Türkmen A C, Zhang X, Mercado P, Shen H, Shchur O, Rangapuram S S, Pineda-Arango S, Kapoor S, Zschiegner J, Maddix D C, Mahoney M W, Torkkola K, Gordon Wilson A, Bohlke-Schneider M, Wang B. Chronos: Learning the language of time series. 2024, arXiv preprint arXiv: 2403.07815

[91]

Lin S, Lin W, Wu W, Zhao F, Mo R, Zhang H . SegRNN: segment recurrent neural network for long-term time-series forecasting. IEEE Internet of Things Journal, 2026, 13( 5): 9861–9871

[92]

Wang S, Wu H, Shi X, Hu T, Luo H, Ma L, Zhang J Y, Zhou J. TimeMixer: decomposable multiscale mixing for time series forecasting. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[93]

Paszke A, Gross S, Massa F, Lerer A, Bradbury J, , et al. PyTorch: an imperative style, high-performance deep learning library. In: Proceedings of the 33rd International Conference on Neural Information Processing Systems. 2019, 721

Rights & permissions

Higher Education Press

PDF (2386KB)

Supplementary files

Highlights

413

Accesses

0

Citation

Detail

Sections
Recommended

/