About the journal
Browse
Collections
Multimedia collections
Authors & reviewers
A comprehensive and practical benchmark for financial time series forecasting
Yifan HU , Yuante LI , Peiyuan LIU , Yuxia ZHU , Naiqi LI , Tao DAI , Shutao XIA , Dawei CHENG , Changjun JIANG
Front. Comput. Sci. ›› 2026, Vol. 20 ›› Issue (10) : 2010629
Financial time series (FinTS) record the behavior of human-brain-augmented decision-making, capturing valuable historical information that can be leveraged for profitable investment strategies. Not surprisingly, this area has attracted considerable attention from researchers, who have proposed a wide range of methods based on various backbones. However, the evaluation of the area often exhibits three systemic limitations: 1) Failure to account for the full spectrum of stock movement patterns observed in dynamic financial markets. (Diversity Gap); 2) The absence of unified assessment protocols undermines the validity of cross-study performance comparisons. (Standardization Deficit); 3) Neglect of critical market structure factors, resulting in inflated performance metrics that lack practical applicability. (Real-World Mismatch). Addressing these limitations, we propose Financial Time Series Benchmark (FinTSB), a comprehensive and practical benchmark for financial time series forecasting (FinTSF). To increase the variety, we categorize movement patterns into four specific parts, tokenize and pre-process the data, and assess the data quality based on some sequence characteristics. To eliminate biases due to different evaluation settings, we standardize the metrics across three dimensions and build a user-friendly, lightweight pipeline incorporating methods from various backbones. To accurately simulate real-world trading scenarios and facilitate practical implementation, we extensively model various regulatory constraints, including transaction fees, among others. Finally, we conduct extensive experiments on FinTSB, highlighting key insights to guide model selection under varying market conditions. Overall, FinTSB provides researchers with a novel and comprehensive platform for improving and evaluating FinTSF methods. The code is available at the website of github.com/TongjiFinLab/FinTSB.
financial time series / computational finance / quantitative trading / benchmark / data mining
| [1] |
Zou J, Zhao Q, Jiao Y, Cao H, Liu Y, Yan Q, Abbasnejad E, Liu L, Shi J Q. Stock market prediction via deep learning techniques: a survey. 2023, arXiv preprint arXiv: 2212.12717v2 |
| [2] |
|
| [3] |
|
| [4] |
|
| [5] |
|
| [6] |
|
| [7] |
|
| [8] |
|
| [9] |
|
| [10] |
|
| [11] |
|
| [12] |
Li Y, Yu Y, Li H, Chen Z, Khashanah K. TradingGPT: multi-agent system with layered memory and distinct characters for enhanced financial trading performance. 2023, arXiv preprint arXiv: 2309.03736 |
| [13] |
Li Y, Yang X, Yang X, Xu M, Wang X, Liu W, Bian J. R&D-Agent-Quant: a multi-agent framework for data-centric factors and model joint optimization. 2025, arXiv preprint arXiv: 2505.15155 |
| [14] |
|
| [15] |
|
| [16] |
|
| [17] |
|
| [18] |
|
| [19] |
Gao Y, Chen H, Wang X, Wang Z, Wang X, Gao J, Ding B. DiffsFormer: a diffusion transformer on stock factor augmentation. 2024, arXiv preprint arXiv: 2402.06656 |
| [20] |
|
| [21] |
Wang J, Zhang Y, Tang K, Wu J, Xiong Z. AlphaStock: a buying-winners-and-selling-losers investment strategy using interpretable deep reinforcement attention networks. In: Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 2019, 1900–1908 |
| [22] |
|
| [23] |
|
| [24] |
|
| [25] |
|
| [26] |
|
| [27] |
|
| [28] |
|
| [29] |
|
| [30] |
|
| [31] |
|
| [32] |
|
| [33] |
|
| [34] |
|
| [35] |
|
| [36] |
|
| [37] |
|
| [38] |
|
| [39] |
|
| [40] |
Schulman J, Wolski F, Dhariwal P, Radford A, Klimov O. Proximal policy optimization algorithms. 2017, arXiv preprint arXiv: 1707.06347 |
| [41] |
|
| [42] |
|
| [43] |
|
| [44] |
|
| [45] |
|
| [46] |
|
| [47] |
Yang X, Liu W, Zhou D, Bian J, Liu T Y. Qlib: an AI-oriented quantitative investment platform. 2020, arXiv preprint arXiv: 2009.11189 |
| [48] |
Wang Y, Wu H, Dong J, Liu Y, Long M, Wang J. Deep time series models: a comprehensive survey and benchmark. 2024, arXiv preprint arXiv: 2407.13278 |
| [49] |
|
| [50] |
|
| [51] |
|
| [52] |
|
| [53] |
|
| [54] |
|
| [55] |
|
| [56] |
Hu Y, Liu P, Li Y, Cheng D, Li N, Dai T, Bao J, Shu-Tao X. FinMamba: market-aware graph enhanced multi-level mamba for stock movement prediction. 2025, arXiv preprint arXiv: 2502.06707 |
| [57] |
|
| [58] |
Hendrycks D, Burns C, Kadavath S, Arora A, Basart S, Tang E, Song D, Steinhardt J. Measuring mathematical problem solving with the math dataset. 2021, arXiv preprint arXiv: 2103.03874 |
| [59] |
Cobbe K, Kosaraju V, Bavarian M, Chen M, Jun H, Kaiser L, Plappert M, Tworek J, Hilton J, Nakano R, Hesse C, Schulman J. Training verifiers to solve math word problems. 2021, arXiv preprint arXiv: 2110.14168 |
| [60] |
Mialon G, Fourrier C, Wolf T, LeCun Y, Scialom T. GAIA: a benchmark for general AI assistants. In: Proceedings of the 12th International Conference on Learning Representations. 2024 |
| [61] |
|
| [62] |
|
| [63] |
|
| [64] |
|
| [65] |
Austin J, Odena A, Nye M, Bosma M, Michalewski H, Dohan D, Jiang E, Cai C, Terry M, Le Q, Sutton C. Program synthesis with large language models. 2021, arXiv preprint arXiv: 2108.07732 |
| [66] |
Chen M, Tworek J, Jun H, Yuan Q, De Oliveira Pinto H P, et al. Evaluating large language models trained on code. 2021, arXiv preprint arXiv: 2107.03374 |
| [67] |
|
| [68] |
Mushtaq R. Augmented dickey fuller test. 2011 |
| [69] |
|
| [70] |
|
| [71] |
|
| [72] |
|
| [73] |
|
| [74] |
Chung J, Gulcehre C, Cho K, Bengio Y. Empirical evaluation of gated recurrent neural networks on sequence modeling. 2014, arXiv preprint arXiv: 1412.3555 |
| [75] |
|
| [76] |
|
| [77] |
Bai S, Kolter J Z, Koltun V. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. 2018, arXiv preprint arXiv: 1803.01271 |
| [78] |
|
| [79] |
|
| [80] |
|
| [81] |
|
| [82] |
|
| [83] |
Zhao Q, Wang Y, Zhou Z, Miao D, Wang L, Qiao Y, Zhao C. Rethinking the zigzag flattening for image reading. 2024, arXiv preprint arXiv: 2202.10240v8 |
| [84] |
Lillicrap T P, Hunt J J, Pritzel A, Heess N, Erez T, Tassa Y, Silver D, Wierstra D. Continuous control with deep reinforcement learning. 2019, arXiv preprint arXiv: 1509.02971v6 |
| [85] |
Haarnoja T, Zhou A, Hartikainen K, Tucker G, Ha S, Tan J, Kumar V, Zhu H, Gupta A, Abbeel P, Levine S. Soft actor-critic algorithms and applications. 2019, arXiv preprint arXiv: 1812.05905v2 |
| [86] |
|
| [87] |
|
| [88] |
|
| [89] |
|
| [90] |
Ansari A F, Stella L, Türkmen A C, Zhang X, Mercado P, Shen H, Shchur O, Rangapuram S S, Pineda-Arango S, Kapoor S, Zschiegner J, Maddix D C, Mahoney M W, Torkkola K, Gordon Wilson A, Bohlke-Schneider M, Wang B. Chronos: Learning the language of time series. 2024, arXiv preprint arXiv: 2403.07815 |
| [91] |
|
| [92] |
|
| [93] |
|
Higher Education Press
/
| 〈 |
|
〉 |