Optimizing metro-integrated freight delivery under time-of-use electricity pricing and non-stationary shipment arrivals

Miaomiao WANG , Lu ZHEN

Eng. Manag ››

PDF (4046KB)
Eng. Manag ›› DOI: 10.1007/s42524-026-6143-x
RESEARCH ARTICLE
Optimizing metro-integrated freight delivery under time-of-use electricity pricing and non-stationary shipment arrivals
Author information +
History +
PDF (4046KB)

Abstract

This study investigates a metro-integrated freight delivery optimization problem under time-of-use electricity pricing, where an entire planning horizon is divided into multiple rolling decision periods. Across these rolling decision periods, shipment arrivals may exhibit different distributions. In each decision period, trains operating along a metro corridor coordinate with external transport modes to deliver shipments. The resulting total energy cost includes both metro train energy cost and external transport energy cost, and is jointly affected by train timetables, speed profiles, shipment access-station choices, and train assignments. For this practical problem, a mixed-integer linear programming model is formulated to minimize the total energy cost by jointly optimizing these interdependent decisions. To improve computational efficiency, we design an online continual reinforcement learning (CRL)-guided algorithm to add cuts to the mixed-integer linear programming model, thereby reducing the solution space. Additionally, we adopt a progress-and-compress-based CRL framework to enable the agent to continually adapt to varying shipment arrival distributions across decision periods. Computational results show that the proposed algorithm obtains high-quality solutions more efficiently than both Gurobi and traditional reinforcement learning. Managerial insights further reveal that (i) minimizing energy cost is not always equivalent to minimizing total energy consumption under time-of-use electricity pricing, especially before the onset of higher electricity prices, and (ii) introducing a small shipment waiting-time penalty can promote just-in-time shipment transfers so as to reduce unnecessary storage pressure at stations.

Graphical abstract

Keywords

multi-mode logistics networks / metro-integrated freight delivery / TOU electricity pricing / continual reinforcement learning / energy optimal

Cite this article

Download citation ▾
Miaomiao WANG, Lu ZHEN. Optimizing metro-integrated freight delivery under time-of-use electricity pricing and non-stationary shipment arrivals. Eng. Manag DOI:10.1007/s42524-026-6143-x

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Abel D, Barreto A, Van Roy B, Precup D, VanHasselt H P, Singh S, (2023). A definition of continual reinforcement learning. Advances in Neural Information Processing Systems, 36: 50377–50407

[2]

Azcuy I, Agatz N, Giesen R, (2021). Designing integrated urban delivery systems using public transport. Transportation Research Part E, Logistics and Transportation Review, 156: 102525

[3]

Bischoff P, Lienkamp B, Rambha T, Schiffer M, (2026). Dynamic capacity allocation of hybrid transportation units for cargo-hitching in urban public transportation systems. Transportation Research Part B: Methodological, 206: 103412

[4]

CAMET (2025). China Association of Metros Urban rail transit annual statistical and analysis report. China. Accessed June 5, 2026 from the website of CAMET

[5]

Cappart Q, Bergman D, Rousseau L M, Prémont-Schwarz I, Parjadis A, (2022). Improving variable orderings of approximate decision diagrams using reinforcement learning. INFORMS Journal on Computing, 34( 5): 2552–2570

[6]

Chen X, Ulmer M W, Thomas B W, (2022). Deep Q-learning for same-day delivery with vehicles and drones. European Journal of Operational Research, 298( 3): 939–952

[7]

Di Z, Yang L, Shi J, Zhou H, Yang K, Gao Z, (2022). Joint optimization of carriage arrangement and flow control in a metro-based underground logistics system. Transportation Research Part B: Methodological, 159: 1–23

[8]

Ding X, Jin J G, Pan H, Wang X, Hu Y, Shi G, (2026). Leveraging the passenger and freight spatiotemporal flow to optimize metro passenger-freight mixed transportation: A new mode for urban logistics systems. Transportation Research Part E, Logistics and Transportation Review, 206: 104560

[9]

Du X, Yang L, Wang X, Zha M, Zhen L, (2025). Pricing and capacity optimization for underground logistics. Transportation Research Part E, Logistics and Transportation Review, 203: 104386

[10]

Hızır A E, Barnhart C, Vaze V, (2026). Large-scale airline crew recovery using mixed-integer optimization and supervised machine learning. Transportation Science, 60( 2): 177–196

[11]

Hörsting L, Cleophas C, (2023). Scheduling shared passenger and freight transport on a fixed infrastructure. European Journal of Operational Research, 306( 3): 1158–1169

[12]

Kaelbling L P, Littman M L, Moore A W, (1996). Reinforcement learning: a survey. Journal of Artificial Intelligence Research, 4: 237–285

[13]

Kirkpatrick J, Pascanu R, Rabinowitz N, Veness J, Desjardins G, Rusu A A, Milan K, Quan J, Ramalho T, Grabska-Barwinska A, Hassabis D, Clopath C, Kumaran D, Hadsell R., (2017). Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences of the United States of America, 114( 13): 3521–3526

[14]

Li K, Liu T, Kumar P N, Han X, (2024a). A reinforcement learning-based hyper-heuristic for AGV task assignment and route planning in parts-to-picker warehouses. Transportation Research Part E, Logistics and Transportation Review, 185: 103518

[15]

Li S, Zhu X, Shang P, Van Woensel T, Yao Y, (2026). Joint optimization of train services and freight delivery in a metro-based underground logistics system. Transportation Research Part B: Methodological, 204: 103377

[16]

Li S, Zhu X, Shang P, Wang L, Li T, (2024b). Scheduling shared passenger and freight transport for an underground logistics system. Transportation Research Part B: Methodological, 183: 102907

[17]

Li Z, Shalaby A, Roorda M J, Mao B, (2021). Urban rail service design for collaborative passenger and freight transport. Transportation Research Part E, Logistics and Transportation Review, 147: 102205

[18]

Liu Q, Hu W, Dong J, Yang K, Ren R, Chen Z, (2025). Cost-benefit analysis of road-underground co-modality strategies for sustainable city logistics. Transportation Research Part D, Transport and Environment, 139: 104585

[19]

Ma M, Zhang F, Liu W, Dixit V, (2022). A game theoretical analysis of metro-integrated city logistics systems. Transportation Research Part B: Methodological, 156: 14–27

[20]

Mo P, Yao Y, D’Ariano A, Liu Z, (2023). The vehicle routing problem with underground logistics: Formulation and algorithm. Transportation Research Part E, Logistics and Transportation Review, 179: 103286

[21]

Nguyen D V A, Gunawan A, Misir M, Hui L K, Vansteenwegen P, (2025). Deep reinforcement learning for solving the stochastic e-waste collection problem. European Journal of Operational Research, 327( 1): 309–325

[22]

Schettini T, Jabali O, Malucelli F, (2022). Demand-driven timetabling for a metro corridor using a short-turning acceleration strategy. Transportation Science, 56( 4): 919–937

[23]

Schwarz J, Czarnecki W, Luketina J, Grabska-Barwinska A, Teh Y W, Pascanu R, Hadsell R, (2018). Progress & compress: a scalable framework for continual learning. International Conference on Machine Learning, 80: 4528–4537

[24]

Shao S, Lin J, Zhang F, (2026). Service network design for a metro-based crowdsourced urban delivery system under demand and supply uncertainty. Transportation Research Part C, Emerging Technologies, 188: 105695

[25]

Shi J, Yang L, Yang J, Gao Z, (2018). Service-oriented train timetabling with collaborative passenger flow control on an oversaturated metro line: An integer linear optimization approach. Transportation Research Part B: Methodological, 110: 26–59

[26]

Simchowitz M, Slivkins A, (2024). Exploration and incentives in reinforcement learning. Operations Research, 72( 3): 983–998

[27]

Sinclair S R, Banerjee S, Yu C L, (2023). Adaptive discretization in online reinforcement learning. Operations Research, 71( 5): 1636–1652

[28]

Stokkink P, Geroliminis N, (2025). On the optimal micro-hub locations in a multi-modal last-mile delivery system. Transportation Research Part E, Logistics and Transportation Review, 203: 104344

[29]

Sun B, Chen S, Meng Q, (2025). Optimizing first-and-last-mile ridesharing services with a heterogeneous vehicle fleet and time-dependent travel times. Transportation Research Part E, Logistics and Transportation Review, 193: 103847

[30]

Syrgkanis V, Zhan R, (2026). Post reinforcement learning inference. Operations Research, 74( 2): 917–957

[31]

Wang H, (2019). Routing and scheduling for a last-mile transportation system. Transportation Science, 53( 1): 131–147

[32]

Wang J, Gao Y, Cheng Y, (2022). On time-dependent critical platforms and tracks in metro systems. Transportation Science, 56( 4): 953–971

[33]

Wu W, Zhu Y, Liu R, (2024). Dynamic scheduling of flexible bus services with hybrid requests and fairness: heuristics-guided multi-agent reinforcement learning with imitation learning. Transportation Research Part B: Methodological, 190: 103069

[34]

Yang R, Xiao Z, Xiong W, Hou Y, (2026). Grid-Interactive and dual-side demand responsive timetable scheduling optimization for flexible urban metro systems. IEEE Transactions on Smart Grid,

[35]

Zhang Y, Chen L, Bao X, Su J, Zhang L, Zheng Y, (2026). Deep reinforcement learning-enhanced branch-and-price algorithm for integrated planning of berth allocation, quay crane assignment, and yard assignment. Computers & Industrial Engineering, 213: 111815

[36]

Zhou J, Xu X, Long J, Ding J, (2021). Integrated optimization approach to metro crew scheduling and rostering. Transportation Research Part C, Emerging Technologies, 123: 102975

RIGHTS & PERMISSIONS

Higher Education Press

PDF (4046KB)

339

Accesses

0

Citation

Detail

Sections
Recommended

/