An adaptive meta-reinforcement learning scheme for maritime dynamic spectrum access and service scheduling

Zhongyang MAO , Jiahuan GENG , Faping LU , Wenbiao TIAN , Jiafang KANG , Yaozong PAN

Eng Inform Technol Electron Eng ›› 2026, Vol. 27 ›› Issue (8) : 260047

PDF (2297KB)
Eng Inform Technol Electron Eng ›› 2026, Vol. 27 ›› Issue (8) :260047 DOI: 10.1631/ENG.ITEE.2026.0047
Research Article
An adaptive meta-reinforcement learning scheme for maritime dynamic spectrum access and service scheduling
Author information +
History +
PDF (2297KB)

Abstract

The rapid proliferation of marine economic activities has led to explosive growth in maritime communication demands. Consequently, the scarcity of spectrum resources, the highly dynamic spectrum environment, and diverse service requirements have become increasingly critical issues. While conventional deep reinforcement learning (DRL) methods perform well for specific training scenarios, they fail to generalize effectively to unknown environments. To address this challenge, this paper investigates a maritime dynamic spectrum access and service scheduling scheme based on adaptive meta-reinforcement learning. First, we construct a channel occupancy model that encompasses both Markov frequency-hopping mode and spread-spectrum frequency-hopping mode. By incorporating a dual-queue mechanism for urgent and normal data packets, we formulate the multi-agent cooperative decision-making problem as a decentralized partially observable Markov decision process (POMDP). Second, to overcome the limited generalization capability of traditional DRL algorithms, we propose a spectrum access and service scheduling algorithm based on a meta-learning paradigm (MetaSASS), which facilitates rapid policy transfer through meta-training across a distribution of tasks. Furthermore, to resolve the adaptation efficiency bottleneck caused by fixed inner-loop learning rates, we design an adaptive meta-reinforcement learning SASS algorithm (AMRLSASS). This algorithm employs a task encoding module to dynamically adjust learning rates, thereby balancing convergence speeds across tasks of varying complexities. Simulation results demonstrate that the proposed AMRLSASS algorithm outperforms baseline algorithms across a series of metrics within unknown environments and effectively validate the superiority of the proposed method in complex maritime environments. The algorithm proposed in this paper provides an effective solution for the rapid adaptation of intelligent spectrum access and service scheduling to unknown scenarios in maritime wireless communication systems.

Keywords

Dynamic spectrum access / Adaptive meta-reinforcement learning / Service scheduling / Maritime communication

Cite this article

Download citation ▾
Zhongyang MAO, Jiahuan GENG, Faping LU, Wenbiao TIAN, Jiafang KANG, Yaozong PAN. An adaptive meta-reinforcement learning scheme for maritime dynamic spectrum access and service scheduling. Eng Inform Technol Electron Eng, 2026, 27 (8) : 260047 DOI:10.1631/ENG.ITEE.2026.0047

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Albinsaid H , Singh K , Biswas S , et al., 2022. Multi-agent reinforcement learning-based distributed dynamic spectrum access. IEEE Trans Cogn Commun Netw, 8 (2): 1174- 1185.

[2]

Antoniou A , Edwards H , Storkey A , 2019. How to train your MAML.https://arxiv.org/abs/1810.09502.

[3]

Atimati E , Crawford D , Stewart R , 2023. Intelligent shared spectrum coordination in heterogeneous networks. IEEE Virtual Conf on Communications, p.252- 257.

[4]

Atimati E , Nyasulu T , Crawford D , et al., 2025. Resource management in dynamic shared spectrum networks. IEEE Int Symp on Dynamic Spectrum Access Networks, p.13- 19.

[5]

Dong L , Qian Y , Xing Y , 2022. Dynamic spectrum access and sharing through actorcritic deep reinforcement learning. EURASIP J Wirel Commun Netw, 2022: 48.

[6]

Feng MJ , Zhang WH , Krunz M , 2023. Dynamic spectrum access in non-stationary environments: a DRL-LSTM integrated approach. Int Conf on Computing, Networking and Communications, p.159- 164.

[7]

Finn C , Abbeel P , Levine S , 2017. Model-agnostic meta-learning for fast adaptation of deep networks. Proc 34th Int Conf on Machine Learning, p.1126- 1135.

[8]

Huang K , Luo ZZ , Liang L , et al., 2022. Fast spectrum sharing in vehicular networks: a meta reinforcement learning approach. IEEE 96th Vehicular Technology Conf, p.1- 5.

[9]

ITU , 2009. Characteristics of VHF Radio Systems and Equipment for the Exchange of Data and Electronic Mail in the Maritime Mobile Service RR Appendix 18 Channels. ITU-R M.1842-1-2009.

[10]

ITU , 2012. Interim Solutions for Improved Efficiency in the Use of the Band 156-174 MHz by Stations in the Maritime Mobile Service. ITU-R M.1084-5-2012.

[11]

ITU , 2024. Table of Transmitting Frequencies in the VHF Maritime Mobile Band. ITU Radio Regulations, Appendix 18.

[12]

Jia XY , Wang T , Du X , 2024. Federated multi-objective meta-reinforcement learning for adaptive edge task offloading. IEEE Int Conf on High Performance Computing and Communications, p.482- 489.

[13]

Kai H , Le L , Shi J , et al., 2025. Meta reinforcement learning for fast spectrum sharing in vehicular networks. China Commun, 22 (9): 320- 332.

[14]

Ke ZY , Wang XM , Du ZY , et al., 2025. Intelligent frequency reuse for dynamic spectrum anti-jamming: a hybrid-reward-based multi-agent deep reinforcement learning approach. IEEE Wirel Commun Lett, 14 (3): 771- 775.

[15]

Khisa S , Elhattab M , Assi C , et al., 2025. Optimizing multi-user uplink cooperative rate-splitting multiple access: efficient user pairing and resource allocation with gradient-based meta learning. IEEE Trans Commun, 73 (9): 7366- 7380.

[16]

Li XH , Zhang YL , Ding HC , et al., 2024. Intelligent spectrum sensing and access with partial observation based on hierarchical multi-agent deep reinforcement learning. IEEE Trans Wirel Commun, 23 (4): 3131- 3145.

[17]

Li YH , Wang Y , Li Y , et al., 2023. Multi-user dynamic spectrum access based on LRQ deep reinforcement learning network. 25th Int Conf on Advanced Communication Technology, p.79- 84.

[18]

Li YZ , Zhang WS , Wang CX , et al., 2020. Deep reinforcement learning for dynamic spectrum sensing and aggregation in multi-channel wireless networks. IEEE Trans Cogn Commun Netw, 6 (2): 464- 475.

[19]

Liu ZY , Wang XJ , Zhang Y , et al., 2023. Meta reinforcement learning for generalized multiple access in heterogeneous wireless networks. 21st Int Symp on Modeling and Optimization in Mobile, Ad Hoc, and Wireless Networks, p.570- 577.

[20]

Liu ZY , Wang XJ , Guo K , et al., 2024. Federated meta-RL based multiple access protocol for diverse heterogeneous wireless networks. Int Conf on Future Communications and Networks, p.1- 6.

[21]

Liu ZY , Wang XJ , Feng CY , et al., 2026. Meta-reinforcement learning with mixture of experts for generalizable multi access in heterogeneous wireless networks. IEEE Trans Commun, 74: 870- 885.

[22]

Lu ZY , Gursoy MC , 2021. Dynamic channel access via meta-reinforcement learning. IEEE Global Communications Conf, p.1- 6.

[23]

Lyu T , Xu HT , Liu FF , et al., 2024. Computing offloading and resource allocation of NOMA-based UAV emergency communication in marine Internet of Things. IEEE Int Things J, 11 (9): 15571- 15586.

[24]

Niu LW , Chen XF , Zhang N , et al., 2023. Multiagent meta-reinforcement learning for optimized task scheduling in heterogeneous edge computing systems. IEEE Int Things J, 10 (12): 10519- 10531.

[25]

Nomikos N , Gkonis PK , Bithas PS , et al., 2023. A survey on UAV-aided maritime communications: deployment considerations, applications, and future challenges. IEEE Open J Commun Soc, 4: 56- 78.

[26]

Rao N , Xu H , Qi ZS , et al., 2024. Fast adaptive jamming resource allocation against frequency-hopping spread spectrum in wireless sensor networks via meta-deepreinforcement-learning. IEEE Trans Aerosp Electron Syst, 60 (6): 7676- 7693.

[27]

Rao N , Xu H , Qi ZS , et al., 2025. Adaptive jamming decision-making against FHSS communications via inexpert demonstrations assisted meta reinforcement learning. IEEE Commun Lett, 29 (1): 105- 109.

[28]

Saggese F , Pasqualini L , Moretti M , et al., 2021. Deep reinforcement learning for URLLC data management on top of scheduled eMBB traffic. IEEE Global Communications Conf, p.1- 6.

[29]

Saggese F , Moretti M , Popovski P , 2022. NOMA power minimization of downlink spectrum slicing for eMBB and URLLC users. IEEE Wireless Communications and Networking Conf, p.1725- 1730.

[30]

Sheng HM , Zhou WJ , Zheng JJ , et al., 2024. Transfer reinforcement learning for dynamic spectrum environment. IEEE Trans Wirel Commun, 23 (2): 1447- 1458.

[31]

Sheng TQ , Zhang WS , Ding WJ , et al., 2022. Dynamic spectrum sharing and aggregation scheme based on deep reinforcement learning. Int Wireless Communications and Mobile Computing, p.290- 294.

[32]

So H , Soya H , 2023. Excluded channel selection scheme for dynamic spectrum sharing in various interference environments. IEEE Access, 11: 69798- 69806.

[33]

Upadhyay D , Upadhyay A , Venu N , et al., 2025. Deep learning-based spectrum sharing for dynamic resource allocation in 6G cognitive radio networks. 4th Int Conf on Power, Control and Computing Technologies, p.1- 6.

[34]

Wang QF , Xu WQ , Chen HH , 2026. A heterogeneous-agent deep reinforcement learning approach for dynamic spectrum access in cognitive wireless networks. IEEE Trans Cogn Commun Netw, 12: 2221- 2235.

[35]

Yang TT , Gao S , Li JB , et al., 2022. Multi-armed bandits learning for task offloading in maritime edge intelligence networks. IEEE Trans Veh Technol, 71 (4): 4212- 4224.

[36]

Yang TT , Zhang WS , Bo YL , et al., 2023. Dynamic spectrum sharing based on federated learning and multi-agent actor-critic reinforcement learning. Int Wireless Communications and Mobile Computing, p.947- 952.

[37]

Yuan L , Zhou FH , Wu QH , et al., 2024. Channel prediction-enhanced intelligent resource allocation for dynamic spectrum-sharing networks. IEEE Int Conf on Communications, p.2767- 2772.

[38]

Zhang SG , Wang Z , Gao GY , et al., 2023. Deep reinforcement learning for UAV-assisted spectrum sharing under partial observability. IEEE 98th Vehicular Technology Conf, p.1- 6.

[39]

Zhang X , Bhuyan A , Kasera SK , et al., 2023. Distributed power allocation for 6-GHz unlicensed spectrum sharing via multi-agent deep reinforcement learning. IEEE Int Conf on Industrial Technology, p.1- 6.

[40]

Zhang YL , Li XH , Ding HC , et al., 2023. A joint scheme on spectrum sensing and access with partial observation: a multi-agent deep reinforcement learning approach. IEEE/CIC Int Conf on Communications in China, p.1- 6.

Rights & permissions

The Authors. Published by Zhejiang University Press Co., Ltd.

PDF (2297KB)

Supplementary files

EITEE20260807-ZYM-ESM

EITEE20260807-ZYM-suppl1

EITEE20260807-ZYM-suppl2

9

Accesses

0

Citation

Detail

Sections
Recommended

/