Path planning of autonomous underwater vehicle for data collection of the Internet of everything

Desheng Chen , Meng Xi , Jiabao Wen , Jingyi He , Huiao Dai , Wenjie Li

›› 2026, Vol. 12 ›› Issue (3) : 520 -527.

PDF (3087KB)
›› 2026, Vol. 12 ›› Issue (3) :520 -527. DOI: 10.1016/j.dcan.2024.10.004
Research article
research-article
Path planning of autonomous underwater vehicle for data collection of the Internet of everything
Author information +
History +
PDF (3087KB)

Abstract

Autonomous Underwater Vehicle (AUV) has become an important tool to accomplish various path planning tasks due to its high intelligence and good maneuverability. Aiming at the problem of data collection at underwater Internet of Everything (IoE) nodes, this paper constructs a complex 3D marine environment based on real marine current data, and proposes a path planning algorithm based on reinforcement learning to ensure that the AUV completes the data collection with a short path length. In particular, in order to address the problem of complex path planning tasks, the Parallel Dense neural Network (PDNet) is proposed to improve the performance of the agent by extracting the core features of the input state. In addition, to simplify the reward shaping, we constructed a marine environment with sparse rewards. Sparse rewards can greatly interfere with the agent’s exploration and learning. To solve the sparse reward problem, the Hindsight Experience Replay (HER) is introduced, which not only solves the sparse reward problem, but also improves the sampling efficiency and convergence of the algorithm.

Keywords

Internet of everything / Autonomous underwater vehicles / Path planning / Deep reinforcement learning

Cite this article

Download citation ▾
Desheng Chen, Meng Xi, Jiabao Wen, Jingyi He, Huiao Dai, Wenjie Li. Path planning of autonomous underwater vehicle for data collection of the Internet of everything. , 2026, 12 (3) : 520-527 DOI:10.1016/j.dcan.2024.10.004

登录浏览全文

4963

注册一个新账户 忘记密码

CRediT authorship contribution statement

Desheng Chen: Methodology, Formal analysis. Meng Xi: Writing – review & editing, Writing – original draft, Methodology, Formal analysis. Jiabao Wen: Methodology, Investigation, Funding acquisition, Conceptualization. Jingyi He: Formal analysis, Conceptualization. Huiao Dai: Investigation, Conceptualization. Wenjie Li: Writing – review & editing, Writing – original draft, Validation, Conceptualization.

Funding

This work was supported by the National Natural Science Foundation of China under Grant 62306211, 62403349, and China Postdoctoral Science Foundation 2023M742608, and Postdoctoral Fellowship Program of CPSF GZC20231919.

Declaration of competing interest

No potential conflict of interest was reported by the authors.

References

[1]

Y. Liu, P. Huang, F. Zhang, Y. Zhao, Distributed formation control using artificial potentials and neural network for constrained multiagent systems, IEEE Trans. Control Syst. Technol. 28 (2) (2018) 697-704.

[2]

P. Khatun, C.M. Bingham, N. Schofield, P. Mellor, Application of fuzzy control algorithms for electric vehicle antilock braking/traction control systems, IEEE Trans. Veh. Technol. 52 (5) (2003) 1356-1364.

[3]

G. Chen, T. Wu, Z. Zhou, Research on ship meteorological route based on a—star algorithm, Math. Probl. Eng. 2021 (2021) 1-8.

[4]

Y.—N. Ma, Y.—J. Gong, C.—F. Xiao, Y. Gao, J. Zhang, Path planning for autonomous underwater vehicles: an ant colony algorithm incorporating alarm pheromone, IEEE Trans. Veh. Technol. 68 (1) (2018) 141-154.

[5]

M. Chen, D. Zhu, Optimal time—consuming path planning for autonomous underwater vehicles based on a dynamic neural network model in ocean current environments, IEEE Trans. Veh. Technol. 69 (12) (2020) 14401-14412.

[6]

V. Roberge, M. Tarbouchi, G. Labonté, Comparison of parallel genetic algorithm and particle swarm optimization for real—time uav path planning, IEEE Trans. Ind. Inform. 9 (1) (2012) 132-141.

[7]

J. Yang, J. Wen, Y. Wang, B. Jiang, H. Wang, H. Song, Fog—based marine environmental information monitoring toward ocean of things, IEEE Int. Things J. 7 (5) (2019) 4238-4247.

[8]

J. Yan, W. Cao, X. Yang, C. Chen, X. Guan, Communication—efficient and collision—free motion planning of underwater vehicles via integral reinforcement learning, IEEE Trans. Neural Netw. Learn. Syst. 35 (6) (2024) 8306-8320, https://doi.org/10.1109/TNNLS.2022.3226776.

[9]

J. Yan, X. Li, X. Yang, X. Luo, C. Hua, X. Guan, Integrated localization and tracking for auv with model uncertainties via scalable sampling—based reinforcement learning approach, IEEE Trans. Syst. Man Cybern. Syst. 52 (11) (2021) 6952-6967.

[10]

M. Samir, C. Assi, S. Sharafeddine, D. Ebrahimi, A. Ghrayeb, Age of information aware trajectory planning of uavs in intelligent transportation systems: a deep learning approach, IEEE Trans. Veh. Technol. 69 (11) (2020) 12382-12395.

[11]

V. Mnih, K. Kavukcuoglu, D. Silver, A.A. Rusu, J. Veness, M.G. Bellemare, A. Graves, M. Riedmiller, A.K. Fidjeland, G. Ostrovski, et al., Human—level control through deep reinforcement learning, Nature 518 (7540) (2015) 529-533.

[12]

J. Yang, S. Xiao, B. Jiang, H. Song, S. Khan, S. Ul Islam, Cache—enabled unmanned aerial vehicles for cooperative cognitive radio networks, IEEE Wirel. Commun. 27 (2) (2020) 155-161.

[13]

X. Li, L. Lu, W. Ni, A. Jamalipour, D. Zhang, H. Du, Federated multi—agent deep reinforcement learning for resource allocation of vehicle—to—vehicle communications, IEEE Trans. Veh. Technol. 71 (8) (2022) 8810-8824.

[14]

K. Li, W. Ni, E. Tovar, A. Jamalipour, On—board deep q—network for uav—assisted online power transfer and data collection, IEEE Trans. Veh. Technol. 68 (12) (2019) 12215-12226.

[15]

J. Yang, J. Huo, M. Xi, J. He, Z. Li, H.H. Song, A time—saving path planning scheme for autonomous underwater vehicles with complex underwater conditions, IEEE Int. Things J. 10 (2) (2022) 1001-1013.

[16]

J. Yang, J. Ni, Y. Li, J. Wen, D. Chen, The intelligent path planning system of agricultural robot via reinforcement learning, Sensors 22 (12) (2022) 4316.

[17]

J. Wang, W. Chi, C. Li, C. Wang, M.Q.—H. Meng, Neural rrt*: learning—based optimal path planning, IEEE Trans. Autom. Sci. Eng. 17 (4) (2020) 1748-1758.

[18]

Y. Li, R. Cui, Z. Li, D. Xu, Neural network approximation based near—optimal motion planning with kinodynamic constraints using rrt, IEEE Trans. Ind. Electron. 65 (11) (2018) 8718-8729.

[19]

J. Wen, J. Yang, T. Wang, Path planning for autonomous underwater vehicles under the influence of ocean currents based on a fusion heuristic algorithm, IEEE Trans. Veh. Technol. 70 (9) (2021) 8529-8544.

[20]

L. Lv, S. Zhang, D. Ding, Y. Wang, Path planning via an improved dqn—based learning policy, IEEE Access 7 (2019) 67319-67330.

[21]

M. Xi, J. Yang, J. Wen, H. Liu, Y. Li, H.H. Song, Comprehensive ocean information—enabled auv path planning via reinforcement learning, IEEE Int. Things J. 9 (18) (2022) 17440-17451.

[22]

H. Xie, D. Yang, L. Xiao, J. Lyu, Connectivity—aware 3d uav path design with deep reinforcement learning, IEEE Trans. Veh. Technol. 70 (12) (2021) 13022-13034.

[23]

D. Hong, S. Lee, Y.H. Cho, D. Baek, J. Kim, N. Chang, Energy—efficient online path planning of multiple drones using reinforcement learning, IEEE Trans. Veh. Technol. 70 (10) (2021) 9725-9740.

[24]

K. Cobbe, O. Klimov, C. Hesse, T. Kim, J. Schulman, Quantifying generalization in reinforcement learning, in: International Conference on Machine Learning, PMLR, 2019, pp. 1282-1289.

[25]

M. Igl, K. Ciosek, Y. Li, S. Tschiatschek, C. Zhang, S. Devlin, K. Hofmann, Generalization in reinforcement learning with selective noise injection and information bottleneck, in: NeurIPS, 2019, pp. 13956-13968.

[26]

Q. Wu, K. Xu, J. Wang, M. Xu, X. Gong, D. Manocha, Reinforcement learning—based visual navigation with information—theoretic regularization, IEEE Robot. Autom. Lett. 6 (2) (2021) 731-738.

[27]

X. Lu, K. Lee, P. Abbeel, S. Tiomkin, Dynamics generalization via information bottleneck in deep reinforcement learning, arXiv preprint arXiv:2008.00614.

[28]

Y. Burda, H. Edwards, A. Storkey, O. Klimov, Exploration by random network distillation, in: Seventh International Conference on Learning Representations, 2019, pp. 1-17.

[29]

A.P. Badia, P. Sprechmann, A. Vitvitskyi, Z.D. Guo, B. Piot, S. Kapturowski, O. Tieleman, M. Arjovsky, A. Pritzel, A. Bolt, C. Blundell, Never give up: learning directed exploration strategies, in: ICLR, OpenReview.net, 2020.

[30]

T.D. Kulkarni, K.R. Narasimhan, A. Saeedi, J.B. Tenenbaum, Hierarchical deep reinforcement learning: integrating temporal abstraction and intrinsic motivation, in: Proceedings of the 30th International Conference on Neural Information Processing Systems, NIPS’16, Curran Associates Inc., 2016, pp. 3682-3690.

[31]

A.S. Vezhnevets, S. Osindero, T. Schaul, N. Heess, M. Jaderberg, D. Silver, K. Kavukcuoglu, Feudal networks for hierarchical reinforcement learning, in: International Conference on Machine Learning, PMLR, 2017, pp. 3540-3549.

[32]

C. Wang, J. Wang, J. Wang, X. Zhang, Deep—reinforcement—learning—based autonomous uav navigation with sparse rewards, IEEE Int. Things J. 7 (7) (2020) 6180-6190.

[33]

H. Song, C.—C. Liu, J. Lawarrée, R.W. Dahlgren, Optimal electricity supply bidding by Markov decision process, IEEE Trans. Power Syst. 15 (2) (2000) 618-624.

[34]

J. Schulman, F. Wolski, P. Dhariwal, A. Radford, O. Klimov, Proximal policy optimization algorithms, arXiv preprint arXiv:1707.06347.

[35]

K. Schulz, L. Sixt, F. Tombari, T. Landgraf, Restricting the flow: information bottlenecks for attribution, in: ICLR, OpenReview.net, 2020.

PDF (3087KB)

2

Accesses

0

Citation

Detail

Sections
Recommended

/