Adaptive optimal consensus over networked multi-agent systems: a digital signal-driven hybrid iteration approach

Jun Li , Jiansong Lu , Lianghao Ji , Huaqing Li

›› 2026, Vol. 12 ›› Issue (5) : 825 -835.

PDF (1859KB)
›› 2026, Vol. 12 ›› Issue (5) :825 -835. DOI: 10.1016/j.dcan.2025.07.010
research-article
Adaptive optimal consensus over networked multi-agent systems: a digital signal-driven hybrid iteration approach
Author information +
History +
PDF (1859KB)

Abstract

The integration of Digital Signal Processing (DSP) and Reinforcement Learning (RL) for optimal consensus control in Networked Multi-Agent Systems (NMASs) has garnered significant research attention. However, prior research encounters some limitations: 1) dependency on initial admissible control policies, and 2) systemic data redundancy arising from ineffective data governance in distributed architectures and slow convergence rates of conventional RL algorithms. To overcome these challenges, this paper proposes a Distributed Collaborative Iteration Adaptive Dynamic Programming (DCIADP) framework. The methodology reformulates the solution of Hamilton-Jacobi-Bellman (HJB) equations by integrating Value Iteration (VI) and Policy Iteration (PI) within a unified architecture, eliminating reliance on prior knowledge of system dynamics. Specifically, a dynamic factor is introduced to synergistically integrate the complementary strengths of VI and PI, achieving accelerated convergence while bypassing the initialization requirement for admissible policies. This innovation significantly mitigates computational overhead in distributed nodes during localized DSP operations. Furthermore, a self-tuning mechanism dynamically optimizes the factor, enhancing adaptability to heterogeneous network conditions. Through rigorous theoretical analysis, the proposed framework is proven to ensure asymptotic convergence and Lyapunov stability. Practical implementation is realized through actor-critic Neural Networks (NNs), incorporating an experience replay mechanism to exploit temporal correlation characteristics in networked data streams. This enables derivation of optimal control policies solely from transmitted network signals, independent of explicit system parameter knowledge. The framework thus establishes a resource-efficient adaptive control paradigm for bandwidth-constrained networked MASs. Finally, several numerical simulations validate the effectiveness and superiority of the proposed approach.

Keywords

Off-policy / Optimal control / Multi-agent systems / Reinforcement learning / Digital signal processing

Cite this article

Download citation ▾
Jun Li, Jiansong Lu, Lianghao Ji, Huaqing Li. Adaptive optimal consensus over networked multi-agent systems: a digital signal-driven hybrid iteration approach. , 2026, 12 (5) : 825-835 DOI:10.1016/j.dcan.2025.07.010

登录浏览全文

4963

注册一个新账户 忘记密码

CRediT authorship contribution statement

Jun Li: Methodology, Investigation, Funding acquisition, Formal analysis, Conceptualization. Jiansong Lu: Methodology, Investigation. Lianghao Ji: Methodology, Investigation, Funding acquisition. Huaqing Li: Methodology, Investigation, Funding acquisition.

Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Acknowledgements

This work was supported in part by the National Natural Science Foundation of China under Grant No. 62276036, in part by the Innovation and Development Joint Fund Project of Chongqing Natural Science Foundation under Grant No. CSTB2024NSCQ-LZX0118, and in part by the National Natural Science Foundation of China under Grant No. 62173278.

References

[1]

R. Olfati—Saber, J.A. Fax, R.M. Murray, Consensus and cooperation in networked multi—agent systems, Proc. IEEE 95 (1) (2007) 215-233.

[2]

L. Zhou, J. Liu, Y. Zheng, F. Xiao, J. Xi, Game—based consensus of hybrid multiagent systems, IEEE Trans. Cybern. 53 (8) (2023) 5346-5357.

[3]

M.—P. Huget, Communication in Multiagent Systems, Springer Press, 2003.

[4]

S. Qiang, L. Liu, The implement of blackboard—based multi—agent intelligent decision support system, in: 2010 Second International Conference on Computer Engineering and Applications, vol. 1, 2010, pp. 572-575.

[5]

K.—Y. Chen, C.—J. Chen, Applying multi—agent technique in multi—section flexible manufacturing system, Expert Syst. Appl. 37 (11) (2010) 7310-7318.

[6]

H. Shi, Z. Zhao, J. Chen, M. Zhou, Y. Liu, Enhancing unmanned aerial vehicle path planning in multi—agent reinforcement learning through adaptive dimensionality reduction, Drones 8 (10) (2024) 521.

[7]

Y. Quan, L. Xi, Smart generation system: a decentralized multi—agent control architecture based on improved consensus algorithm for generation command dispatch of sustainable energy systems, Appl. Energy 365 (2024) 123209.

[8]

A. Perrusquía, W. Guo, Uncovering reward goals in distributed drone swarms using physics—informed multiagent inverse reinforcement learning, IEEE Trans. Cybern. 55 (1) (2025) 14-23.

[9]

P.H. Luzolo, Z. Elrawashdeh, I. Tchappi, S. Galland, F. Outay, Combining multi—agent systems and artificial intelligence of things: technical challenges and gains, Internet of Things 28 (2024) 101364.

[10]

C. Hua—Min, W. Shou—Feng, P. Wang, S. Lin, F. Chao, Deep q—learning for intelligent band coordination in 5g heterogeneous network supporting v2x communication, Wirel. Commun. Mob. Comput. (2022), https://doi.org/10.1155/2022/9653334.

[11]

X. Jin, S. , C. Deng, M. Chadli, Distributed adaptive security consensus control for a class of multi—agent systems under network decay and intermittent attacks, Inf. Sci. 547 (2021) 88-102.

[12]

A. Hu, Y. Wang, J. Cao, A. Alsaedi, Event—triggered bipartite consensus of multiagent systems with switching partial couplings and topologies, Inf. Sci. 521 (2020) 1-13.

[13]

L. Ji, Z. Lin, C. Zhang, S. Yang, J. Li, H. Li, Data—based optimal consensus control for multiagent systems with time delays: using prioritized experience replay, IEEE Trans. Syst. Man Cybern. Syst. 54 (5) (2024) 3244-3256.

[14]

C. Chen, F.L. Lewis, K. Xie, S. Xie, Y. Liu, Off—policy learning for adaptive optimal output synchronization of heterogeneous multi—agent systems, Automatica 119 (2020) 109081.

[15]

D. Liang, Y. Dong, C. Wang, G. Zhai, Data—driven cooperative output regulation of linear discrete—time multiagent systems with unknown dynamics, IEEE Trans. Syst. Man Cybern. Syst. 54 (8) (2024) 5025-5034.

[16]

J. Zhao, K. Zhu, H. Hu, X. Yu, X. Li, H. Wang, Formation control of networked mobile robots with unknown reference orientation, IEEE/ASME Trans. Mechatron. 28 (4) (2023) 2200-2212.

[17]

J. Yang, J. Dai, H.B. Gooi, H.D. Nguyen, P. Wang, Hierarchical blockchain design for distributed control and energy trading within microgrids, IEEE Trans. Smart Grid 13 (4) (2022) 3133-3144.

[18]

R. Yang, H. Zhang, G. Feng, H. Yan, Z. Wang, Robust cooperative output regulation of multi—agent systems via adaptive event—triggered control, Automatica 102 (2019) 129-136.

[19]

K.G. Vamvoudakis, F.L. Lewis, G.R. Hudas, Multi—agent differential graphical games: online adaptive learning solution for synchronization with optimality, Automatica 48 (8) (2012) 1598-1611.

[20]

W. Wang, H. Zhang, A new and effective nonparametric variable step—size normalized least—mean—square algorithm and its performance analysis, Signal Process. 210 (2023) 109060.

[21]

H. Zhu, M. Zhang, Y. Suo, T.D. Tran, J. Van der Spiegel, Design of a digital address event triggered compressive acquisition image sensor, IEEE Trans. Circuits Syst. I, Regul. Pap. 63 (2) (2016) 191-199.

[22]

E. Twahirwa, J. Rwigema, R. Datta, Design and deployment of vehicular Internet of things for smart city applications, Sustainability 14 (1) (2022) 176.

[23]

K.I. Seong, S.O. Hwang, Load balancing in decentralized smart grid trade system using blockchain, J. Intell. Fuzzy Syst. 35 (6) (2018) 5901-5911.

[24]

Y. Yan, B. Zhang, C. Li, J. Bai, Z. Yao, A novel model—assisted decentralized multi—agent reinforcement learning for joint optimization of hybrid beamforming in massive mimo mmwave systems, IEEE Trans. Veh. Technol. 72 (11) (2023) 14743-14755.

[25]

A. Ornelas—Gutierrez, C. Vargas—Rosales, R. Villalpando—Hernandez, J. Zuniga—Mejia, Sinr management in roundabout vehicular ad hoc networks: a control and reinforcement learning with digital beamforming approach, IEEE Access 12 (2024) 85580-85600.

[26]

R.S. Sutton, A.G. Barto, Reinforcement Learning: An Introduction, MIT Press, 2018.

[27]

V. Mnih, K. Kavukcuoglu, D. Silver, A.A. Rusu, J. Veness, M.G. Bellemare, A. Graves, M. Riedmiller, A.K. Fidjeland, G. Ostrovski, et al., Human—level control through deep reinforcement learning, Nature 518 (7540) (2015) 529-533.

[28]

D. Silver, A. Huang, C.J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al., Mastering the game of go with deep neural networks and tree search, Nature 529 (7587) (2016) 484-489.

[29]

H. Jiang, H. Zhang, G. Xiao, X. Cui, Data—based approximate optimal control for nonzero—sum games of multi—player systems using adaptive dynamic programming, Neurocomputing 275 (2018) 192-199.

[30]

H. Jiang, H. Zhang, Y. Luo, X. Cui, H control with constrained input for completely unknown nonlinear systems using data—driven reinforcement learning method , Neurocomputing 237 (2017) 226-234.

[31]

X. Bu, Z. Hou, H. Zhang, Data—driven multiagent systems consensus tracking using model free adaptive control, IEEE Trans. Neural Netw. Learn. Syst. 29 (5) (2017) 1514-1524.

[32]

B. Hu, Z.—H. Guan, F.L. Lewis, C.P. Chen, Adaptive tracking control of cooperative robot manipulators with Markovian switched couplings, IEEE Trans. Ind. Electron. 68 (3) (2020) 2427-2436.

[33]

B. Luo, D. Liu, H. Wu, D. Wang, F.L. Lewis, Policy gradient adaptive dynamic programming for data—based optimal control, IEEE Trans. Cybern. 47 (10) (2016) 3341-3354.

[34]

X. Yang, H. Zhang, Z. Wang, Data—based optimal consensus control for multiagent systems with policy gradient reinforcement learning, IEEE Trans. Neural Netw. Learn. Syst. 33 (8) (2021) 3872-3883.

[35]

X. Xu, R. Li, Z. Zhao, H. Zhang, The gradient convergence bound of federated multiagent reinforcement learning with efficient communication, IEEE Trans. Wirel. Commun. 23 (1) (2023) 507-528.

[36]

C. Chen, F.L. Lewis, K. Xie, S. Xie, Adaptive optimal control of unknown nonlinear systems via homotopy—based policy iteration, IEEE Trans. Autom. Control 69 (5) (2023) 3396-3403.

[37]

H. Jiang, B. Zhou, G.R. Duan, Modified 𝜆—policy iteration based adaptive dynamic programming for unknown discrete—time linear systems , IEEE Trans. Neural Netw. Learn. Syst. 35 (3) (2023) 3291-3301.

[38]

M.I. Abouheaf, F.L. Lewis, Multi—agent differential graphical games: Nash online adaptive learning solutions, in: Proc. 52nd IEEE Conf. Decis, Control (CDC)., IEEE, 2013, pp. 5803-5809.

[39]

C. Xiong, Q. Ma, J. Guo, F.L. Lewis, Data—based optimal synchronization of heterogeneous multiagent systems in graphical games via reinforcement learning, IEEE Trans. Neural Netw. Learn. Syst. (2023) 1-9.

[40]

B. Luo, H. Wu, T. Huang, D. Liu, Data—based approximate policy iteration for affine nonlinear continuous—time optimal control design, Automatica 50 (12) (2014) 3281-3290.

[41]

A. Al—Tamimi, F.L. Lewis, M. Abu—Khalaf, Discrete—time nonlinear HJB solution using approximate dynamic programming: convergence proof, IEEE Trans. Syst. Man Cybern., Part B, Cybern. 38 (4) (2008) 943-949.

[42]

P. Li, W. Zou, J. Guo, Z. Xiang, Optimal consensus of a class of discrete—time linear multi—agent systems via value iteration with guaranteed admissibility, Neurocomputing 516 (2023) 1-10.

[43]

L. Ji, C. Wang, C. Zhang, H. Wang, H. Li, Optimal consensus model—free control for multi—agent systems subject to input delays and switching topologies, Inf. Sci. 589 (2022) 497-515.

[44]

O. Qasem, W. Gao, T. Bian, Adaptive optimal control of continuous—time linear systems via hybrid iteration, in: Proceeding of 2021 IEEE Symposium Series on Computational Intelligence (SSCI), IEEE, 2021, pp. 1-07.

[45]

B. Luo, Y. Yang, H. Wu, T. Huang, Balancing value iteration and policy iteration for discrete—time control, IEEE Trans. Syst. Man Cybern. Syst. 50 (11) (2019) 3948-3958.

[46]

B. Piot, M. Geist, O. Pietquin, Bridging the gap between imitation learning and inverse reinforcement learning, IEEE Trans. Neural Netw. Learn. Syst. 28 (8) (2016) 1814-1826.

[47]

A. Oroojlooy, D. Hajinezhad, A review of cooperative multi—agent deep reinforcement learning, Appl. Intell. (2022) 1-46.

[48]

Q. Wang, H.E. Psillakis, C. Sun, Cooperative control of multiple agents with unknown high—frequency gain signs under unbalanced and switching topologies, IEEE Trans. Autom. Control 64 (6) (2018) 2495-2501.

[49]

S. Changyin, M. Chao Xu, Important scientific problems of multi—agent deep reinforcement learning, Acta Autom. Sin. 46 (7) (2020) 1301-1312.

[50]

J. Fan, D. Li, R. Li, T. Yang, Q. Wang, Analysis for cooperative combat system of manned—unmanned aerial vehicles and combat simulation, in: 2017 IEEE International Conference on Unmanned Systems (ICUS), 2017, pp. 204-209.

[51]

C.—M. Yu, M.—L. Ku, L.—C. Wang, Dtc—hsr: distributed topology control and hierarchical self—routing for bluetooth load balancing networks, IEEE Internet Things J. 9 (20) (2022) 19545-19560.

[52]

Y. Zheng, Q. Zhao, J. Ma, L. Wang, Second—order consensus of hybrid multi—agent systems, Syst. Control Lett. 125 (2019) 51-58.

[53]

J. Li, L. Ji, Optimal consensus control for second—order discrete—time multi—agent systems: using online policy iteration algorithm, in: Proceeding of 2020 IEEE Symposium Series on Computational Intelligence (SSCI), IEEE, 2020, pp. 2210-2216.

[54]

N. An, Q. Wang, X. Zhao, Q. Wang, Discrete—time nonlinear optimal control using multi—step reinforcement learning, IEEE Trans. Circuits Syst. II, Express Briefs 71 (4) (2023) 2279-2283.

[55]

H. Zhang, D. Yue, C. Dou, W. Zhao, X. Xie, Data—driven distributed optimal consensus control for unknown multiagent systems with input—delay, IEEE Trans. Cybern. 49 (6) (2018) 2095-2105.

PDF (1859KB)

2

Accesses

0

Citation

Detail

Sections
Recommended

/