Adaptive optimal consensus over networked multi-agent systems: a digital signal-driven hybrid iteration approach✩
Jun Li , Jiansong Lu , Lianghao Ji , Huaqing Li
›› 2026, Vol. 12 ›› Issue (5) : 825 -835.
The integration of Digital Signal Processing (DSP) and Reinforcement Learning (RL) for optimal consensus control in Networked Multi-Agent Systems (NMASs) has garnered significant research attention. However, prior research encounters some limitations: 1) dependency on initial admissible control policies, and 2) systemic data redundancy arising from ineffective data governance in distributed architectures and slow convergence rates of conventional RL algorithms. To overcome these challenges, this paper proposes a Distributed Collaborative Iteration Adaptive Dynamic Programming (DCIADP) framework. The methodology reformulates the solution of Hamilton-Jacobi-Bellman (HJB) equations by integrating Value Iteration (VI) and Policy Iteration (PI) within a unified architecture, eliminating reliance on prior knowledge of system dynamics. Specifically, a dynamic factor is introduced to synergistically integrate the complementary strengths of VI and PI, achieving accelerated convergence while bypassing the initialization requirement for admissible policies. This innovation significantly mitigates computational overhead in distributed nodes during localized DSP operations. Furthermore, a self-tuning mechanism dynamically optimizes the factor, enhancing adaptability to heterogeneous network conditions. Through rigorous theoretical analysis, the proposed framework is proven to ensure asymptotic convergence and Lyapunov stability. Practical implementation is realized through actor-critic Neural Networks (NNs), incorporating an experience replay mechanism to exploit temporal correlation characteristics in networked data streams. This enables derivation of optimal control policies solely from transmitted network signals, independent of explicit system parameter knowledge. The framework thus establishes a resource-efficient adaptive control paradigm for bandwidth-constrained networked MASs. Finally, several numerical simulations validate the effectiveness and superiority of the proposed approach.
Off-policy / Optimal control / Multi-agent systems / Reinforcement learning / Digital signal processing
| [1] |
|
| [2] |
|
| [3] |
|
| [4] |
|
| [5] |
|
| [6] |
|
| [7] |
|
| [8] |
|
| [9] |
|
| [10] |
|
| [11] |
|
| [12] |
|
| [13] |
|
| [14] |
|
| [15] |
|
| [16] |
|
| [17] |
|
| [18] |
|
| [19] |
|
| [20] |
|
| [21] |
|
| [22] |
|
| [23] |
|
| [24] |
|
| [25] |
|
| [26] |
|
| [27] |
|
| [28] |
|
| [29] |
|
| [30] |
|
| [31] |
|
| [32] |
|
| [33] |
|
| [34] |
|
| [35] |
|
| [36] |
|
| [37] |
|
| [38] |
|
| [39] |
|
| [40] |
|
| [41] |
|
| [42] |
|
| [43] |
|
| [44] |
|
| [45] |
|
| [46] |
|
| [47] |
|
| [48] |
|
| [49] |
|
| [50] |
|
| [51] |
|
| [52] |
|
| [53] |
|
| [54] |
|
| [55] |
|
/
| 〈 |
|
〉 |