TY - GEN
T1 - Multi-Agent Cooperative Attitude Takeover of Defunct Satellites with Unknown Dynamics via Off-Policy Reinforcement Learning
AU - Huang, Xintong
AU - Shen, Ganghui
AU - Zhang, Yizhai
AU - Ma, Zhiqiang
AU - Ma, Tianze
AU - Huang, Panfeng
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - This paper proposes a model-free Off-Policy reinforcement learning cooperative control method to address the attitude stabilization problem during the cooperative takeover of a defunct large satellite by multiple microsatellites. Upon attachment, the inertia matrix of the combined assembly undergoes significant changes with high uncertainty, rendering traditional model-based strategies such as computed torque control ineffective. The problem is thus formulated as a multi-satellite team cooperative game, and a data-driven off-policy algorithm solves the multi-input coupled Hamilton-Jacobi-Bellman (HJB) equation. Using Integral Reinforcement Learning (IRL), the dependence on system drift dynamics and control input matrices is eliminated through integral operations along system trajectories. This transforms the differential HJB equation, which involves unknown dynamic parameters, into an integral algebraic equation solvable using only input-output data. By introducing exploration noise to satisfy the persistence of excitation condition, the algorithm learns optimal cooperative policies offline from historical data, mitigating safety risks of on-policy trial-and-error and improving data utilization. Simulation results demonstrate that, without any prior knowledge of the assembly's inertial parameters, the proposed method achieves optimal cooperative cost minimization, stabilizing the defunct satellite's attitude with rapid convergence and smooth transient response.
AB - This paper proposes a model-free Off-Policy reinforcement learning cooperative control method to address the attitude stabilization problem during the cooperative takeover of a defunct large satellite by multiple microsatellites. Upon attachment, the inertia matrix of the combined assembly undergoes significant changes with high uncertainty, rendering traditional model-based strategies such as computed torque control ineffective. The problem is thus formulated as a multi-satellite team cooperative game, and a data-driven off-policy algorithm solves the multi-input coupled Hamilton-Jacobi-Bellman (HJB) equation. Using Integral Reinforcement Learning (IRL), the dependence on system drift dynamics and control input matrices is eliminated through integral operations along system trajectories. This transforms the differential HJB equation, which involves unknown dynamic parameters, into an integral algebraic equation solvable using only input-output data. By introducing exploration noise to satisfy the persistence of excitation condition, the algorithm learns optimal cooperative policies offline from historical data, mitigating safety risks of on-policy trial-and-error and improving data utilization. Simulation results demonstrate that, without any prior knowledge of the assembly's inertial parameters, the proposed method achieves optimal cooperative cost minimization, stabilizing the defunct satellite's attitude with rapid convergence and smooth transient response.
KW - Attitude control
KW - Defunct satellite takeover
KW - Hamilton-Jacobi-Bellman equation
KW - Model-free control
KW - Multi-agent systems
KW - Off-policy reinforcement learning
UR - https://www.scopus.com/pages/publications/105044213179
U2 - 10.1109/ICAISISAS68969.2026.11567712
DO - 10.1109/ICAISISAS68969.2026.11567712
M3 - 会议稿件
AN - SCOPUS:105044213179
T3 - 2026 Joint International Conference on Automation-Intelligence-Safety, ICAIS 2026 and International Symposium on Autonomous Systems, ISAS 2026
BT - 2026 Joint International Conference on Automation-Intelligence-Safety, ICAIS 2026 and International Symposium on Autonomous Systems, ISAS 2026
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2026 Joint International Conference on Automation-Intelligence-Safety, ICAIS 2026 and International Symposium on Autonomous Systems, ISAS 2026
Y2 - 8 May 2026 through 10 May 2026
ER -