TY - JOUR
T1 - Spacecraft on-orbit observation maneuver decision-making method based on multi-policy learning
AU - Jia, Zhenshuai
AU - Xiao, Bing
AU - Qian, Hanyu
AU - Zhang, Zheyu
N1 - Publisher Copyright:
© 2026, Chinese Institute of Electronics. All rights reserved.
PY - 2026/5/27
Y1 - 2026/5/27
N2 - A spacecraft multi-stage maneuver decision-making method based on deep reinforcement learning is proposed to address the maneuver decision problem for spacecraft approaching space targets during on-orbit observation service. Firstly, the on-orbit observation task is divided into target approach-observation preparation-continuous observation three stages, establishing the multi-stage task model and constraint set to enhance task solvability. Secondly, the multi-stage policy learning algorithm is proposed, constructing the multi-stage training environment and task reward function, integrating predictive guidance and rule-coupled maneuver guidance mechanisms to enhance algorithm exploration capability and convergence stability. Finally, simulations demonstrate that compared to classical reinforcement learning algorithms, this algorithm reduces convergence time by 30.9, increases average task cumulative reward by 9.28, and decreases average pulse consumption by 13.91. Moreover, compared to the traditional optimization method, it effectively enhances core task indicators, validating its effectiveness.
AB - A spacecraft multi-stage maneuver decision-making method based on deep reinforcement learning is proposed to address the maneuver decision problem for spacecraft approaching space targets during on-orbit observation service. Firstly, the on-orbit observation task is divided into target approach-observation preparation-continuous observation three stages, establishing the multi-stage task model and constraint set to enhance task solvability. Secondly, the multi-stage policy learning algorithm is proposed, constructing the multi-stage training environment and task reward function, integrating predictive guidance and rule-coupled maneuver guidance mechanisms to enhance algorithm exploration capability and convergence stability. Finally, simulations demonstrate that compared to classical reinforcement learning algorithms, this algorithm reduces convergence time by 30.9, increases average task cumulative reward by 9.28, and decreases average pulse consumption by 13.91. Moreover, compared to the traditional optimization method, it effectively enhances core task indicators, validating its effectiveness.
KW - deep reinforcement learning
KW - intelligent decision-making
KW - on-orbit observation
KW - spacecraft maneuvering
UR - https://www.scopus.com/pages/publications/105039992000
U2 - 10.12305/j.issn.1001-506X.2026.05.15
DO - 10.12305/j.issn.1001-506X.2026.05.15
M3 - 文章
AN - SCOPUS:105039992000
SN - 1001-506X
VL - 48
SP - 1590
EP - 1598
JO - Xi Tong Gong Cheng Yu Dian Zi Ji Shu/Systems Engineering and Electronics
JF - Xi Tong Gong Cheng Yu Dian Zi Ji Shu/Systems Engineering and Electronics
IS - 5
ER -