Abstract
This study presents a reinforcement learning framework for solving constrained spacecraft targeta"attackera"defender (TAD) problem in 3-D orbital environment. We formulate the TAD problem using state transition of relative orbital dynamics under impulsive maneuver and multiple practical constraints are considered. To solve the complex TAD game problem, a Hamilton-regularized multiagent twin-delayed deep deterministic policy gradient algorithm is proposed, which achieves an efficiently training by guiding the spacecrafts exploring the optimal strategy effectively. An adaptive curriculum learning methodology is designed, which is driven automatically such that the convergence of strategies is improved progressively. The reward decomposition combined with nonlinear process-oriented and terminal condition rewards can solve the sparse reward issue and therefore make the strategy avoid converging to local optima. A sun-elevation angle adaptation method is introduced to amend the strategy that makes the constraint be satisfied in TAD game. Numerical simulations and comparison results demonstrate the effectiveness of the proposed method.
| Original language | English |
|---|---|
| Pages (from-to) | 11404-11417 |
| Number of pages | 14 |
| Journal | IEEE Transactions on Aerospace and Electronic Systems |
| Volume | 62 |
| DOIs | |
| State | Published - 2026 |
Keywords
- Impulsive maneuver
- attackera
- defender (TAD) game
- reinforcement learning (RL)
- spacecraft
- targeta
Fingerprint
Dive into the research topics of 'Impulsive Maneuver Strategy Design for Target-Attacker-Defender Game of Spacecrafts via a Curriculum Learning-Based Hamilton-Regularized MATD3 Algorithm'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver