Abstract
This paper investigates an improved Proximity Policy Optimization(PPO)-based intelligent adaptive control algorithm for spacecraft swarm to achieve rapid rounding-up and high-precision tracking of targets in elliptical orbits. Firstly, an improved reinforcement learning algorithm is proposed. And a stability-constrained hybrid learning-control framework is developed, which can learn controller parameters by the improved algorithm. By analytically constraining the output space of the actor network through theoretical derivation, it can be guaranteed that all learned policies satisfy Lyapunov stability conditions a priori. Secondly, considering various constraints, a spacecraft swarm round-up algorithm is designed based on the learning-control framework. A hierarchical reward function is designed to integrate parameters with significantly different operational characteristics into a unified training framework. The composite reward function incorporates multiple constraints to guide the agent in learning optimal rounding-up strategies while achieving high-precision tracking control. Finally, numerical simulations with environment disturbances and sensor measurement noises demonstrate the effectiveness and robustness of the proposed controller.
| Original language | English |
|---|---|
| Pages (from-to) | 543-555 |
| Number of pages | 13 |
| Journal | Acta Astronautica |
| Volume | 249 |
| DOIs | |
| State | Published - Dec 2026 |
Keywords
- Adaptive control
- Pursuit-evasion game
- Reinforcement learning
- Rounding-up
- Spacecraft swarm
Fingerprint
Dive into the research topics of 'Improved learning-based intelligent adaptive control for spacecraft swarm rounding-up in elliptical orbits'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver