TY - JOUR
T1 - Behavior Cloning-TD3 Based End-to-End Autonomous Flight of Quadrotor UAVs
AU - Wang, Yuanshun
AU - Li, Bo
AU - Xiao, Bing
AU - Gao, Yufeng
AU - Shen, Qiang
N1 - Publisher Copyright:
© 1965-2011 IEEE.
PY - 2026
Y1 - 2026
N2 - This study investigates an end-to-end autonomous flight approach for quadrotor unmanned aerial vehicles (UAVs) under local environmental perception. To address the limitations of traditional hierarchical architectures such as the decoupling of planning and control, heavy reliance on global perception a end to-end flight control method, BC-TD3, is proposed by integrating Behavior Cloning (BC) with Twin Delayed Deep Deterministic Policy Gradient (TD3). Firstly, onboard LiDAR data are used to perceive the environment and construct a collision risk model, which is combined with positional data to form the state representation. Then, two independent experience storage mechanisms are designed: a Prioritized Experience Replay (PER) buffer and a successful trajectory experience pool. After each update of the policy network, the behavior cloning model, trained on successful trajectories, is employed to further refine the policy. In addition, a composite reward function with four components dynamic distance reward, heading alignment reward, collision penalty, and time efficiency penalty is designed to improve the stability and safety of the learned autonomous flight policy. Finally, simulation experiments are conducted to validate the effectiveness and superiority of the proposed approach.
AB - This study investigates an end-to-end autonomous flight approach for quadrotor unmanned aerial vehicles (UAVs) under local environmental perception. To address the limitations of traditional hierarchical architectures such as the decoupling of planning and control, heavy reliance on global perception a end to-end flight control method, BC-TD3, is proposed by integrating Behavior Cloning (BC) with Twin Delayed Deep Deterministic Policy Gradient (TD3). Firstly, onboard LiDAR data are used to perceive the environment and construct a collision risk model, which is combined with positional data to form the state representation. Then, two independent experience storage mechanisms are designed: a Prioritized Experience Replay (PER) buffer and a successful trajectory experience pool. After each update of the policy network, the behavior cloning model, trained on successful trajectories, is employed to further refine the policy. In addition, a composite reward function with four components dynamic distance reward, heading alignment reward, collision penalty, and time efficiency penalty is designed to improve the stability and safety of the learned autonomous flight policy. Finally, simulation experiments are conducted to validate the effectiveness and superiority of the proposed approach.
KW - Autonomous flight
KW - Behavior cloning and TD3
KW - End-to-end
KW - Local Perception
KW - Quadrotor UAV
UR - https://www.scopus.com/pages/publications/105045160194
U2 - 10.1109/TAES.2026.3712715
DO - 10.1109/TAES.2026.3712715
M3 - 文章
AN - SCOPUS:105045160194
SN - 0018-9251
JO - IEEE Transactions on Aerospace and Electronic Systems
JF - IEEE Transactions on Aerospace and Electronic Systems
ER -