TY - GEN
T1 - An Improved RT-DETR Detection Algorithm for Small Targets in UAV Aerial Images
AU - Zhang, Danni
AU - Zhang, Hao
AU - Zou, Weijun
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - The small-size target detection in UAV aerial images faces challenges including small size, complex backgrounds, as well as limited computational resources. In response to them, this paper provides an improved RT-DETR algorithm. A DualConv-Block structure is introduced into the backbone network by incorporating Dual Conv, which integrates the advantages of group convolutions and heterogeneous convolutions. This reduces computational cost while preserving the original information and promoting information sharing. Additionally, a redesigned feature fusion structure is proposed, consisting of a Scale Sequence Feature Fusion (SSFF) module and a Triple Feature Encoder (TFE) module, which enhance multi-scale information extraction and feature integration capabilities, thereby improving small target detection accuracy. Furthermore, we propose that Inner-Focaler-IoU loss its function, combining the ideas of Inner-IoU and Focaler-IoU with adaptively concentrate on samples in various levels of difficulty to seek to improve boundary regression accuracy. Experiments on the VisDrone-2019 dataset demonstrate that the improved model achieves mAP0.5 scores of 49.8% and 39.5% on the validation and test sets, respectively, showing an improvement of 2.4% and 1.3% compared to the baseline model. Moreover, the parameter count is reduced by 11.0%. This algorithm effectively balances model size and detection accuracy, providing an efficient and lightweight solution for small target detection in UAV aerial images.
AB - The small-size target detection in UAV aerial images faces challenges including small size, complex backgrounds, as well as limited computational resources. In response to them, this paper provides an improved RT-DETR algorithm. A DualConv-Block structure is introduced into the backbone network by incorporating Dual Conv, which integrates the advantages of group convolutions and heterogeneous convolutions. This reduces computational cost while preserving the original information and promoting information sharing. Additionally, a redesigned feature fusion structure is proposed, consisting of a Scale Sequence Feature Fusion (SSFF) module and a Triple Feature Encoder (TFE) module, which enhance multi-scale information extraction and feature integration capabilities, thereby improving small target detection accuracy. Furthermore, we propose that Inner-Focaler-IoU loss its function, combining the ideas of Inner-IoU and Focaler-IoU with adaptively concentrate on samples in various levels of difficulty to seek to improve boundary regression accuracy. Experiments on the VisDrone-2019 dataset demonstrate that the improved model achieves mAP0.5 scores of 49.8% and 39.5% on the validation and test sets, respectively, showing an improvement of 2.4% and 1.3% compared to the baseline model. Moreover, the parameter count is reduced by 11.0%. This algorithm effectively balances model size and detection accuracy, providing an efficient and lightweight solution for small target detection in UAV aerial images.
KW - Dual Conv
KW - RT-DETR
KW - Scale Sequence Fusion
KW - UAV images
KW - loss function
KW - small target detection
UR - https://www.scopus.com/pages/publications/105040926070
U2 - 10.1109/CAC67268.2025.11487056
DO - 10.1109/CAC67268.2025.11487056
M3 - 会议稿件
AN - SCOPUS:105040926070
T3 - Proceedings - 2025 China Automation Congress, CAC 2025
SP - 1153
EP - 1158
BT - Proceedings - 2025 China Automation Congress, CAC 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2025 China Automation Congress, CAC 2025
Y2 - 26 September 2025 through 28 September 2025
ER -