跳到主要导航 跳到搜索 跳到主要内容

基 于 频 域 和 跨 尺 度 特 征 融 合 的 无 人 机 灰 度 小 目 标检 测 算 法

  • Pengyao Zhou
  • , Meibo Lü
  • , Xin Ning
  • , Zhe Fu
  • Northwestern Polytechnical University Xian

科研成果: 期刊稿件文章同行评审

摘要

Objective Due to the blurred fine details of targets, low contrast between targets and background, and variations in viewing angles, detecting small objects in UAV grayscale images remains highly challenging. These factors lead to severe loss of detailed information in shallow features and insufficient texture representation in deep features, thereby increasing the likelihood of false positives. To address these issues, we propose the ICA-YOLO detector. ICA-YOLO enhances the extraction of target details in shallow layers by generating precise positional encodings, while in deep layers, it adopts an adaptive cross-layer feature fusion strategy and establishes long-range dependencies between local channel features and global spatial features, thereby improving the representation of target texture information. Methods ICA-YOLO is developed based on the YOLOv8 framework and incorporates three core modules: frequency slice-assisted module (FSAM), cross-scale feature fusion network (CFFNet), and feature enhancement module (FEM). In the backbone network, ICA-YOLO introduces the newly designed FSAM. FSAM enlarges small targets at the image level through a slicing operation, thereby introducing more global detail information in the frequency domain. By computing the information entropy of semantic features within each sliced region, FSAM guides the encoder to generate accurate positional encodings for small-object regions, enriching their spatial positional information. In the Neck part, ICA-YOLO replaces the PAN structure in YOLOv8 with the newly designed CFFNet. CFFNet employs an adaptive weighted fusion strategy and cross-layer connection mechanism to retain more texture and positional information. By fusing deep semantic features with shallow detailed features, CFFNet effectively enhances the feature representation capability for multi-scale objects. To overcome the limitations of a single receptive field in semantic feature expression, the FEM establishes long-range dependencies between local features of small targets and global contextual features. In the Head part, ICA-YOLO adds a dedicated small-object detection head to better utilize shallow detailed features, thereby improving small-object detection performance. For training, the stochastic gradient descent (SGD) optimizer is adopted with a batch size of 96 for 300 epochs. The input image size is set to 640×640, and the initial learning rate is 0.001. Evaluation metrics include mAP@0.5∶0.95, mAP@ 0.5, the number of parameters (Params), and detection frame rate. Results and Discussions To comprehensively verify the effectiveness of the proposed algorithm, a series of experiments are conducted on both the self-constructed HG-UAV dataset and the VisDrone2019 dataset. As shown in visualization results, the ablation studies on the HG-UAV dataset demonstrate that the proposed FSAM, CFFNet, and FEM each enhance the overall performance of ICA-YOLO at different levels. Compared with the baseline YOLOv8 model, ICA-YOLO achieves improvements of 7.5 percentage points and 6.1 percentage points in mAP@0.5 and mAP@0.5∶0.95, respectively, which effectively validates the superior performance and robustness of ICA-YOLO in high-altitude small object detection tasks. The comparative experiments conducted to analyze the impact of different slicing strategies on the FSAM indicate that the proposed {4×4, 5×5, 2×2, 2×4, 4× 2} slicing configurations significantly enhance the representation of fine-grained object features, thereby improving detection accuracy. To further validate the superiority of CFFNet under different weighting coefficients, additional comparative experiments are performed. To evaluate the robustness and adaptability of ICA-YOLO under varying test environments, illumination conditions, and viewing angles, the VisDrone dataset is employed for comparison with several state-of-the-art detection algorithms, including mainstream models such as YOLOv5s, YOLOv8s, and YOLOv11s, as well as improved variants such as Drone-YOLO, Modified-YOLOv8, ACAM-YOLO, and SOD-YOLO. Furthermore, considering that the proposed model integrates Transformer-based design principles, ICA-YOLO is also compared against Transformer-based or hybrid architectures, including TPH-YOLOv5, EMSD-DETR, and FM-RTDETR-M. The results reveal that ICA-YOLO achieves an optimal balance between detection accuracy and computational efficiency. Despite a slight reduction in inference speed, ICA-YOLO attains the highest detection precision among all compared models. To validate the practical applicability of ICA-YOLO, deployment experiments are further conducted on Rockchip RK 3588 and NVIDIA Orin Nano embedded platforms, using 2000 images from the HG-UAV dataset for benchmarking. As shown in Tab.7, ICA-YOLO achieves the best trade-off between detection accuracy and real-time performance across different embedded platforms, confirming its effectiveness and deployability for UAV-based embedded detection applications. Conclusions To address the limitations of existing YOLO-based detectors in UAV grayscale imagery, such as blurred detail features, low contrast between targets and background, and accuracy degradation caused by viewpoint variations—this paper proposes an efficient object detector named ICA-YOLO for UAV image detection tasks. ICA-YOLO consists of three core modules: FSAM, CFFNet, and FEM. Specifically, FSAM employs a parameter-free Fourier transform to project feature maps into the frequency domain and measures the texture information entropy differences among frequency sub-blocks to guide positional encoding, thereby enhancing the extraction of fine-grained target details. CFFNet introduces a weight-adaptive cross-layer feature fusion strategy based on adaptive weights (AWs) to effectively mitigate texture and positional information loss caused by upsampling operations during feature fusion. FEM establishes long-range dependencies between local channel features and global spatial features, further strengthening the representation of target texture information.Experimental results demonstrate that ICA-YOLO outperforms baseline models and existing state-of-the-art methods on both the HG-UAV and VisDrone2019 datasets in terms of mAP@0.5 and mAP@ 0.5∶0.95. Moreover, ICA-YOLO has been successfully deployed on embedded platforms such as Rockchip RK 3588 and NVIDIA Orin Nano, achieving good real-time performance while maintaining high detection accuracy. Although ICA-YOLO exhibits excellent overall performance in UAV small-object detection and shows promising potential for embedded deployment, it still faces certain limitations when running on domestic embedded platforms with limited computing power (e. g., RK 3588). To address this issue, future work will explore lightweight strategies such as group convolution to further optimize the memory utilization and computational efficiency of ICA-YOLO. In addition, considering that complex factors such as object occlusion and viewpoint variation can constrain the mAP@0.5, subsequent research will focus on optimizing the regression loss function to further enhance the model′s engineering applicability in UAV grayscale small-object detection scenarios.

投稿的翻译标题UAV-Grayscale Small Object Detection Algorithm Based on Frequency-Domain and Cross-Scale Feature Fusion
源语言繁体中文
文章编号1228003
期刊Laser and Optoelectronics Progress
63
12
DOI
出版状态已出版 - 6月 2026

关键词

  • feature enhancement
  • feature fusion
  • Fourier transform
  • grayscale image small object detection
  • positional encoding

指纹

探究 '基 于 频 域 和 跨 尺 度 特 征 融 合 的 无 人 机 灰 度 小 目 标检 测 算 法' 的科研主题。它们共同构成独一无二的指纹。

引用此