跳到主要导航 跳到搜索 跳到主要内容

Pointwise Frequency Attention and Coarse-to-Fine Learning With Image-to-Instance Level Supervision for Fine-Grained Object Detection in Remote Sensing Images

  • Zhengzhou University of Light Industry

科研成果: 期刊稿件文章同行评审

摘要

Fine-grained object detection (FGOD) models aim to distinguish different subcategories within the same coarse-grained category. Although two-stage object detection frameworks are extensively employed in current FGOD models due to their high detection accuracy, they still face two major problems. First, insufficient discriminability of the backbone features results in frequent misclassification among similar subcategory objects. Second, the limited capability of the backbone features to represent object direction leads to localization deviation. To address the first problem, a coarse-to-fine image-level supervised feature discriminability enhancement (C2FIm-FDE) module is proposed. Guided by the generated coarse- and fine-grained image-level category labels, the C2FIm-FDE module progressively strengthens the discriminability of backbone features at the image level. To tackle the second problem, a pointwise frequency attention based on the discrete wavelet transform (PFA-DWT) is proposed. By adaptively combining multiband frequency features and leveraging the dual localization capability of the DWT in the spatial and frequency domains, the PFA-DWT module achieves accurate fusion of frequency-domain features encoding local object orientation information with spatial features via a pointwise attention mechanism. To further enhance the discriminability of similar subcategories, a coarse-to-fine instance-level supervised detection (C2FIn) head is proposed, which gradually focuses on differences among subcategories within the same coarse-grained category by sequentially introducing coarse- and fine-grained instance-level labels. Ablation studies validate the effectiveness of the C2FIm-FDE, PFA-DWT modules, and C2FIn head, as well as their arbitrary combinations. Quantitative comparisons with popular FGOD models on the two large-scale public datasets, FAIR1M-1.0 and FAIR1M-2.0, demonstrate the excellent detection accuracy and generalization capability of our FGOD model. Source code is available at https://github.com/PFA-DWT/PFA-C2FI.

源语言英语
文章编号5619812
期刊IEEE Transactions on Geoscience and Remote Sensing
64
DOI
出版状态已出版 - 2026

指纹

探究 'Pointwise Frequency Attention and Coarse-to-Fine Learning With Image-to-Instance Level Supervision for Fine-Grained Object Detection in Remote Sensing Images' 的科研主题。它们共同构成独一无二的指纹。

引用此