Skip to main navigation Skip to search Skip to main content

Pointwise Frequency Attention and Coarse-to-Fine Learning With Image-to-Instance Level Supervision for Fine-Grained Object Detection in Remote Sensing Images

  • Zhengzhou University of Light Industry

Research output: Contribution to journalArticlepeer-review

Abstract

Fine-grained object detection (FGOD) models aim to distinguish different subcategories within the same coarse-grained category. Although two-stage object detection frameworks are extensively employed in current FGOD models due to their high detection accuracy, they still face two major problems. First, insufficient discriminability of the backbone features results in frequent misclassification among similar subcategory objects. Second, the limited capability of the backbone features to represent object direction leads to localization deviation. To address the first problem, a coarse-to-fine image-level supervised feature discriminability enhancement (C2FIm-FDE) module is proposed. Guided by the generated coarse- and fine-grained image-level category labels, the C2FIm-FDE module progressively strengthens the discriminability of backbone features at the image level. To tackle the second problem, a pointwise frequency attention based on the discrete wavelet transform (PFA-DWT) is proposed. By adaptively combining multiband frequency features and leveraging the dual localization capability of the DWT in the spatial and frequency domains, the PFA-DWT module achieves accurate fusion of frequency-domain features encoding local object orientation information with spatial features via a pointwise attention mechanism. To further enhance the discriminability of similar subcategories, a coarse-to-fine instance-level supervised detection (C2FIn) head is proposed, which gradually focuses on differences among subcategories within the same coarse-grained category by sequentially introducing coarse- and fine-grained instance-level labels. Ablation studies validate the effectiveness of the C2FIm-FDE, PFA-DWT modules, and C2FIn head, as well as their arbitrary combinations. Quantitative comparisons with popular FGOD models on the two large-scale public datasets, FAIR1M-1.0 and FAIR1M-2.0, demonstrate the excellent detection accuracy and generalization capability of our FGOD model. Source code is available at https://github.com/PFA-DWT/PFA-C2FI.

Original languageEnglish
Article number5619812
JournalIEEE Transactions on Geoscience and Remote Sensing
Volume64
DOIs
StatePublished - 2026

Keywords

  • C2FIn head
  • Coarse-to-fine image-level supervised feature discriminability enhancement (C2FIm-FDE) module
  • fine-grained object detection (FGOD)
  • pointwise frequency attention based on discrete wavelet transform (PFA-DWT) module
  • remote sensing images (RSIs)

Fingerprint

Dive into the research topics of 'Pointwise Frequency Attention and Coarse-to-Fine Learning With Image-to-Instance Level Supervision for Fine-Grained Object Detection in Remote Sensing Images'. Together they form a unique fingerprint.

Cite this