TY - JOUR
T1 - Pointwise Frequency Attention and Coarse-to-Fine Learning With Image-to-Instance Level Supervision for Fine-Grained Object Detection in Remote Sensing Images
AU - Qian, Xiaoliang
AU - Li, Xiaobin
AU - Wang, Wei
AU - Yao, Xiwen
AU - Cheng, Gong
N1 - Publisher Copyright:
© 1980-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Fine-grained object detection (FGOD) models aim to distinguish different subcategories within the same coarse-grained category. Although two-stage object detection frameworks are extensively employed in current FGOD models due to their high detection accuracy, they still face two major problems. First, insufficient discriminability of the backbone features results in frequent misclassification among similar subcategory objects. Second, the limited capability of the backbone features to represent object direction leads to localization deviation. To address the first problem, a coarse-to-fine image-level supervised feature discriminability enhancement (C2FIm-FDE) module is proposed. Guided by the generated coarse- and fine-grained image-level category labels, the C2FIm-FDE module progressively strengthens the discriminability of backbone features at the image level. To tackle the second problem, a pointwise frequency attention based on the discrete wavelet transform (PFA-DWT) is proposed. By adaptively combining multiband frequency features and leveraging the dual localization capability of the DWT in the spatial and frequency domains, the PFA-DWT module achieves accurate fusion of frequency-domain features encoding local object orientation information with spatial features via a pointwise attention mechanism. To further enhance the discriminability of similar subcategories, a coarse-to-fine instance-level supervised detection (C2FIn) head is proposed, which gradually focuses on differences among subcategories within the same coarse-grained category by sequentially introducing coarse- and fine-grained instance-level labels. Ablation studies validate the effectiveness of the C2FIm-FDE, PFA-DWT modules, and C2FIn head, as well as their arbitrary combinations. Quantitative comparisons with popular FGOD models on the two large-scale public datasets, FAIR1M-1.0 and FAIR1M-2.0, demonstrate the excellent detection accuracy and generalization capability of our FGOD model. Source code is available at https://github.com/PFA-DWT/PFA-C2FI.
AB - Fine-grained object detection (FGOD) models aim to distinguish different subcategories within the same coarse-grained category. Although two-stage object detection frameworks are extensively employed in current FGOD models due to their high detection accuracy, they still face two major problems. First, insufficient discriminability of the backbone features results in frequent misclassification among similar subcategory objects. Second, the limited capability of the backbone features to represent object direction leads to localization deviation. To address the first problem, a coarse-to-fine image-level supervised feature discriminability enhancement (C2FIm-FDE) module is proposed. Guided by the generated coarse- and fine-grained image-level category labels, the C2FIm-FDE module progressively strengthens the discriminability of backbone features at the image level. To tackle the second problem, a pointwise frequency attention based on the discrete wavelet transform (PFA-DWT) is proposed. By adaptively combining multiband frequency features and leveraging the dual localization capability of the DWT in the spatial and frequency domains, the PFA-DWT module achieves accurate fusion of frequency-domain features encoding local object orientation information with spatial features via a pointwise attention mechanism. To further enhance the discriminability of similar subcategories, a coarse-to-fine instance-level supervised detection (C2FIn) head is proposed, which gradually focuses on differences among subcategories within the same coarse-grained category by sequentially introducing coarse- and fine-grained instance-level labels. Ablation studies validate the effectiveness of the C2FIm-FDE, PFA-DWT modules, and C2FIn head, as well as their arbitrary combinations. Quantitative comparisons with popular FGOD models on the two large-scale public datasets, FAIR1M-1.0 and FAIR1M-2.0, demonstrate the excellent detection accuracy and generalization capability of our FGOD model. Source code is available at https://github.com/PFA-DWT/PFA-C2FI.
KW - C2FIn head
KW - Coarse-to-fine image-level supervised feature discriminability enhancement (C2FIm-FDE) module
KW - fine-grained object detection (FGOD)
KW - pointwise frequency attention based on discrete wavelet transform (PFA-DWT) module
KW - remote sensing images (RSIs)
UR - https://www.scopus.com/pages/publications/105037900139
U2 - 10.1109/TGRS.2026.3688254
DO - 10.1109/TGRS.2026.3688254
M3 - 文章
AN - SCOPUS:105037900139
SN - 0196-2892
VL - 64
JO - IEEE Transactions on Geoscience and Remote Sensing
JF - IEEE Transactions on Geoscience and Remote Sensing
M1 - 5619812
ER -