跳到主要导航 跳到搜索 跳到主要内容

A unified detection model for multimodal aerospace remote sensing images based on mixture of experts

投稿的翻译标题: 一 种 基 于 混 合 专 家 组 的 多 模 态 航 天 遥 感 图 像统 一 检 测 模 型
  • Yuanjie ZHI
  • , Xin GE
  • , Fan ZHANG
  • , Zhi YANG
  • , Mingyang MA
  • , Shaohui MEI
  • Northwestern Polytechnical University Xian
  • China Aerospace Science and Technology Corporation
  • State Grid Electric Power Engineering Research Institute Company Ltd.

科研成果: 期刊稿件文章同行评审

摘要

With the increasing number of remote sensing satellites deployed in orbit in China,the quantity of aerospace remote sensing images,represented by Synthetic Aperture Radar(SAR)and optical(RGB)images,is rapidly growing,along with the demand for tasks such as object detection from these massive datasets. However,due to objective factors such as differences in imaging mechanisms and resolutions,images from different satellites exhibit significant modality feature differences. These differences are particularly pronounced between SAR and RGB remote sensing images,making it difficult for a single model to learn feature information across different types of remote sensing images. As a result,each satellite typically requires a dedicated model for detection tasks,which has become a major obstacle to collaborative recognition and relay detection applications in satellite remote sensing. To address this issue,this paper innovatively proposes a self-distillation multimodal detection model based on a Mixture of Experts (MoE). First,a modality-aware MoE structure is constructed,employing a small number of high-quality experts as teachers to guide other experts,while simultaneously incorporating modality-invariant constraints to further reduce cross-modality feature shifts. Second, a Fourier-enhanced diffusion detection head is developed, combining frequency-domain feature enhancement to improve the capability of capturing detailed information of detection targets. To evaluate the model performance,aerospace images were selected and cropped from the public datasets FAIR1M and SARDet_100K,resulting in a dataset of 68 983 aerospace remote sensing images for object detection under different backgrounds and imaging mechanisms. Experimental results demonstrate that,compared with existing single-modality detection methods,the proposed model performs better in detection tasks across both modalities,with a significant improvement in mean Average Precision(mAP). This fully demonstrates that the proposed model possesses significant application value in multimodal aerospace remote sensing image object detection,and exhibits good adaptability to various types of satellite remote sensing images.

投稿的翻译标题一 种 基 于 混 合 专 家 组 的 多 模 态 航 天 遥 感 图 像统 一 检 测 模 型
源语言英语
文章编号532864
期刊Hangkong Xuebao/Acta Aeronautica et Astronautica Sinica
47
10
DOI
出版状态已出版 - 2026

学术指纹

探究 '一 种 基 于 混 合 专 家 组 的 多 模 态 航 天 遥 感 图 像统 一 检 测 模 型' 的科研主题。它们共同构成独一无二的学术指纹。

引用此