跳到主要导航 跳到搜索 跳到主要内容

Exploring Multi-Level Attention and Semantic Relationship for Remote Sensing Image Captioning

  • Zhenghang Yuan
  • , Xuelong Li
  • , Qi Wang
  • Northwestern Polytechnical University Xian

科研成果: 期刊稿件文章同行评审

52 引用 (Scopus)

摘要

Remote sensing image captioning, which aims to understand high-level semantic information and interactions of different ground objects, is a new emerging research topic in recent years. Though image captioning has developed rapidly with convolutional neural networks (CNNs) and recurrent neural networks (RNNs), the image captioning task for remote sensing images still suffers from two main limitations. One limitation is that the scales of objects in remote sensing images vary dramatically, which makes it difficult to obtain an effective image representation. Another limitation is that the visual relationship in remote sensing images is still underused, which should have great potential to improve the final performance. In order to deal with these two limitations, an effective framework for captioning the remote sensing image is proposed in this paper. The framework is based on multi-level attention and multi-label attribute graph convolution. Specifically, the proposed multi-level attention module can adaptively focus not only on specific spatial features, but also on features of specific scales. Moreover, the designed attribute graph convolution module can employ the attribute-graph to learn more effective attribute features for image captioning. Extensive experiments are conducted and the proposed method achieves superior performance on UCM-captions, Sydney-captions and RSICD dataset.

源语言英语
文章编号8943170
页(从-至)2608-2620
页数13
期刊IEEE Access
8
DOI
出版状态已出版 - 2020

学术指纹

探究 'Exploring Multi-Level Attention and Semantic Relationship for Remote Sensing Image Captioning' 的科研主题。它们共同构成独一无二的学术指纹。

引用此