Skip to main navigation Skip to search Skip to main content

MTCNet: Multimodal time-sensitive target fast concealment network for UAV aerial imagery via real-time mask generation

  • Ruitao Lu
  • , Zhanhong Zhuo
  • , Siyu Wang
  • , Dingwen Zhang
  • , Bobo Xi
  • , Xiangtao Zheng
  • , Zhiqiang Hou
  • , Xiaogang Yang
  • Rocket Force University of Engineering
  • Xuzhou Institute of Technology
  • State Key Laboratory of Integrated Services Networks
  • Fuzhou University
  • Xi'an Institute of Posts and Telecommunications

Research output: Contribution to journalArticlepeer-review

Abstract

The rapid concealment of time-sensitive targets in UAV-based aerial imagery holds significant importance for public security and asset protection. However, it remains a key challenge in current research under the complex real-world aerial scenes and modality variations. To address this issue, a novel multimodal time-sensitive target fast concealment Network for UAV aerial imagery via real-time mask generation (MTCNet) is proposed. MTCNet employs a two-stage framework consisting of real-time target mask generation and image inpainting to achieve efficient target concealment. In the mask generation stage, we construct SPD-YOLODA, a novel real-time mask generation network based on a convolutional neural network architecture for domain adaptation. The network integrates the Space-to-Depth Convolution (SPD-Conv) module and the Multi-Scale Dilated Attention (MSDA) module to improve the segmentation accuracy of small multimodal targets and enhance edge detail preservation in complex scenes. In the object concealment stage, we design a cascade inpainting network combining convolutional and transformer networks based on mask updating (CTM). By employing a convolutional downsampling-Transformer-upsampling coarse generation network enhanced by the Multi-Head Contextual Attention (MCA) module and an adaptive mask update mechanism, along with a Conv-U-Net refinement network, CTM fills in the missing background of target areas effectively and conceals the targets across various resolutions rapidly. A specialized multimodal time-sensitive target dataset, named MTT-10K, is constructed from the perspective of UAV aerial imagery. Extensive experimental results on the MTT-10K demonstrate that MTCNet can efficiently generate effective masks in real-time and achieve rapid target concealment through the inpainting network. Furthermore, generalization experiments highlight the strong adaptability and robustness of MTCNet compared with existing methods.

Original languageEnglish
Article number134088
JournalExpert Systems with Applications
Volume333
DOIs
StatePublished - 15 Jan 2027

Keywords

  • Deep learning
  • Image inpainting
  • Mask generation
  • Multimodal
  • YOLO

Fingerprint

Dive into the research topics of 'MTCNet: Multimodal time-sensitive target fast concealment network for UAV aerial imagery via real-time mask generation'. Together they form a unique fingerprint.

Cite this