Abstract
The rapid concealment of time-sensitive targets in UAV-based aerial imagery holds significant importance for public security and asset protection. However, it remains a key challenge in current research under the complex real-world aerial scenes and modality variations. To address this issue, a novel multimodal time-sensitive target fast concealment Network for UAV aerial imagery via real-time mask generation (MTCNet) is proposed. MTCNet employs a two-stage framework consisting of real-time target mask generation and image inpainting to achieve efficient target concealment. In the mask generation stage, we construct SPD-YOLODA, a novel real-time mask generation network based on a convolutional neural network architecture for domain adaptation. The network integrates the Space-to-Depth Convolution (SPD-Conv) module and the Multi-Scale Dilated Attention (MSDA) module to improve the segmentation accuracy of small multimodal targets and enhance edge detail preservation in complex scenes. In the object concealment stage, we design a cascade inpainting network combining convolutional and transformer networks based on mask updating (CTM). By employing a convolutional downsampling-Transformer-upsampling coarse generation network enhanced by the Multi-Head Contextual Attention (MCA) module and an adaptive mask update mechanism, along with a Conv-U-Net refinement network, CTM fills in the missing background of target areas effectively and conceals the targets across various resolutions rapidly. A specialized multimodal time-sensitive target dataset, named MTT-10K, is constructed from the perspective of UAV aerial imagery. Extensive experimental results on the MTT-10K demonstrate that MTCNet can efficiently generate effective masks in real-time and achieve rapid target concealment through the inpainting network. Furthermore, generalization experiments highlight the strong adaptability and robustness of MTCNet compared with existing methods.
| Original language | English |
|---|---|
| Article number | 134088 |
| Journal | Expert Systems with Applications |
| Volume | 333 |
| DOIs | |
| State | Published - 15 Jan 2027 |
Keywords
- Deep learning
- Image inpainting
- Mask generation
- Multimodal
- YOLO
Fingerprint
Dive into the research topics of 'MTCNet: Multimodal time-sensitive target fast concealment network for UAV aerial imagery via real-time mask generation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver