Abstract
Adversarial attacks expose the vulnerabilities of deep learning models, and effective adversarial attack methods aid in uncovering potential weaknesses. Existing gradient-based adversarial attacks overfit the characteristics of attacked white-box models, resulting in poor transferability for black-box models. This paper investigates transferable adversarial attacks for black-box models, proposing a multi-perspective confidence fusion-based method to enhance the transferability of adversarial examples. This approach is integrated as a universal component into gradient-based adversarial attack processes to improve transferability. Specifically, a multi-perspective transformation strategy based on dual-pixel space is designed to guide the model in perceiving image information across different channels and spatial scales, thereby expanding the model’s areas of focus, generating a differentiated attention distribution for the image, and enabling multi-view perception of image information. To model conflicts and uncertainties across different perspectives, a conflict-aware confidence fusion method based on the evidence theory framework is proposed. This method extracts common predictive information from the confidence outputs of the model across multiple perspectives, thereby avoiding decision-making interference caused by perspective-specific biases and effectively enhancing the reliability of multi-perspective decision fusion. A bidirectional loss optimization function is designed to optimize the deviation of adversarial examples from the correct model decision boundary, guiding them to lie in the shared vulnerable regions across different views and models, thereby improving the transfer attack performance against black-box models. Experiments show that the proposed method can effectively improve the transferability of adversarial examples in cross-model architecture attack scenarios. After integrating the existing gradient-based adversarial attacks with the multi-perspective confidence fusion method, the transfer attack success rate is improved by an average of 21.15% and 13.02% for conventionally trained convolutional neural networks (CNNs) and Transformer models, respectively, by 13.84% for defense models, and by 16.14% for ensemble models.
| Translated title of the contribution | 基于多视角置信度融合的对抗样本迁移性提升方法 |
|---|---|
| Original language | English |
| Pages (from-to) | 1062-1077 |
| Number of pages | 16 |
| Journal | Tien Tzu Hsueh Pao/Acta Electronica Sinica |
| Volume | 54 |
| Issue number | 3 |
| DOIs | |
| State | Published - Mar 2026 |
Keywords
- adversarial attack
- confidence score
- deep learning
- image classification
- transferability
- 图像分类
- 对抗攻击
- 深度学习
- 置信度
- 迁移性
Fingerprint
Dive into the research topics of 'A Method for Enhancing the Transferability of Adversarial Examples Based on Multi-Perspective Confidence Fusion'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver