- Research Article
12
- 10.1016/j.cose.2022.102749
Transferable adversarial examples can efficiently fool topic models
- May 02, 2022
- Computers & Security
- Zhen Wang + 4 more +4
Transferable adversarial examples can efficiently fool topic models
Deep learning models have emerged as strong and efficient tools that can be applied to a broad spectrum of complex learning problems and many real-world applications. However, more and more works show that deep models are vulnerable to adversarial examples. Compared to vanilla attack settings, this paper advocates a more practical setting of data-free black-box attack, for which the attackers can completely not access the structures and parameters of the target model, as well as the intermediate features and any training data associated with the model. To tackle this task, previous methods generate transferable adversarial examples from a transparent substitute model to the target model. However, we found that these works have the limitations of taking static substitute model structure for different targets, only using hard synthesized examples once, and still relying on data statistics of the target model. This may potentially harm the performance of attacking the target model. To this end, we propose a novel Dynamic Routing and Knowledge Re-Learning framework (DraKe) to effectively learn a dynamic substitute model from the target model. Specifically, given synthesized training samples, a dynamic substitute structure learning strategy is proposed to adaptively generate optimal substitute model structure via a policy network according to different target models and tasks. To facilitate the substitute training, we present a graph-based structure information learning to capture the structural knowledge learned from the target model. For the inherent limitation that online data generation can only be learned once, a dynamic knowledge re-learning strategy is proposed to adjust the weights of optimization objectives and re-learn hard samples. Extensive experiments on four public image classification datasets and one face recognition benchmark are conducted to evaluate the efficacy of our Drake. We can obtain significant improvement compared with state-of-the-art competitors. More importantly, our DraKe consistently achieves attack superiority for different target models (e.g., residual networks, and vision transformers), showing great potential for complex real-world applications.
Transferable adversarial examples can efficiently fool topic models
Transferable adversarial examples can efficiently fool topic models
How Robust is Your Automatic Diagnosis Model?
Automatic diagnosis based on clinical notes has become a popular research field recently, and many proposed deep learning models have achieved competitive performance in diseases inference. However, previous research reveals that deep learning models are susceptible to negligibly perturbed inputs named adversarial examples, which contradicts with the safety and reliability requirements of the medical domain. To analyze the vulnerability and robustness of current automatic diagnosis models, we investigate in the generation of adversarial text examples. The main challenges for generating adversarial text examples are divided into three parts. First, the word embedding space is discrete, which makes it hard to perturb as small as adversarial image examples generation. Second, previous adversarial example generation methods focus mainly on multi-class classification models, while automatic diagnosis is a multi-label classification task. Third, the semantic and medical meaning of clinical notes are vital in disease inference, and even small perturbations can change them to a large extent. In this paper, we address the three main challenges and propose Clinical-Attacker, a general framework for both white-box and black-box adversarial text examples generation against automatic diagnosis models. Experimental results on MIMIC-III dataset demonstrate that our framework can easily alter the predictions of automatic diagnosis models with the semantic and medical meaning preserved.
Read moreCrafting Transferable Adversarial Examples Against Face Recognition via Gradient Eroding
In recent years, deep neural networks (DNNs) have made significant progress on face recognition (FR). However, DNNs have been found to be vulnerable to adversarial examples, leading to fatal consequences in real-world applications. This article focuses on improving the transferability of adversarial examples against FR models. We propose gradient eroding (GE) to make the gradient of the residual blocks more diverse, by eroding the back-propagation dynamically. We also propose a novel black-box adversarial attack named corrasion attack based on GE. Extensive experiments demonstrate that our approach can effectively improve the transferability of adversarial attacks against FR models. Our approach overperforms 29.35% in fooling rate than state-of-the-art black-box attacks. Leveraging adversarial training with adversarial examples generated by us, the robustness of models can be improved by up to 43.2%. Besides, corrasion attack successfully breaks two online FR systems, achieving a highest fooling rate of 89.8%.
Read moreAssessing the Threat of Adversarial Examples on Deep Neural Networks for Remote Sensing Scene Classification: Attacks and Defenses
Deep neural networks, which can learn the representative and discriminative features from data in a hierarchical manner, have achieved state-of-the-art performance in the remote sensing scene classification task. Despite the great success that deep learning algorithms have obtained, their vulnerability toward adversarial examples deserves our special attention. In this article, we systematically analyze the threat of adversarial examples on deep neural networks for remote sensing scene classification. Both targeted and untargeted attacks are performed to generate subtle adversarial perturbations, which are imperceptible to a human observer but may easily fool the deep learning models. Simply adding these perturbations to the original high-resolution remote sensing (HRRS) images, adversarial examples can be generated, and there are only slight differences between the adversarial examples and the original ones. An intriguing discovery in our study shows that most of these adversarial examples may be misclassified into the wrong category by the state-of-the-art deep neural networks with very high confidence. This phenomenon, undoubtedly, may limit the practical deployment of these deep learning models in the safety-critical remote sensing field. To address this problem, the adversarial training strategy is further investigated in this article, which significantly increases the resistibility of deep models toward adversarial examples. Extensive experiments on three benchmark HRRS image data sets demonstrate that while most of the well-known deep neural networks are sensitive to adversarial perturbations, the adversarial training strategy helps to alleviate their vulnerability toward adversarial examples.
Read moreEfficient Generation of Targeted and Transferable Adversarial Examples for Vision-Language Models via Diffusion Models
Adversarial attacks, particularly targeted transfer-based attacks, can be used to assess the adversarial robustness of large visual-language models (VLMs), allowing for a more thorough examination of potential security flaws before deployment. However, previous transfer-based adversarial attacks incur high costs due to high iteration counts and complex method structure. Furthermore, due to the unnaturalness of adversarial semantics, the generated adversarial examples have low transferability. These issues limit the utility of existing methods for assessing robustness. To address these issues, we propose AdvDiffVLM, which uses diffusion models to generate natural, unrestricted and targeted adversarial examples via score matching. Specifically, AdvDiffVLM uses Adaptive Ensemble Gradient Estimation (AEGE) to modify the score during the diffusion model’s reverse generation process, ensuring that the produced adversarial examples have natural adversarial targeted semantics, which improves their transferability. Simultaneously, to improve the quality of adversarial examples, we use the GradCAM-guided Mask Generation (GCMG) to disperse adversarial semantics throughout the image rather than concentrating them in a single area. Finally, AdvDiffVLM embeds more target semantics into adversarial examples after multiple iterations. Experimental results show that our method generates adversarial examples 5x to 10x faster than state-of-the-art (SOTA) transfer-based adversarial attacks while maintaining higher quality adversarial examples. Furthermore, compared to previous transfer-based adversarial attacks, the adversarial examples generated by our method have better transferability. Notably, AdvDiffVLM can successfully attack a variety of commercial VLMs in a black-box environment, including GPT-4V. The code is available at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/gq-max/AdvDiffVLM</uri>
Read moreClustering Approach for Detecting Multiple Types of Adversarial Examples
With intentional feature perturbations to a deep learning model, the adversary generates an adversarial example to deceive the deep learning model. As an adversarial example has recently been considered in the most severe problem of deep learning technology, its defense methods have been actively studied. Such effective defense methods against adversarial examples are categorized into one of the three architectures: (1) model retraining architecture; (2) input transformation architecture; and (3) adversarial example detection architecture. Especially, defense methods using adversarial example detection architecture have been actively studied. This is because defense methods using adversarial example detection architecture do not make wrong decisions for the legitimate input data while others do. In this paper, we note that current defense methods using adversarial example detection architecture can classify the input data into only either a legitimate one or an adversarial one. That is, the current defense methods using adversarial example detection architecture can only detect the adversarial examples and cannot classify the input data into multiple classes of data, i.e., legitimate input data and various types of adversarial examples. To classify the input data into multiple classes of data while increasing the accuracy of the clustering model, we propose an advanced defense method using adversarial example detection architecture, which extracts the key features from the input data and feeds the extracted features into a clustering model. From the experimental results under various application datasets, we show that the proposed method can detect the adversarial examples while classifying the types of adversarial examples. We also show that the accuracy of the proposed method outperforms the accuracy of recent defense methods using adversarial example detection architecture.
Read moreOn the Salience of Adversarial Examples
Adversarial examples are beginning to evolve as rapidly as the deep learning models they are designed to attack. These intentionally-manipulated inputs attempt to mislead the targeted model while maintaining the appearance of innocuous input data. Countermeasures against these attacks that take a global approach tend to be lossy to the original data, or ineffective in removing the perturbations. Localized approaches have proven effective, however it is difficult to identify affected areas in the data in order to apply a targeted cleaning algorithm. For image data, visual saliency estimation models identify important features in an image, and provide a targeting mechanism for countering adversarial examples. In this work, we examine the effectiveness of state-of-the-art saliency models on complex scenes, in their original and perturbed forms. In a thorough range of standard metrics, we compare performance on clean image data with adversarial examples to demonstrate the vulnerability of deep learning-based saliency models to adversarial examples.
Read moreFE-DaST: Fast and effective data-free substitute training for black-box adversarial attacks
FE-DaST: Fast and effective data-free substitute training for black-box adversarial attacks
Robust Token Gradient and Frequency-Aware Transferable Adversarial Attacks on Vision Transformers
Vision Transformers (ViTs) have achieved remarkable performance in computer vision tasks but are vulnerable to adversarial attacks. Recent studies have demonstrated the feasibility of crafting transferable adversarial examples based on ViT models. However, the adversarial examples generated by ViTs exhibit poor generalization, primarily due to structural differences between models and the tendency to overfit, which significantly hinders cross-architecture transferability. In this paper, we propose a novel framework to improve the generalization and transferability of adversarial attacks across diverse models, focusing on two key strategies: Token Gradient Divergence (TGD) and Multi-level Frequency-aware Attack (MFA). TGD, as a gradient regularization method, addresses the structural gradient issue of surrogate models, which is one of the causes of overfitting. By increasing the gradient divergence between tokens and eliminating the influence of the class token gradient, TGD enhances the transferability of adversarial examples across models. Meanwhile, MFA employs an implicit ensemble approach to enhance attack generalization. Through multiple spectral augmentations, it increases input diversity and simulates ensemble learning. By targeting critical frequency regions across models, MFA enhances adversarial example adaptability to different architectures, significantly boosting cross-architecture transferability. Extensive experiments on both ViTs and CNNs demonstrate that TGD-MFA significantly outperforms state-of-the-art transfer-based attacks, achieving substantial improvements in adversarial transferability and robustness.
Read moreDRHA-SR: Dual-Region Hierarchical Attack for Stealthy Black-Box Adversarial Examples in Remote Sensing
With the rapid development of Deep Neural Networks (DNNs), remote sensing image analysis has achieved significant progress in scene classification, object detection, and semantic segmentation. However, the increasing deployment of DNNs in real-world remote sensing applications exposes them to adversarial risks. A critical challenge in this domain is the poor visual stealth of black-box adversarial examples, perturbations are often easily perceived by humans, limiting their practicality. This issue stems from the distribution mismatch between visually salient regions and model-sensitive regions in remote sensing images. The current approaches predominantly focus on improving attack success rates while neglecting such heterogeneity, resulting in redundant and conspicuous perturbations. To address this, we propose a Dual-Region Hierarchical Attack (DRHA) that improves stealth and attack performance by using the different roles of visually salient and model-sensitive regions. Shallow perturbations are applied to salient areas to preserve perceptual similarity, while directional perturbations guided by integrated gradients target high-contribution regions to enhance model deception. Our method leverages superpixel segmentation and attribution analysis to localize perturbation regions precisely. Experiments on the UCM and AID datasets show that DRHA improves average success rates by 10%–32% over baseline methods, while maintaining superior stealth and perturbation sparsity. These results demonstrate the effectiveness of region-aware attack design in constructing imperceptible and transferable adversarial examples for remote sensing tasks.
Read moreOn the Adversarial Transferability of Generalized "Skip Connections".
Skip connection is an essential ingredient for modern deep models to be deeper and more powerful. Despite their huge success in normal scenarios (state-of-the-art classification performance on natural examples), we investigate and identify an interesting property of skip connections under adversarial scenarios, namely, the use of skip connections allows easier generation of highly transferable adversarial examples. Specifically, in ResNet-like models (with skip connections), we find that biasing backpropagation to favor gradients from skip connections-while suppressing those from residual modules via a decay factor-allows one to craft adversarial examples with high transferability. Based on this insight, we propose the Skip Gradient Method (SGM). Although starting from ResNet-like models in vision domains, we further extend SGM to more advanced architectures, including Vision Transformers (ViTs), models with varying-length paths, and other domains such as natural language processing. We conduct comprehensive transfer-based attacks against diverse model families, including ResNets, Transformers, Inceptions, Neural Architecture Search-based models, and Large Language Models (LLMs). The results demonstrate that employing SGM can greatly improve the transferability of crafted attacks in almost all cases. Furthermore, we demonstrate that SGM can still be effective under more challenging settings such as ensemble-based attacks, targeted attacks, and against defense equipped models. At last, we provide theoretical explanations and empirical insights on how SGM works. Our findings not only motivate new adversarial research into the architectural characteristics of models but also open up further challenges for secure model architecture design.
Read moreFrom Image to Code
Recent years, Machine Learning has been widely used in malware analysis and achieved unprecedented success. However, deep learning models are found to be highly vulnerable to adversarial examples, which leads to the machine learning-based malware analysis methods vulnerable to malware makers. Exploring the attack algorithm can not only promote the generation of more effective malware analysis methods, but also can promote the development of the defense algorithm. Different machine learning models use different malware features as their classification basis, and accordingly there will be different attack methods against them. For malware visualization method, corresponding effective adversarial attack has not yet appeared. Most existing malware adversarial examples for malware visualization are generated at the feature level, and do not consider whether the generated adversarial examples can be executed and complete their original functions. In this paper, we explored how to modify an Android executable file without affecting its original functions and made it become an adversarial example. We proposed an executable adversarial examples attack strategy for machine learning-based malware visualization analysis. Experimental result shows that the executable adversarial examples we generated can be normally run on Android devices without affecting its original functions, and can confuse the malware family classifier with 93% success rate. We explored possible defense methods and hope to contribute to building a more robust malware classification method.
Read moreParallel Rectangle Flip Attack: A Query-based Black-box Attack against Object Detection
Object detection has been widely used in many safety- critical tasks, such as autonomous driving. However, its vulnerability to adversarial examples has not been sufficiently studied, especially under the practical scenario of black-box attacks, where the attacker can only access the query feedback of predicted bounding-boxes and top- 1 scores returned by the attacked model. Compared with black-box attack to image classification, there are two main challenges in black-box attack to detection. Firstly, even if one bounding-box is successfully attacked, another sub- optimal bounding-box may be detected near the attacked bounding-box. Secondly, there are multiple bounding- boxes, leading to very high attack cost. To address these challenges, we propose a Parallel Rectangle Flip Attack (PRFA) via random search. We explain the difference between our method with other attacks in Fig. 1. Specifically, we generate perturbations in each rectangle patch to avoid sub-optimal detection near the attacked region. Besides, utilizing the observation that adversarial perturbations mainly locate around objects’ contours and critical points under white-box attacks, the search space of attacked rectangles is reduced to improve the attack efficiency. Moreover, we develop a parallel mechanism of attacking multiple rectangles simultaneously to further accelerate the attack process. Extensive experiments demonstrate that our method can effectively and efficiently attack various popular object detectors, including anchor-based and anchor- free, and generate transferable adversarial examples.
Read moreGM-Attack: Improving the Transferability of Adversarial Attacks
In the real world, blackbox attacks seem to be widely existed due to the lack of detailed information of models to be attacked. Hence, it is desirable to obtain adversarial examples with high transferability which will facilitate practical adversarial attacks. Instead of adopting traditional input transformation approaches, we propose a mechanism to derive masked images through removing some regions from the initial input images. In this manuscript, the removed regions are spatially uniformly distributed squares. For comparison, several transferable attack methods are adopted as the baselines. Eventually, extensive empirical evaluations are conducted on the standard ImageNet dataset to validate the effectiveness of GM-Attack. As indicated, our GM-Attack can craft more transferable adversarial examples compared with other input transformation methods and attack success rate on Inc-v4 has been improved by 6.5% over state-of-the-art methods.KeywordsDeep neural networksAdversarial attackAdversarial examplesData augmentationWhite-box/black-box attackTransferability
Read moreSelf-Attention Context Network: Addressing the Threat of Adversarial Attacks for Hyperspectral Image Classification.
Deep learning models have shown their great capability for the hyperspectral image (HSI) classification task in recent years. Nevertheless, their vulnerability towards adversarial attacks could not be neglected. In this study, we systematically analyze the influence of adversarial attacks on the HSI classification task for the first time. While existing research of adversarial attacks focuses on the generation of adversarial examples in the RGB domain, the experiments in this study show such adversarial examples could also exist in the hyperspectral domain. Although the difference between the generated adversarial image and the original hyperspectral data is imperceptible to the human visual system, most of the existing state-of-the-art deep learning models could be fooled by the adversarial image to make wrong predictions. To address this challenge, a novel self-attention context network (SACNet) is further proposed. We discover that the global context information contained in HSI can significantly improve the robustness of deep neural networks when confronted with adversarial attacks. Extensive experiments on three benchmark HSI datasets demonstrate that the proposed SACNet possesses stronger resistibility towards adversarial examples compared with the existing state-of-the-art deep learning models.
Read more