FBFS: fog backdoor attack based on feature similarity
1 School of Mathematics and Statistics, Lanzhou University, Lanzhou 730000, China
2 School of Mathematics and Computer Science, Guangdong Ocean University, Zhanjiang 524088, China
Abstract

Most existing backdoor attacks embed triggers in images in ways that are either conspicuous to human observers or easily detected by feature-space defenses, thereby sacrificing stealthiness or robustness. To address these issues, we propose a fog backdoor attack based on feature similarity (FBFS), which enhances the visual concealment of backdoor triggers as well as the feature-space homogeneity between poisoned and clean samples. Specifically, FBFS employs a standard optical model to simulate natural fog as a visually plausible trigger and injects it into image samples. Additionally, a feature similarity penalty term is incorporated into the loss function to enforce consistency in the feature representations of poisoned and clean samples, thereby evading defenses that rely on latent separability. Experiments conducted on Canadian Institute for Advanced Research 10-class dataset (CIFAR-10), German Traffic Sign Recognition Benchmark (GTSRB), and a subset of ImageNet demonstrate that, under a 10% poisoning rate, FBFS achieves over 90% attack success rate while maintaining clean sample accuracy above 85%. Moreover, detection rates under representative feature-space defense methods, including activation clustering and spectral signature analysis, remain below 40%, demonstrating that the proposed method effectively balances attack performance and resistance to detection, exhibiting both stealth and robustness.

Keywords

backdoor attack; feature similarity; visual stealth; feature space; latent separability; model robustness

Preview