01 — The principle

A small perturbation can cross the model’s decision boundary.

An adversarial example is an input deliberately modified to cause a model error. The perturbation is generally small and sometimes imperceptible to humans. Mathematical norms quantify its size: the L0 norm counts modified pixels, the L2 norm measures Euclidean distance between images, and the L∞ norm bounds the maximum change to each pixel.

A targeted attack seeks a specific class; an untargeted attack seeks any error. In a white-box setting, the attacker knows the architecture and parameters. In a black-box setting, access to the model is limited. Adversarial transfer describes the ability of an example designed for one network to fool another.

02 — Two reference attacks

Fast Gradient Sign Method and Projected Gradient Descent: two optimization strategies.

The Fast Gradient Sign Method (FGSM) uses the sign of the loss gradient to move each pixel in a direction that rapidly increases error. As a single-step method, it provides a fast and inexpensive initial robustness test.

Projected Gradient Descent (PGD) repeats this operation and projects the result back into the allowed perturbation region. This iterative optimization generally finds stronger examples. A robust evaluation therefore combines simple, iterative and defense-adaptive attacks.

03 — Why it works

The network may exploit predictive cues that are difficult for humans to perceive.

My research connects adversarial sensitivity with convolutional neural network (CNN) reliance on textures and high spatial frequencies. Such information may be highly predictive in training data while carrying little semantic meaning for a human observer. A small optimized change can then move the decision.

This view also partly explains transfer: multiple architectures learn similar non-robust features. It encourages complementing average accuracy with an analysis of learned shortcuts and behavior under distribution shift.

04 — Evaluating defenses

Test each protection against several adaptive attacks.

My research dissertation reviews four families: data preprocessing or augmentation, adversarial detectors, gradient masking and adversarial training. Adversarial training remains a strong approach, but it is costly and can remain vulnerable to adaptive or out-of-distribution attacks. Gradient masking can give a false sense of security: an attack that bypasses the gradient problem may still fool the model. A defense therefore needs evaluation with adaptive attacks.

In my research dissertation simulations, inhibiting supposedly non-robust neurons did not improve performance. Transfer learning from frequency-filtered adversarial examples, however, raised accuracy against PGD on standard images from about 13.6% to 59.1%. This result remains contextual: a limited dataset, no statistical test and a need for black-box evaluation. Frequency structure offers a measurable direction for robustness.