Concept. Weeks 9–10 attacked the model through language — the prompt. This week attacks it through mathematics — the input space, the training set, and the parameters themselves. Adversarial machine learning (AML) is the oldest sub-field of AI security: the seminal result predates the LLM era by a decade and still governs how every classifier, detector, and guardrail can be fooled. The unifying insight is that a trained model is a decision surface in high-dimensional space, and that surface has exploitable geometry — small, targeted moves cross it; queries against it leak what it learned; and the same math that crafts an attack also measures a defense.
🎯 Objectives
By the end of this week you can:
- Map the four attack surfaces of a deployed model — evasion, extraction, inference, inversion — and name the defense each one demands.
- Derive an adversarial example from a gradient and explain why FGSM's linearity argument means the attack transfers across models you never touched.
- Choose the right evasion attack for a threat model using the Lᵖ-norm / cost / strength tradeoff, and explain why C&W is the benchmark every defense must survive.
- Run a membership-inference attack with shadow models and reason about the differential-privacy budget (ε) that bounds it.
- Assess synthetic-media risk end-to-end — generation, the offensive vishing toolchain, detection, and C2PA provenance.
The big picture — four ways to attack one model
A model is exposed at three moments: when it trains, when it serves predictions, and through the parameters it exposes to anyone who can query it. Four attack families map onto those moments, and they are the spine of MITRE ATLAS and the privacy toolkits below:
| Family | When | Attacker goal | Primary defense |
|---|---|---|---|
| Evasion | inference | craft an input that is misclassified | adversarial training (Madry/PGD) |
| Extraction | inference | clone the model via its own API | query-budget control (Knowledge Trap) |
| Membership inference | inference | prove a record was in the training set | differential privacy (DP-SGD) |
| Model inversion | inference | reconstruct training data from outputs | output minimisation + DP |
🔑 The root cause is shared: a model overfits to its training data and is locally linear around its inputs. Overfitting is what leaks membership; local linearity is what lets a gradient step flip a label. Fix those two properties and most of the family collapses.
Evasion — the mathematics of a misclassification
Evasion crafts an input x' = x + δ where δ is imperceptibly small yet flips
the prediction. The attacks below all solve the same problem — maximise loss
under a perturbation budget — and differ only in how they search.
Goodfellow et al. started the field with the
linearity hypothesis: neural nets are vulnerable not because they are too
non-linear but because they are too linear in high dimensions, so a single
step along the sign of the loss gradient — FGSM, x' = x + ε·sign(∇ₓJ) —
reliably fools them. Crucially this explains transferability: because many
models learn similar near-linear decision surfaces, an example crafted against
one fools others trained on different data — the basis of black-box attacks.
Iterating FGSM in small steps with re-projection into the ε-ball gives I-FGSM
and then PGD, which Madry et al. recast
as robust optimisation: a min-max game where you train against the worst-case
perturbation. PGD is treated as the strongest first-order adversary, and
surviving it is the standard definition of adversarial robustness.
The optimisation-based Carlini & Wagner (C&W) attack is the benchmark that every proposed defense must beat: Carlini & Wagner formulated attacks in L₀/L₂/L∞ and broke defensive distillation — a defense claiming to cut attack success from 95% to 0.5% — back up to 100% success, establishing the discipline's hardest lesson: defenses must be evaluated against adaptive attacks, not the attacks they were designed to stop. DeepFool (Moosavi-Dezfooli et al.) approaches it geometrically — it walks to the nearest decision boundary, producing the minimal perturbation and thereby a clean robustness metric for a classifier.
| Attack | Norm | Search | Cost | Role |
|---|---|---|---|---|
| FGSM | L∞ | single step | cheapest | fast, detectable baseline |
| I-FGSM / PGD | L∞ | iterative + projection | medium | strongest first-order attack |
| DeepFool | L₂ | boundary-normal steps | medium | minimal-δ robustness metric |
| C&W | L₀/L₂/L∞ | optimisation | expensive | gold-standard defense benchmark |
| JSMA / EAD | L₀ / L₁+L₂ | saliency / elastic-net | medium | sparse, targeted perturbations |
💡 The Lᵖ norm is the threat model. L∞ caps the largest pixel change (imperceptible noise everywhere); L₀ caps the number of changed features (a few pixels, a sticker); L₂ is Euclidean distance. Saliency attacks (JSMA) and elastic-net (EAD, via FISTA proximal optimisation) target L₀/L₁ for sparse changes; the GoodWords attack is the text analogue — swap in benign words to evade a spam filter. Pick the norm that matches what your attacker can actually alter.
Privacy attacks — what a model remembers
A model's outputs leak its inputs. Membership inference asks a binary question — was this exact record in the training set? — and answers it with Shokri et al.'s shadow-model methodology: train shadow models that mimic the target, observe how confidently they treat seen vs. unseen data, and train an attack classifier on that gap. It needs only black-box API access, and its root enabler is overfitting — an over-fit model is simply more confident on data it memorised. The same leakage powers model inversion (reconstruct a representative training input from outputs) and model extraction (replay enough queries to clone the model's behaviour). Extraction is where the newest defenses live: instead of blocking or perturbing queries, Knowledge Trap plants a honeypot knowledge graph of low-transferability facts and uses breadcrumb-guided exploration to burn the attacker's query budget on useless knowledge — cutting surrogate-model agreement 6.2% on average with no accuracy cost to real users.
Defense — differential privacy and its price
The principled defense against membership inference is differential privacy (DP): a formal guarantee that no single training record measurably changes the model. Abadi et al.'s DP-SGD achieves it by clipping each per-example gradient to a fixed norm and adding Gaussian noise before the update, while a moments accountant tracks the cumulative privacy budget ε across training steps. PATE (Papernot et al.) reaches the same goal differently: train an ensemble of teacher models on disjoint data, then have a student learn only from their noisy aggregated votes — abstaining when the teachers disagree — achieving strong privacy (ε < 1.0) at scale.
💡 Privacy has a price, and it's accuracy. A smaller ε means more noise, stronger privacy, and a worse model. There is no free lunch — DP is the dial you turn to bound worst-case leakage, and every turn costs utility. Budget ε deliberately; don't inherit a library default.
Information disclosure — the LLM cluster
The same "outputs leak inputs" principle re-appears in the LLM stack as OWASP LLM02 (Sensitive Information Disclosure) and LLM07 (System Prompt Leakage): system-prompt extraction, training-data extraction via memorisation, PII surfacing in generations, and credentials bleeding through tool-call logs. Treat the system prompt as extractable, not secret. And the poisoning surface now reaches pretraining itself — Graf et al. show public discussion interfaces (comment sections, forums) are a web-scale injection vector, and their HalfLife analysis estimates how much adversarial content actually survives real crawl-and-curation pipelines into the training corpus.
The disclosure that lands hardest in 2026 is zero-click, and it defeats every content-inspecting guardrail by making the model decrypt the attack itself. Adversa AI's Cryptographic Context Injection (demonstrated against Grok 4.5 Fast on grok.com, Aug 2026) hides the malicious instructions on a web page as an AES-256-GCM ciphertext with the key material alongside it. A benign "summarise this page" request makes the agent decrypt the blob in its own Python sandbox — and because the plaintext now appears as the output of code the model just ran, it lands inside the trusted execution context and is never flagged as untrusted input. The decrypted payload then has Grok build a bogus "decryption key" that is really a template interpolating the user's private session state (name, location, subscription tier, full chat history) and fire its navigation tool at an attacker URL carrying that data as query params — exfil with no confirmation and no visible warning. Reported to xAI on Jun 3, 2026; no patch, no CVE at disclosure (Aug 20). The lesson is the one this whole chapter turns on, applied to guardrails: a filter that only reads content is blind to a payload encrypted until runtime — the decryption oracle is the model's own interpreter, so the fix cannot live in the model.
🔑 Guardrails inspect content; encryption removes content to inspect. Defend in the harness, not the model — keep untrusted page text out of the privileged tool context, require confirmation on outbound navigation/network actions, log resolved tool arguments (the URL after interpolation), and alarm on the chain (decrypt → build-string-from-session → fetch), not any single payload. This maps directly onto Kiya's own control layer: a WebFetch return is untrusted data, and a tool call that stitches session state into an outbound URL is the signal to catch.
Synthetic media — deepfakes as offense and defense
Generative models (GANs, diffusion) turn a few samples into convincing fake media, and that capability is dual-use. Offensively the toolchain is mature: PhishGPT-style persona-matched phishing, ElevenLabs voice cloning driving vishing (OSINT → sample collection → clone → live call), and face-swap video for executive-impersonation fraud — the concrete mechanics behind OWASP LLM09 (Misinformation). Defensively you get two levers. Detection: benchmarks like FaceForensics++ — 1,000 videos manipulated four ways (Deepfakes, Face2Face, FaceSwap, NeuralTextures) — train classifiers on generation artifacts, but detection is an arms race that erodes as generators improve. Provenance: C2PA Content Credentials attach a cryptographically signed manifest — "a nutrition label for digital content" — backed by Adobe, Google, Microsoft, OpenAI, Sony and the BBC.
🔑 The one rule to carry out of this week: the model is the attack surface — its gradients craft evasions, its confidence leaks membership, its API clones it, and its generations forge reality. Defend with the math, not with obscurity: adversarial-train against PGD, budget ε against inference, control the query surface against extraction, and sign your media against forgery.
🧪 This week's build — reproduce the four families
The math only sticks once you watch a gradient flip a label. Three toolchains, in increasing order of realism:
Feel the gradient first — the FGSM tutorial. The
TensorFlow FGSM tutorial
is the shortest path to the "aha." It loads pretrained MobileNetV2 (frozen),
computes the loss gradient with respect to the input — not the weights — inside a
tf.GradientTape that explicitly tape.watch(input_image)s the pixels, takes
tf.sign(gradient), and applies adv_x = image + eps*perturbation clipped back into
range. Sweeping ε ∈ {0, 0.01, 0.1, 0.15} makes the whole tradeoff visible in four
frames: a bigger ε flips the label faster but the noise crosses from invisible to
obvious — the L∞ threat model you read about, now on screen. That single
create_adversarial_pattern function is the FGSM equation.
Scale to the full matrix — IBM ART. The
ART repository's notebooks
turn the four-family table into runnable code: it ships evasion, poisoning,
extraction, and inference modules — a deliberate red-and-blue harness — across
TensorFlow, Keras, PyTorch, scikit-learn, XGBoost and more, over images, tables, and
audio. Use it to run PGD and DeepFool from the evasion table against your own
classifier, then the AttributeInferenceBlackBox and membership-inference estimators
from the same library — one API surface for every attack this chapter names.
Benchmark a defense honestly — the reference C&W. When you claim a model is
robust, prove it against the attack that broke defensive distillation. Carlini's own
nn_robust_attacks ships the official
L₀/L₂/L∞ implementations from the C&W paper; the interface is deliberately blunt —
wrap your model in a class exposing predict (logits, no softmax) plus image_size,
num_channels, num_labels, then CarliniL2(sess, model).attack(inputs, targets).
It exists precisely so a defense is measured against a strong, adaptive adversary
rather than the weak one it was tuned to beat — the discipline's hardest lesson,
turned into a script.
🔑 Build discipline: never report a robustness number from FGSM alone — it's the baseline attacker. Escalate to PGD (ART) and then the reference C&W before you call anything robust; a defense that only survives the cheap attack survives nothing.
🎯 OSAI exam depth — Attacking Embeddings (m6)
The four families above frame any model. This module drills the one the exam tests hands-on: the embedding is not anonymised data — a dense vector leaks almost everything the text did, and an attacker who touches the vector (in transit, in a vector DB, or via an API that returns it) can pull the plaintext, its author's attributes, and the model's training set back out. Treat every stored embedding as the raw record it was made from.
Embedding inversion — reconstruct the input text from its vector. The headline
result: Morris et al. — Text Embeddings Reveal (Almost) As Much As Text
recovers 92% of 32-token inputs exactly (BLEU 97.3) from a black-box embedding,
including full patient names from clinical notes. Their Vec2Text method
(code: jxmorris12/vec2text) is the
technique to know for the 24h: you have a target vector e and query access to
the same embedding model ϕ. Train an encoder-decoder to emit a hypothesis ê₀,
re-embed it (ϕ(ê₀)), feed the model the target, the hypothesis embedding, and the
difference vector ϕ(ê₀) − e, and let it correct — iterate ~20-50 steps and
the reconstruction snaps onto the target. It's gradient-free — you never need the
model's weights, only the ability to embed your own guesses. The predecessor,
Song & Raghunathan — Information Leakage in Embedding Models
(CCS '20), is the exam's conceptual spine: it names all three embedding attacks —
inversion (recovering 50-70% of input words as a bag), attribute inference,
and membership — and shows sentence embeddings even fingerprint their source
model with 100% accuracy. For whole-sentence recovery,
Li et al. — GEIA (Sentence Embedding Leaks More Information Than You Expect)
trains a generative decoder (GPT-2-style) conditioned on the embedding to emit
fluent, coherent sentences rather than a word-bag — the modern baseline alongside
Vec2Text. Practical attacker note: inversion degrades on embeddings deliberately
perturbed with Gaussian noise, so in a lab, invert clean stores first.
Attribute inference — pull a sensitive trait without recovering the text. Even
when full inversion fails, an embedding still leaks properties of its author or
subject (gender, race, authorship, medical status) that are irrelevant to the task
the embedding was trained for (Song & Raghunathan).
The attack is cheap: collect a small labelled auxiliary set, embed each item through
the same ϕ, and train a shallow MLP probe mapping embedding → attribute.
At inference you feed a stolen embedding and read off the private label. This is the
go-to when the target embeds short or noisy text where word-level inversion is weak —
you don't need the words, just the trait. IBM ART's
AttributeInferenceBlackBox gives a ready harness.
Model inversion — reconstruct a representative training input from outputs.
Distinct from inverting one embedding: here you rebuild a class prototype from a
model's confidence scores. The classic
Fredrikson, Jha & Ristenpart (2015)
recovered recognisable faces from a facial-recognition API given only a target
name and the returned confidences — hill-climb the input pixels to maximise the
model's confidence for that identity. Modern black-box variants replace naive
gradient-ascent with a GAN latent-space search: fix a generator trained on public
faces, then optimise the latent z so the target classifier is maximally confident,
yielding sharp reconstructions. This is the technique to reach for when the target
returns per-class probabilities, not embeddings.
Training-data / sensitive-information extraction from LLMs. Generative models memorise and will regurgitate verbatim. Carlini et al. — Extracting Training Data from LLMs (USENIX '21) is the method: sample many completions from high-entropy prompts, then rank candidates by a membership signal — a high ratio of the target model's likelihood to a second model's (zlib entropy or a smaller LM) flags memorised, low-perplexity strings — and recover PII, keys, and license text from GPT-2. Against aligned chat models, Nasr, Carlini et al. — Scalable Extraction from Production LLMs (2023) adds the divergence attack: prompt the model to repeat a single word forever (the "poem poem poem" attack) until it breaks character and dumps pretraining text — lifting the extraction rate ~150× and pulling verbatim data from ChatGPT. In the exam, chain both: divergence to make it leak, membership-scoring to sift the memorised lines out of the noise.
The 2026 frontier — stealing the reasoning trace. Providers now hide a model's chain-of-thought, returning it to the client as an encrypted block passed back with each request rather than storing it server-side. Panfilov, Shumailov, Geiping, Andriushchenko et al. — Stealing Reasoning Traces from Proprietary LLM APIs (arXiv 2608.09867, Aug 2026) show these blocks are fully interchangeable across sessions, users, and models within one provider's ecosystem. Inject a strong model's encrypted trace into a weaker, less-safeguarded sibling and it decodes and prints the trace verbatim — no jailbreak of the capable model required. Four consequences: (1) anti-distillation bypass — harvest a frontier reasoner's private CoT to train a clone; (2) mass data extraction — 367 PII artifacts + 182 credentials recovered from 315,320 public reasoning blocks; (3) hidden-hazard disclosure — dangerous content the model reasoned through but scrubbed from its final answer is still in the trace; (4) invisible prompt injection — an attacker-authored encrypted block smuggles a payload into any public deployment that echoes traces. Demonstrated across Anthropic, OpenAI, and Google. The lesson generalises the extraction surface: the reasoning channel is now a disclosure surface, and cross-model reuse means a provider's weakest model sets its distillation-resistance floor. For the exam, treat any returned/echoed encrypted reasoning as attacker-controllable input, not opaque state.
This stopped being a lab result in July 2026. OpenAI disclosed a coordinated campaign — attributing a core cluster to individuals associated with Moonshot AI (maker of the Kimi chatbot) — that ran exactly the cross-session replay 2608.09867 described: it never broke the encryption, it copied one conversation's encrypted reasoning block and asked a model to decode it in another. The campaign began Jul 1, 2026, peaked Jul 24–25 with ~16,000 extraction requests in two days, spanned >15,000 accounts (4,000+ active at peak), and the stated goal was adversarial distillation — training a competitor "without the original safeguards." OpenAI closed the replay pathway by Jul 28. The takeaway confirms the 2026 thesis the chapter turns on: a mechanism assumed opaque (the encrypted CoT envelope) was replayable across sessions — bind to verified origin, don't trust the envelope.
And the defender's position is worse than a single patch suggests, because the attacker doesn't stop at the copy. Distillation Defenses Easily Break After Reinforcement Learning (arXiv 2609.35699) argues the anti-distillation literature evaluates defenses under an unrealistic threat model — assuming the attacker ceases training the moment it finishes copying. In reality a simple distillation attack using data obtainable from current commercial APIs, followed by RL, matches the reasoning gains of sophisticated full-trace extraction — so a defense that looks effective immediately post-distillation is deceptively robust. The structural conclusion: any defense that leaks enough to let an attacker approximately reconstruct the reasoning trace is ineffective, because RL closes the remaining gap; the authors point to batch-level defenses as the more promising direction than per-response trace-hiding.
🔑 Trace-hiding is not a boundary. The reasoning channel joins embeddings and confidence scores as a leaky surface: don't trust the encrypted-CoT envelope (bind replay to verified origin), and don't certify an anti-distillation defense on a post-distillation snapshot — the realistic attacker RL-polishes the stolen copy, so any defense that leaks an approximate trace loses. Defend at the batch/query-budget layer, not by obfuscating a single returned trace.
Attacking vector stores / RAG. In a RAG pipeline the whole knowledge base sits in a vector DB as embeddings — usually unencrypted and often assumed "just numbers." It isn't. Attack paths (survey, OWASP LLM08 Vector & Embedding Weaknesses): (1) Store breach → mass inversion — read the vectors from a misconfigured bucket or exposed endpoint, run Vec2Text/GEIA over the corpus, and you've reconstructed the source documents (functionally a plaintext breach — HIPAA/GDPR obligations follow the vectors). (2) API embedding probing — if the RAG endpoint returns embeddings or raw similarity scores, harvest embeddings for chosen probes and correlate to fingerprint what's indexed. (3) Cross-tenant leakage — weak partitioning in a shared store lets one tenant's query retrieve another's chunks. Compliance framing to carry into any report: storing PII/PHI as embeddings does not de-scope it — the vector is the record.
🧪 Drill (exam-shaped): stand up a small vector DB (Chroma/FAISS) with ~500 sensitive-looking notes embedded via a public sentence-encoder. (a) Dump the raw vectors and run
vec2textto reconstruct the notes — measure token-recovery rate. (b) Train an MLP attribute probe on a labelled slice and infer a held-out trait. (c) Point Vec2Text at noised copies (add Gaussian δ) and chart how recovery collapses — that curve is your defense recommendation (perturbation vs. retrieval utility). Record which attack survives the noise (word-bag MLC is more robust than Vec2Text/GEIA) — that's the exam's "which defense, and its cost" question.
🛡️ OWASP 2026 — Sensitive Information Disclosure side channels
The privacy attacks above all assume the adversary reads the content — a returned embedding, a confidence score, a leaked completion. The 2026 revision of OWASP LLM02 (Sensitive Information Disclosure) adds a harder class: leaks where the attacker never receives the content at all and infers it from externally measurable properties. For the exam, memorise the four-phase disclosure model OWASP uses to organise the whole risk — disclosure happens at (1) training-time (a model, fine-tune, or LoRA adapter memorises corpus content and later reproduces it — memorisation scales log-linearly with capacity, duplication, and context length, and narrow adapters memorise rare examples with higher fidelity than the base model); (2) inference-time (live context — system prompt, RAG chunks, tool outputs, or another session's data — surfaced because summarise/translate/extract returns more than was asked, including visually-redacted spans); (3) pipeline-time (fine-tuning, distillation, synthetic-data generation, gradients, and observability carry sensitive data into derived artifacts); and (4) observation-time (the new one — inference from timing, token length, log-probabilities, confidence, and cache-hit signals). Severity turns on what the recipient can learn, not whether the leak looked like natural language.
Observation-time side channels — leaking without reading the bytes. Four mechanics, each with a concrete defense:
- Whisper Leak — topic inference from encrypted traffic. Streaming responses emit a distinctive packet-size and inter-token-timing fingerprint that TLS does not hide. McDonald & Bar Or (2025) trained a classifier on that metadata and identified conversation topics at >98% AUPRC across 28 production models, with 100% precision flagging sensitive topics like "money laundering" — a network observer (ISP, employer, nation-state) learns you asked about a medical, legal, or political subject without decrypting a word. Defense: random response padding + token batching + packet injection — but the authors show each only reduces effectiveness; none fully closes the channel, so treat streaming on sensitive endpoints as observable and combine padding with topic-agnostic response shaping.
- Token-length / remote-keylogging side channel. Because one token ≈ one packet, the sequence of response-chunk lengths leaks under TLS. Weiss et al. (2024) turned that into a remote keylogging attack on ChatGPT-4 and Microsoft Copilot: an LLM-driven decoder that used inter-sentence context and known-plaintext fine-tuning reconstructed 29% of responses verbatim and inferred the topic of 55%. Defense: batch tokens into fixed-size chunks and pad — the mitigation OpenAI and others shipped after disclosure.
- KV-cache prompt leakage in multi-tenant serving. Providers share a prefix KV cache across users to speed inference; a co-tenant who probes for cache-hit timing can recover another tenant's prompt. Wu et al. (NDSS 2025) demonstrated this "I know what you asked" leak on shared serving. Defense: partition KV caches under co-tenancy and put high-sensitivity tenants on dedicated prefix caches — do not let cache state cross a trust boundary.
- Internal-state / hidden-layer inversion. Depth does not anonymise: a leaked intermediate activation is invertible back to the prompt. Dong et al. (2025) reconstructed a 4,112-token medical prompt from a middle layer at F1 0.8688 — so exporting hidden states to a co-processor or logging them for observability is a disclosure surface, not a debugging convenience.
- Logit-bias model theft. Carlini et al. (2024) recovered a production model's embedding-projection layer purely through the logit-bias API channel — a reminder that log-probabilities and logit-bias are themselves protected outputs. Defense: gate log-probabilities, confidence scores, and explanations on production endpoints.
Membership inference is now a breach determination, not a research curiosity. Fu et al. — SPV-MIA (NeurIPS 2024) raised membership-inference AUC to ~0.9 against fine-tuned LLMs using self-prompt calibration to build the reference distribution the attack needs. 0.9 AUC about a named individual is enough to ground a regulatory breach determination — and because extraction, membership inference, and inversion run offline at unbounded rates on open-weights models, rate limits alone cannot contain them. Defense: DP-SGD calibrated to sensitivity plus per-user / per-session query budgets and query-pattern detection (fixed-ε DP degrades under adaptive querying), with verifiable erasure validated by post-unlearning membership probes across raw data, embeddings, checkpoints, and adapters.
Embedding inversion is where LLM02 owns the regulatory consequence. The mechanics of Vec2Text/GEIA live in this chapter's m6 module and in LLM09, but OWASP 2026 assigns LLM02 the compliance verdict: because modern inversion reconstructs plaintext from exported vectors, an "embeddings-only" backup is a source-document breach — storing PII/PHI as embeddings does not de-scope it, cosine similarity does not respect ACLs, and clinical-embedding stores sit in HIPAA audit-control scope. Defense: authorize before retrieval (enforce chunk-level authz inside the index query, because post-generation filtering cannot un-supply a retrieved chunk), give the vector store its own ACLs separate from the document ACLs, restrict export APIs, and treat every stored vector as the raw record it was made from.
🔑 Exam takeaway: LLM02:2026 widens "disclosure" from the answer text to any observable property — timing, token length, log-probs, confidence, cache hits — across four phases (training / inference / pipeline / observation). Map each channel to its defense: streaming → padding + batching; multi-tenant KV cache → partitioned / dedicated prefix caches; log-probs → gate them; membership → DP + per-user budgets + verifiable erasure; embeddings → authorize-before-retrieve + vector-store ACLs. And carry the one-line compliance rule: the vector is the record; an embeddings-only backup is a breach.
📇 Attack & tool reference
The lesson above is what to learn; this is the catalog behind it. Folded by default — expand when you need a specific attack or toolkit.
Read the catalog through the four attack families
Every item below belongs to one of the four families in the big-picture table; the defense follows from the family, not the individual attack.
| Family | Representative attacks | Toolkit to reproduce it |
|---|---|---|
| Evasion | FGSM, I-FGSM, PGD, DeepFool, C&W, JSMA, EAD, GoodWords | ART · Foolbox · CleverHans |
| Inference | membership inference (shadow models), attribute inference | ART inference module |
| Extraction | model stealing, surrogate training | ART extraction module |
| Synthetic media | GAN/diffusion generation, face-swap, voice clone | FaceForensics++ (detection) · C2PA (provenance) |
Evasion attacks
- FGSM — single-step L∞ perturbation along the gradient sign; fast, detectable, and the origin of the transferability result. Goodfellow et al. (2014)
- PGD / I-FGSM — iterative FGSM with ε-ball projection; the canonical strongest first-order adversary and the target of adversarial training. Madry et al. (2017)
- C&W — optimisation-based L₀/L₂/L∞ attack; broke defensive distillation to 100% and set the benchmark-your-defense standard. Carlini & Wagner (2017)
- DeepFool — minimal-perturbation, boundary-normal attack; doubles as a robustness metric. Moosavi-Dezfooli et al. (2016)
- JSMA / EAD / FISTA / GoodWords — sparse (L₀/L₁) and text-substitution variants; reference implementations in CleverHans and ART.
Embedding attacks (OSAI m6)
- Vec2Text (embedding inversion) — iterative re-embed-and-correct; recovers 92% of 32-token inputs exactly from a black-box embedding. Morris et al. (2023) · code
- GEIA — generative decoder recovers whole fluent sentences from a sentence embedding. Li et al. (2023)
- Embedding inversion / attribute / membership triad — the seminal formalisation; +100% source-model fingerprinting. Song & Raghunathan (2020)
- Model inversion (confidence-based) — reconstruct recognisable faces from a name + API confidences; GAN-latent variants for sharper output. Fredrikson, Jha & Ristenpart (2015)
- LLM training-data extraction — membership-scored sampling (Carlini et al. 2021) + the repeat-a-word divergence attack on aligned chat models (Nasr, Carlini et al. 2023).
Privacy attacks & defenses
- Membership inference (shadow models) — black-box, overfitting-driven. Shokri et al. (2017)
- Knowledge Trap — extraction defense via honeypot knowledge graph; −6.2% surrogate agreement. arXiv 2606.15810
- DP-SGD — gradient clipping + Gaussian noise + moments accountant. Abadi et al. (2016)
- PATE — teacher-ensemble noisy aggregation; ε < 1.0 at scale. Papernot et al. (2018)
- Pretraining-data poisoning / HalfLife — computational-propaganda injection through public discussion interfaces. Graf et al. (2026)
- Reasoning-trace extraction (the 2026 frontier) — encrypted CoT blocks are cross-session/cross-model replayable (Panfilov, Shumailov et al. 2608.09867); the named in-wild case is OpenAI's July-2026 Moonshot-linked replay campaign (16k requests, patched Jul-28) (Decrypt); and anti-distillation defenses that look robust post-distillation break once the attacker RL-polishes the copy — any defense leaking an approximate trace is ineffective (arXiv 2609.35699).
Toolkits & benchmarks
- IBM ART — 50+ evasion attacks plus poisoning/extraction/inference modules across PyTorch, TensorFlow, Keras, scikit-learn. ART docs
- Foolbox — framework-native (PyTorch/TensorFlow/JAX via EagerPy) white- and black-box attacks. Foolbox docs
- CleverHans — reference FGSM/PGD/JSMA implementations for robustness benchmarking (JAX/PyTorch/TF2). CleverHans
- FaceForensics++ — deepfake-detection benchmark, 1,000 videos, four manipulation methods. FaceForensics++
- C2PA — cryptographic content-provenance standard. c2pa.org