Skip to content
Phase 2, week 11

Adversarial Machine Learning

0 of 57 items done. ~14h30m estimated.

Concept. Weeks 9–10 attacked the model through language — the prompt. This week attacks it through mathematics — the input space, the training set, and the parameters themselves. Adversarial machine learning (AML) is the oldest sub-field of AI security: the seminal result predates the LLM era by a decade and still governs how every classifier, detector, and guardrail can be fooled. The unifying insight is that a trained model is a decision surface in high-dimensional space, and that surface has exploitable geometry — small, targeted moves cross it; queries against it leak what it learned; and the same math that crafts an attack also measures a defense.

🎯 Objectives

By the end of this week you can:

  • Map the four attack surfaces of a deployed model — evasion, extraction, inference, inversion — and name the defense each one demands.
  • Derive an adversarial example from a gradient and explain why FGSM's linearity argument means the attack transfers across models you never touched.
  • Choose the right evasion attack for a threat model using the Lᵖ-norm / cost / strength tradeoff, and explain why C&W is the benchmark every defense must survive.
  • Run a membership-inference attack with shadow models and reason about the differential-privacy budget (ε) that bounds it.
  • Assess synthetic-media risk end-to-end — generation, the offensive vishing toolchain, detection, and C2PA provenance.

The big picture — four ways to attack one model

A model is exposed at three moments: when it trains, when it serves predictions, and through the parameters it exposes to anyone who can query it. Four attack families map onto those moments, and they are the spine of MITRE ATLAS and the privacy toolkits below:

Family When Attacker goal Primary defense
Evasion inference craft an input that is misclassified adversarial training (Madry/PGD)
Extraction inference clone the model via its own API query-budget control (Knowledge Trap)
Membership inference inference prove a record was in the training set differential privacy (DP-SGD)
Model inversion inference reconstruct training data from outputs output minimisation + DP

🔑 The root cause is shared: a model overfits to its training data and is locally linear around its inputs. Overfitting is what leaks membership; local linearity is what lets a gradient step flip a label. Fix those two properties and most of the family collapses.

Evasion — the mathematics of a misclassification

Evasion crafts an input x' = x + δ where δ is imperceptibly small yet flips the prediction. The attacks below all solve the same problem — maximise loss under a perturbation budget — and differ only in how they search.

Goodfellow et al. started the field with the linearity hypothesis: neural nets are vulnerable not because they are too non-linear but because they are too linear in high dimensions, so a single step along the sign of the loss gradient — FGSM, x' = x + ε·sign(∇ₓJ) — reliably fools them. Crucially this explains transferability: because many models learn similar near-linear decision surfaces, an example crafted against one fools others trained on different data — the basis of black-box attacks. Iterating FGSM in small steps with re-projection into the ε-ball gives I-FGSM and then PGD, which Madry et al. recast as robust optimisation: a min-max game where you train against the worst-case perturbation. PGD is treated as the strongest first-order adversary, and surviving it is the standard definition of adversarial robustness.

The optimisation-based Carlini & Wagner (C&W) attack is the benchmark that every proposed defense must beat: Carlini & Wagner formulated attacks in L₀/L₂/L∞ and broke defensive distillation — a defense claiming to cut attack success from 95% to 0.5% — back up to 100% success, establishing the discipline's hardest lesson: defenses must be evaluated against adaptive attacks, not the attacks they were designed to stop. DeepFool (Moosavi-Dezfooli et al.) approaches it geometrically — it walks to the nearest decision boundary, producing the minimal perturbation and thereby a clean robustness metric for a classifier.

Attack Norm Search Cost Role
FGSM L∞ single step cheapest fast, detectable baseline
I-FGSM / PGD L∞ iterative + projection medium strongest first-order attack
DeepFool L₂ boundary-normal steps medium minimal-δ robustness metric
C&W L₀/L₂/L∞ optimisation expensive gold-standard defense benchmark
JSMA / EAD L₀ / L₁+L₂ saliency / elastic-net medium sparse, targeted perturbations

💡 The Lᵖ norm is the threat model. L∞ caps the largest pixel change (imperceptible noise everywhere); L₀ caps the number of changed features (a few pixels, a sticker); L₂ is Euclidean distance. Saliency attacks (JSMA) and elastic-net (EAD, via FISTA proximal optimisation) target L₀/L₁ for sparse changes; the GoodWords attack is the text analogue — swap in benign words to evade a spam filter. Pick the norm that matches what your attacker can actually alter.

Privacy attacks — what a model remembers

A model's outputs leak its inputs. Membership inference asks a binary question — was this exact record in the training set? — and answers it with Shokri et al.'s shadow-model methodology: train shadow models that mimic the target, observe how confidently they treat seen vs. unseen data, and train an attack classifier on that gap. It needs only black-box API access, and its root enabler is overfitting — an over-fit model is simply more confident on data it memorised. The same leakage powers model inversion (reconstruct a representative training input from outputs) and model extraction (replay enough queries to clone the model's behaviour). Extraction is where the newest defenses live: instead of blocking or perturbing queries, Knowledge Trap plants a honeypot knowledge graph of low-transferability facts and uses breadcrumb-guided exploration to burn the attacker's query budget on useless knowledge — cutting surrogate-model agreement 6.2% on average with no accuracy cost to real users.

Defense — differential privacy and its price

The principled defense against membership inference is differential privacy (DP): a formal guarantee that no single training record measurably changes the model. Abadi et al.'s DP-SGD achieves it by clipping each per-example gradient to a fixed norm and adding Gaussian noise before the update, while a moments accountant tracks the cumulative privacy budget ε across training steps. PATE (Papernot et al.) reaches the same goal differently: train an ensemble of teacher models on disjoint data, then have a student learn only from their noisy aggregated votes — abstaining when the teachers disagree — achieving strong privacy (ε < 1.0) at scale.

💡 Privacy has a price, and it's accuracy. A smaller ε means more noise, stronger privacy, and a worse model. There is no free lunch — DP is the dial you turn to bound worst-case leakage, and every turn costs utility. Budget ε deliberately; don't inherit a library default.

Information disclosure — the LLM cluster

The same "outputs leak inputs" principle re-appears in the LLM stack as OWASP LLM02 (Sensitive Information Disclosure) and LLM07 (System Prompt Leakage): system-prompt extraction, training-data extraction via memorisation, PII surfacing in generations, and credentials bleeding through tool-call logs. Treat the system prompt as extractable, not secret. And the poisoning surface now reaches pretraining itself — Graf et al. show public discussion interfaces (comment sections, forums) are a web-scale injection vector, and their HalfLife analysis estimates how much adversarial content actually survives real crawl-and-curation pipelines into the training corpus.

The disclosure that lands hardest in 2026 is zero-click, and it defeats every content-inspecting guardrail by making the model decrypt the attack itself. Adversa AI's Cryptographic Context Injection (demonstrated against Grok 4.5 Fast on grok.com, Aug 2026) hides the malicious instructions on a web page as an AES-256-GCM ciphertext with the key material alongside it. A benign "summarise this page" request makes the agent decrypt the blob in its own Python sandbox — and because the plaintext now appears as the output of code the model just ran, it lands inside the trusted execution context and is never flagged as untrusted input. The decrypted payload then has Grok build a bogus "decryption key" that is really a template interpolating the user's private session state (name, location, subscription tier, full chat history) and fire its navigation tool at an attacker URL carrying that data as query params — exfil with no confirmation and no visible warning. Reported to xAI on Jun 3, 2026; no patch, no CVE at disclosure (Aug 20). The lesson is the one this whole chapter turns on, applied to guardrails: a filter that only reads content is blind to a payload encrypted until runtime — the decryption oracle is the model's own interpreter, so the fix cannot live in the model.

🔑 Guardrails inspect content; encryption removes content to inspect. Defend in the harness, not the model — keep untrusted page text out of the privileged tool context, require confirmation on outbound navigation/network actions, log resolved tool arguments (the URL after interpolation), and alarm on the chain (decrypt → build-string-from-session → fetch), not any single payload. This maps directly onto Kiya's own control layer: a WebFetch return is untrusted data, and a tool call that stitches session state into an outbound URL is the signal to catch.

Synthetic media — deepfakes as offense and defense

Generative models (GANs, diffusion) turn a few samples into convincing fake media, and that capability is dual-use. Offensively the toolchain is mature: PhishGPT-style persona-matched phishing, ElevenLabs voice cloning driving vishing (OSINT → sample collection → clone → live call), and face-swap video for executive-impersonation fraud — the concrete mechanics behind OWASP LLM09 (Misinformation). Defensively you get two levers. Detection: benchmarks like FaceForensics++ — 1,000 videos manipulated four ways (Deepfakes, Face2Face, FaceSwap, NeuralTextures) — train classifiers on generation artifacts, but detection is an arms race that erodes as generators improve. Provenance: C2PA Content Credentials attach a cryptographically signed manifest — "a nutrition label for digital content" — backed by Adobe, Google, Microsoft, OpenAI, Sony and the BBC.

🔑 The one rule to carry out of this week: the model is the attack surface — its gradients craft evasions, its confidence leaks membership, its API clones it, and its generations forge reality. Defend with the math, not with obscurity: adversarial-train against PGD, budget ε against inference, control the query surface against extraction, and sign your media against forgery.

🧪 This week's build — reproduce the four families

The math only sticks once you watch a gradient flip a label. Three toolchains, in increasing order of realism:

Feel the gradient first — the FGSM tutorial. The TensorFlow FGSM tutorial is the shortest path to the "aha." It loads pretrained MobileNetV2 (frozen), computes the loss gradient with respect to the input — not the weights — inside a tf.GradientTape that explicitly tape.watch(input_image)s the pixels, takes tf.sign(gradient), and applies adv_x = image + eps*perturbation clipped back into range. Sweeping ε ∈ {0, 0.01, 0.1, 0.15} makes the whole tradeoff visible in four frames: a bigger ε flips the label faster but the noise crosses from invisible to obvious — the L∞ threat model you read about, now on screen. That single create_adversarial_pattern function is the FGSM equation.

Scale to the full matrix — IBM ART. The ART repository's notebooks turn the four-family table into runnable code: it ships evasion, poisoning, extraction, and inference modules — a deliberate red-and-blue harness — across TensorFlow, Keras, PyTorch, scikit-learn, XGBoost and more, over images, tables, and audio. Use it to run PGD and DeepFool from the evasion table against your own classifier, then the AttributeInferenceBlackBox and membership-inference estimators from the same library — one API surface for every attack this chapter names.

Benchmark a defense honestly — the reference C&W. When you claim a model is robust, prove it against the attack that broke defensive distillation. Carlini's own nn_robust_attacks ships the official L₀/L₂/L∞ implementations from the C&W paper; the interface is deliberately blunt — wrap your model in a class exposing predict (logits, no softmax) plus image_size, num_channels, num_labels, then CarliniL2(sess, model).attack(inputs, targets). It exists precisely so a defense is measured against a strong, adaptive adversary rather than the weak one it was tuned to beat — the discipline's hardest lesson, turned into a script.

🔑 Build discipline: never report a robustness number from FGSM alone — it's the baseline attacker. Escalate to PGD (ART) and then the reference C&W before you call anything robust; a defense that only survives the cheap attack survives nothing.

🎯 OSAI exam depth — Attacking Embeddings (m6)

The four families above frame any model. This module drills the one the exam tests hands-on: the embedding is not anonymised data — a dense vector leaks almost everything the text did, and an attacker who touches the vector (in transit, in a vector DB, or via an API that returns it) can pull the plaintext, its author's attributes, and the model's training set back out. Treat every stored embedding as the raw record it was made from.

Embedding inversion — reconstruct the input text from its vector. The headline result: Morris et al. — Text Embeddings Reveal (Almost) As Much As Text recovers 92% of 32-token inputs exactly (BLEU 97.3) from a black-box embedding, including full patient names from clinical notes. Their Vec2Text method (code: jxmorris12/vec2text) is the technique to know for the 24h: you have a target vector e and query access to the same embedding model ϕ. Train an encoder-decoder to emit a hypothesis ê₀, re-embed it (ϕ(ê₀)), feed the model the target, the hypothesis embedding, and the difference vector ϕ(ê₀) − e, and let it correct — iterate ~20-50 steps and the reconstruction snaps onto the target. It's gradient-free — you never need the model's weights, only the ability to embed your own guesses. The predecessor, Song & Raghunathan — Information Leakage in Embedding Models (CCS '20), is the exam's conceptual spine: it names all three embedding attacks — inversion (recovering 50-70% of input words as a bag), attribute inference, and membership — and shows sentence embeddings even fingerprint their source model with 100% accuracy. For whole-sentence recovery, Li et al. — GEIA (Sentence Embedding Leaks More Information Than You Expect) trains a generative decoder (GPT-2-style) conditioned on the embedding to emit fluent, coherent sentences rather than a word-bag — the modern baseline alongside Vec2Text. Practical attacker note: inversion degrades on embeddings deliberately perturbed with Gaussian noise, so in a lab, invert clean stores first.

Attribute inference — pull a sensitive trait without recovering the text. Even when full inversion fails, an embedding still leaks properties of its author or subject (gender, race, authorship, medical status) that are irrelevant to the task the embedding was trained for (Song & Raghunathan). The attack is cheap: collect a small labelled auxiliary set, embed each item through the same ϕ, and train a shallow MLP probe mapping embedding → attribute. At inference you feed a stolen embedding and read off the private label. This is the go-to when the target embeds short or noisy text where word-level inversion is weak — you don't need the words, just the trait. IBM ART's AttributeInferenceBlackBox gives a ready harness.

Model inversion — reconstruct a representative training input from outputs. Distinct from inverting one embedding: here you rebuild a class prototype from a model's confidence scores. The classic Fredrikson, Jha & Ristenpart (2015) recovered recognisable faces from a facial-recognition API given only a target name and the returned confidences — hill-climb the input pixels to maximise the model's confidence for that identity. Modern black-box variants replace naive gradient-ascent with a GAN latent-space search: fix a generator trained on public faces, then optimise the latent z so the target classifier is maximally confident, yielding sharp reconstructions. This is the technique to reach for when the target returns per-class probabilities, not embeddings.

Training-data / sensitive-information extraction from LLMs. Generative models memorise and will regurgitate verbatim. Carlini et al. — Extracting Training Data from LLMs (USENIX '21) is the method: sample many completions from high-entropy prompts, then rank candidates by a membership signal — a high ratio of the target model's likelihood to a second model's (zlib entropy or a smaller LM) flags memorised, low-perplexity strings — and recover PII, keys, and license text from GPT-2. Against aligned chat models, Nasr, Carlini et al. — Scalable Extraction from Production LLMs (2023) adds the divergence attack: prompt the model to repeat a single word forever (the "poem poem poem" attack) until it breaks character and dumps pretraining text — lifting the extraction rate ~150× and pulling verbatim data from ChatGPT. In the exam, chain both: divergence to make it leak, membership-scoring to sift the memorised lines out of the noise.

The 2026 frontier — stealing the reasoning trace. Providers now hide a model's chain-of-thought, returning it to the client as an encrypted block passed back with each request rather than storing it server-side. Panfilov, Shumailov, Geiping, Andriushchenko et al. — Stealing Reasoning Traces from Proprietary LLM APIs (arXiv 2608.09867, Aug 2026) show these blocks are fully interchangeable across sessions, users, and models within one provider's ecosystem. Inject a strong model's encrypted trace into a weaker, less-safeguarded sibling and it decodes and prints the trace verbatim — no jailbreak of the capable model required. Four consequences: (1) anti-distillation bypass — harvest a frontier reasoner's private CoT to train a clone; (2) mass data extraction — 367 PII artifacts + 182 credentials recovered from 315,320 public reasoning blocks; (3) hidden-hazard disclosure — dangerous content the model reasoned through but scrubbed from its final answer is still in the trace; (4) invisible prompt injection — an attacker-authored encrypted block smuggles a payload into any public deployment that echoes traces. Demonstrated across Anthropic, OpenAI, and Google. The lesson generalises the extraction surface: the reasoning channel is now a disclosure surface, and cross-model reuse means a provider's weakest model sets its distillation-resistance floor. For the exam, treat any returned/echoed encrypted reasoning as attacker-controllable input, not opaque state.

This stopped being a lab result in July 2026. OpenAI disclosed a coordinated campaign — attributing a core cluster to individuals associated with Moonshot AI (maker of the Kimi chatbot) — that ran exactly the cross-session replay 2608.09867 described: it never broke the encryption, it copied one conversation's encrypted reasoning block and asked a model to decode it in another. The campaign began Jul 1, 2026, peaked Jul 24–25 with ~16,000 extraction requests in two days, spanned >15,000 accounts (4,000+ active at peak), and the stated goal was adversarial distillation — training a competitor "without the original safeguards." OpenAI closed the replay pathway by Jul 28. The takeaway confirms the 2026 thesis the chapter turns on: a mechanism assumed opaque (the encrypted CoT envelope) was replayable across sessions — bind to verified origin, don't trust the envelope.

And the defender's position is worse than a single patch suggests, because the attacker doesn't stop at the copy. Distillation Defenses Easily Break After Reinforcement Learning (arXiv 2609.35699) argues the anti-distillation literature evaluates defenses under an unrealistic threat model — assuming the attacker ceases training the moment it finishes copying. In reality a simple distillation attack using data obtainable from current commercial APIs, followed by RL, matches the reasoning gains of sophisticated full-trace extraction — so a defense that looks effective immediately post-distillation is deceptively robust. The structural conclusion: any defense that leaks enough to let an attacker approximately reconstruct the reasoning trace is ineffective, because RL closes the remaining gap; the authors point to batch-level defenses as the more promising direction than per-response trace-hiding.

🔑 Trace-hiding is not a boundary. The reasoning channel joins embeddings and confidence scores as a leaky surface: don't trust the encrypted-CoT envelope (bind replay to verified origin), and don't certify an anti-distillation defense on a post-distillation snapshot — the realistic attacker RL-polishes the stolen copy, so any defense that leaks an approximate trace loses. Defend at the batch/query-budget layer, not by obfuscating a single returned trace.

Attacking vector stores / RAG. In a RAG pipeline the whole knowledge base sits in a vector DB as embeddings — usually unencrypted and often assumed "just numbers." It isn't. Attack paths (survey, OWASP LLM08 Vector & Embedding Weaknesses): (1) Store breach → mass inversion — read the vectors from a misconfigured bucket or exposed endpoint, run Vec2Text/GEIA over the corpus, and you've reconstructed the source documents (functionally a plaintext breach — HIPAA/GDPR obligations follow the vectors). (2) API embedding probing — if the RAG endpoint returns embeddings or raw similarity scores, harvest embeddings for chosen probes and correlate to fingerprint what's indexed. (3) Cross-tenant leakage — weak partitioning in a shared store lets one tenant's query retrieve another's chunks. Compliance framing to carry into any report: storing PII/PHI as embeddings does not de-scope it — the vector is the record.

🧪 Drill (exam-shaped): stand up a small vector DB (Chroma/FAISS) with ~500 sensitive-looking notes embedded via a public sentence-encoder. (a) Dump the raw vectors and run vec2text to reconstruct the notes — measure token-recovery rate. (b) Train an MLP attribute probe on a labelled slice and infer a held-out trait. (c) Point Vec2Text at noised copies (add Gaussian δ) and chart how recovery collapses — that curve is your defense recommendation (perturbation vs. retrieval utility). Record which attack survives the noise (word-bag MLC is more robust than Vec2Text/GEIA) — that's the exam's "which defense, and its cost" question.

🛡️ OWASP 2026 — Sensitive Information Disclosure side channels

The privacy attacks above all assume the adversary reads the content — a returned embedding, a confidence score, a leaked completion. The 2026 revision of OWASP LLM02 (Sensitive Information Disclosure) adds a harder class: leaks where the attacker never receives the content at all and infers it from externally measurable properties. For the exam, memorise the four-phase disclosure model OWASP uses to organise the whole risk — disclosure happens at (1) training-time (a model, fine-tune, or LoRA adapter memorises corpus content and later reproduces it — memorisation scales log-linearly with capacity, duplication, and context length, and narrow adapters memorise rare examples with higher fidelity than the base model); (2) inference-time (live context — system prompt, RAG chunks, tool outputs, or another session's data — surfaced because summarise/translate/extract returns more than was asked, including visually-redacted spans); (3) pipeline-time (fine-tuning, distillation, synthetic-data generation, gradients, and observability carry sensitive data into derived artifacts); and (4) observation-time (the new one — inference from timing, token length, log-probabilities, confidence, and cache-hit signals). Severity turns on what the recipient can learn, not whether the leak looked like natural language.

Observation-time side channels — leaking without reading the bytes. Four mechanics, each with a concrete defense:

  • Whisper Leak — topic inference from encrypted traffic. Streaming responses emit a distinctive packet-size and inter-token-timing fingerprint that TLS does not hide. McDonald & Bar Or (2025) trained a classifier on that metadata and identified conversation topics at >98% AUPRC across 28 production models, with 100% precision flagging sensitive topics like "money laundering" — a network observer (ISP, employer, nation-state) learns you asked about a medical, legal, or political subject without decrypting a word. Defense: random response padding + token batching + packet injection — but the authors show each only reduces effectiveness; none fully closes the channel, so treat streaming on sensitive endpoints as observable and combine padding with topic-agnostic response shaping.
  • Token-length / remote-keylogging side channel. Because one token ≈ one packet, the sequence of response-chunk lengths leaks under TLS. Weiss et al. (2024) turned that into a remote keylogging attack on ChatGPT-4 and Microsoft Copilot: an LLM-driven decoder that used inter-sentence context and known-plaintext fine-tuning reconstructed 29% of responses verbatim and inferred the topic of 55%. Defense: batch tokens into fixed-size chunks and pad — the mitigation OpenAI and others shipped after disclosure.
  • KV-cache prompt leakage in multi-tenant serving. Providers share a prefix KV cache across users to speed inference; a co-tenant who probes for cache-hit timing can recover another tenant's prompt. Wu et al. (NDSS 2025) demonstrated this "I know what you asked" leak on shared serving. Defense: partition KV caches under co-tenancy and put high-sensitivity tenants on dedicated prefix caches — do not let cache state cross a trust boundary.
  • Internal-state / hidden-layer inversion. Depth does not anonymise: a leaked intermediate activation is invertible back to the prompt. Dong et al. (2025) reconstructed a 4,112-token medical prompt from a middle layer at F1 0.8688 — so exporting hidden states to a co-processor or logging them for observability is a disclosure surface, not a debugging convenience.
  • Logit-bias model theft. Carlini et al. (2024) recovered a production model's embedding-projection layer purely through the logit-bias API channel — a reminder that log-probabilities and logit-bias are themselves protected outputs. Defense: gate log-probabilities, confidence scores, and explanations on production endpoints.

Membership inference is now a breach determination, not a research curiosity. Fu et al. — SPV-MIA (NeurIPS 2024) raised membership-inference AUC to ~0.9 against fine-tuned LLMs using self-prompt calibration to build the reference distribution the attack needs. 0.9 AUC about a named individual is enough to ground a regulatory breach determination — and because extraction, membership inference, and inversion run offline at unbounded rates on open-weights models, rate limits alone cannot contain them. Defense: DP-SGD calibrated to sensitivity plus per-user / per-session query budgets and query-pattern detection (fixed-ε DP degrades under adaptive querying), with verifiable erasure validated by post-unlearning membership probes across raw data, embeddings, checkpoints, and adapters.

Embedding inversion is where LLM02 owns the regulatory consequence. The mechanics of Vec2Text/GEIA live in this chapter's m6 module and in LLM09, but OWASP 2026 assigns LLM02 the compliance verdict: because modern inversion reconstructs plaintext from exported vectors, an "embeddings-only" backup is a source-document breach — storing PII/PHI as embeddings does not de-scope it, cosine similarity does not respect ACLs, and clinical-embedding stores sit in HIPAA audit-control scope. Defense: authorize before retrieval (enforce chunk-level authz inside the index query, because post-generation filtering cannot un-supply a retrieved chunk), give the vector store its own ACLs separate from the document ACLs, restrict export APIs, and treat every stored vector as the raw record it was made from.

🔑 Exam takeaway: LLM02:2026 widens "disclosure" from the answer text to any observable property — timing, token length, log-probs, confidence, cache hits — across four phases (training / inference / pipeline / observation). Map each channel to its defense: streaming → padding + batching; multi-tenant KV cache → partitioned / dedicated prefix caches; log-probs → gate them; membership → DP + per-user budgets + verifiable erasure; embeddings → authorize-before-retrieve + vector-store ACLs. And carry the one-line compliance rule: the vector is the record; an embeddings-only backup is a breach.

📇 Attack & tool reference

The lesson above is what to learn; this is the catalog behind it. Folded by default — expand when you need a specific attack or toolkit.

Read the catalog through the four attack families

Every item below belongs to one of the four families in the big-picture table; the defense follows from the family, not the individual attack.

Family Representative attacks Toolkit to reproduce it
Evasion FGSM, I-FGSM, PGD, DeepFool, C&W, JSMA, EAD, GoodWords ART · Foolbox · CleverHans
Inference membership inference (shadow models), attribute inference ART inference module
Extraction model stealing, surrogate training ART extraction module
Synthetic media GAN/diffusion generation, face-swap, voice clone FaceForensics++ (detection) · C2PA (provenance)
Evasion attacks
  • FGSM — single-step L∞ perturbation along the gradient sign; fast, detectable, and the origin of the transferability result. Goodfellow et al. (2014)
  • PGD / I-FGSM — iterative FGSM with ε-ball projection; the canonical strongest first-order adversary and the target of adversarial training. Madry et al. (2017)
  • C&W — optimisation-based L₀/L₂/L∞ attack; broke defensive distillation to 100% and set the benchmark-your-defense standard. Carlini & Wagner (2017)
  • DeepFool — minimal-perturbation, boundary-normal attack; doubles as a robustness metric. Moosavi-Dezfooli et al. (2016)
  • JSMA / EAD / FISTA / GoodWords — sparse (L₀/L₁) and text-substitution variants; reference implementations in CleverHans and ART.
Embedding attacks (OSAI m6)
  • Vec2Text (embedding inversion) — iterative re-embed-and-correct; recovers 92% of 32-token inputs exactly from a black-box embedding. Morris et al. (2023) · code
  • GEIA — generative decoder recovers whole fluent sentences from a sentence embedding. Li et al. (2023)
  • Embedding inversion / attribute / membership triad — the seminal formalisation; +100% source-model fingerprinting. Song & Raghunathan (2020)
  • Model inversion (confidence-based) — reconstruct recognisable faces from a name + API confidences; GAN-latent variants for sharper output. Fredrikson, Jha & Ristenpart (2015)
  • LLM training-data extraction — membership-scored sampling (Carlini et al. 2021) + the repeat-a-word divergence attack on aligned chat models (Nasr, Carlini et al. 2023).
Privacy attacks & defenses
  • Membership inference (shadow models) — black-box, overfitting-driven. Shokri et al. (2017)
  • Knowledge Trap — extraction defense via honeypot knowledge graph; −6.2% surrogate agreement. arXiv 2606.15810
  • DP-SGD — gradient clipping + Gaussian noise + moments accountant. Abadi et al. (2016)
  • PATE — teacher-ensemble noisy aggregation; ε < 1.0 at scale. Papernot et al. (2018)
  • Pretraining-data poisoning / HalfLife — computational-propaganda injection through public discussion interfaces. Graf et al. (2026)
  • Reasoning-trace extraction (the 2026 frontier) — encrypted CoT blocks are cross-session/cross-model replayable (Panfilov, Shumailov et al. 2608.09867); the named in-wild case is OpenAI's July-2026 Moonshot-linked replay campaign (16k requests, patched Jul-28) (Decrypt); and anti-distillation defenses that look robust post-distillation break once the attacker RL-polishes the copy — any defense leaking an approximate trace is ineffective (arXiv 2609.35699).
Toolkits & benchmarks
  • IBM ART — 50+ evasion attacks plus poisoning/extraction/inference modules across PyTorch, TensorFlow, Keras, scikit-learn. ART docs
  • Foolbox — framework-native (PyTorch/TensorFlow/JAX via EagerPy) white- and black-box attacks. Foolbox docs
  • CleverHans — reference FGSM/PGD/JSMA implementations for robustness benchmarking (JAX/PyTorch/TF2). CleverHans
  • FaceForensics++ — deepfake-detection benchmark, 1,000 videos, four manipulation methods. FaceForensics++
  • C2PA — cryptographic content-provenance standard. c2pa.org

Recommended resources0/40

Sign in to tick items off and track your progress.

Show

📖 Core Path

📚 Further Reading

Core AML Papers
Talks & Strategic Perspectives
Hands-on AML Tools
  • 🔧 IBM ART GitHub + Notebooks — Jupyter notebooks with working examples across attack categories
  • 🧪 TensorFlow FGSM Tutorial — Step-by-step adversarial example generation with TensorFlow
  • 🔧 Foolbox Documentation — Framework-native (PyTorch/TensorFlow/JAX via EagerPy) white- and black-box attacks; clean unified API
  • 🔧 C&W Attack Implementation — Official TensorFlow implementations of L0, L2, Linf attacks
  • 🔧 CleverHans — Reference FGSM/PGD/JSMA implementations for robustness benchmarking (JAX/PyTorch/TF2)
Attacking Embeddings (OSAI m6)
Privacy Attacks & Defenses
  • 📄 Knowledge Trap — arXiv 2606.15810 (Jun 2026) — Extraction defense that plants a Honeypot Knowledge Graph + breadcrumb-guided exploration to burn the attacker's query budget on low-transferability knowledge; cuts surrogate Agreement 6.2% avg with no legit-user accuracy hit. Reframes extraction defense as knowledge-space traversal control (~30 min)
  • 📄 Pretraining Data Can Be Poisoned through Computational Propaganda — arXiv 2607.15267 (Jul 2026) — Graf/Hajishirzi/Smith/Kohlbrenner/Lo (AI2/UW). Public discussion interfaces are a web-scale content-injection vector that survives real crawl + curation; introduces HalfLife to estimate how much adversarial content reaches the training corpus. Training-time supply-chain threat to pretraining itself (~30 min)
  • 📄 PATE — Papernot et al. (2018) — Private Aggregation of Teacher Ensembles; noisy-vote knowledge transfer to a student, ε < 1.0 at scale; alternative to DP-SGD
OWASP 2026 — LLM02 Sensitive Information Disclosure (side channels)
Deepfakes & Synthetic Media
  • 🔧 FaceForensics++ — Deepfake-detection benchmark: 1,000 videos manipulated four ways (Deepfakes, Face2Face, FaceSwap, NeuralTextures)
  • 📄 C2PA (Content Provenance) — Cryptographically signed content-authenticity manifests ("a nutrition label for digital content"); backed by Adobe, Google, Microsoft, OpenAI, Sony, BBC
  • 🎥 Two Minute Papers — Deepfake Detection — Visual overview of detection methods

📡 From the Resources feed

  • 📄 TRACE — Temporal Gradient Inversion for Private Trajectory Reconstruction in Embodied RL (arXiv 2609.30258) — autoregressively reconstructs private observation-action trajectories from per-step policy-learning gradients in distributed embodied RL by exploiting temporal correlation between successive gradients, beating prior single-frame gradient-inversion attacks — extends the gradient-leakage privacy threat (cf. Shokri membership inference above) into the federated/distributed-training setting; NeurIPS 2026 (in Trove since 2026-09-26 (security/ai-security)) 📡
  • 📄 Distillation Defenses Easily Break After Reinforcement Learning (arXiv 2609.35699) — argues existing anti-distillation defenses for closed-source LLMs are evaluated under an unrealistic threat model (no further training after distillation). A simple distillation attack using data obtainable from current commercial APIs, followed by RL, matches the reasoning gains of sophisticated full-trace-extraction attacks — breaking defenses that looked effective immediately post-distillation. Conclusion: any defense that leaks enough to approximately reconstruct reasoning traces is ineffective; batch-level defenses are the more promising direction. Extends the 2026 cross-model-extraction frontier (cf. reasoning-trace theft 2608.09867) — the attacker doesn't stop at the copy, they RL-polish it. [Sep-30 daily-pulse]
  • 📰 OpenAI: Moonshot-linked campaign replayed encrypted chain-of-thought to distill it (Decrypt) — the named real-world incident behind the reasoning-trace-theft research line above (2608.09867, single-provider-key finding). OpenAI disclosed a campaign starting Jul-1-2026, peaking Jul-24/25 with 16,000 extraction requests from 4,000+ users (broader cluster >15,000 accounts), that did not break encryption but manipulated interactions to replay one conversation's encrypted CoT and decode it in another → adversarial distillation of a rival's reasoning without its safety guardrails. OpenAI attributes a core cluster to individuals associated with Moonshot AI (Kimi) and patched the replay pathway by Jul-28. Confirms the 2026 thesis: the mechanism assumed opaque (encrypted CoT block) is replayable across sessions — bind to verified origin, don't trust the envelope. [Oct-02 daily-pulse]
  • 📄 Do Defenses Against LLM Extraction Work Across Attacks? A Lifecycle Benchmark of Black-Box Model Extraction (arXiv 2610.00839) — a unified benchmark pitting six model-extraction attacks against ten defenses (plus two adaptive paraphrase/back-translation evasions) under matched query budgets — standardized evidence for which extraction defenses actually hold across attack types (in Trove since 2026-10-03 (security/ai-security)) 📡
  • 📄 Guess My Weight: Profiled Side-Channel Recovery of Floating-Point Neural-Network Weights (arXiv 2610.04436) — a profiled power-analysis template attack that recovers exact 32-bit float neural-network weights from an ARM Cortex-M4 at 99% bit-exact in ~171 traces (100% at 263), via coarse-to-fine search over IEEE-754 space — a physical side-channel model-theft vector beyond query-based extraction (in Trove since 2026-10-06 (security/ai-security)) 📡

Study checklist

↪ See roadmap.md → Phase 2 → Week 11

  • Map the four attack families (evasion, extraction, membership inference, inversion) to their defenses
  • Derive FGSM from a loss gradient; explain the linearity hypothesis and transferability
  • Run PGD (iterative) and DeepFool (minimal-δ) evasion attacks with ART or Foolbox
  • Reproduce a C&W attack and explain why it is the defense-evaluation benchmark
  • Choose an evasion attack by Lᵖ norm (L0/L2/L∞); note JSMA (saliency) and EAD (elastic-net) sparse variants
  • Build path: FGSM on MobileNetV2 (TF tutorial, ε-sweep) → PGD/DeepFool + inference modules in ART → benchmark a defense against the reference C&W; never report robustness from FGSM alone
  • Perform a membership-inference attack via shadow-model methodology
  • Understand model extraction and inversion; study Knowledge Trap query-budget defense
  • Treat echoed/encrypted reasoning traces as attacker-controllable: study cross-model trace reuse (anti-distillation bypass + PII/cred extraction)
  • Trace the OpenAI/Moonshot July-2026 CoT-replay campaign (named in-wild case); explain why RL-after-distillation breaks post-distillation defense evals → defend at batch/query-budget, not by hiding one trace
  • Understand DP-SGD (clipping + noise + moments accountant) and the ε privacy–utility tradeoff
  • Study PATE (teacher-ensemble noisy aggregation) as an alternative privacy defense
  • Cluster the LLM info-disclosure surface: LLM07 prompt leakage, LLM02 PII/data extraction, pretraining poisoning (HalfLife)
  • Trace Cryptographic Context Injection (Grok): encrypted payload → model-runtime decryption oracle → session-state exfil via URL params; defend in the harness, not the model
  • Study synthetic-media generation (GANs vs diffusion) and detection via FaceForensics++
  • Study the offensive deepfake toolchain: PhishGPT, ElevenLabs voice cloning, vishing, video impersonation
  • Assess AI-disinformation risk (LLM09) and C2PA provenance as a defense

Study notes

Sign in to take notes.