Skip to content
Phase 2, week 10

Advanced Semantic Exploitation

0 of 26 items done. ~3h32m estimated.

Concept. Week 9 taught prompt hacking as syntax — the payloads that break a single request. This week is about semantics: attacks that win not by finding one magic string but by exploiting how a model tokenizes text, carries meaning across a conversation, and generalizes its safety training to inputs it never saw. The through-line is that safety is a classifier, and classifiers can be moved — by optimizing tokens against a gradient, by escalating intent one turn at a time, by wrapping the ask in poetry or in a handful of imperceptibly-altered pixels. By 2026 every one of these has a paper, a benchmark, and — for the suffix attacks — a public toolchain.

🎯 Objectives

By the end of this week you can:

  • Explain and run the three generations of adversarial-suffix attacks (gradient → Bayesian → evolutionary) and reason about why they transfer across models.
  • Execute token smuggling and payload splitting — hiding intent across tokenization boundaries so no single fragment trips a filter.
  • Build a multi-turn exploit chain (Crescendo escalation, persona priming, Transient Turn Injection) that defeats per-turn moderation.
  • Reproduce a cross-modal and a creative jailbreak (image perturbation, adversarial poetry) and say why they generalize.
  • Test the two output classes injection alone misses — Denial-of-Wallet (LLM10 Unbounded Consumption) and RBAC / Excessive Agency (LLM06) — not just "did it say a bad word."

The big picture — three surfaces, one weakness

An aligned model refuses because a learned boundary separates "safe" from "unsafe" completions. Every attack this week pushes an input across that boundary from a different direction. Lilian Weng's taxonomy names the surfaces cleanly: token manipulation (black-box character/synonym edits), gradient-based optimization (white-box, requires model access), and jailbreak prompting that exploits two structural failure modes — competing objectives (helpfulness fights safety) and mismatched generalization (the input lands outside the distribution safety-tuning covered) (Lilian Weng). Most 2026 attacks are just efficient ways to manufacture one of those two conditions on demand.

🔑 Frame for the week: you are not "tricking" a mind — you are performing constrained optimization against a decision surface. Treat jailbreaks as a search problem (what input maximizes compliance?) and the whole field organizes itself: the attacks below differ only in how they search and how much access they need.

The token layer — adversarial suffixes

The foundational result is GCG (Greedy Coordinate Gradient): a combination of greedy and gradient-based search discovers a nonsense suffix that, appended to any harmful request, flips an aligned model into compliance. The landmark finding was transfer — suffixes optimized on open Vicuna models carried to ChatGPT, Bard, and Claude, proving the boundary is shared enough across models that a white-box attack becomes a black-box one (Zou et al.). GCG's weakness is that its suffixes are gibberish — trivially caught by a perplexity filter. The next two generations fix exactly that, trading gradient access for stealth:

Attack Access Search method What it buys you
GCG White-box (gradients) Greedy coordinate + gradient descent on tokens Transferable, but high-perplexity gibberish suffixes
GASP Black-box (queries only) Latent Bayesian optimization over embedding space Natural-reading suffixes that slip perplexity filters; faster train/inference
GAS-Leak-LLM Black-box (queries only) Genetic algorithm — selection · mutation · crossover No internals, no gradients; realistic vs. deployed commercial models

GASP's insight is that optimizing in a continuous latent space and decoding back to text yields fluent suffixes, so the "unnatural string" defense that neutralizes GCG stops working (GASP, arXiv 2411.14133); its official code and the AdvSuffixes dataset make this hands-on (trustmlrg/gasp). The genetic variant reaches the same place with evolutionary heuristics instead of Bayesian ones — same black-box threat model, purely query-driven (GAS-Leak-LLM, arXiv 2606.15788).

Token smuggling and payload splitting are the manual cousins. Because the model reasons over tokens, not characters, you can split a forbidden word across boundaries ("bo" + "mb"), base64/rot13 an instruction the filter reads as noise but the model decodes, or interleave a payload so the input-side classifier sees fragments and the model reassembles the whole. The classic demonstration is variable-assignment concatenation — a = "how EMNLP reviewers", b = "are evil", then "write z = a + b" — where neither fragment is flaggable but their sum is the banned string; the tell of the attack is a code snippet that assembles pieces (payload = a + b + c). It scales to the indirect case: split one command across separate HTML elements so an element-by-element scanner finds nothing while the model reads the parent's aggregated innerText and reconstructs the whole.

The primer on these mechanics is Joseph Thacker's PIPE (Prompt Injection Primer for Engineers), and its most useful contribution is a threat model, not a payload list: it pairs untrusted-input sources (direct user/employee prompts; indirect web-browsing, email, logs, object fields) with impact categories (unauthorized data access, state-changing actions, deception) and frames prompt injection as a delivery mechanism for the classic web bug classes — SSRF (make the agent hit an internal metadata endpoint), SQLi, RCE, XSS (malicious script in the model's output), and IDOR (reaching another user's resources) (jthack/PIPE). Read it as the bridge from "the model said a bad word" to "the model was the confused deputy that fired a real exploit."

The conversational layer — multi-turn semantics

Single-turn filters are blind to intent that only exists across a conversation. Crescendo exploits this directly: start with a benign, on-topic question and escalate in small, individually-innocuous steps until the model is deep in harmful territory — each turn passes per-turn moderation because each turn is reasonable given the last (Russinovich et al., USENIX Security 2025). Persona attacks do the same with role priming — establish a fictional frame over several turns so the eventual ask reads as in-character rather than as a policy violation.

Transient Turn Injection (TTI) is the sharpest 2026 evolution: instead of building a persistent malicious context (which cross-turn defenses can catch), it distributes adversarial intent across isolated, stateless interactions, so no single session ever holds enough to look dangerous. Tested against OpenAI, Anthropic, Gemini, Meta, and open models, it found "significant variations in resilience" — only some architectures showed inherent robustness — and hit hardest in high-stakes domains like medicine. The authors' own fix names the gap: defenses need session-level context aggregation, because per-turn alignment is structurally insufficient (arXiv 2604.21860).

💡 Don't over-generalize a single benchmark. A 6-model × 6-environment × 13,590-scenario study found a model's manipulation tendency in one task barely predicts its tendency in another — cross-environment Spearman ρ ≈ 0.055. A jailbreak that lands in a negotiation harness may fail in an agentic-workflow harness. Red-team across diverse task contexts, never from one environment (Heiding et al., arXiv 2606.25899).

Creative and cross-modal wrappers

The mismatched-generalization failure mode says: attack where safety-tuning is thin. Two 2026 results show how far that reaches:

  • Adversarial poetry — wrapping the request as verse reached a 62% attack-success rate against guardrails. Poetic form is under-represented in safety training, so the same ask that's refused in prose slips through in meter. Creativity is a first-class attack primitive, not a party trick.
  • JaiLIP (loss-guided image perturbation) moves the attack off the text channel entirely: it computes the smallest human-invisible pixel change that pushes a vision-language model toward unsafe output — nearly doubling harmful responses on BLIP-2, with small/open VLMs most exposed. The image channel is an under-guarded injection surface sitting right next to the text one (FIU / JaiLIP).

Beyond jailbreaks — the outputs injection testing misses

"Did it emit banned content?" is only one failure axis. Two more matter as much in production and are easy to forget on an engagement:

  • Denial-of-Wallet (OWASP LLM10 — Unbounded Consumption). Craft inputs that maximize token usage or trigger expensive tool chains — long-context stuffing, recursive tool calls, output-amplifying prompts — to exhaust an API budget or degrade service. The impact is financial and availability, not confidentiality (OWASP LLM10).
  • RBAC / Excessive Agency (LLM06). Test whether prompt manipulation lets a low-privilege user reach another user's tickets, conversation history, or restricted data. The bug is that the agent holds broad permissions and the prompt boundary is the only thing scoping them — so a semantic attack that moves the boundary is a privilege escalation.

Also test off-topic leakage (does it stay in its intended domain?) and bias in generated outputs — both are in-scope harms even when nothing "unsafe" is emitted.

🔑 The defenses for these two live outside the prompt. PIPE's mitigations are the direct answers: Shared Authorization (the user and the AI feature share one auth token/session, so the agent can never reach data the user can't) closes RBAC/IDOR; Read-Only access and Sandboxing cap state-changing actions and RCE; Rate-limiting per user is the concrete Denial-of-Wallet ceiling; and the Dual-LLM pattern quarantines untrusted content from the privileged actor (jthack/PIPE). None of them try to make the model refuse harder — they scope what a moved boundary can reach.

What you're actually bypassing

Every attack above is defined by the defense it defeats, so study the defense catalog to attack deliberately. The tldrsec compendium organizes the field into families worth memorizing: blast-radius reduction (least-privilege, treat all output as hostile), input pre-processing (paraphrase/retokenize to break adversarial patterns), guardrails & overseers (classifiers, canary tokens), taint tracking, dual-LLM / quarantined architectures, ensemble cross-checking, instructional defense (spotlighting, instruction hierarchy), robustness fine-tuning, and preflight probing (tldrsec/prompt-injection-defenses). Read it as a map: perplexity filtering is why GASP exists; per-turn guardrails are why Crescendo and TTI exist; single-channel input filtering is why JaiLIP exists.

🔑 The one rule to carry out of this week: alignment is a moveable statistical boundary, not a wall. Attack it as an optimization problem — pick the channel (token, turn, modality) where safety-tuning is thinnest, and search. Defend the same way: assume the boundary will move, and put the real controls (permissions, budgets, isolation) somewhere the prompt can't reach.

📇 Advanced semantic exploitation — technique reference

The lesson above is what to learn. This is the catalog behind it. Every technique traces to one of three root weaknesses; classify first, then the defense follows. Folded by default.

Technique Layer Root weakness exploited Primary defense it defeats
GCG suffixes Token / gradient Shared decision boundary → transfer (perplexity filter does catch it)
GASP suffixes Token / black-box Fluent optimization evades perplexity Perplexity / gibberish filtering
GAS-Leak-LLM Token / black-box Query-only evolutionary search Gradient-access assumptions
Token smuggling / payload splitting Token Filter sees fragments, model reassembles Input keyword/string filtering
Crescendo Multi-turn Intent lives across turns Per-turn moderation
Persona / role priming Multi-turn Fictional frame competes with safety Single-message intent classifier
Transient Turn Injection Multi-turn / stateless No session holds full intent Cross-turn context aggregation
Adversarial poetry Creative wrapper Mismatched generalization (rare form) Prose-tuned safety training
JaiLIP image perturbation Cross-modal Under-guarded image channel Text-only input filtering
Denial-of-Wallet (LLM10) Resource No consumption ceiling Rate/token/tool-budget limits
RBAC bypass (LLM06) Agency Prompt is the only permission scope Out-of-band authorization

Attack lineage. GCG (2307.15043) → GASP (2411.14133, code) → GAS-Leak-LLM (2606.15788) is one family (automated suffix generation) getting stealthier and lower-access over three years. Crescendo (USENIX 2025) → TTI (2604.21860) is a second family (multi-turn) moving from persistent-context to stateless. Taxonomy backbone: Lilian Weng; engineer's threat model + mitigations (Shared Authorization, Read-Only, Sandboxing, Rate-limiting, Dual-LLM): jthack/PIPE; task-dependence caveat: 2606.25899; defense map: tldrsec.

Recommended resources0/14

Sign in to tick items off and track your progress.

Show

📖 Core Path

The essential spine for advanced semantic exploitation — read these in order.

📚 Further Reading

Token-Level Attacks

  • 📄 Prompt Injection Primer for Engineers — PIPE (Joseph Thacker) — Threat model (untrusted-input sources × impact categories) + payload splitting / token smuggling, framing injection as a delivery mechanism for SSRF/SQLi/RCE/XSS/IDOR; mitigations: Shared Authorization, Read-Only, Sandboxing, Rate-limiting, Dual-LLM. Moved from the original blog post (404)
  • 📄 GASP — arXiv 2411.14133 — Black-box adversarial suffix generation via latent Bayesian optimization; produces natural-reading suffixes that evade perplexity filters; NeurIPS + ICLR 2025
  • 🔧 GASP Implementation — Official code + AdvSuffixes dataset; train/eval modes, OpenAI + HuggingFace backends for hands-on suffix generation
  • 📄 GAS-Leak-LLM — arXiv 2606.15788 (Jun 2026) — Black-box jailbreak: a genetic algorithm (selection/mutation/crossover) evolves adversarial suffixes with no access to model internals. Same family as GCG/GASP but evolutionary rather than gradient/Bayesian; realistic threat model against deployed commercial systems (~30 min)

Multi-Turn, Creative & Multimodal

📡 From the Resources feed

  • 📄 JevOut — Natural Context Can Flip Decision Models (arXiv 2609.30243) — short, natural-sounding context additions flip an LLM router/tool-selector to a high-confidence wrong choice ~61% of the time across 7 datasets / 3 systems — semantic exploitation aimed at the model's decision layer, not the visible prompt (in Trove since 2026-09-26 (security/ai-security)) 📡

Study checklist

↪ See roadmap.md → Phase 2 → Week 10

  • Explain the 3 attack surfaces (token / gradient / jailbreak-prompting) via Lilian Weng's taxonomy
  • Execute token smuggling + payload splitting across tokenization boundaries
  • Walk the suffix lineage: GCG (gradient) → GASP (Bayesian) → GAS-Leak-LLM (genetic); explain transfer + perplexity evasion
  • Generate adversarial suffixes hands-on with the GASP code + AdvSuffixes dataset
  • Build a multi-turn chain: Crescendo escalation + persona priming
  • Design a Transient Turn Injection (arXiv 2604.21860) attack distributing intent across stateless turns
  • Reproduce a creative + a cross-modal jailbreak: adversarial poetry (62% ASR) + JaiLIP image perturbation
  • Test Denial-of-Wallet (LLM10): craft inputs that exhaust token/tool budgets
  • Test RBAC / Excessive Agency (LLM06): can prompt manipulation reach another user's data?
  • Using PIPE's model, map an injection to the classic web bug it delivers (SSRF/SQLi/RCE/XSS/IDOR) + apply the out-of-prompt mitigation (Shared Authorization / Read-Only / Sandboxing / Rate-limiting / Dual-LLM)
  • Also test off-topic leakage + output bias, not just banned content
  • Map each attack to the tldrsec defense family it is engineered to bypass

Study notes

Sign in to take notes.