Concept. A transformer is a next-token predictor built almost entirely out of one operation — attention — and everything that makes it powerful also makes it exploitable. Attention weighs every input token against every other, which is exactly why an instruction buried in retrieved text or a tool description gets the same consideration as your system prompt. This week you learn the architecture from the inside so the attacks in Phase 2 read as consequences, not surprises: prompt injection is attention doing its job, token smuggling is tokenization doing its job, and weaponized model variants are fine-tuning doing its job.
🎯 Objectives
By the end of this week you can:
- Trace the NLP lineage — word2vec embeddings → RNN/LSTM → transformer — and explain why self-attention replaced recurrence (Illustrated Word2Vec, Colah on LSTMs).
- Walk the query/key/value mechanism, multi-head attention, and positional encoding, and name the data/instruction conflation that attention creates (Illustrated Transformer).
- Explain how BPE / byte-level BPE tokenization maps text to tokens, and how token boundaries enable token smuggling past naive filters (HF Tokenizers).
- Reason about temperature / top-k / top-p and why non-determinism matters for both testing and probing (Raschka).
- Explain how LoRA/PEFT makes weaponized model variants economically feasible, and cite the evidence that a handful of examples unaligns a safety-tuned model (Qi et al. 2023).
- Build a basic chatbot against a transformer API to feel the inference loop end-to-end.
From word vectors to attention
Before transformers, meaning was already geometric. word2vec learns to
place each word at a point in a dense vector space by predicting a word from
its neighbours (CBOW) or its neighbours from the word (skip-gram), so that
proximity encodes semantics and vector arithmetic captures relationships —
the famous king − man + woman ≈ queen (Illustrated
Word2Vec). That embedding
space is not a historical footnote: it is the same space RAG retrieval
(Week 4) searches and the same space adversarial perturbations (Week 12)
navigate.
The problem was sequence. RNNs loop hidden state forward token by token, but as the gap between related words grows they "become unable to learn to connect the information" — the vanishing-gradient problem. LSTMs patched this with a cell-state "conveyor belt" gated by forget / input / output gates, buying longer memory (Colah on LSTMs). But both remained strictly sequential — token n waits for token n−1 — which caps parallelism and still strains on long dependencies.
🔑 The transformer's one big idea: drop recurrence entirely. Because every token attends to every other in parallel, training is far more parallelizable and long-range dependencies are one hop away, not n hops (Attention Is All You Need).
Inside the transformer
Self-attention lets each token gather context from the whole sequence. For every token the model derives three vectors — Query, Key, Value — and computes, for each pair, a score as the dot product of one token's Query with another's Key, scaled and softmax-normalized; the output is the Value vectors summed by those weights (Illustrated Transformer). When the model encodes "it" in "the animal didn't cross the street because it was tired," attention is what lets "it" bind to "animal."
Two refinements matter for security reasoning:
- Multi-head attention runs several attention mechanisms in parallel (eight in the original), each with its own Q/K/V projections, so different heads specialize in different relationships before being concatenated (Illustrated Transformer).
- Positional encoding is added to token embeddings because bare attention is order-blind — it has no inherent notion of sequence, so position must be injected via sinusoidal patterns (Vaswani et al.).
If you want to see this rather than read it, the Transformer Explainer (Georgia Tech) runs a live GPT-2-small in the browser: type a prompt and hover a token to watch which earlier tokens it attends to across all 12 heads, then trace the embedding → Q/K/V → MLP → softmax path end-to-end. It exposes the same three surfaces this chapter names as security-relevant — the token split at the input, the attention weights in the middle, and a temperature slider on the output distribution — in one view.
💡 The security reading of attention. Attention does not distinguish trusted instruction from untrusted data — both are just tokens competing for weight. There is no architectural "instruction channel." That conflation is the mechanism behind prompt injection: text retrieved from a webpage, an email, or a tool output can out-compete your system prompt simply by being phrased as a stronger instruction. Every defense in Phase 3 (spotlighting, delimiters, input sanitization) is an attempt to re-impose a boundary the architecture never had.
Tokenization: where text becomes tokens
Models don't see characters or words — they see tokens. Subword
tokenizers keep the vocabulary compact while still representing unseen
words: common words stay whole, rare ones decompose (annoyingly →
["annoying", "ly"]) (HF
Tokenizers).
The dominant algorithms differ in how they build the vocabulary:
| Algorithm | Merge rule | Used by | Note |
|---|---|---|---|
| BPE | merge the most frequent adjacent pair, iteratively | GPT, Llama, Qwen | deterministic merge list |
| Byte-level BPE | BPE over 256 byte values, not Unicode chars | GPT-2, RoBERTa | no <unk> — every byte is representable |
| WordPiece | merge the pair that most raises training likelihood | BERT family | favors informative merges, not just frequent |
| Unigram | prune a large vocab by loss contribution; probabilistic | T5, Pegasus | can sample alternate tokenizations |
| SentencePiece | BPE/Unigram on raw bytes incl. space (▁) |
multilingual | works for space-free languages |
The security angle: tokenization is a filter-evasion surface. Because a
string can split into tokens more than one way, an attacker can fracture a
blocked word across token boundaries — token smuggling — so a naive
substring filter never sees the forbidden term while the model still
reconstructs its meaning. Standard (non-byte-level) BPE also collapses
out-of-vocabulary characters into <unk>, and homoglyph / character
substitution can produce "ambiguous tokenization patterns" that bypass
content filters (HF LLM
Course). Byte-level
BPE removes the <unk> hole but not the multiple-encodings problem — which
is why you'll spend time in the OpenAI
Tokenizer watching payloads split.
🔑 Rule of thumb: any input filter that operates on characters is blind to the model's token view. Filter on normalized, canonicalized text and validate at the token level, or the boundary you think you drew isn't the one the model sees.
Sampling: temperature, top-k, top-p
Given the model's probability distribution over the next token, the decoding strategy decides what actually comes out — and it is a knob you must understand to test AI systems honestly. Three parameters shape it (Raschka, HF blog):
- Temperature reshapes the distribution before sampling:
0is greedy/deterministic (same output every time), higher values flatten it toward randomness (Cohere). - Top-k keeps only the k highest-probability candidates — a fixed-size pool — and discards the rest.
- Top-p (nucleus) keeps the smallest set whose cumulative probability reaches
p(e.g.0.9) — an adaptive pool that shrinks when the model is confident and grows when it isn't.
💡 Why this matters for security. Non-determinism is a double-edged evaluator. A jailbreak that fails at temperature
0may succeed at1.0on one of many samples — so red-team results must report the sampling config, and a single "it refused" run proves nothing. Conversely, determinism probing (Week 9) exploits temperature-0reproducibility to reverse-engineer behavior. Greedy and beam search are consistent but repetitive; sampling is diverse but can wander incoherent.
Fine-tuning as attack surface
Full retraining of a 70B model is out of reach for most attackers.
PEFT — parameter-efficient fine-tuning — changes that by freezing the
base model and training only a tiny set of extra parameters, "significantly
decreasing computational and storage costs" while matching full fine-tuning
quality (HF PEFT). LoRA is the
canonical method: it decomposes each weight update into two low-rank
matrices ΔW = B·A injected into the attention projections, so the forward
pass becomes y = Wx + BAx. Trainable parameters drop by roughly two orders
of magnitude and the resulting adapter is a few MB, not GB (HF LoRA
guide).
That economics is the threat. A cheap, shareable adapter means:
- Weaponized variants are feasible. An attacker can produce and distribute a malicious LoRA the way you'd share a config file.
- Adapters are supply-chain payloads. A LoRA pulled from an untrusted hub can carry a backdoor; you merge it and inherit its behavior (Phase 2, Week 16).
- Alignment is fragile under fine-tuning. Qi et al. jailbroke GPT-3.5 Turbo's safety guardrails by fine-tuning on only ~10 adversarial examples for under
$0.20— and even benign fine-tuning "can inadvertently degrade the safety alignment." Safety does not survive customization (Qi et al. 2023).
🔑 Carry this into Phase 2: the same three mechanisms you learned as architecture are the same three you'll exploit as attack surface — attention (prompt injection), tokenization (smuggling), and fine-tuning (unalignment + backdoored adapters). Understand them here and the offensive weeks become applied, not arcane.
🧪 Hands-on
Build a minimal chatbot against a transformer API and instrument it so you
see the plumbing: log the tokenization of each turn (watch a word split
into multiple tokens), sweep temperature from 0 to 1.2 on the same
prompt and observe determinism collapse, and try to slip a benign "blocked"
word past a naive substring filter by fragmenting it. You are not building a
product — you are building intuition for the primitives every later attack
stands on.