Skip to content
Phase 1, week 3

Deep Learning & Transformer Architectures

0 of 36 items done. ~17h22m estimated.

Concept. A transformer is a next-token predictor built almost entirely out of one operation — attention — and everything that makes it powerful also makes it exploitable. Attention weighs every input token against every other, which is exactly why an instruction buried in retrieved text or a tool description gets the same consideration as your system prompt. This week you learn the architecture from the inside so the attacks in Phase 2 read as consequences, not surprises: prompt injection is attention doing its job, token smuggling is tokenization doing its job, and weaponized model variants are fine-tuning doing its job.

🎯 Objectives

By the end of this week you can:

  • Trace the NLP lineage — word2vec embeddings → RNN/LSTM → transformer — and explain why self-attention replaced recurrence (Illustrated Word2Vec, Colah on LSTMs).
  • Walk the query/key/value mechanism, multi-head attention, and positional encoding, and name the data/instruction conflation that attention creates (Illustrated Transformer).
  • Explain how BPE / byte-level BPE tokenization maps text to tokens, and how token boundaries enable token smuggling past naive filters (HF Tokenizers).
  • Reason about temperature / top-k / top-p and why non-determinism matters for both testing and probing (Raschka).
  • Explain how LoRA/PEFT makes weaponized model variants economically feasible, and cite the evidence that a handful of examples unaligns a safety-tuned model (Qi et al. 2023).
  • Build a basic chatbot against a transformer API to feel the inference loop end-to-end.

From word vectors to attention

Before transformers, meaning was already geometric. word2vec learns to place each word at a point in a dense vector space by predicting a word from its neighbours (CBOW) or its neighbours from the word (skip-gram), so that proximity encodes semantics and vector arithmetic captures relationships — the famous king − man + woman ≈ queen (Illustrated Word2Vec). That embedding space is not a historical footnote: it is the same space RAG retrieval (Week 4) searches and the same space adversarial perturbations (Week 12) navigate.

The problem was sequence. RNNs loop hidden state forward token by token, but as the gap between related words grows they "become unable to learn to connect the information" — the vanishing-gradient problem. LSTMs patched this with a cell-state "conveyor belt" gated by forget / input / output gates, buying longer memory (Colah on LSTMs). But both remained strictly sequential — token n waits for token n−1 — which caps parallelism and still strains on long dependencies.

🔑 The transformer's one big idea: drop recurrence entirely. Because every token attends to every other in parallel, training is far more parallelizable and long-range dependencies are one hop away, not n hops (Attention Is All You Need).

Inside the transformer

Self-attention lets each token gather context from the whole sequence. For every token the model derives three vectors — Query, Key, Value — and computes, for each pair, a score as the dot product of one token's Query with another's Key, scaled and softmax-normalized; the output is the Value vectors summed by those weights (Illustrated Transformer). When the model encodes "it" in "the animal didn't cross the street because it was tired," attention is what lets "it" bind to "animal."

Two refinements matter for security reasoning:

  • Multi-head attention runs several attention mechanisms in parallel (eight in the original), each with its own Q/K/V projections, so different heads specialize in different relationships before being concatenated (Illustrated Transformer).
  • Positional encoding is added to token embeddings because bare attention is order-blind — it has no inherent notion of sequence, so position must be injected via sinusoidal patterns (Vaswani et al.).

If you want to see this rather than read it, the Transformer Explainer (Georgia Tech) runs a live GPT-2-small in the browser: type a prompt and hover a token to watch which earlier tokens it attends to across all 12 heads, then trace the embedding → Q/K/V → MLP → softmax path end-to-end. It exposes the same three surfaces this chapter names as security-relevant — the token split at the input, the attention weights in the middle, and a temperature slider on the output distribution — in one view.

💡 The security reading of attention. Attention does not distinguish trusted instruction from untrusted data — both are just tokens competing for weight. There is no architectural "instruction channel." That conflation is the mechanism behind prompt injection: text retrieved from a webpage, an email, or a tool output can out-compete your system prompt simply by being phrased as a stronger instruction. Every defense in Phase 3 (spotlighting, delimiters, input sanitization) is an attempt to re-impose a boundary the architecture never had.

Tokenization: where text becomes tokens

Models don't see characters or words — they see tokens. Subword tokenizers keep the vocabulary compact while still representing unseen words: common words stay whole, rare ones decompose (annoyingly → ["annoying", "ly"]) (HF Tokenizers). The dominant algorithms differ in how they build the vocabulary:

Algorithm Merge rule Used by Note
BPE merge the most frequent adjacent pair, iteratively GPT, Llama, Qwen deterministic merge list
Byte-level BPE BPE over 256 byte values, not Unicode chars GPT-2, RoBERTa no <unk> — every byte is representable
WordPiece merge the pair that most raises training likelihood BERT family favors informative merges, not just frequent
Unigram prune a large vocab by loss contribution; probabilistic T5, Pegasus can sample alternate tokenizations
SentencePiece BPE/Unigram on raw bytes incl. space (▁) multilingual works for space-free languages

The security angle: tokenization is a filter-evasion surface. Because a string can split into tokens more than one way, an attacker can fracture a blocked word across token boundaries — token smuggling — so a naive substring filter never sees the forbidden term while the model still reconstructs its meaning. Standard (non-byte-level) BPE also collapses out-of-vocabulary characters into <unk>, and homoglyph / character substitution can produce "ambiguous tokenization patterns" that bypass content filters (HF LLM Course). Byte-level BPE removes the <unk> hole but not the multiple-encodings problem — which is why you'll spend time in the OpenAI Tokenizer watching payloads split.

🔑 Rule of thumb: any input filter that operates on characters is blind to the model's token view. Filter on normalized, canonicalized text and validate at the token level, or the boundary you think you drew isn't the one the model sees.

Sampling: temperature, top-k, top-p

Given the model's probability distribution over the next token, the decoding strategy decides what actually comes out — and it is a knob you must understand to test AI systems honestly. Three parameters shape it (Raschka, HF blog):

  • Temperature reshapes the distribution before sampling: 0 is greedy/deterministic (same output every time), higher values flatten it toward randomness (Cohere).
  • Top-k keeps only the k highest-probability candidates — a fixed-size pool — and discards the rest.
  • Top-p (nucleus) keeps the smallest set whose cumulative probability reaches p (e.g. 0.9) — an adaptive pool that shrinks when the model is confident and grows when it isn't.

💡 Why this matters for security. Non-determinism is a double-edged evaluator. A jailbreak that fails at temperature 0 may succeed at 1.0 on one of many samples — so red-team results must report the sampling config, and a single "it refused" run proves nothing. Conversely, determinism probing (Week 9) exploits temperature-0 reproducibility to reverse-engineer behavior. Greedy and beam search are consistent but repetitive; sampling is diverse but can wander incoherent.

Fine-tuning as attack surface

Full retraining of a 70B model is out of reach for most attackers. PEFT — parameter-efficient fine-tuning — changes that by freezing the base model and training only a tiny set of extra parameters, "significantly decreasing computational and storage costs" while matching full fine-tuning quality (HF PEFT). LoRA is the canonical method: it decomposes each weight update into two low-rank matrices ΔW = B·A injected into the attention projections, so the forward pass becomes y = Wx + BAx. Trainable parameters drop by roughly two orders of magnitude and the resulting adapter is a few MB, not GB (HF LoRA guide).

That economics is the threat. A cheap, shareable adapter means:

  • Weaponized variants are feasible. An attacker can produce and distribute a malicious LoRA the way you'd share a config file.
  • Adapters are supply-chain payloads. A LoRA pulled from an untrusted hub can carry a backdoor; you merge it and inherit its behavior (Phase 2, Week 16).
  • Alignment is fragile under fine-tuning. Qi et al. jailbroke GPT-3.5 Turbo's safety guardrails by fine-tuning on only ~10 adversarial examples for under $0.20 — and even benign fine-tuning "can inadvertently degrade the safety alignment." Safety does not survive customization (Qi et al. 2023).

🔑 Carry this into Phase 2: the same three mechanisms you learned as architecture are the same three you'll exploit as attack surface — attention (prompt injection), tokenization (smuggling), and fine-tuning (unalignment + backdoored adapters). Understand them here and the offensive weeks become applied, not arcane.

🧪 Hands-on

Build a minimal chatbot against a transformer API and instrument it so you see the plumbing: log the tokenization of each turn (watch a word split into multiple tokens), sweep temperature from 0 to 1.2 on the same prompt and observe determinism collapse, and try to slip a benign "blocked" word past a naive substring filter by fragmenting it. You are not building a product — you are building intuition for the primitives every later attack stands on.

Recommended resources0/28

Sign in to tick items off and track your progress.

Show

📖 Core Path

The essential six — read these and you have the week.

📚 Further Reading

The Transformer Architecture
NLP Lineage (word2vec → RNNs → Transformers)
Tokenization & Security
RLHF & Fine-Tuning
LoRA & PEFT (Fine-Tuning as Attack Surface)
Temperature & Sampling

Study checklist

↪ See roadmap.md → Phase 1 → Week 3

  • Trace the NLP lineage: word2vec embedding space → RNN/LSTM (vanishing gradients) → why self-attention replaced recurrence
  • Walk the transformer: Q/K/V self-attention, multi-head, positional encoding
  • Name the data/instruction conflation and connect it to prompt injection
  • Explain BPE / byte-level BPE and how token boundaries enable token smuggling
  • Reason about temperature / top-k / top-p and why non-determinism affects testing
  • Explain LoRA/PEFT economics + cite that ~10 examples unalign a safety-tuned model
  • 🧪 Build a basic transformer-API chatbot; log tokenization + sweep temperature
  • 🔧 Use the Transformer Explainer to watch attention + temperature live on GPT-2

Study notes

Sign in to take notes.