Skip to content
Phase 1, week 1

Mathematical Literacy for AI Security

0 of 23 items done. ~114h42m estimated.

Concept. Every AI-security attack and defense you will study for the next 25 weeks bottoms out in one of three pieces of mathematics: the geometry of vector spaces (what a RAG system retrieves, and how it can be poisoned), the probability distribution over tokens (what an adversarial suffix bends toward an unsafe answer), and gradient descent (how a model learns — and how poisoned data corrupts that learning). You do not need to become a mathematician. You need enough fluency that when Week 12 says "embedding poisoning" or Week 11 says "gradient-based jailbreak," the attack is obvious rather than mysterious. That fluency is the whole job of Week 1.

🎯 Objectives

By the end of this week you can:

  • Compute cosine similarity and a dot product by hand on small vectors, and explain why they — not raw distance — decide what a retrieval system returns.
  • Read a row of logits and turn it into a probability distribution with softmax, and explain what temperature does to that distribution.
  • Explain gradient descent and backpropagation well enough to see how a single poisoned training example nudges a model's weights.
  • Map each primitive to the specific later-week attack that exploits it (embedding poisoning, adversarial suffixes, data poisoning).
  • Recognise where probability distributions show up on the defensive side — confidence, uncertainty, and out-of-distribution detection.

Why a security course opens with linear algebra

It is tempting to skip to the exploits. Resist it. The reason is that modern AI systems represent meaning as position in a high-dimensional space, and "position" is a purely mathematical object with purely mathematical weaknesses. An attacker who understands the geometry can place a crafted document at exactly the coordinates that a victim's query will reach for — that is embedding poisoning, and it is indistinguishable from geometry until you can do the geometry. The three foundational specialisations for this material — Imperial College's Mathematics for Machine Learning (Coursera) and DeepLearning.AI's Mathematics for ML and Data Science (Coursera) — both organise themselves the same way: linear algebra to represent data as vectors and matrices, calculus to optimise, probability to quantify uncertainty. That ordering is not accidental; it is the anatomy of every model you will attack.

Vectors, dot products, and the geometry of meaning

An embedding is just a list of numbers — a vector — that a model assigns to a piece of text so that similar meanings land near each other. "Near" is measured, almost universally, by cosine similarity: the cosine of the angle between two vectors, computed as their dot product divided by the product of their lengths.

💡 Cosine similarity by hand. For a = [1, 0, 1] and b = [1, 1, 0]: dot product = 1·1 + 0·1 + 1·0 = 1; each length = √2; so cos = 1 / (√2·√2) = 0.5. A cosine of 1.0 means identical direction, 0 means orthogonal (unrelated), -1 means opposite. Do this by hand a few times — the intuition is what makes Week 12 legible.

Why cosine and not plain distance? Because it ignores magnitude and keeps only direction — two documents about the same topic are "similar" even if one is far longer. That choice is exactly the seam an attacker works: to poison a retrieval index you do not need your malicious document to be close in every dimension, only to point the same direction as the target query. The three canonical metrics differ in what they are blind to:

Metric Formula (intuition) Blind to Where it bites in security
Dot product Σ aᵢbᵢ nothing — magnitude and direction count rewards long, keyword-stuffed poison docs
Cosine similarity dot product ÷ both lengths magnitude the default RAG metric → embedding poisoning (Week 12)
Euclidean distance √Σ(aᵢ−bᵢ)² direction alone clustering / anomaly detection on defense side

The visual intuition for all of this — vectors as arrows, the dot product as projection, matrices as transformations of space — is best absorbed from 3Blue1Brown's Essence of Linear Algebra (3Blue1Brown) and drilled with Khan Academy's exercises (Khan Academy).

🔑 Rule to carry forward: in a vector store, similarity is a weapon. Anything that can write to the index can choose coordinates, and choosing coordinates near a victim query is the entire mechanism of RAG poisoning.

Softmax and logits — the distribution attackers steer

A language model does not emit a word; it emits a vector of raw scores called logits, one per token in its vocabulary. Softmax turns those scores into a probability distribution: exponentiate each logit and divide by the sum of all the exponentials, so the outputs are positive and sum to exactly 1.0 (Machine Learning Mastery). For logits [1, 3, 2], softmax gives [0.090, 0.665, 0.245] — the largest logit dominates but the others keep a share.

The exponentiation is the load-bearing step: it amplifies differences, so a small lead in logit space becomes a large lead in probability space. Softmax is a "soft" argmax — where argmax would return a hard [0, 1, 0], softmax spreads probability to reflect uncertainty (Machine Learning Mastery). Temperature rescales the logits before the exponential: low temperature sharpens the distribution toward the top token (more deterministic), high temperature flattens it (more random). The clear walkthrough of logits → probabilities is StatQuest's Softmax (StatQuest).

💡 Why this is a security concept. An adversarial suffix (Weeks 9–11) is a string of tokens chosen so that, after softmax, the probability mass shifts onto a harmful continuation the model would normally refuse. The attacker is not "tricking" the model in some fuzzy sense — they are literally editing the logits that feed softmax. If you can read a softmax output, you can read the objective an attacker is optimising.

Gradient descent and backpropagation — how learning is corrupted

Models learn by minimising a loss — a single number measuring how wrong they are. Gradient descent repeatedly nudges every weight a little in the direction that most reduces that loss; the gradient is the vector of those directions. Backpropagation is how the gradient is computed efficiently: the network is a computational graph, and the chain rule lets you flow the gradient backward from the loss to every weight, multiplying local gradients along the way (Stanford CS231n).

The mechanics are worth internalising because attacks live inside them. Each operation ("gate") in the graph has a simple local rule (Stanford CS231n):

  • add gate — passes the incoming gradient through unchanged to both inputs.
  • multiply gate — routes each input's gradient scaled by the other input's value, so mismatched magnitudes create outsized, non-obvious effects.
  • max gate — sends the entire gradient to the input that won the forward pass; the losers get zero.

The best from-scratch build of this intuition is Karpathy's micrograd walkthrough, which implements an autograd engine and trains a neuron line by line (Andrej Karpathy); the visual companion is 3Blue1Brown's Neural Networks series (3Blue1Brown, extended as a Manning liveVideo), and the 10-hour hands-on version is freeCodeCamp's Linear Algebra for ML.

🔑 The security payoff: because training follows the gradient, whoever influences the training data influences the gradient. Data poisoning (Week 11–12) works by contributing examples whose gradient tugs the weights toward the attacker's goal. And a gradient-based jailbreak (e.g. GCG) simply runs gradient descent on the input instead of the weights — same math, aimed at the prompt.

Probability and statistics — the defender's language

The distributions above are not only an attack surface; they are the defender's primary instrument. Confidence scores, anomaly detection, out-of-distribution flags, and the statistical tests behind red-team evaluations all rest on the same foundations: random variables, expectation and variance, conditional probability, and Bayes' theorem. Brown University's interactive Seeing Theory covers exactly this spine — basic and compound probability, distributions and the central limit theorem, then frequentist and Bayesian inference and regression — as visual, manipulable widgets (Seeing Theory), and StatQuest's Statistics Fundamentals playlist is the video counterpart (StatQuest). When Week 17 builds a semantic firewall that flags "unusual" inputs, "unusual" will mean low probability under a distribution — and you will already know what that sentence means.

Building the fluency — watch to understand, then drill to retain

Mathematical intuition does not survive a single viewing; it comes from alternating seeing the geometry and doing it. Split your week's resources by which job they do rather than working top to bottom:

  • Watch to understand — 3Blue1Brown's visuals-first series build the mental picture of vectors-as-arrows and gradient-descent-as-rolling-downhill that everything else hangs on; the Neural Networks set is also sold as a Manning liveVideo with 22 graded exercises if you want the enhanced, checkpointed version.
  • Drill to retain — passive video is where intuition decays. Cement it against immediate feedback: Khan Academy's Linear Algebra unit runs vectors → matrix transformations → eigenvalues as interactive, auto-graded exercises, and freeCodeCamp's 33-chapter Linear Algebra for ML (10h 48m) is the code-along counterpart that turns each concept into runnable Python. For the probability spine — random variables, distributions, Bayes — StatQuest's Statistics Fundamentals playlist is the drill track behind the defender's-language section above.
  • Do by hand — the single highest-leverage exercise of the week is not on a screen at all: compute a dot product, two lengths, and a cosine similarity on paper (see the callout below), then perturb one vector and watch the cosine move. That is the exact reasoning you will reuse for embedding poisoning in Week 12, and it takes five minutes.

🔑 Watch-then-drill, not watch-then-move-on. A concept you have only seen animated is a concept you will recognise but cannot use. Every primitive in this week reappears as an attack in Phase 2 — pay the drill cost now so the offensive weeks read as applications, not new material.

📇 Primitive → attack-surface reference

This is the one-screen map to keep. Every math object on the left is a real attack surface on the right; the week is where you exploit or defend it.

Math primitive What it computes How it becomes an attack surface Later week
Cosine similarity / dot product semantic closeness of two embeddings craft a document that points at a victim query's direction → it gets retrieved RAG poisoning, Wk 12
Softmax over logits token probability distribution append a suffix that shifts probability onto an unsafe token Prompt hacking, Wks 9–11
Temperature scaling sharpness of the token distribution exploit high-temperature sampling to widen the space of "reachable" outputs Wks 9–11
Gradient descent / backprop how weights move to cut loss contribute training data whose gradient steers the weights Adversarial ML, Wks 11–12
Gradient on the input sensitivity of loss to each input token optimise a jailbreak string directly (GCG-style) Wk 11
Probability distribution / variance model uncertainty flag low-probability (out-of-distribution) inputs as hostile Guardrails, Wks 17–18

💡 Practical this week (do it, don't just read it): take two 3–5 dimensional vectors, compute their dot product, their lengths, and their cosine similarity by hand. Then perturb one vector slightly and watch the cosine change. That five-minute exercise is the exact mental model you will reuse to reason about embedding poisoning — the geometry is stable, so this is a foundation that does not go stale.

Verified status. Stable foundation — the mathematics does not change, and these primitives underpin every attack and defense in the remaining 25 weeks. The security framings above are the load-bearing addition: learn the math through the lens of what it exposes, and Phase 2's offensive weeks read like applications rather than surprises.

Recommended resources0/14

Sign in to tick items off and track your progress.

Show

📖 Core Path

The six essentials — do these in order. Linear algebra first (the geometry behind embeddings), then how models learn (backprop), then softmax (the token distribution attackers steer).

📚 Further Reading

Linear Algebra & Vector Spaces
Probability & Statistics for ML
Gradient Descent & Backpropagation
Softmax & Logits

Study checklist

↪ See roadmap.md → Phase 1 → Week 1

  • Compute cosine similarity + dot product by hand on 3-5 dim vectors
  • Explain why cosine (not Euclidean) drives RAG retrieval — and how that enables embedding poisoning (Week 12)
  • Turn a row of logits into a probability distribution with softmax; explain temperature
  • Map softmax steering to adversarial suffixes (Weeks 9-11)
  • Understand gradient descent + backpropagation basics (chain rule, add/multiply/max gates)
  • Connect gradient flow to data poisoning and gradient-based jailbreaks (Weeks 11-12)
  • Complete a linear algebra refresher (vectors, matrices, dot products)
  • Explain where probability shows up on defense — confidence, anomaly/OOD detection as low-probability-under-a-distribution (Week 17)
  • Run the watch-then-drill path: 3B1B visuals → Khan/freeCodeCamp/StatQuest drills → cosine-by-hand

Study notes

Sign in to take notes.