Concept. You cannot attack — or defend — a system whose learning paradigm you can't name. This week is the engineering floor for everything that follows: the three ways machines learn, the Python toolchain that builds them, and a local model runtime that lets you poke at an LLM without an API key sitting between you and the failure. Verified status: stable foundation — none of this is bleeding-edge, and that's the point.
🎯 Objectives
By the end of this week you can:
- Name supervised, unsupervised, and reinforcement learning, and map each paradigm to the attack class it invites.
- Build and evaluate a real classifier in scikit-learn — the same shape as the spam/malware detectors you'll later try to evade (scikit-learn).
- Explain overfitting, generalization, and data leakage as security properties, not just accuracy problems.
- Read PyTorch code — tensors, autograd, the training loop — well enough to follow an adversarial-example paper (Raschka).
- Run a local LLM with Ollama so the model is the variable under test, not a remote black box (Ollama).
The big picture: three ways to learn, three ways to break
Every model you will attack in Phase 2 was trained under one of three paradigms, and the paradigm is the threat model. Supervised learning maps labelled inputs to outputs — spam/not-spam, malware/benign — and its signature weakness is the adversarial example: a perturbation too small for a human to notice that flips the label. Unsupervised learning finds structure in unlabelled data — clustering, anomaly detection — and breaks when an attacker either hides inside a dense cluster or poisons the notion of "normal." Reinforcement learning optimises a reward signal through trial and error, and fails through reward hacking: the agent maximises the proxy you wrote down instead of the goal you meant. Google's crash course frames the same split around classification, regression, and the generalisation problem that sits under all of it (Google ML Crash Course).
🔑 The frame for the week: the learning paradigm predicts the attack. Supervised → adversarial examples. Unsupervised → poisoning the baseline. Reinforcement → reward hacking. Learn to name the paradigm on sight.
| Paradigm | Learns from | Canonical task | Signature attack | You'll exploit it in |
|---|---|---|---|---|
| Supervised | Labelled examples | Spam / malware classification | Adversarial examples, evasion | Week 11 |
| Unsupervised | Unlabelled structure | Clustering, anomaly detection | Poisoning the "normal" baseline | Week 11–12 |
| Reinforcement | Reward signal | Agent control, recommenders | Reward hacking, spec gaming | Week 5, Week 21 |
Building a classifier — and why security people build one first
The fastest way to internalise supervised learning is to build the exact
artefact you'll later attack. scikit-learn's text-classification tutorial
walks the full pipeline end to end: load a labelled corpus, turn documents
into a bag-of-words count matrix with CountVectorizer, reweight it
with tf-idf so common words stop dominating, fit a classifier
(Multinomial Naive Bayes or a linear SGDClassifier), and wrap the whole
chain in a Pipeline so preprocessing and model travel together
(scikit-learn).
The user guide is the reference behind it — supervised models (SVM, Naive
Bayes, trees, ensembles), unsupervised methods (K-means, DBSCAN, one-class
SVM for outlier detection), and, crucially, the model-selection machinery:
cross-validation, grid search, and metrics beyond accuracy
(scikit-learn User Guide).
The security payoff is that once you've built a spam filter, evasion stops being abstract. You can see which tokens the linear model weights, which means you can see which words an attacker would add to flip the decision. A classifier you built is a classifier you can reason about breaking.
💡 Metrics are a security control. Accuracy hides a filter that passes every attack when attacks are rare. For any detector, read precision and recall separately — recall is your miss rate on the malicious class, and that's the number an attacker optimises against (scikit-learn User Guide).
Generalisation, overfitting, and data leakage as attack surface
A model that memorised its training set scores beautifully in the lab and collapses on anything an attacker crafts. Overfitting is therefore not just a quality bug — it's brittleness an adversary can probe. The defensive discipline is generalisation: hold out data, cross-validate, and treat the validation split as sacred (Google ML Crash Course).
The subtler failure is data leakage — when information from the test set or the future sneaks into training (fitting a scaler on the whole dataset before splitting is the classic case). scikit-learn calls this out as a top "common pitfall" precisely because it produces models that look strong and are secretly worthless (scikit-learn User Guide). In security terms, leakage is a supply-chain integrity failure in your own pipeline: your evaluation lied to you, so your risk assessment is wrong. Pipelines exist partly to make leakage hard — fit inside the split, never across it.
The fastest way to feel this rather than memorise it is Kaggle's hands-on
micro-course, which walks the whole loop on real data: split with
train_test_split, score with Mean Absolute Error, then sweep a decision
tree's max_leaf_nodes to watch the error trace out a U — shallow trees
underfit, deep trees overfit, and the minimum is the generalisation sweet spot
before random forests average the variance away
(Kaggle — Intro to ML).
That U is the same curve an attacker reads in reverse: the deeper into
overfitting a defender's detector sits, the more brittle decision boundary
there is to probe.
Reading PyTorch: the language of the attacks to come
scikit-learn covers classical ML; every modern adversarial-ML and LLM paper
speaks PyTorch. You don't need to train transformers this week — you need
to read the code. Four concepts unlock most of it. Tensors are NumPy
arrays that also live on a GPU and remember their history. Autograd is
automatic differentiation: call .backward() on a loss and PyTorch computes
every gradient — the same gradients an attacker uses to craft an
adversarial example against the input. Models subclass nn.Module with a
forward() method, and the training loop is an unchanging ritual:
forward pass → loss → optimizer.zero_grad() → loss.backward() →
optimizer.step() (Raschka).
The official tutorials walk the same workflow — tensors, datasets, autograd,
build/train/save — and their computer-vision track already includes an
adversarial-examples recipe, a direct preview of Week 11
(PyTorch Tutorials). If you'd
rather learn by typing than by reading, two free code-first tracks cover the
same ground: the Zero to Mastery course is ten chapters of "I write PyTorch
code, you write PyTorch code" from fundamentals through classification, custom
datasets, and deployment (learnpytorch.io), and
Udacity's Intro to Deep Learning with PyTorch teaches NNs, CNNs, and RNNs with
a lesson from Soumith Chintala, PyTorch's creator, on how the framework was
built (Udacity).
You don't need all of it — you need enough fluency that .backward() reads as
"compute every gradient" without stopping to think.
🔑 Gradients cut both ways. The
.backward()call that trains a model is the same primitive that attacks it — FGSM and PGD are just gradient ascent on the input instead of the weights. Understand autograd once and both directions come free (Raschka).
Reinforcement learning and reward hacking
RL is the odd paradigm: no labelled answer, just an agent taking actions in an environment, collecting reward, and improving a policy over many episodes. Hugging Face's Deep RL course teaches this hands-on with Stable Baselines3 and Colab notebooks, training agents in toy environments before real ones (Hugging Face). Kaggle's game-AI course builds the intuition in a tighter arc on Connect Four: start with a one-step-lookahead heuristic (score the board one move ahead), watch it walk into a loss because it can't see the opponent's reply, then fix it with N-step lookahead / minimax — you maximise your score assuming the opponent minimises it — before finally handing the whole problem to a learned policy trained with Stable Baselines (Kaggle — Game AI & RL). The progression is the security lesson in miniature: a hand-written objective (the heuristic) is only as good as the futures it can see, and the moment you replace it with a learned reward the agent starts optimising the number, not your intent. The security-relevant lesson is reward hacking: agents relentlessly exploit any gap between the reward you specified and the behaviour you wanted — a boat-racing agent that farms bonus points in a loop instead of finishing the race. This is the direct ancestor of spec gaming and goal misgeneralisation in agentic AI (Week 5, Week 21): the agent isn't broken, your objective was, and it found the hole faster than you did.
Local LLMs with Ollama: removing the black box
When you start attacking language models, a remote API is a confounder — you
can't tell a model's behaviour from a provider's guardrail, rate limit, or
silent A/B test. Ollama removes that variable: install it, ollama pull
a model like Llama 3 or Mistral, and ollama run it entirely on your own
hardware (Ollama; Ollama Model
Library). It ships a local REST API on
localhost:11434 and an official Python client, so you can script probes
against a model you fully control (KDnuggets).
Modelfiles let you pin a system prompt, temperature, and stop tokens
without retraining — "think Docker for LLMs" — which is exactly the harness
you want for reproducible prompt-injection experiments
(freeCodeCamp).
💡 Local isn't automatically safe. Ollama's default listener and any Modelfile you import are attack surface in their own right (the same "binds a port, trusts local input" pattern you'll meet in the MCP CVE wave). Run it deliberately: bind loopback, and treat imported Modelfiles as untrusted config (freeCodeCamp).
Privacy and cost are the usual pitches — no prompts leaving your machine, no per-token bill — but for a security learner the real win is repeatability: same weights, same seed, same failure, every time.
🧪 This week, hands-on
A graded path — pick the rung that matches where you are; each maps to one paradigm in the table above.
- Supervised, from scratch. Build a spam classifier with scikit-learn's text pipeline (CountVectorizer → tf-idf → Naive Bayes), then look at the top-weighted tokens and reason about how you'd evade it. If ML itself is new, do Kaggle's Intro to ML first — its
train_test_split→ MAE →max_leaf_nodessweep makes the overfitting/validation lesson concrete before you weaponise it (Kaggle — Intro to ML). - Split honestly. Fit everything inside the training split, report precision and recall, and deliberately introduce a leak to watch the fake accuracy appear.
- Reinforcement, by playing. Work through Kaggle's Connect Four arc — one-step heuristic → minimax → Stable Baselines — to feel where a hand-written objective ends and a learned reward (and its hacks) begins (Kaggle — Game AI & RL).
- Read PyTorch fluently. If the training loop still looks foreign, run one code-first track end to end (learnpytorch.io or Udacity ud188) until
.backward()is second nature — you'll need it to follow every Week 11 attack. - Control the runtime. Run one model locally with Ollama and hit its
localhost:11434API from a five-line Python script — this is the harness you'll reuse from Week 9 onward.
🔑 The one rule to carry out of this week: name the paradigm, build the artefact, control the runtime. Everything in Phase 2 is an attack on a supervised classifier, an unsupervised baseline, an RL reward, or an LLM — and you now have all four on your own machine.