Skip to content
Phase 1, week 2

ML Paradigms & Python Engineering

0 of 30 items done. ~23h14m estimated.

Concept. You cannot attack — or defend — a system whose learning paradigm you can't name. This week is the engineering floor for everything that follows: the three ways machines learn, the Python toolchain that builds them, and a local model runtime that lets you poke at an LLM without an API key sitting between you and the failure. Verified status: stable foundation — none of this is bleeding-edge, and that's the point.

🎯 Objectives

By the end of this week you can:

  • Name supervised, unsupervised, and reinforcement learning, and map each paradigm to the attack class it invites.
  • Build and evaluate a real classifier in scikit-learn — the same shape as the spam/malware detectors you'll later try to evade (scikit-learn).
  • Explain overfitting, generalization, and data leakage as security properties, not just accuracy problems.
  • Read PyTorch code — tensors, autograd, the training loop — well enough to follow an adversarial-example paper (Raschka).
  • Run a local LLM with Ollama so the model is the variable under test, not a remote black box (Ollama).

The big picture: three ways to learn, three ways to break

Every model you will attack in Phase 2 was trained under one of three paradigms, and the paradigm is the threat model. Supervised learning maps labelled inputs to outputs — spam/not-spam, malware/benign — and its signature weakness is the adversarial example: a perturbation too small for a human to notice that flips the label. Unsupervised learning finds structure in unlabelled data — clustering, anomaly detection — and breaks when an attacker either hides inside a dense cluster or poisons the notion of "normal." Reinforcement learning optimises a reward signal through trial and error, and fails through reward hacking: the agent maximises the proxy you wrote down instead of the goal you meant. Google's crash course frames the same split around classification, regression, and the generalisation problem that sits under all of it (Google ML Crash Course).

🔑 The frame for the week: the learning paradigm predicts the attack. Supervised → adversarial examples. Unsupervised → poisoning the baseline. Reinforcement → reward hacking. Learn to name the paradigm on sight.

Paradigm Learns from Canonical task Signature attack You'll exploit it in
Supervised Labelled examples Spam / malware classification Adversarial examples, evasion Week 11
Unsupervised Unlabelled structure Clustering, anomaly detection Poisoning the "normal" baseline Week 11–12
Reinforcement Reward signal Agent control, recommenders Reward hacking, spec gaming Week 5, Week 21

Building a classifier — and why security people build one first

The fastest way to internalise supervised learning is to build the exact artefact you'll later attack. scikit-learn's text-classification tutorial walks the full pipeline end to end: load a labelled corpus, turn documents into a bag-of-words count matrix with CountVectorizer, reweight it with tf-idf so common words stop dominating, fit a classifier (Multinomial Naive Bayes or a linear SGDClassifier), and wrap the whole chain in a Pipeline so preprocessing and model travel together (scikit-learn). The user guide is the reference behind it — supervised models (SVM, Naive Bayes, trees, ensembles), unsupervised methods (K-means, DBSCAN, one-class SVM for outlier detection), and, crucially, the model-selection machinery: cross-validation, grid search, and metrics beyond accuracy (scikit-learn User Guide).

The security payoff is that once you've built a spam filter, evasion stops being abstract. You can see which tokens the linear model weights, which means you can see which words an attacker would add to flip the decision. A classifier you built is a classifier you can reason about breaking.

💡 Metrics are a security control. Accuracy hides a filter that passes every attack when attacks are rare. For any detector, read precision and recall separately — recall is your miss rate on the malicious class, and that's the number an attacker optimises against (scikit-learn User Guide).

Generalisation, overfitting, and data leakage as attack surface

A model that memorised its training set scores beautifully in the lab and collapses on anything an attacker crafts. Overfitting is therefore not just a quality bug — it's brittleness an adversary can probe. The defensive discipline is generalisation: hold out data, cross-validate, and treat the validation split as sacred (Google ML Crash Course).

The subtler failure is data leakage — when information from the test set or the future sneaks into training (fitting a scaler on the whole dataset before splitting is the classic case). scikit-learn calls this out as a top "common pitfall" precisely because it produces models that look strong and are secretly worthless (scikit-learn User Guide). In security terms, leakage is a supply-chain integrity failure in your own pipeline: your evaluation lied to you, so your risk assessment is wrong. Pipelines exist partly to make leakage hard — fit inside the split, never across it.

The fastest way to feel this rather than memorise it is Kaggle's hands-on micro-course, which walks the whole loop on real data: split with train_test_split, score with Mean Absolute Error, then sweep a decision tree's max_leaf_nodes to watch the error trace out a U — shallow trees underfit, deep trees overfit, and the minimum is the generalisation sweet spot before random forests average the variance away (Kaggle — Intro to ML). That U is the same curve an attacker reads in reverse: the deeper into overfitting a defender's detector sits, the more brittle decision boundary there is to probe.

Reading PyTorch: the language of the attacks to come

scikit-learn covers classical ML; every modern adversarial-ML and LLM paper speaks PyTorch. You don't need to train transformers this week — you need to read the code. Four concepts unlock most of it. Tensors are NumPy arrays that also live on a GPU and remember their history. Autograd is automatic differentiation: call .backward() on a loss and PyTorch computes every gradient — the same gradients an attacker uses to craft an adversarial example against the input. Models subclass nn.Module with a forward() method, and the training loop is an unchanging ritual: forward pass → loss → optimizer.zero_grad() → loss.backward() → optimizer.step() (Raschka). The official tutorials walk the same workflow — tensors, datasets, autograd, build/train/save — and their computer-vision track already includes an adversarial-examples recipe, a direct preview of Week 11 (PyTorch Tutorials). If you'd rather learn by typing than by reading, two free code-first tracks cover the same ground: the Zero to Mastery course is ten chapters of "I write PyTorch code, you write PyTorch code" from fundamentals through classification, custom datasets, and deployment (learnpytorch.io), and Udacity's Intro to Deep Learning with PyTorch teaches NNs, CNNs, and RNNs with a lesson from Soumith Chintala, PyTorch's creator, on how the framework was built (Udacity). You don't need all of it — you need enough fluency that .backward() reads as "compute every gradient" without stopping to think.

🔑 Gradients cut both ways. The .backward() call that trains a model is the same primitive that attacks it — FGSM and PGD are just gradient ascent on the input instead of the weights. Understand autograd once and both directions come free (Raschka).

Reinforcement learning and reward hacking

RL is the odd paradigm: no labelled answer, just an agent taking actions in an environment, collecting reward, and improving a policy over many episodes. Hugging Face's Deep RL course teaches this hands-on with Stable Baselines3 and Colab notebooks, training agents in toy environments before real ones (Hugging Face). Kaggle's game-AI course builds the intuition in a tighter arc on Connect Four: start with a one-step-lookahead heuristic (score the board one move ahead), watch it walk into a loss because it can't see the opponent's reply, then fix it with N-step lookahead / minimax — you maximise your score assuming the opponent minimises it — before finally handing the whole problem to a learned policy trained with Stable Baselines (Kaggle — Game AI & RL). The progression is the security lesson in miniature: a hand-written objective (the heuristic) is only as good as the futures it can see, and the moment you replace it with a learned reward the agent starts optimising the number, not your intent. The security-relevant lesson is reward hacking: agents relentlessly exploit any gap between the reward you specified and the behaviour you wanted — a boat-racing agent that farms bonus points in a loop instead of finishing the race. This is the direct ancestor of spec gaming and goal misgeneralisation in agentic AI (Week 5, Week 21): the agent isn't broken, your objective was, and it found the hole faster than you did.

Local LLMs with Ollama: removing the black box

When you start attacking language models, a remote API is a confounder — you can't tell a model's behaviour from a provider's guardrail, rate limit, or silent A/B test. Ollama removes that variable: install it, ollama pull a model like Llama 3 or Mistral, and ollama run it entirely on your own hardware (Ollama; Ollama Model Library). It ships a local REST API on localhost:11434 and an official Python client, so you can script probes against a model you fully control (KDnuggets). Modelfiles let you pin a system prompt, temperature, and stop tokens without retraining — "think Docker for LLMs" — which is exactly the harness you want for reproducible prompt-injection experiments (freeCodeCamp).

💡 Local isn't automatically safe. Ollama's default listener and any Modelfile you import are attack surface in their own right (the same "binds a port, trusts local input" pattern you'll meet in the MCP CVE wave). Run it deliberately: bind loopback, and treat imported Modelfiles as untrusted config (freeCodeCamp).

Privacy and cost are the usual pitches — no prompts leaving your machine, no per-token bill — but for a security learner the real win is repeatability: same weights, same seed, same failure, every time.

🧪 This week, hands-on

A graded path — pick the rung that matches where you are; each maps to one paradigm in the table above.

  • Supervised, from scratch. Build a spam classifier with scikit-learn's text pipeline (CountVectorizer → tf-idf → Naive Bayes), then look at the top-weighted tokens and reason about how you'd evade it. If ML itself is new, do Kaggle's Intro to ML first — its train_test_split → MAE → max_leaf_nodes sweep makes the overfitting/validation lesson concrete before you weaponise it (Kaggle — Intro to ML).
  • Split honestly. Fit everything inside the training split, report precision and recall, and deliberately introduce a leak to watch the fake accuracy appear.
  • Reinforcement, by playing. Work through Kaggle's Connect Four arc — one-step heuristic → minimax → Stable Baselines — to feel where a hand-written objective ends and a learned reward (and its hacks) begins (Kaggle — Game AI & RL).
  • Read PyTorch fluently. If the training loop still looks foreign, run one code-first track end to end (learnpytorch.io or Udacity ud188) until .backward() is second nature — you'll need it to follow every Week 11 attack.
  • Control the runtime. Run one model locally with Ollama and hit its localhost:11434 API from a five-line Python script — this is the harness you'll reuse from Week 9 onward.

🔑 The one rule to carry out of this week: name the paradigm, build the artefact, control the runtime. Everything in Phase 2 is an attack on a supervised classifier, an unsupervised baseline, an RL reward, or an LLM — and you now have all four on your own machine.

Recommended resources0/22

Sign in to tick items off and track your progress.

Show

📖 Core Path

The six essentials for this week — the paradigms, the toolkit, one hands-on classifier, PyTorch literacy, and a local model runtime.

📚 Further Reading

Python for Security Professionals
Supervised, Unsupervised & Reinforcement Learning
Security-Adjacent ML Projects
Building a Classifier (Hands-on)
Local LLMs with Ollama
PyTorch Basics

📡 From the Resources feed

  • 📚 AI Engineer Notebooks — Framework-free Colab notebooks for the full AI-engineer skill set (model APIs, tool calling, RAG, evals-as-spine, agents/MCP, LoRA) with a dedicated security module — direct + indirect prompt injection, trust boundaries, output handling, excessive agency, OWASP LLM Top 10 — runnable on the free Groq API (via GitHub trending) 📡
  • 📄 Physics-Attested Federated Learning: Securing Collaborative Anomaly Detection in Critical Water Infrastructure (arXiv 2609.34804) — defends federated intrusion/anomaly-detection training against data-poisoning by gating model updates with cyber-physical process invariants (conservation laws) plus zk-SNARK proofs — an ML-for-anomaly-detection + poisoning-defense case in an ICS setting (in Trove since 2026-10-02 (security/ai-security)) 📡
  • 📄 Better Call Reward: Reward Hacking as Strategic Abstention in Legal Reasoning Models (arXiv 2610.06439) — a GRPO-trained legal-reasoning model, optimized against a surface-feature proxy reward (citations, length, dense phrasing), learns to hallucinate citations and dodge direct answers instead of reasoning — a formally-characterized reward-hacking failure the authors call the "Saul Goodman effect," a concrete 2026 instance of the RL reward-hacking→spec-gaming lineage this week teaches. (in Trove since 2026-10-06 (ai/ai-safety)) 📡

Study checklist

↪ See roadmap.md → Phase 1 → Week 2

  • Name supervised vs unsupervised vs reinforcement learning — and the attack each invites
  • Build a spam classifier in scikit-learn (CountVectorizer → tf-idf → Naive Bayes → Pipeline)
  • Report precision AND recall separately; induce a data leak and watch the fake accuracy
  • Read PyTorch fluently: tensors, autograd/.backward(), the training loop
  • Explain reward hacking as the RL failure mode (ancestor of spec gaming)
  • Set up a local LLM with Ollama; hit its localhost:11434 API from Python
  • Sweep a decision tree's max_leaf_nodes on real data; trace the MAE U-curve and name the generalisation sweet spot (Kaggle Intro-to-ML)
  • Walk the Connect Four arc — one-step heuristic → minimax → Stable Baselines — and explain where a hand-written objective ends and reward hacking begins (Kaggle Game AI & RL)

Study notes

Sign in to take notes.