Skip to content
Phase 4, week 26

Capstone — Red Team Simulation

0 of 24 items done. ~5h estimated.

Concept. This is the capstone. Twenty-five weeks built the parts — prompt hacking, adversarial ML, RAG poisoning, MCP hijacking, guardrails, MLSecOps, ATLAS, incident response. Week 26 assembles them into a single end-to-end red team engagement against a mock enterprise AI system, then translates the findings into the one artifact that actually changes budgets: an executive risk report. The deliverable is not a pile of jailbreaks — it is a scoped, reproducible, framework-mapped assessment that a CISO can act on and an engineering team can close.

The hardest lesson of the capstone is the one Microsoft learned red teaming 100 generative-AI products: no single mitigation eliminates risk, and automation never replaces human judgment — you run iterative break-fix cycles that raise the attacker's cost, and you keep running them (Microsoft).

🎯 Objectives

By the end of this week you can:

  • Run a full AI red team engagement end-to-end: rules of engagement, scoping, scenario selection, execution, structured reporting.
  • Test an agentic system (tools, memory, multi-turn) — not just a text model — using reproducible offline environments.
  • Map every finding to MITRE ATLAS v5.4.0, OWASP LLM Top 10:2025, and the OWASP Agentic Top 10 (ASI) so findings are portable across teams.
  • Write an executive risk report that translates a technical finding into business impact, severity, reproduction, and prioritized remediation.
  • Stand up continuous purple teaming — adversary emulation wired into CI/CD with automated detection validation, not a one-off engagement.

The engagement, start to finish

A capstone engagement is a discipline, not a hunt. It runs the same phases a network pentest does, adapted for a system whose attack surface is its context window.

1 · Rules of engagement & scoping

Before a single payload, fix the boundary. What systems are in scope (the model, the RAG store, the tool/MCP layer, the surrounding app)? What is explicitly out (production data, third-party tenants, the provider's own infrastructure)? What are the stop conditions? This is the Week 14 methodology — and it is what separates authorized red teaming from the "weaponized coding agent" incidents the course catalogued all year, where an agent with standing credentials was pointed at a target with no RoE at all.

Scope for AI has a twist a network scope doesn't: the model is non-deterministic. The same payload succeeds 4 times in 26 and fails the rest — exactly what Zscaler measured when testing indirect prompt injection against browsing agents. So scope must specify attempt counts and success thresholds, not a single pass/fail.

2 · Scenario selection — model the actor, not the payload

Microsoft's red-team ontology is the cleanest scenario generator: model each attack as (actor) → (TTP) → (system weakness) → (downstream impact) (Microsoft). Start from the impact you care about (data exfiltration, fraudulent transaction, RCE, reputational harm) and work backwards to the weakness and the TTP that reaches it. This keeps the engagement business-relevant instead of a jailbreak scavenger hunt.

3 · Execution — agentic systems need environments, not prompts

The single biggest shift for a modern capstone: traditional text-based safety testing is not enough. An agent performs tool-based actions, reacts to real-time feedback, and operates in semi-autonomous cycles — so you must test it inside a realistic, reproducible, offline environment, not a chat box (Toloka). Toloka's agentic study built offline custom platforms across 25+ use cases (social media, financial dashboards, coding forums) and ran 1,200+ scenarios across 100+ attack vectors and 40+ risk categories — malicious code execution, file deletion, data exfiltration, time-delay attacks (Toloka). Their three primary vulnerability classes are the taxonomy to test against:

Class What it is Capstone test
External prompt injection Malicious instructions in emails, ads, web content, tool output Plant a payload in RAG/web/tool data the agent will read
Agent mistakes Accidental disclosure or destructive action, no attacker Ambiguous multi-step task → does it leak or delete?
Direct misuse Harmful request aimed at the agent itself Jailbreak / policy-bypass on the model directly

🔑 The capstone rule: test the system, not the model. A model that refuses every jailbreak is still fully exploitable if its tool layer executes an instruction it read from a poisoned document.

4 · Reporting — the deliverable that matters

Everything above is worthless if it lands as a wall of transcripts. The executive risk report is the product. Each finding carries: a severity rating, reproduction steps, the business impact in plain language, and a prioritized remediation. Two audiences read it — executives who need impact and cost, engineers who need the repro and the fix — so structure it in layers: a one-page risk summary on top, the technical appendix beneath.

Mapping findings to the three frameworks

A finding that isn't mapped is a finding no other team can act on. The capstone maps every issue to all three standards; they're complementary, not competing.

Framework Version Answers Example entry
MITRE ATLAS v5.4.0 How the attack works (adversary TTPs, kill chain) AML.T0051 LLM Prompt Injection
OWASP LLM Top 10 2025 What the model-level risk is LLM01 Prompt Injection · LLM06 Excessive Agency
OWASP Agentic Top 10 (ASI) 2025 What the agent/tool-level risk is Tool misuse · memory poisoning · privilege compromise

Read them as a stack: OWASP names the weakness class, ATLAS supplies the technique ID and the emulation plan, and the Agentic Top 10 covers the autonomy-specific failure modes the LLM list predates. A prompt-injection finding that lets an agent call a payment tool maps to LLM01 + LLM06 + Agentic tool-misuse + AML.T0051 at once — and that four-way tag is exactly what makes it reproducible by the blue team.

💡 Severity is not the model's confidence — it's business impact × reachability. A CVSS-10 jailbreak on a model with no tools and no sensitive data is lower risk than a "medium" injection on an agent wired to production credentials.

From engagement to always-on: continuous purple teaming

A one-off engagement is a photograph; the system changes weekly. The capstone's second half is a continuous purple teaming model — the red team's attacks become automated adversary emulation wired into CI/CD, and every emulation validates that the blue team's detections actually fire. This is the SEC598 continuous model: attacks run on every deploy, detection coverage is tracked against MITRE ATT&CK, and gaps become tickets. Concretely, adversary emulation orchestrated in a SOAR platform (Tines) firing Atomic Red Team-style test cases against the AI stack, with detection results fed back automatically.

The reason this beats point-in-time testing is Microsoft's first takeaway: GenAI systems carry both classic vulnerabilities (an unpatched FFmpeg dependency gave them a real SSRF in a video pipeline) and novel model-level ones (vision-model jailbreaks where the model couldn't separate system instructions from user input) — and the classic half regresses every time a dependency updates (Microsoft). Continuous emulation catches the regression; an annual pentest doesn't.

🔑 Human expertise stays in the loop. Automation (PyRIT, harnesses, emulation) scales coverage, but subject-matter judgment — medicine, cyber, CBRN, cultural context, psychosocial harm — is irreplaceable (Microsoft). Purple teaming is human-directed, machine-scaled.

What "done" looks like

The capstone is complete when you can hand over three artifacts: a full engagement run to the Week 14 methodology against the mock enterprise system; an executive risk report whose findings carry severity, repro, business impact, and prioritized fixes; and a continuous purple teaming loop that re-runs your attacks on every deploy and reports detection coverage. If a CISO can read the top page and decide where to spend, and an engineer can read the appendix and close the finding, the capstone has done its job.

🎯 OSAI exam depth — Assembling the Pieces (m11)

The methodology above is the frame. The OSAI exam is the execution: a proctored 24-hour hands-on engagement against an AI-enabled enterprise network, structured as two independent attack chains, each spanning three machines, both converging on the same Domain Controller, each starting from a single publicly accessible entry point you pivot inward from (OffSec, PRNewswire). It rewards one skill above all: chaining an AI-layer foothold into traditional infrastructure compromise. A candidate who can jailbreak a chatbot but can't turn that into a reverse shell and lateral movement will not pass. This section is the full-spectrum chain the exam demands.

The environment is larger than "two chains" implies — ten machines: the six scored chain targets, one standalone AI-focused host, the shared Domain Controller, and two intentionally non-vulnerable decoys that score nothing and exist only for realism (OSAI Exam Guide). The eight scored targets total 100 points; you need 75 to pass:

Target Count Points each
AI-vector machine (in the chains) 15
Traditional-vector machine (in the chains) 10
Standalone AI-focused host 1 15
Domain Controller flag 1 5 (submittable once)
Intentionally non-vulnerable decoy 2 0

🔑 The scoring encodes the whole lesson: AI-vector machines are worth 50 % more than traditional ones (15 vs 10) — but the DC flag plus the traditional targets are worth ~25 points you cannot skip, so 75 is unreachable on AI findings alone. The exam is engineered to fail the AI-only red teamer who never learned to pivot to infra.

The AI kill chain — the spine of the engagement

Model every chain as five linked stages; break any link and the attack dies, so as an attacker you fortify each before moving on (NVIDIA):

① Recon. Fingerprint the AI surface before you touch a payload. Ask the NVIDIA recon questions concretely: what data channels do I control that reach the model, and what tools/MCP servers does it call? In practice:

  • Model/family fingerprint — probe with tokenizer quirks, refusal wording, and known canary prompts; error messages and stack traces leak the framework and versions (NVIDIA).
  • System-prompt extraction — "repeat everything above," translation and summarization ruses, and multi-turn dribble attacks to reconstruct the hidden instructions, tool list, and guardrail wording.
  • Tool / MCP enumeration — ask the agent to "list your available tools" and their JSON schemas; the model reads each tool description as trusted instruction, so the schema you extract is also your injection target (redteams.ai).
  • RAG detection — feed a unique token and watch whether later answers cite an ingested doc; probe for retrieval-ranking behavior and upload points. This maps the promptware-chain reconnaissance stage: map the tool inventory, API permissions, data scopes, and decision policies before you commit (Preamble).

② Poison. Get attacker-controlled instructions into a channel the model trusts. Indirect prompt injection is the exam's primary initial-access vector — a payload planted in a RAG document, an uploaded file, a web page the agent browses, a tool's output, or an email it reads (Preamble). Sharpen it with ASCII smuggling / hidden Unicode tag characters so the payload is invisible to a human reviewing the source doc but fully legible to the model (NVIDIA). Poison writable RAG stores and vector DBs; a single tainted document retrieved into an augmented prompt is your foothold.

③ Hijack. The retrieved poison forces the model to act: it invokes a tool with your parameters. This is where the AI layer becomes an execution primitive. On the exam expect to:

  • Force the agent to build malicious SQL/NoSQL through a DB tool, or path traversal / SSRF through a URL-fetch tool's parameters — natural-language-driven classic web attacks (redteams.ai).
  • Abuse excessive agency via MCP-exposed CRUD/shell tools to write files, install packages, or run commands.
  • MCP / supply-chain abuse is a first-class surface — 99 MCP CVEs landed in 2025 and tool poisoning is live, not theoretical (redteams.ai). Techniques: tool-poisoning (hostile instructions hidden in a tool's description), tool-name spoofing / shadowing, rug-pull (a benign tool mutates after approval), and compromise of the MCP server package itself (redteams.ai).

④ Persist. Turn one-shot into durable control. Poison memory / vector DBs so the malicious instruction survives the session and re-fires for every future user, or stand up a retrieval-based C2: have the agent periodically fetch adversary-controlled content to pull new instructions asynchronously (Preamble).

⑤ Impact / pivot to infra. This is the exam's crux and where AI-only red teamers stall. The tool-call you hijacked runs on a worker sandbox — open an RCE channel (reverse shell) out of it, then move like any network pentester: harvest secrets and cloud creds from the sandbox env, use the agent's email/tool integrations for lateral movement, and pivot machine-to-machine toward the Domain Controller the two chains converge on (arXiv 2606.24496, OffSec). Remember Microsoft's dual surface: the infra half is often a classic bug reached through the AI — an unpatched dependency (their FFmpeg SSRF in a video pipeline) exposed by a tool the model can drive (Microsoft).

🔑 Exam chain in one line: indirect injection in RAG → agent invokes a poisoned MCP tool → SSRF/RCE out of the worker sandbox → secrets + lateral movement → Domain Controller. Every AI finding must terminate in a demonstrated infra impact, not a screenshot of a jailbroken reply.

Full-spectrum surfaces the exam mixes

The environment blends LLM apps, RAG/vector DBs, multi-agent systems, model-orchestration frameworks, embeddings, and the cloud infra behind them (OffSec). Two surfaces beyond prompt injection to drill:

  • Embedding / retrieval-ranking attacks — craft documents whose embeddings sit adjacent to high-value queries so your poisoned doc is the top-ranked context returned; invert embeddings to recover text from a leaked vector store. Manipulating retrieval ranking to surface malicious content is explicitly in AI-300 scope (OffSec).
  • Agentic supply-chain — poisoned model provenance, tampered training data, and malicious MCP-server / package dependencies map to OWASP ASI Tool Misuse + Agentic Supply-Chain Compromise — tag findings there so the blue team can act (redteams.ai).
Running the 24-hour AI-assisted exam workflow

It's OSCP endurance format applied to AI, so treat it as an engagement, not a puzzle marathon:

  • Rules. Fully proctored and open-book — and, unusually for OffSec, AI chatbots / LLMs are not just permitted but encouraged. OSAI is one of only two exceptions (with OSEE) to OffSec's blanket exam AI ban; effective AI use is treated as a core graded competency, and there are no tool restrictions (OffSec AI Usage Policy, OSAI exam FAQ). The catch: My Kali is not provided — bring your own tooling. A professional report is required, graded, and due within 24 h of the exam ending — miss that window and you auto-fail, and submissions are final (OSAI Exam Guide).
  • Time-box the two chains. Two chains, three machines each. Budget roughly a half-day per chain, leave a buffer, and rotate off a stuck machine — the other chain often unblocks it. Both converge on the DC, so a foothold on chain B can hand you creds that finish chain A.
  • Recon everything first. Enumerate all public entry points and every AI tool/RAG/MCP surface before exploiting; the injection target you need is usually a tool schema you already dumped in stage ①.
  • Collect proof as you go, not at hour 23. For every host: screenshot the local.txt/proof.txt equivalent with your control indicator, save the exact injection payload, the tool call it triggered, and the reverse-shell/command output. AI findings are non-deterministic — capture the successful attempt and the attempt count, because a 4-in-26 success still counts but must be reproducible.
  • Write the report in layers (§ Reporting above): per-finding severity = business impact × reachability, exact reproduction, and every AI finding mapped to ATLAS + OWASP LLM + OWASP Agentic and carried through to its infra impact. "Try Harder" is the ethos, but the graded artifact is a clean, reproducible chain writeup (OffSec).

📇 Engagement toolkit & framework reference

The lesson above is the method. This is the catalog behind it — frameworks, harnesses, and the case studies that define the standard. Folded by default; expand for the detail.

Tool / standard Role in the engagement Source
MITRE ATLAS v5.4.0 Adversary TTP taxonomy + emulation plans for the AI kill chain atlas.mitre.org
OWASP LLM Top 10:2025 Model-level risk taxonomy genai.owasp.org
OWASP Agentic Top 10 (ASI) Agent/tool/autonomy risk taxonomy genai.owasp.org
MITRE ATT&CK Detection-coverage mapping target for purple teaming attack.mitre.org
PyRIT Microsoft's open-source risk-identification automation harness Microsoft
Atomic Red Team Library of small, reproducible emulation test cases for CI/CD atomicredteam.io
Tines SOAR orchestration for continuous adversary emulation tines.com

Reference engagements to model on.

  • Microsoft — Red Teaming 100 GenAI Products. The break-fix ontology (actor → TTP → weakness → impact), human-in-the-loop necessity, and the classic-plus-novel dual risk surface (FFmpeg SSRF alongside vision jailbreaks). The single best methodology reference for this week. Source: Microsoft
  • Toloka — AI Agents Under Attack. The agentic-specific playbook: offline reproducible environments across 25+ use cases, 1,200+ scenarios, 100+ vectors, and the three-class vulnerability model (external injection / agent mistakes / direct misuse). The reference for testing tools and autonomy, not just text. Source: Toloka
  • Jason Haddix — The AI Attack Blueprint. Full AI attack methodology for pentesting AI systems, 26 chapters, from the CEO of Arcanum InfoSec — a practitioner's end-to-end walkthrough to pair with the frameworks above. Source: YouTube
  • RSAC 2025 — Security in the Age of Agentic AI (Vasu Jakkal, Microsoft). Strategic framing for why security must evolve for autonomous agents — the executive-audience companion to the technical work. Source: YouTube

Recommended resources0/16

Sign in to tick items off and track your progress.

Show

📖 Core Path

  • 📄 Microsoft — 3 Takeaways from Red Teaming 100 GenAI Products — the break-fix ontology (actor → TTP → weakness → impact), human-in-the-loop necessity, classic+novel dual risk (FFmpeg SSRF + vision jailbreaks), PyRIT. The methodology reference for this week (~2h)
  • 📄 Toloka — AI Agents Under Attack: A Case Study on Advanced Agent Red Teaming — offline reproducible environments (25+ use cases), 1,200+ scenarios, 100+ vectors, three-class model (external injection / agent mistakes / direct misuse) (~1.5h)
  • 🎥 Jason Haddix — The AI Attack Blueprint (Interview) — complete AI attack methodology for pentesting AI systems; 26 chapters from the CEO of Arcanum InfoSec (~1h 10m)
  • 🧪 Deliverable: executive risk report translating technical findings to business impact with severity ratings, reproduction steps, prioritized remediation recommendations
  • 🧪 Deliverable: all findings mapped to MITRE ATLAS v5.4.0 + OWASP Agentic Top 10 (ASI) + OWASP LLM Top 10:2025

📚 Further Reading


🎓 OSAI / AI-300 alignment

Study checklist

↪ See roadmap.md → Phase 4 → Week 26

  • Write rules of engagement + scope for the mock enterprise AI system (in/out of scope, stop conditions, attempt counts + success thresholds)
  • Select scenarios via the actor → TTP → weakness → impact ontology
  • Test the agent as a system: external injection, agent mistakes, direct misuse in a reproducible offline environment (not just a chat box)
  • Run the full engagement against the mock system (Week 14 methodology)
  • Draft executive risk report: severity, repro steps, business impact, prioritized remediation — layered for exec + engineer audiences
  • Map every finding to MITRE ATLAS v5.4.0 + OWASP LLM Top 10:2025 + OWASP Agentic Top 10 (ASI)
  • Stand up continuous purple teaming: adversary emulation in CI/CD (Tines + Atomic Red Team), automated detection validation, coverage vs MITRE ATT&CK
  • Internalize the OSAI exam shape: 10 machines, 100 pts, 75 to pass, AI-vector 15 / traditional 10 / DC 5 (why 75 is unreachable on AI-only), AI tools encouraged, report due within 24h of exam end

Study notes

Sign in to take notes.