AI Security Mastery Roadmap
Canonical reference document. This is what you read when you're studying a topic. The companion
study-checklist.mdis the execution layer (tick boxes) and links into the sections here.Last verified against trends: 2026-10-01 (see
retro-2026-10.mdfor structural changes;reconciliation-2026-09-28.mdfor content diff). OWASP LLM Top 10 pins to the 2026 edition (reordered Aug-2026, 8 of 10 categories moved). MITRE ATLAS pins to v2026.09 (15 Sep 2026). OWASP Agentic Top 10 (ASI) pins to the 2026 release, now a ranked list (ASI01 Agent Goal Hijack → ASI10 Rogue Agents).
How to use this document
This roadmap is structured as 4 phases × 26 weeks, mirroring the checklist. For each week:
- Concept — what the topic is and why it matters (the "what")
- Key references — primary sources, papers, CVEs, frameworks
- Verified status — fresh / aging / stale; last reviewed date
- Practical — what you do with it (links to checklist tick boxes)
The top of the document (this file) covers cross-cutting frames: the lethal trifecta, the OWASP/ATLAS/ASI taxonomies, and the governance stack. Read these before phase-specific sections — they're the vocabulary used throughout.
Cross-cutting frames
The "lethal trifecta"
An agent system becomes fundamentally compromisable via indirect prompt injection when all three of these are present:
- Privileged access — the agent can read or modify sensitive resources
- Untrusted input — the agent processes content from sources outside the trust boundary (web pages, emails, RAG documents, tool outputs)
- Ability to publish/exfiltrate — the agent can send data outward (network calls, writing to shared storage, posting to channels)
Remove any one leg and the trifecta collapses. This is the architectural test you apply at design time. Source: Lakera, Indirect Prompt Injection.
Multi-stage agent kill chains (new pattern, Apr 2026)
Individual CVEs are no longer the unit of analysis. Attackers now chain vulnerabilities through AI agents as the exploitation mechanism:
Repo → Agent → Container escape. Prompt injection seeded in a GitHub repo → AI coding agent processes it → constructs a Docker auth bypass (CVE-2026-34040: body > 1MB silently drops AuthZ) → creates privileged container → mounts host filesystem. The developer only ran
git clone. (Cyera Research, Mar 2026; patched Docker 29.3.1)Document → MCP → Secrets. Malicious Google Doc → AI agent fetches attacker MCP server instructions → runs Python payload → harvests secrets. Same pattern as CVE-2025-59944 (one-char typo in config path).
PR → CI Agent → Key leak. "Comment and Control" — PR title injection caused Claude Code, Gemini CLI Action, and GitHub Copilot Agent to leak API keys as GitHub comments. (VentureBeat, Apr 2026)
Security scanner → CI → API key harvest (TeamPCP, Mar 2026). Compromised Trivy GitHub Action (CVE-2026-33634, CVSS 9.4) → stolen CI/CD creds → self-propagating npm worm (47 packages) → backdoored LiteLLM + Telnyx SDK on PyPI → harvested API keys for 100+ LLM providers simultaneously. Used WAV steganography for C2 and ICP canisters for censorship-resistant command-and-control. Cisco breached as downstream victim.
[Apr-29 research]npm worm → OIDC token theft → AI agent config persistence (Shai-Hulud/Miasma May–Jun 2026). Self-propagating npm worm compromised 373 package-version entries across 169 npm packages (@tanstack, @mistralai, @squawk, @uipath). Orphaned commit technique on GitHub forks minted valid OIDC publish tokens despite 2FA — packages carried valid SLSA provenance attestations. Novel persistence: injects
.claude/settings.jsonSessionStarthook +.vscode/tasks.jsonrunOn:folderOpeninto every accessible repo via GraphQL mutations impersonating the Claude Code bot. The infected repo becomes the propagation vector — no npm install needed on the next victim. Dead-man's switch: payload destroys home directory if persistence hooks removed before credential rotation. Exfiltration via Session P2P network (indistinguishable from encrypted messaging). First supply chain attack to weaponize AI coding agent configurations as a propagation mechanism.[May-12]Jun 1 update: "Miasma" variant built on same codebase hit 32 @redhat-cloud-services npm packages (Red Hat RHSB-2026-006). TeamPCP/UNC6780 reusing Mini Shai-Hulud + OIDC provenance defeat. Sources: StepSecurity, Socket, Snyk, Wiz (Miasma)
Defensive implication: Patching individual CVEs doesn't break the chain. Defense requires sandboxing (Week 19), intent-based access control (Week 18), and kill switches (Week 21) working together.
Ontological shift — probabilistic vs deterministic security
Traditional security treats vulnerabilities as unintended execution paths in code. AI systems blur the data/instruction boundary at the model level — the very capability that makes LLMs useful (following natural language instructions) is what makes them vulnerable to instruction injection. No equivalent to SQL prepared statements exists for natural language. This is why prompt injection is a class, not a bug, and why defense-in-depth is mandatory.
The three current taxonomies (use all three)
| Framework | Scope | Version | Why |
|---|---|---|---|
| OWASP LLM Top 10 | Single-model LLM apps | 2026 (Aug-2026 reorder) | Ranked by incident data for the first time; 8 of 10 categories moved. Full detail + rationale: week-08 chapter. |
| OWASP Agentic Top 10 (ASI) | Agent systems | 2026 | Now a ranked list — ASI01 (Agent Goal Hijack) → ASI10 (Rogue Agents). ASI prefix replaces earlier "AAI" prefix. |
| MITRE ATLAS | Adversarial TTPs | v2026.09 (15 Sep 2026) | TTPs catalog; 16 tactics / 120 techniques / 88 sub-techniques / 40 mitigations / 73 case studies. Key agent technique IDs: AML.T0100 (AI Agent Clickbait), AML.T0110 (AI Agent Tool Poisoning), AML.T0096 (AI Service API for C2). Case study AML.CS0042 (SesameOp — OpenAI Assistants API as backdoor C2 channel). |
OWASP LLM Top 10 (2026) — current numbering, only LLM01/LLM02 held their spots:
- LLM01 Prompt Injection · 2. LLM02 Sensitive Information Disclosure ·
- LLM03 Excessive Agency
[UP from LLM06]· 4. LLM04 Data and Model Poisoning · - LLM05 Supply Chain · 6. LLM06 Unbounded Consumption
[UP from LLM10]· - LLM07 Misinformation
[UP from LLM09]· - LLM08 Hidden Context Exposure
[renamed from System Prompt Leakage, broadened]· - LLM09 Vector and Embedding Weaknesses ·
- LLM10 Improper Output Handling
[DOWN from LLM05]
Full rationale for the reorder (incident-data blend, 0.75 expert/0.25 data) and the companion OWASP Agentic Top 10 2026 list: week-08 chapter.
OWASP AI Exchange — the umbrella project that houses the LLM Top 10, Agentic Top 10 (ASI), and the older ML Security Top 10 (2023). The ML Top 10 covers traditional ML attacks (adversarial examples, data poisoning, model theft) at the scikit-learn/PyTorch level — distinct from LLM-specific prompt injection. All three lists share the AI Exchange taxonomy.
Sources: OWASP AI Exchange · OWASP LLM Top 10 · OWASP Agentic AI Threats & Mitigations · MITRE ATLAS
Governance stack (Phase 4)
NIST AI RMF (risk-thinking) ←→ ISO/IEC 42001 (operational AIMS controls) ←→ EU AI Act (binding law). Each layer feeds the next: NIST identifies, ISO operationalizes, EU AI Act demonstrates legal compliance.
NIST Cyber AI Profile (IR 8596, draft Dec 2025). Overlays
AI-specific considerations onto CSF 2.0 — three focus areas: securing
AI systems, AI-enabled cyber defense, thwarting AI-enabled attacks.
Subcategories tiered High/Moderate/Foundational. Companion effort:
SP 800-53 COSAiS (Control Overlays for Securing AI Systems) —
implementation-level guidance. 6,500+ community contributors.
Final release expected mid-2026. [Apr-29 research]
U.S. AI Security Executive Order (Jun 2, 2026). Voluntary model
testing (submit ≤30 days before release), federal AI cybersecurity
benchmarks, "AI cybersecurity clearinghouse" for vulnerability sharing.
Explicitly NOT a mandatory licensing/preclearance regime. Triggered by
Anthropic limiting Mythos Preview release over offensive cyber
capabilities. Sources: White House,
NPR [Jun-11]
EU AI Act enforcement deadlines:
Aug 2, 2026— Commission enforcement powers over GPAI model providers activate; high-risk systems in Annex III must comply. Log retention minimum: 6 months (Articles 19 & 26). Digital Omnibus package could delay Annex III to Dec 2027 (in trilogue).Aug 2, 2027— pre-existing models + regulated products- Penalties up to €35M or 7% of worldwide turnover
Source: EU AI Act 2026 timeline
U.S. state frontier-model law trio — the de-facto national floor. With no binding federal statute (the Jun-2 EO is voluntary), three states now set enforceable frontier-model standards, together ≈40% of the U.S. AI market:
- California SB-53 and New York RAISE Act — first-wave transparency + safety-framework disclosure for large developers.
- Illinois AI Safety Measures Act (SB 315), signed Jul-06, effective
Jan-1-2028— first-in-nation third-party audit mandate + 72h incident reporting + $1M / $3M penalties; targets developers >$500M revenue. Source: Gov. Pritzker newsroom - Colorado Chatbot Safety Act (HB 26-1263), effective
Jan-2027— consumer chatbot disclosure + age checks + $20K/violation (unverified primary — held).
Applicability to Kiya: none of these bind us (revenue/scale thresholds, and Kiya
is a private family system) — but they are the emerging agentic-audit vocabulary
(third-party audit, incident reporting, disclosure) that the Phase-4 governance
self-audit rubrics (TrustX ARC, SovereignPA-Bench) operationalize. [Jul-12 daily-pulse]
Stack-affecting CVEs (our production tools)
These affect tools we actually run — audit priority, not just study.
| CVE | Component | CVSS | Status | Notes |
|---|---|---|---|---|
| CVE-2026-35022 | Claude Code CLI (auth.ts) |
7.8 (9.9 in CI/CD) | Disputed — Anthropic closed as "Informative" | Auth helper values execute via execa(shell:true) in -p mode. Audit all apiKeyHelper/awsAuthRefresh in .claude/settings.json. [Apr-29] |
| CVE-2026-35020 | Claude Code CLI (which.ts) |
8.4 | Disputed | Unsanitized string interpolation into shell. Confirmed on v2.1.91. |
| CVE-2026-35021 | Claude Code CLI (promptEditor.ts) |
7.8 | Disputed | POSIX shell command substitution in editor launch. |
| CVE-2026-33068 | Claude Code CLI (trust dialog) | 7.7 | Fixed in v2.1.53 | .claude/settings.json bypass. We are safe (v2.1.118+). |
| CVE-2025-6514 | mcp-remote | 9.6 | Fixed in v0.1.16 | Arbitrary command exec connecting to untrusted servers. Check if installed. |
| CVE-2025-68143/44/45 | mcp-server-git | High | Fixed | Path bypass + argument injection. Verify our version. |
| CVE-2026-24887 | Claude Code CLI (find command) | High | Fixed in v2.0.72 | Parsing error let attackers bypass user approval prompt via find command; untrusted content in context window → arbitrary system command execution. [May-05] |
| CVE-2026-41686 | anthropic-sdk-typescript | Medium | Fixed in v0.91.1 | BetaLocalFilesystemMemoryTool used 0o666/0o777 defaults; local attackers on shared hosts could read/modify agent memory files. Affects v0.79.0–0.91.0. [May-05] |
| CVE-2026-39861 | Claude Code CLI (sandbox) | 10.0 NVD / 7.7 CNA | Fixed in v2.1.64 | Sandbox escape via symlink following — sandboxed process creates symlink inside workspace pointing outside; unsandboxed write follows symlink → arbitrary file write → RCE. All platforms. [May-12] |
| CVE-2026-31431 | Linux kernel (algif_aead) |
7.8 | Patch available | Copy Fail — 732-byte splice() exploit overwrites page-cache of setuid binaries → deterministic root. Container escape via shared page cache. Affects all kernels since 4.14. Our VPS (6.8.0-106) has module unloaded but autoloadable. Reboot to 6.8.0-111 needed. [May-12] |
| Claude Code deny-rule bypass | Claude Code CLI (bashPermissions.ts) |
High | Fixed in v2.1.90 | Adversa AI: 50+ subcommands in a single Bash call causes deny-rule enforcement to skip entirely (performance cap). Attacker hides malicious command at position 51 in a benign-looking build script. Fixed with tree-sitter parser. [Apr-01, Adversa] |
| Claude Code GitHub Action | claude-code-action |
7.8 (CVSS v4) | Fixed in v2.1.128 + action v1.0.94 | Microsoft Threat Intel (Jun 5): Read tool bypassed env scrubbing → /proc/self/environ exfiltration of ANTHROPIC_API_KEY, GITHUB_TOKEN. Separate [bot] suffix trust bypass (RyotaK, Jun 1). $4,800 bounty. [Jun-05] |
| Claude Code MCP self-approval | Claude Code CLI (claude mcp list/get) |
Hardening | Fixed in v2.1.196 | A cloned repo could self-approve its own .mcp.json servers via committed .claude/settings.json (enableAllProjectMcpServers/enabledMcpjsonServers) → auto-spawn without consent. 2.1.196 ignores those flags in untrusted workspaces; servers stay ⏸ Pending until you trust the folder. VPS on 2.1.183 — behind; update via make update-claude (latest 2.1.197, Jul 1, ships Sonnet 5 default). [Jun-30] |
Source: Phoenix Security, Microsoft Security Blog (Jun 5), Adversa AI
Phase 1: Engineering Foundation (Weeks 1–8)
Before you can attack or defend AI, you need to know how it's built. Phase 1 covers the math, the architecture, and the application layer that everything else attacks.