Skip to content
Phase 4, week 22

Risk Management Frameworks

0 of 48 items done. ~16h49m estimated.

Risk Management Frameworks

Phase 4 · Governance & Advanced Strategy. Phases 2 and 3 taught you to break agents and to defend them one control at a time. This week zooms out to the question a security engineer eventually has to answer to a board, an auditor, or a regulator: how do you know your AI system's risk is being managed — deliberately, repeatably, and provably — rather than patched reactively? The answer is a framework: a shared vocabulary and a set of control objectives that turn "we think it's fine" into "here is our risk posture, mapped to named functions, with evidence."

🎯 Objectives

By the end of this week you can:

  • Explain the four core functions of the NIST AI RMF (Govern, Map, Measure, Manage) and run a lightweight assessment against a real system.
  • Distinguish a voluntary risk framework (NIST AI RMF, SAIF) from a certifiable management standard (ISO/IEC 42001) — and know which one an auditor accepts.
  • Map Google SAIF's controls across the Data / Infrastructure / Model / Application surface and connect them to the AI-specific risks from earlier phases.
  • Read the agentic-AI governance frontier — Five Eyes guidance, the Cisco readiness gap, the IETF MCP-security draft — and locate where the standards are ahead of, or behind, deployed reality.
  • Translate a top-down framework into an agent-level, auditable spec using the 2026 research layer (ARC risk-tiering, AJR/ADP delegation policy, dynamic capability scoping).

The big picture: three kinds of instrument

Newcomers collapse "framework," "standard," and "guidance" into one blur. They are three different instruments with three different jobs, and the most common governance mistake is reaching for the wrong one.

  • A risk framework gives you a vocabulary and a process for reasoning about risk. It is voluntary; you borrow what applies. NIST AI RMF is the canonical example.
  • A management standard is certifiable: an external auditor checks you against fixed clauses and issues a certificate a customer can demand in a contract. ISO/IEC 42001 is the one that matters for AI.
  • Guidance / advisories (Five Eyes, NSA, SAIF) are technical opinion — authoritative, often government-issued, but non-binding. They tell you what good looks like without grading you.

💡 The distinction that trips people up: NIST AI RMF and SAIF describe what to think about; ISO 42001 is the only one here you can be audited and certified against. If a customer asks "are you certified," a framework won't satisfy them — a standard will.

Instrument Type Binding? Scope Use it for
NIST AI RMF 1.0 Risk framework Voluntary Org-wide AI risk A shared risk vocabulary + assessment process
ISO/IEC 42001 Management standard Certifiable AI Management System Auditable operational controls, customer-facing certification
Google SAIF Technical guidance Voluntary ML/AI security controls Mapping controls to AI-specific threats across the stack
Five Eyes agentic guidance Government advisory Non-binding Agentic AI A best-practice baseline for autonomous agents
Int'l AI Safety Report 2026 Scientific consensus None Frontier / systemic The "stacked safety" / defense-in-depth vocabulary regulators borrow

NIST AI RMF — the four functions

The AI RMF (NIST AI 100-1, released Jan 2023) organizes everything around four functions that form a loop, not a checklist (NIST hub):

  • Govern — the cross-cutting culture function: policies, roles, accountability, and risk tolerance. It wraps the other three; without it, Map/Measure/Manage are one-off exercises that decay.
  • Map — establish context and enumerate risks: what is the system for, who is affected, where can it fail. You cannot manage a risk you have not named.
  • Measure — quantify and track what you mapped: metrics, red-team results, evals, drift monitoring.
  • Manage — act on what you measured: prioritize, mitigate, accept, or retire; allocate resources to the highest risks.

The framework is deliberately domain-agnostic, so NIST layers profiles on top — the same four functions, re-populated for a narrower domain. Two matter here:

  • The GenAI Profile (NIST AI 600-1) (Jul 2024, under EO 14110) names 12 risk categories unique to or amplified by generative AI and pins 200+ suggested actions to them, each tagged by actor type (developer / deployer / user): CBRN information, confabulation (confidently-stated false content — the RMF's chosen term over "hallucination"), dangerous/violent/hateful content, data privacy, environmental impact, harmful bias & homogenization, human-AI configuration (over-reliance / automation bias), information integrity (deepfakes/misinfo), information security (it lists direct and indirect prompt injection + data poisoning here), intellectual property, obscene/degrading content, and value-chain & component integration. Read it as the checklist that turns "Map your risks" from a blank page into 12 named buckets.
  • The draft Cyber AI Profile (IR 8596) (preliminary draft Dec 2025, comments to Jan 30 2026) overlays AI onto CSF 2.0 rather than the AI RMF, and is organized around three focus areas that every org eventually meets: Secure (securing the AI you build/adopt), Defend (AI-enabled cyber defense, with human oversight kept in the loop), and Thwart (resisting AI-enabled attacks). Each of CSF's six functions (Govern/Identify/Protect/Detect/Respond/Recover) gets a table of AI-specific considerations with a 1–3 priority per subcategory.
  • IR 8587 — Protecting Tokens and Assertions from Forgery, Theft, and Misuse (NIST + CISA, finalized 15 Sep 2026) is the token-plumbing layer under all of this and the one that speaks directly to agent identity. Its controls — shorter-lived tokens, audience restrictions, cryptographically binding a token to the holder's key (kills replay), CAEP/RISC shared-signals for revocation, and keeping tokens out of logs/CI/build artifacts — apply to AI-agent signed tokens the same as human ones. The load-bearing caveat for us: it explicitly excludes agent authorization and API keys from scope — a compromised token proves authentication, never that the action the agent takes was intended. So IR 8587 hardens the credential but leaves the "token ≠ authorization" gap that policy gateways, purpose-bound permissions, and out-of-band approval must fill (the exact seam Week 21's control patterns cover). Companion: NIST's Feb-2026 NCCoE concept paper on accelerating AI-agent identity & authorization demonstrations.

The Playbook turns each sub-category into suggested actions — voluntary, borrow what applies.

🔑 The RMF's core discipline: Map before you Manage. Most AI incidents trace back to a risk nobody wrote down — an unmapped tool permission, an unmeasured drift, an ungoverned self-modification. The loop only works if Govern holds the other three accountable.

ISO/IEC 42001 — the AI Management System

Where the RMF is a way of thinking, ISO/IEC 42001 is a way of operating — the first certifiable AI Management System (AIMS) standard, built in the same clause structure as ISO 27001 so the two integrate. Implementation guides describe a six-phase roadmap — scope, assessment, objectives, execution, monitoring, improvement — that folds into existing 27001/GDPR programs rather than replacing them (AIGL guide · BSI guide · clause walkthrough). The payoff is external: a certificate you can put in front of a customer or regulator. The mental model to carry: RMF tells you what risks exist; 42001 proves you have a managed system for handling them.

Google SAIF — controls mapped to the stack

SAIF is the ML-specific technical layer. Its six core elements extend familiar security practice into AI: expand strong security foundations to AI, extend detection and response to AI threats, automate defenses, harmonize platform-level controls, adapt controls to AI-specific threats, and contextualize AI risk within surrounding business processes. The framework detail organizes controls across four surfaces — Data · Infrastructure · Model · Application — and its risk catalog maps named threats (prompt injection, data poisoning, rogue actions) to specific mitigations, with a dedicated agent diagram in SAIF 2.0. This is where the abstract frameworks meet the concrete attacks you already studied: SAIF's "rogue actions" is the agentic-misuse class from Phase 3, given a control mapping. (HTB's COAE certification is aligned to SAIF; the new Microsoft SC-500 beta cert covers the same agentic threat-model ground.)

The agentic frontier: where standards lag reality

The frameworks above were written for AI systems; autonomous agents strain them, and 2026 is where the governance world admitted it. Three signals define the gap:

  • Five Eyes, Careful Adoption of Agentic AI Services (May 2026) — the first coordinated multi-government advisory on agentic AI (CISA, NSA + AU/CA/NZ/UK). 28 pages, 23 risks, 100+ best practices, five risk categories (privilege, design/configuration, behavioral, structural, accountability). Its headline verdict: prompt injection is "the most persistent and difficult-to-fix threat — no confirmed solution." Core recommendations: zero trust, defense-in-depth, least privilege, fail-safe by default, "never grant broad/unrestricted access."
  • Cisco, State of AI Security 2026 — the readiness gap in two numbers: 83% of organizations are deploying agentic AI; only 29% feel ready to secure it. This is the single citation to reach for when arguing the control gap is industry-wide, not project-specific.
  • IETF draft-mohiuddin-mcp-security-considerations-00 (Jun 2026) — the first Internet-Draft toward normative MCP security requirements. It catalogues the recurring vuln classes from Week 15, ships an open detector (mcp-safeguard), and names "Protocol Pivoting": an injected instruction crossing from MCP tool-invocation into A2A delegation, turning one poisoned context into lateral movement.

⚠️ Honesty caveat on the IETF draft: it is an individual submission — not IETF-endorsed, no standards-track status, production validation still future work. Treat it as the direction of travel, not a ratified control. The required reading that anchors the whole phase is the International AI Safety Report 2026 (Bengio et al., 100+ authors, 30+ countries), the source of the "stacked safety" defense-in-depth vocabulary regulators now use.

From framework to auditable agent spec

Top-down frameworks stop at the org boundary; they don't tell you how to govern one agent's authority. A wave of 2026 research fills that gap, and it maps almost one-to-one onto the problems in this very system. Read these as the missing spec layer under NIST/ISO:

  • Risk-tiering — TrustX ARC. A repeatable rubric for classifying internally-built agents: 12-dimension scoring, a 5-level autonomy model, and 7 agent types (including a coding-assistant extension) feeding a 3-tier governance output with mapped controls. Grounded in NIST AI RMF / ISO 42001. The tool you run before granting an agent new autonomy.
  • The delegation spec — AJR + ADP. Argues that decisions we bury in prompts and tool schemas are really requirements-level commitments, and names the delegated-autonomy boundary: what may be delegated, under what graduated (not binary) authority, with what oversight, and how control returns. Two artifacts: an Agency Justification Record (is an agent even warranted over a simpler alternative?) and an Agentic Delegation Policy capturing purpose, authority, information access, coordination, assurance, and evolution.
  • Prevention-first least privilege — Dynamic Capability Scoping. The sharpest thesis of the three: "a credential that does not exist in an agent's context cannot be misused regardless of the agent's reasoning or evasion sophistication." Don't hand an agent every tool its role might need — scope per task at admission time via three combined sources: role-based ceilings, a task-context classifier, and policy-derived combination prohibitions. On a 600-prompt synthetic dataset (human agreement κ=0.967), iterating dataset↔policy cut ceiling violations 93% (46 → 3), with an observe-only mode that logs out-of-context requests as a misalignment signal.

🔑 The one rule to carry out of this week: governance is layered, not chosen. Use NIST AI RMF for the vocabulary, ISO 42001 for the auditable system, SAIF for the technical controls, and the agentic research layer (ARC → AJR/ADP → capability scoping) to push least-privilege down to the individual agent. No single instrument is sufficient; the stack is the control.

Applying it to this system

This framework stack is not academic here — it grades our own posture, and the grade is mixed. The bypassPermissions mode this builder runs under directly contradicts the Five Eyes "never grant broad/unrestricted access" rule; our only partial mitigations are workspace confinement and behavioral guardrails. The dynamic-capability-scoping model is the concrete counterweight — narrow which tools even exist per task rather than approving them all. The AJR/ADP pair is the missing spec under our one-way/two-way decision rule, which today lives as CLAUDE.md prose rather than an auditable per-agent policy. The honest self-audit exercise for this week: run ARC's tiering on this multi-agent orchestration, then check where our authority model would fail a 42001 auditor.

📇 Frameworks, profiles & governance-research reference

The lesson above is what to learn. This is the catalog behind it — the primary framework documents plus the 2026 agent-governance research, grouped by where each fits the stack. Folded by default; expand for the detail.

Primary frameworks & standards
Instrument Document Role
NIST AI RMF AI 100-1 · Playbook Four-function risk process + suggested actions
NIST profiles GenAI 600-1 · Cyber AI IR 8596 GenAI-specific and CSF-overlay risk profiles
ISO/IEC 42001 Standard · BSI · AIGL guide Certifiable AI Management System
Google SAIF Details · Risks Control mapping across Data/Infra/Model/App
Safety science Int'l AI Safety Report 2026 "Stacked safety" defense-in-depth vocabulary
Agent-governance research (the spec layer under the frameworks)
  • CAGE-1 — deployment-readiness eval and its companion AGL-1 governance layer — score authority, policy enforcement, retrieval/memory integrity, tool safety, decision replay, and a kill-switch ("can it be stopped before business impact"). A rubric to grade an agent's own governance posture.
  • SovereignPA-Bench — an executable benchmark for user-owned personal agents, scoring 8 axes including privacy, consent, evidence-grounding, manipulation-resistance, and auditability. The closest published rubric to what a family assistant is.
  • LOGOS — a living logic for evolving agent teams — treats every learned prompt/memory/skill/tool as an untrusted release candidate until held-out evidence + human authorization permit promotion, with fail-closed verification and portable audit traces. The maturity target for any agent-self-modification pipeline.
  • Gold Eagle — federal AI-discovered-vulnerability clearinghouse — stood up by a Jun 2 2026 executive order (Treasury, CISA, NSA, ONCD, in voluntary collaboration with AI developers + critical-infra operators; announced Jul 14) to coordinate and deconflict the flood of AI-found bugs — validate, prioritize, and route "actionable remediation information." Notably it reuses CMU's existing VINCE / CERT-CC coordination platform rather than building new pipes; no staffing/funding/selection-criteria detail disclosed. The disclosure pipeline itself becoming an AI-scaling problem; a Phase 4 governance marker to watch.
Framework introductions (video)

Recommended resources0/35

Sign in to tick items off and track your progress.

Show

📖 Core Path

The six essentials for Week 22 — read these to own the framework stack.

📚 Further Reading

Agentic-AI Risk-Tiering & Delegation
  • 📄 Specifying the Delegated-Autonomy Boundary: Requirements Engineering for Agentic AI — arXiv 2607.17225 — names the delegated-autonomy boundary: what may be delegated, under what graduated authority, with what oversight, and how control returns. Two artifacts: an Agency Justification Record (AJR) — is an agent even warranted over a simpler alternative — and an Agentic Delegation Policy (ADP) capturing purpose, authority, information access, coordination, assurance, evolution. The missing spec layer under a one-way/two-way authority rule (~40 min) [Jul-19]
  • 📄 Dynamic Capability Scoping for Enterprise AI Agents — arXiv 2607.22445 (Noyan; ICML 2026 Agents-in-the-Wild) — prevention-first: "a credential that does not exist in an agent's context cannot be misused." Three-source least-privilege ceiling — (1) role-based ceilings, (2) task-context classifier, (3) policy-derived combination prohibitions. 600-prompt synthetic dataset (human κ=0.967); iterating dataset↔policy cut ceiling violations 93% (46→3); observe-only mode logs out-of-context requests as a misalignment signal. The counterweight to over-privilege (~40 min) [Jul-27 daily-pulse]
  • 📄 LOGOS: A Living Logic for AI Agent Teams That Evolve With Humans — arXiv 2607.10878 — self-evolution governance: treats every learned prompt/memory/skill/tool/workflow as an untrusted release candidate until held-out evidence + human-controlled policy + explicit authorization permit promotion; portable auditable event traces with fail-closed verification. The maturity target for an agent-self-modification pipeline (~40 min) [Jul-14 daily-pulse]
AI-Discovered-Vulnerability Governance
  • 🌐 Socket — White House Launches "Gold Eagle" Initiative to Manage Surge in AI-Discovered Vulnerabilities (Jul-14) — Federal clearinghouse to coordinate scanning, validate findings, prioritize patches for AI-discovered vulns across critical infra, federal systems, and open source — response to AI surfacing bugs faster than maintainers can validate/disclose/patch. Reuses CMU's VINCE platform; no staffing/funding detail disclosed. A Phase 4 governance marker — the disclosure pipeline is now itself an AI-scaling problem. Stood up by a Jun 2 2026 executive order (Treasury/CISA/NSA/ONCD); reuses CMU's VINCE / CERT-CC coordination platform; announced Jul 14 (~15 min) [Jul-20 daily-pulse]
NIST AI RMF
  • 🎥 NIST — AI Risk Management Framework Explainer Video — Official video intro to the four core functions (Govern, Map, Measure, Manage) (~15 min)
  • 🌐 NIST AI Risk Management Framework — Hub page: framework, playbook, profiles, crosswalks, use cases; GenAI + critical-infrastructure profiles
  • 📄 NIST AI RMF Playbook — Suggested actions per sub-category; voluntary, borrow what applies (~2h)
  • 📄 NIST AI 600-1 (GenAI Profile) — 12 GenAI risk categories (CBRN, confabulation, info-integrity, info-security incl. direct+indirect prompt injection, value-chain, …) with 200+ suggested actions tagged by actor type; layered on top of AI RMF 1.0
  • 📄 NIST Cyber AI Profile (IR 8596) — AI-specific overlay onto CSF 2.0 (not the AI RMF); three focus areas Secure / Defend / Thwart, each mapped across CSF's six functions with 1–3 priority per subcategory. Preliminary draft Dec 2025, comments to Jan 30 2026 (csrc draft page)
ISO/IEC 42001
AI Governance Frameworks & Advisories
Certifications — AI Security
Google SAIF & International AI Safety

📡 From the Resources feed

  • 🌐 Wiz — From Concept to Context Engine: AI-Powered Data Discovery — DSPM architecture for sensitive-data discovery: small models triage/select at volume, large models do deep analysis, deterministic pattern libraries kept separate from the AI for cost/reliability, sandboxed multi-format parsing (PDF/archive/disk image) — a reusable secure multi-agent design (cheap-triage → expensive-reason, determinism outside the model). (via vendor blog) 📡
  • 📄 OpenClaw Threat Model with MAESTRO (Ken Huang) — worked example applying the MAESTRO seven-layer agentic-AI threat model to a real system (foundation model → data ops → agent framework → deployment → observability → compliance → ecosystem), surfacing attack chains like channel-borne prompt injection and plaintext credential storage with mapped mitigations (via Kiya discovery) 📡
  • 🌐 Orca — Securing Shadow AI: Detecting Unapproved LLMs in Your Cloud — three-step cloud detection (NHI/OAuth mapping, SaaS & CI/CD auditing, read-vs-write risk assessment) for inventorying unapproved AI — key gap: IDE extensions and MCP servers bypass web gateways entirely (via vendor blog) 📡
  • 🌐 Orca — Data Security Posture Management (DSPM) for AI — legacy data tools are blind to embeddings, RAG corpora and model weights; once sensitive data is in the weights, GDPR erasure is impossible without retraining (EU AI Act fines up to €35M) — a four-pillar discover/lineage/govern/remediate framework (via vendor blog) 📡
  • 🌐 JFrog — Why Uniform Governance Fails with Enterprise AI Agents — argues for proportional, artifact-centric governance: treat MCPs, plugins, skills and models as versioned software artifacts with controls scaled to each agent's autonomy and trust boundary, not one uniform policy; cites Gartner's warning that 40% of enterprises will decommission agents by 2027 over binary governance failures (via vendor blog) 📡
  • 🌐 Orca — 2026 State of AI Security Report — telemetry from 1,200+ production orgs — 81% running AI packages have a known vuln (avg CVSS 8.79), exploit availability jumped 250x (0.2%->50.1%) in two years, 99.9% of fixable vulns unpatched, 56% deployed agents without controls; the posture baseline for governance planning (via vendor blog) 📡
  • 🌐 NIST — RFI: Modernizing the National Vulnerability Database in the Age of AI — the federal governance signal that AI is reshaping vuln management: NIST opened a 62-day comment window (Aug 12 → Oct 13, docket NIST-2026-0100) seeking a blueprint for a continuous/contextual/automated NVD, citing AI-assisted discovery/triage/exploitation and a 72% YoY vuln surge (50,340 by Aug) that manual periodic-scan workflows can't scale to — the policy backdrop for risk-based prioritization (CISA-KEV-first triage) (via Kiya discovery) 📡

Study checklist

↪ See roadmap.md → Phase 4 → Week 22

  • Distinguish the three instruments: risk framework (NIST AI RMF) vs certifiable standard (ISO 42001) vs advisory (SAIF/Five Eyes)
  • Run a NIST AI RMF assessment on a mock system — Govern / Map / Measure / Manage, Map-before-Manage
  • Map a mock GenAI system against NIST AI 600-1's 12 risk categories — which buckets fire (confabulation, info-security/prompt-injection, value-chain)?
  • Note IR 8596's Secure / Defend / Thwart triad over CSF 2.0's six functions — which focus area(s) does this system live in?
  • Understand ISO/IEC 42001 AIMS requirements and the six-phase implementation roadmap
  • Complete an AI asset inventory + risk classification exercise
  • Study Google SAIF — 6 core elements; map controls across Data/Infra/Model/App; connect "rogue actions" to Phase 3 attacks
  • Read Five Eyes "Careful Adoption of Agentic AI Services" — assess a bypassPermissions-style agent against "never grant broad/unrestricted access"
  • Skim IETF draft-mohiuddin-mcp-security-considerations-00 — note "Protocol Pivoting"; remember it's an individual submission, not IETF-endorsed
  • Note the Cisco readiness-gap stat (83% deploy / 29% ready) as the industry-wide control-gap citation
  • Read the International AI Safety Report 2026 (Bengio et al.) — "stacked safety" vocabulary
  • Push framework down to the agent: run ARC risk-tiering, draft an AJR/ADP delegation policy, apply prevention-first capability scoping
  • Apply NIST IR 8587 token controls (short-lived + audience-restricted + key-bound) to an agent's credentials — then name the gap it leaves: a valid token ≠ an intended action, so add an out-of-band authorization gate

Study notes

Sign in to take notes.