Risk Management Frameworks
Phase 4 · Governance & Advanced Strategy. Phases 2 and 3 taught you to break agents and to defend them one control at a time. This week zooms out to the question a security engineer eventually has to answer to a board, an auditor, or a regulator: how do you know your AI system's risk is being managed — deliberately, repeatably, and provably — rather than patched reactively? The answer is a framework: a shared vocabulary and a set of control objectives that turn "we think it's fine" into "here is our risk posture, mapped to named functions, with evidence."
🎯 Objectives
By the end of this week you can:
- Explain the four core functions of the NIST AI RMF (Govern, Map, Measure, Manage) and run a lightweight assessment against a real system.
- Distinguish a voluntary risk framework (NIST AI RMF, SAIF) from a certifiable management standard (ISO/IEC 42001) — and know which one an auditor accepts.
- Map Google SAIF's controls across the Data / Infrastructure / Model / Application surface and connect them to the AI-specific risks from earlier phases.
- Read the agentic-AI governance frontier — Five Eyes guidance, the Cisco readiness gap, the IETF MCP-security draft — and locate where the standards are ahead of, or behind, deployed reality.
- Translate a top-down framework into an agent-level, auditable spec using the 2026 research layer (ARC risk-tiering, AJR/ADP delegation policy, dynamic capability scoping).
The big picture: three kinds of instrument
Newcomers collapse "framework," "standard," and "guidance" into one blur. They are three different instruments with three different jobs, and the most common governance mistake is reaching for the wrong one.
- A risk framework gives you a vocabulary and a process for reasoning about risk. It is voluntary; you borrow what applies. NIST AI RMF is the canonical example.
- A management standard is certifiable: an external auditor checks you against fixed clauses and issues a certificate a customer can demand in a contract. ISO/IEC 42001 is the one that matters for AI.
- Guidance / advisories (Five Eyes, NSA, SAIF) are technical opinion — authoritative, often government-issued, but non-binding. They tell you what good looks like without grading you.
💡 The distinction that trips people up: NIST AI RMF and SAIF describe what to think about; ISO 42001 is the only one here you can be audited and certified against. If a customer asks "are you certified," a framework won't satisfy them — a standard will.
| Instrument | Type | Binding? | Scope | Use it for |
|---|---|---|---|---|
| NIST AI RMF 1.0 | Risk framework | Voluntary | Org-wide AI risk | A shared risk vocabulary + assessment process |
| ISO/IEC 42001 | Management standard | Certifiable | AI Management System | Auditable operational controls, customer-facing certification |
| Google SAIF | Technical guidance | Voluntary | ML/AI security controls | Mapping controls to AI-specific threats across the stack |
| Five Eyes agentic guidance | Government advisory | Non-binding | Agentic AI | A best-practice baseline for autonomous agents |
| Int'l AI Safety Report 2026 | Scientific consensus | None | Frontier / systemic | The "stacked safety" / defense-in-depth vocabulary regulators borrow |
NIST AI RMF — the four functions
The AI RMF (NIST AI 100-1, released Jan 2023) organizes everything around four functions that form a loop, not a checklist (NIST hub):
- Govern — the cross-cutting culture function: policies, roles, accountability, and risk tolerance. It wraps the other three; without it, Map/Measure/Manage are one-off exercises that decay.
- Map — establish context and enumerate risks: what is the system for, who is affected, where can it fail. You cannot manage a risk you have not named.
- Measure — quantify and track what you mapped: metrics, red-team results, evals, drift monitoring.
- Manage — act on what you measured: prioritize, mitigate, accept, or retire; allocate resources to the highest risks.
The framework is deliberately domain-agnostic, so NIST layers profiles on top — the same four functions, re-populated for a narrower domain. Two matter here:
- The GenAI Profile (NIST AI 600-1) (Jul 2024, under EO 14110) names 12 risk categories unique to or amplified by generative AI and pins 200+ suggested actions to them, each tagged by actor type (developer / deployer / user): CBRN information, confabulation (confidently-stated false content — the RMF's chosen term over "hallucination"), dangerous/violent/hateful content, data privacy, environmental impact, harmful bias & homogenization, human-AI configuration (over-reliance / automation bias), information integrity (deepfakes/misinfo), information security (it lists direct and indirect prompt injection + data poisoning here), intellectual property, obscene/degrading content, and value-chain & component integration. Read it as the checklist that turns "Map your risks" from a blank page into 12 named buckets.
- The draft Cyber AI Profile (IR 8596) (preliminary draft Dec 2025, comments to Jan 30 2026) overlays AI onto CSF 2.0 rather than the AI RMF, and is organized around three focus areas that every org eventually meets: Secure (securing the AI you build/adopt), Defend (AI-enabled cyber defense, with human oversight kept in the loop), and Thwart (resisting AI-enabled attacks). Each of CSF's six functions (Govern/Identify/Protect/Detect/Respond/Recover) gets a table of AI-specific considerations with a 1–3 priority per subcategory.
- IR 8587 — Protecting Tokens and Assertions from Forgery, Theft, and Misuse (NIST + CISA, finalized 15 Sep 2026) is the token-plumbing layer under all of this and the one that speaks directly to agent identity. Its controls — shorter-lived tokens, audience restrictions, cryptographically binding a token to the holder's key (kills replay), CAEP/RISC shared-signals for revocation, and keeping tokens out of logs/CI/build artifacts — apply to AI-agent signed tokens the same as human ones. The load-bearing caveat for us: it explicitly excludes agent authorization and API keys from scope — a compromised token proves authentication, never that the action the agent takes was intended. So IR 8587 hardens the credential but leaves the "token ≠ authorization" gap that policy gateways, purpose-bound permissions, and out-of-band approval must fill (the exact seam Week 21's control patterns cover). Companion: NIST's Feb-2026 NCCoE concept paper on accelerating AI-agent identity & authorization demonstrations.
The Playbook turns each sub-category into suggested actions — voluntary, borrow what applies.
🔑 The RMF's core discipline: Map before you Manage. Most AI incidents trace back to a risk nobody wrote down — an unmapped tool permission, an unmeasured drift, an ungoverned self-modification. The loop only works if Govern holds the other three accountable.
ISO/IEC 42001 — the AI Management System
Where the RMF is a way of thinking, ISO/IEC 42001 is a way of operating — the first certifiable AI Management System (AIMS) standard, built in the same clause structure as ISO 27001 so the two integrate. Implementation guides describe a six-phase roadmap — scope, assessment, objectives, execution, monitoring, improvement — that folds into existing 27001/GDPR programs rather than replacing them (AIGL guide · BSI guide · clause walkthrough). The payoff is external: a certificate you can put in front of a customer or regulator. The mental model to carry: RMF tells you what risks exist; 42001 proves you have a managed system for handling them.
Google SAIF — controls mapped to the stack
SAIF is the ML-specific technical layer. Its six core elements extend familiar security practice into AI: expand strong security foundations to AI, extend detection and response to AI threats, automate defenses, harmonize platform-level controls, adapt controls to AI-specific threats, and contextualize AI risk within surrounding business processes. The framework detail organizes controls across four surfaces — Data · Infrastructure · Model · Application — and its risk catalog maps named threats (prompt injection, data poisoning, rogue actions) to specific mitigations, with a dedicated agent diagram in SAIF 2.0. This is where the abstract frameworks meet the concrete attacks you already studied: SAIF's "rogue actions" is the agentic-misuse class from Phase 3, given a control mapping. (HTB's COAE certification is aligned to SAIF; the new Microsoft SC-500 beta cert covers the same agentic threat-model ground.)
The agentic frontier: where standards lag reality
The frameworks above were written for AI systems; autonomous agents strain them, and 2026 is where the governance world admitted it. Three signals define the gap:
- Five Eyes, Careful Adoption of Agentic AI Services (May 2026) — the first coordinated multi-government advisory on agentic AI (CISA, NSA + AU/CA/NZ/UK). 28 pages, 23 risks, 100+ best practices, five risk categories (privilege, design/configuration, behavioral, structural, accountability). Its headline verdict: prompt injection is "the most persistent and difficult-to-fix threat — no confirmed solution." Core recommendations: zero trust, defense-in-depth, least privilege, fail-safe by default, "never grant broad/unrestricted access."
- Cisco, State of AI Security 2026 — the readiness gap in two numbers: 83% of organizations are deploying agentic AI; only 29% feel ready to secure it. This is the single citation to reach for when arguing the control gap is industry-wide, not project-specific.
- IETF draft-mohiuddin-mcp-security-considerations-00 (Jun 2026) — the first Internet-Draft toward normative MCP security requirements. It catalogues the recurring vuln classes from Week 15, ships an open detector (
mcp-safeguard), and names "Protocol Pivoting": an injected instruction crossing from MCP tool-invocation into A2A delegation, turning one poisoned context into lateral movement.
⚠️ Honesty caveat on the IETF draft: it is an individual submission — not IETF-endorsed, no standards-track status, production validation still future work. Treat it as the direction of travel, not a ratified control. The required reading that anchors the whole phase is the International AI Safety Report 2026 (Bengio et al., 100+ authors, 30+ countries), the source of the "stacked safety" defense-in-depth vocabulary regulators now use.
From framework to auditable agent spec
Top-down frameworks stop at the org boundary; they don't tell you how to govern one agent's authority. A wave of 2026 research fills that gap, and it maps almost one-to-one onto the problems in this very system. Read these as the missing spec layer under NIST/ISO:
- Risk-tiering — TrustX ARC. A repeatable rubric for classifying internally-built agents: 12-dimension scoring, a 5-level autonomy model, and 7 agent types (including a coding-assistant extension) feeding a 3-tier governance output with mapped controls. Grounded in NIST AI RMF / ISO 42001. The tool you run before granting an agent new autonomy.
- The delegation spec — AJR + ADP. Argues that decisions we bury in prompts and tool schemas are really requirements-level commitments, and names the delegated-autonomy boundary: what may be delegated, under what graduated (not binary) authority, with what oversight, and how control returns. Two artifacts: an Agency Justification Record (is an agent even warranted over a simpler alternative?) and an Agentic Delegation Policy capturing purpose, authority, information access, coordination, assurance, and evolution.
- Prevention-first least privilege — Dynamic Capability Scoping. The sharpest thesis of the three: "a credential that does not exist in an agent's context cannot be misused regardless of the agent's reasoning or evasion sophistication." Don't hand an agent every tool its role might need — scope per task at admission time via three combined sources: role-based ceilings, a task-context classifier, and policy-derived combination prohibitions. On a 600-prompt synthetic dataset (human agreement κ=0.967), iterating dataset↔policy cut ceiling violations 93% (46 → 3), with an observe-only mode that logs out-of-context requests as a misalignment signal.
🔑 The one rule to carry out of this week: governance is layered, not chosen. Use NIST AI RMF for the vocabulary, ISO 42001 for the auditable system, SAIF for the technical controls, and the agentic research layer (ARC → AJR/ADP → capability scoping) to push least-privilege down to the individual agent. No single instrument is sufficient; the stack is the control.
Applying it to this system
This framework stack is not academic here — it grades our own posture, and the
grade is mixed. The bypassPermissions mode this builder runs under directly
contradicts the Five Eyes "never grant broad/unrestricted access" rule; our
only partial mitigations are workspace confinement and behavioral guardrails.
The dynamic-capability-scoping model is the concrete counterweight — narrow
which tools even exist per task rather than approving them all. The AJR/ADP
pair is the missing spec under our one-way/two-way decision rule, which today
lives as CLAUDE.md prose rather than an auditable per-agent policy. The honest
self-audit exercise for this week: run ARC's tiering on this multi-agent
orchestration, then check where our authority model would fail a 42001 auditor.
📇 Frameworks, profiles & governance-research reference
The lesson above is what to learn. This is the catalog behind it — the primary framework documents plus the 2026 agent-governance research, grouped by where each fits the stack. Folded by default; expand for the detail.
Primary frameworks & standards
| Instrument | Document | Role |
|---|---|---|
| NIST AI RMF | AI 100-1 · Playbook | Four-function risk process + suggested actions |
| NIST profiles | GenAI 600-1 · Cyber AI IR 8596 | GenAI-specific and CSF-overlay risk profiles |
| ISO/IEC 42001 | Standard · BSI · AIGL guide | Certifiable AI Management System |
| Google SAIF | Details · Risks | Control mapping across Data/Infra/Model/App |
| Safety science | Int'l AI Safety Report 2026 | "Stacked safety" defense-in-depth vocabulary |
Agent-governance research (the spec layer under the frameworks)
- CAGE-1 — deployment-readiness eval and its companion AGL-1 governance layer — score authority, policy enforcement, retrieval/memory integrity, tool safety, decision replay, and a kill-switch ("can it be stopped before business impact"). A rubric to grade an agent's own governance posture.
- SovereignPA-Bench — an executable benchmark for user-owned personal agents, scoring 8 axes including privacy, consent, evidence-grounding, manipulation-resistance, and auditability. The closest published rubric to what a family assistant is.
- LOGOS — a living logic for evolving agent teams — treats every learned prompt/memory/skill/tool as an untrusted release candidate until held-out evidence + human authorization permit promotion, with fail-closed verification and portable audit traces. The maturity target for any agent-self-modification pipeline.
- Gold Eagle — federal AI-discovered-vulnerability clearinghouse — stood up by a Jun 2 2026 executive order (Treasury, CISA, NSA, ONCD, in voluntary collaboration with AI developers + critical-infra operators; announced Jul 14) to coordinate and deconflict the flood of AI-found bugs — validate, prioritize, and route "actionable remediation information." Notably it reuses CMU's existing VINCE / CERT-CC coordination platform rather than building new pipes; no staffing/funding/selection-criteria detail disclosed. The disclosure pipeline itself becoming an AI-scaling problem; a Phase 4 governance marker to watch.
Framework introductions (video)
- NIST AI RMF explainer · SANS — Five Must-Haves of an AI Governance Framework — the four functions and governance-framework essentials in ~15–30 min each.