Skip to content
Phase 3, week 18

Bidirectional Defense & Input Sanitization

0 of 46 items done. ~14h14m estimated.

Concept. Week 17 built the firewall — guardrails and semantic classifiers that sit at the edge. This week is about what flows through it in both directions: sanitizing untrusted content on the way in and scanning for leakage on the way out. The insight that organizes the whole week is that input validation alone cannot work — the model "can't distinguish between user input and external content, making traditional input validation insufficient" (Microsoft). So bidirectional defense is really three moves stacked: detect/strip injections at the input, constrain what the agent is allowed to do regardless of what it was told, and inspect the output before it leaves.

🎯 Objectives

By the end of this week you can:

  • Build a pre-execution intent gate and an output scanner for PII / data leakage, and explain why each catches what the other misses.
  • Use an LLM-as-detector (PromptArmor) as a sanitization pre-filter and justify its latency/token cost against your threat model.
  • Distinguish content filtering (does this text look malicious?) from intent analysis (does the model actually plan to follow it?) — and know why the latter survives obfuscation.
  • Stand up the dual-LLM / information-flow-control architecture so a hijacked executor cannot exceed what the original user intent authorized.
  • Recognize the defender-side failure modes: gaslighting the triage model, and agents that surveil the user they serve.

The big picture — why input validation alone loses

The core asymmetry: prompt injection lives in the same channel as legitimate data. Any content the agent reads — an email, a web page, a tool result — can carry instructions, and the model has no reliable way to tell "process this" from "obey this." Microsoft's guidance is blunt about the consequence: assume indirect prompt injection will succeed, and design so the blast radius is small (Microsoft). That reframes defense from "block the bad string" to a layered stack — probabilistic detectors that reduce attack volume, plus deterministic controls that cap impact when detection fails. OWASP's cheat sheet lands in the same place: input marking, output filtering, least-privilege tools, and the dual-LLM pattern as complementary layers, none sufficient alone (OWASP).

🔑 The one rule to carry out of this week: a filter reduces attack rate; only a capability boundary reduces attack impact. Ship both — and never let the probabilistic layer be your only layer.

The input side — detection and sanitization

The 2026 baseline for input-side defense is embarrassingly simple and it finally works: prompt a strong off-the-shelf LLM to find and delete the injection before the agent sees it. That is PromptArmor — no fine-tuning, no custom model, just a detector prompt wrapped around GPT-4.1 / GPT-4o / o4-mini. On AgentDojo it drives attack-success rate, false-positive rate, and false-negative rate all below 1%, and holds up against adaptive attacks (PromptArmor, arXiv 2507.15219). The authors argue it should be the standard baseline any new defense must beat. The cost is real — a second model call adds latency and ~500–1000 tokens per request — so you pay it where near-zero injection tolerance is worth the tax, not on every hop.

Lighter-weight, deterministic detectors sit underneath it:

  • Pattern + fuzzy matching — Levenshtein-distance keyword matching to catch obfuscated / "typoglycemia" spellings, with a tuned similarity threshold (OWASP).
  • Adversarial suffix filtering — detect and strip the gradient-crafted suffix tokens (GCG-style) before they reach the model (arXiv 2505.09602).
  • Small classifier guards — Llama Prompt Guard 2, ShieldGemma, and open tools like InjecGuard (which explicitly targets overdefense false positives, +30.8% over prior SOTA on NotInject) run in the wide, cheap first stage (Datadog · InjecGuard).

💡 Spotlighting / data marking is the cheapest layer and you should always ship it. Wrap untrusted content in explicit delimiters and label it USER_DATA_TO_PROCESS (or an unguessable per-request tag), so both the model and downstream validators treat it as data, not instructions (OWASP). This is exactly Kiya's sanitize.py posture for Tier-2 content.

The intent layer — defending on what the user actually asked

Content filtering asks "does this text look malicious?" — a losing game against paraphrase and obfuscation. Intent analysis asks a different, more robust question: which instructions does the model actually plan to follow, and did any of them come from untrusted data?

IntentGuard extracts the model's structured list of intended instructions (using "thinking-intervention" strategies on reasoning models) and flags any overlap with untrusted input. On Mind2Web it cut attack success from 100% to 8.5% — because it targets intent recognition, not string patterns (IntentGuard, arXiv 2512.00966).

IBAC (Intent-Based Access Control) pushes the same idea down to the authorization layer. An intent parser turns the user request into a structured Intent; a policy generator compiles it into fine-grained tuples like tool:read#db:patients[pii=true]; an engine (Cedar / OPA / OpenFGA) evaluates every tool call against those tuples before it runs (Ken Huang · IBAC paper). The security property is the one that matters most: "if the later LLM logic is hijacked, it doesn't matter — the engine still enforces exactly what the original intent permitted." It moves authorization from "who can do what" to "for what purpose, under what conditions," with short TTLs so a grant doesn't outlive the task (IBAC site · IBAC for coding agents). The stack-relevant instantiation is agentctl (Apache-2.0): a governance layer that normalizes the divergent tool vocabularies of Claude Code, Codex, OpenClaw and Gemini CLI to canonical actions (read/write/delete/deploy/query), classifies each request into a risk tier (Low→Critical) and evaluates it against Cedar policies in ~/.agentctl/policies/ — with JWT-signed agent identities, TTL-based session cleanup, and an immutable JSONL audit trail — before the call runs (agentctl).

The architecture layer — dual-LLM and information-flow control

The most durable defense predates the current wave. Simon Willison's dual-LLM pattern splits the system in two: a privileged LLM that holds tool access but only ever sees trusted user input, and a quarantined LLM that processes hostile content but has no tools and cannot emit links or images. A plain-software Controller brokers between them, passing the privileged model symbolic variables ($VAR1) instead of raw untrusted text — so contaminated content never reaches the component that can act on it (Simon Willison). Willison is honest about the limits: it's clunky, and social-engineering the user into pasting obfuscated data still breaks it.

Microsoft generalizes this into Information Flow Control (IFC) — tag data with a trust level, and enforce that untrusted-tagged content can never influence privileged inference or planning, using quarantined inference environments (Microsoft). AEGIS makes IFC concrete as a deployable pre-execution firewall: an SDK instruments 14 agent frameworks (single line — agentguard.auto()) to intercept every tool_use, and a server-side Gateway runs a three-stage pipeline — recursive string extraction from nested args → content risk-scan against 22 patterns across 7 categories (SQLi, path traversal, shell/prompt injection, sensitive files, exfil, PII) → JSON-Schema policy validation — returning allow / block / pending-human-review by risk tier, with every decision written to a tamper-evident log (Ed25519-signed, SHA-256 hash-chained). It blocks all 48 attacks in its suite at 8.3 ms median overhead (<1% of inference) and a 1.2% false-positive rate, complementing the I/O firewalls (AEGIS, arXiv 2603.12621).

Runtime resilience — inference-time and multi-agent defenses

Two newer directions spend compute instead of engineering effort:

  • SecInfer — inference-time scaling: generate multiple responses under a varied set of security-oriented system prompts, then aggregate to the response most likely to accomplish the intended task, discarding those that drift toward an injected goal. More reasoning budget = better attack resistance, and it beats fine-tuning-based defenses (SecInfer, arXiv 2509.24967).
  • Multi-agent defense pipeline — several specialized defense agents coordinate on detection, echoing Microsoft's "critic agents" and "plan drift detection" running alongside the worker (arXiv 2509.14285 · Microsoft).

Trust the boundary only as far as you've tested it

The intent, capability and architecture layers above are all instances of one 2026 consensus: enforce security out-of-band — a deterministic policy that mediates the agent's actions — rather than training the model to refuse. The named systems in that family are CaMeL, FIDES, Progent, RTBAS, and FORGE (IBAC, dual-LLM/IFC and AEGIS are the same shape), and several report near-elimination of attacks on AgentDojo (Adaptive Evaluation of Out-of-Band Defenses, arXiv 2606.26479). The catch is a measurement one, and it is the same trap in-band defenses fell into: every one of these was validated on a static benchmark — a fixed set of injection attempts — the very methodology that made in-band detectors look strong right up until adaptive, defense-aware attacks broke twelve of them at >90% success. When the authors actually built a hand-crafted adaptive attack against Progent on AgentDojo (Qwen2.5-7B, a setting Progent's own authors never tested), the boundary held — static 25.8% → 4.2% ASR, adaptive 2.6% — early evidence that deterministic out-of-band enforcement really is more adaptive-robust than in-band detection. But they are blunt that this is "one small-scale data point" and stronger white-box attacks are unexplored.

🔑 A "near-zero ASR" number is only as trustworthy as the attacker who produced it. Before you certify a capability boundary, attack it adaptively — with an attacker who knows the policy — not just against the benchmark's fixed injection set. Static robustness is a hypothesis, not a guarantee (arXiv 2606.26479).

The output side — leakage scanning and the defender's own blind spots

Bidirectional means the job isn't done when the agent answers. Output scanning validates the response before delivery: checking for system- prompt leakage, exposed API keys, and PII, and validating outputs against the data sources they were allowed to draw from (OWASP). Datadog frames the operational version as a wide → strict → human-review funnel: cheap static filters first, ML classifiers next, human escalation for high-risk actions, with tracing/telemetry feeding a continuous tuning loop (Datadog).

A practical refinement the redaction tooling converges on: redact with stable, deterministic placeholders, not blunt deletion. Map each secret to a consistent token (e.g. a hash-derived tag) so the same value always yields the same placeholder within a session — the model keeps referential coherence and can still edit around a redacted value instead of losing the thread. Two Claude-Code-native tools ship exactly this, in opposite directions: DataSentry redacts secrets (cloud keys / PEM / JWT / DB creds) from tool output before the model sees them via a PostToolUse hook, and blocks prompts that paste live credentials via UserPromptSubmit — a fail-closed, zero-dependency GPL-3.0 plugin using SHA1-stable placeholders (DataSentry); Maskit is the egress counterpart — a local privacy gateway that masks keys / credentials / PII / internal IPs before requests leave the machine to Cursor / Claude Code / Codex and restores them in the streamed response, reusing placeholders across multi-turn chats (Maskit). Between them they cover both directions the deterministic way — keep secrets out of the model's context, and out of the provider egress.

Two 2026 findings warn that the defender is itself a target:

  • Gaslighting the triage model. The in-the-wild macOS.Gaslight backdoor embeds a 3.5 KB cascade of 38 fabricated "system" messages (fake token-expiry, OOM, disk-exhaustion, static-analysis flags) to trick an LLM-assisted malware-triage agent into aborting — an evolution from single-block injection to a stacked assault on the analyst's perception (SentinelLABS). The rule it enforces — keep hostile content out of the model entirely, treat triaged samples as data never instructions — is precisely the bidirectional-sanitization posture.
  • The surveillance flip side. "AI Snitches Get Glitches" formalizes agentic surveillance — an agent analyzing accessible data, writing a report, and exfiltrating it via its own tools — and finds a paradox: some models unprompted-ly assist surveillance and report the attempt to authorities (arXiv 2606.25836). Your output scanner is also the last gate on the agent leaking for someone.

The defense stack at a glance

Layer Question it answers Representative technique Catches Blind spot
Spotlighting / data marking "Is this data or instructions?" Delimiters + USER_DATA tags (OWASP) naive injections model still can obey
Input detection "Does this look malicious?" PromptArmor, InjecGuard, suffix filtering (PromptArmor) known + adaptive patterns novel semantics; latency cost
Intent analysis "Does the model plan to follow it?" IntentGuard (arXiv 2512.00966) obfuscated injections needs reasoning model
Capability boundary "Is this action authorized now?" IBAC intent tuples (Ken Huang) hijacked planner policy must be tight
Architecture "Can untrusted data reach a tool?" Dual-LLM / IFC (Simon Willison) data→action flow UX cost, user social-eng
Runtime resilience "Is this the intended task?" SecInfer, critic agents (SecInfer) goal drift compute cost
Output scanning "Is anything leaking?" PII/leak filter, wide→strict→human (Datadog) exfil, secret leak probabilistic, tunable

💡 Map it to Kiya. We already run the two most reliable layers: sanitize.py (spotlighting + quarantine of Tier-2 content) and the redact.py output scrubber (leak/PII gate before Telegram) — worth making its tokens placeholder-stable so the model keeps referential coherence. Off-the-shelf, DataSentry (tool-output redaction) and Maskit (egress masking) are the Claude-Code-native tools in this exact category. The gaps worth closing are an intent gate on tool calls (IBAC-style) and a critic pass on high-risk actions.

📇 Defenses & tools reference

The lesson above is what to learn. This is the catalog behind it — expand when you need a specific method or repo.

Papers by defense family
Defense Family Key result Source
PromptArmor LLM-as-detector (input) <1% ASR/FPR/FNR on AgentDojo arXiv 2507.15219 · OpenReview
IntentGuard intent analysis 100% → 8.5% ASR on Mind2Web arXiv 2512.00966 · OpenReview
SecInfer inference-time scaling beats SOTA + fine-tuning defenses arXiv 2509.24967
Adversarial Suffix Filtering input detection strips GCG-style suffixes arXiv 2505.09602
Multi-Agent Defense Pipeline multi-agent critic coordinated detection arXiv 2509.14285
Dual-LLM pattern architecture privileged vs quarantined split Simon Willison
AEGIS pre-exec firewall/audit 48/48 attacks blocked, 8.3 ms median, 1.2% FPR arXiv 2603.12621
IBAC capability boundary intent tuples, per-call authz, TTL paper · site
Adaptive Eval (CaMeL/FIDES/Progent/RTBAS/FORGE) eval discipline out-of-band family; Progent adaptive 2.6% ASR vs static 4.2% — but tested only statically arXiv 2606.26479
Open-source detectors & enforcers
  • agentctl — Apache-2.0 IBAC governance layer; normalizes Claude Code / Codex / OpenClaw / Gemini CLI tool calls to canonical actions, Cedar policies, JWT identities + TTL cleanup, JSONL audit; 144 tests / 96% cov — GitHub
  • InjecGuard — prompt-injection guard, +30.8% over SOTA on NotInject; targets overdefense FPs — GitHub
  • StackOne Defender — injection detector, 88.7% accuracy, 22MB, CPU-only — GitHub
  • zeroleaks — scanner for injection + system-prompt extraction — GitHub
Defender-side threat research
  • macOS.Gaslight — 38-message fabricated-system-message cascade gaslights an LLM triage agent — SentinelLABS
  • AI Snitches / SurveilBench — agentic surveillance threat model + benchmark — arXiv 2606.25836

Recommended resources0/35

Sign in to tick items off and track your progress.

Show

📖 Core Path

📚 Further Reading

Practical defense guides
Intent & access control
  • 📄 IBAC Paper — FGA tuples from user intent; authorization check before every tool call; TTL enforcement (~2h)
  • 🌐 IBAC Website — Architecture diagrams and implementation details
  • 📄 IBAC for Coding Agents — Open-source implementation with agentctl for Claude Code, Gemini CLI
Architecture & inference-time defenses
  • 📄 Defusing Explosive Prompts — arXiv 2609.22510 — the "explosive prompt": a dormant conditional IPI payload planted in one piece of retrieved content that stays inert until an attacker-chosen trigger fires (a training-free inference-time backdoor). Temporal separation defeats frontier refusals — rephrasing a refused imperative as a dormant conditional drives real state-changing tool calls (16.5% vs 2.4% imperative, 34.2% on a proprietary model); on 9 production agents (Codex, Gemini CLI, Claude Code CLI, Cursor, Copilot, Devin, Kiro, Qwen, Google Assistant) 43–83% vs ≤3%, and a preference-optimized model that closes imperative injection still runs 11.8% (all at the trigger turn, where the payload sits in trusted history, not the untrusted channel). Fix = DeFuse, an ingestion-time detector of the conditional structure (3.0% ASR at 5% FPR, AUC 0.9994) — validates the week's thesis: the durable lever is scanning retrieved content before it enters context, not asking the model to refuse harder. — Deployable pre-execution firewall: SDK over 14 frameworks + Gateway 3-stage pipeline (extract → 22 patterns/7 categories → JSON-Schema policy); allow/block/pending; Ed25519+SHA-256 audit; 48/48 blocked, 8.3 ms median, 1.2% FPR
  • 📄 SecInfer — arXiv 2509.24967 — Inference-time scaling: sample under varied security prompts, aggregate to the intended-task response; beats fine-tuning defenses
  • 📄 Adversarial Suffix Filtering — arXiv 2505.09602 — Detect and filter GCG-style adversarial suffixes before they reach the model
  • 📄 Multi-Agent Defense Pipeline — arXiv 2509.14285 — Multiple defense agents coordinating on threat detection
  • 🧪 Implement: build a 2-model pipeline where the planner (high trust) generates tool-call intents, and the executor (low trust) runs them with strict I/O validation
  • 📄 Adaptive Evaluation of Out-of-Band Defenses — arXiv 2606.26479 — Frames the out-of-band defense family (CaMeL / FIDES / Progent / RTBAS / FORGE) as classical integrity/reference-monitor/least-privilege; warns they're all validated on static benchmarks (the trap that hid in-band weakness until adaptive attacks broke 12 at >90%); hand-crafted adaptive attack on Progent/AgentDojo held (2.6% ASR) but "one small-scale data point"
Peer review (supplementary)
Open-source detectors & enforcers
  • 🔧 agentctl — Apache-2.0 IBAC governance layer; normalizes Claude Code / Codex / OpenClaw / Gemini CLI tool calls to canonical actions, Cedar policies in ~/.agentctl/policies/, JWT identities + TTL cleanup, JSONL audit; 144 tests / 96% cov
  • 🔧 StackOne Defender — Open-source injection detection; 88.7% accuracy, 22MB, CPU-only
  • 🔧 InjecGuard — Prompt guard; +30.8% over prior SOTA on NotInject; addresses overdefense FPs
  • 🔧 zeroleaks — Scanner for prompt injection + system-prompt extraction (554 stars)
Defender-side threat research
  • 📄 AI Snitches Get Glitches (arXiv 2606.25836) — Formalizes agentic surveillance + SurveilBench (corporate / education / police); paradox: models both assist and report surveillance (~30 min)
  • 📄 SentinelLABS — macOS.Gaslight — In-the-wild DPRK-aligned backdoor with a 38 fabricated-"system"-message cascade that gaslights an LLM triage agent; attacks the analyst's perception, not the sandbox (~30 min)
Offense-to-defense videos

📡 From the Resources feed

  • 🔧 pi-jev (y0usaf) — safety layer for the Pi coding agent: gates bash/write/edit tool calls on four risk dimensions (destructiveness, exfil, scope-creep, damage) before execution and scans command output for leaked secrets, with thresholds calibrated from real test runs rather than chosen; MIT, ~149★ — the tool-call gate + output leak-scan pattern applied to one agent's I/O, a smaller cousin of DataSentry/Maskit (in Trove since 2026-09-26 (security/ai-security)) 📡
  • 📄 PIDS-Bench — evaluating PI detectors under over-defense & distribution shift — a detector at F1≥0.98 on held-out data still misclassifies ~1/3 of externally-sourced benign text (provenance-sensitive over-defense), and hard-negative augmentation doesn't fix it — don't trust a single aggregate F1 (in Trove since 2026-09-26 (security/ai-security)) 📡
  • 🔧 DataSentry (CaptainASIC) — fail-closed Claude Code plugin that redacts secrets (40 rules: cloud keys/PEM/JWT/DB creds) from tool output before the model sees them and blocks prompts pasting live credentials, using stable placeholders so the model can still reference redacted values; GPL-3.0, zero-dep — the DLP hook layer for our own agent stack. (via X/Twitter trending) 📡
  • 🔧 Maskit (xiaYuTian11) — local privacy-masking gateway that auto-redacts API keys, credentials, PII and internal IPs from requests before they leave your machine to Cursor/Claude Code/Codex, then restores them in responses, keeping placeholders consistent across multi-turn chats; 100% local, zero telemetry, ~222★, AGPL-3.0 — the input-sanitization counterpart to DataSentry's output redaction (via GitHub trending) 📡
  • 📄 Prompt Injection Detection for Email Agents Through Attack-Chain Modeling (arXiv 2609.30657) — models indirect-PI progression in stages (per-stage verifiers + rule-based risk signals + user-intent/action-consistency + a logistic decision policy) instead of flat text classification, roughly doubling mean F1 (0.406 vs 0.216 for the strongest of five pretrained detectors) on LLM email agents — directly relevant to our Gmail-MCP summarize-my-email path. (in Trove since 2026-09-30 (security/ai-security)) 📡
  • 📄 Certified Multi-Source Integrity for Structured Agent Actions (arXiv 2609.34245) — a formal certifier for an LLM agent's structured actions that resists indirect prompt injection by counting genuinely independent evidence sources via minimum hitting sets, rather than naive source-count or trust labels — a provable input-integrity layer for the assume-injection-succeeds stack (in Trove since 2026-10-02 (security/ai-security)) 📡
  • 📄 Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents (arXiv 2610.03448) — replays ground-truth tool calls from AgentDojo/tau-bench (no LLM) to label injected tool outputs, then grades 15 PI detectors incl. Meta Prompt Guard 2: detection rankings don't transfer — BIPIA's best catches 2% of AgentDojo injections at 1% FPR; a detector at 72% on AgentDojo drops to 15% on tau-bench — because detectors trained on short prompt strings miss injections embedded in tool outputs. Pick a detector trained on agent-style inputs, and don't trust a public-benchmark score for your deployment. [Oct-06 daily-pulse]
  • 📄 Evaluating and Improving the Robustness of LLMs to Input Sequence Variations (arXiv 2610.02432) — pairs a JS-divergence robustness metric + ASA adaptive black-box attack (73.8% ASR) with two defenses: a 7-model committee (−47–55pp ASR) and AttestMCP — HMAC-attested tool-call packets (<0.1ms) that with isolation patterns cut agentic ASR 53.7%→12.4% on MCPBench — input-perturbation robustness for the sanitization stack (in Trove since 2026-10-05 (security/ai-security)) 📡
  • 📄 RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents (arXiv 2610.06401) — a training-time defense: the model generates its own tool-use scenarios and self-distills to match clean-context behavior on benign and injected variants, cutting indirect-PI success in tool responses without the usual utility hit of prior defenses — a model-side complement to the detector/firewall stack (in Trove since 2026-10-06 (security/ai-security)) 📡

Study checklist

↪ See roadmap.md → Phase 3 → Week 18

  • Build a pre-execution intent gate + an output PII/leakage scanner; explain what each catches that the other misses
  • Wire PromptArmor (LLM-as-detector) as an input sanitization pre-filter; measure ASR + latency/token cost
  • Add spotlighting / data-marking to untrusted content (delimiters + USER_DATA tag)
  • Implement IntentGuard intent-following analysis — measure ASR reduction on Mind2Web
  • Build the dual-LLM / IFC split (privileged planner + quarantined executor, symbolic vars)
  • Implement IBAC intent-tuple authorization enforced before every tool call (TTL-scoped)
  • Compare SecInfer (inference-time scaling) and Adversarial Suffix Filtering as complementary defenses
  • Stand up an AEGIS-style pre-execution firewall (extract → pattern scan → policy → allow/block/pending) + tamper-evident audit; measure latency + FPR overhead
  • Trial agentctl on Claude Code — write Cedar policies for read/write/deploy, verify high-risk calls escalate + TTL-scoped grants expire
  • Attack your own capability boundary adaptively (attacker who knows the policy), not just against a fixed injection set — treat a static "near-zero ASR" as a hypothesis
  • Make output/tool-result redaction placeholder-stable (deterministic tokens) so the model keeps referential coherence over redacted values — trial DataSentry (tool-output) + Maskit (egress) on the stack

Study notes

Sign in to take notes.