Concept. Week 17 built the firewall — guardrails and semantic classifiers that sit at the edge. This week is about what flows through it in both directions: sanitizing untrusted content on the way in and scanning for leakage on the way out. The insight that organizes the whole week is that input validation alone cannot work — the model "can't distinguish between user input and external content, making traditional input validation insufficient" (Microsoft). So bidirectional defense is really three moves stacked: detect/strip injections at the input, constrain what the agent is allowed to do regardless of what it was told, and inspect the output before it leaves.
🎯 Objectives
By the end of this week you can:
- Build a pre-execution intent gate and an output scanner for PII / data leakage, and explain why each catches what the other misses.
- Use an LLM-as-detector (PromptArmor) as a sanitization pre-filter and justify its latency/token cost against your threat model.
- Distinguish content filtering (does this text look malicious?) from intent analysis (does the model actually plan to follow it?) — and know why the latter survives obfuscation.
- Stand up the dual-LLM / information-flow-control architecture so a hijacked executor cannot exceed what the original user intent authorized.
- Recognize the defender-side failure modes: gaslighting the triage model, and agents that surveil the user they serve.
The big picture — why input validation alone loses
The core asymmetry: prompt injection lives in the same channel as legitimate data. Any content the agent reads — an email, a web page, a tool result — can carry instructions, and the model has no reliable way to tell "process this" from "obey this." Microsoft's guidance is blunt about the consequence: assume indirect prompt injection will succeed, and design so the blast radius is small (Microsoft). That reframes defense from "block the bad string" to a layered stack — probabilistic detectors that reduce attack volume, plus deterministic controls that cap impact when detection fails. OWASP's cheat sheet lands in the same place: input marking, output filtering, least-privilege tools, and the dual-LLM pattern as complementary layers, none sufficient alone (OWASP).
🔑 The one rule to carry out of this week: a filter reduces attack rate; only a capability boundary reduces attack impact. Ship both — and never let the probabilistic layer be your only layer.
The input side — detection and sanitization
The 2026 baseline for input-side defense is embarrassingly simple and it finally works: prompt a strong off-the-shelf LLM to find and delete the injection before the agent sees it. That is PromptArmor — no fine-tuning, no custom model, just a detector prompt wrapped around GPT-4.1 / GPT-4o / o4-mini. On AgentDojo it drives attack-success rate, false-positive rate, and false-negative rate all below 1%, and holds up against adaptive attacks (PromptArmor, arXiv 2507.15219). The authors argue it should be the standard baseline any new defense must beat. The cost is real — a second model call adds latency and ~500–1000 tokens per request — so you pay it where near-zero injection tolerance is worth the tax, not on every hop.
Lighter-weight, deterministic detectors sit underneath it:
- Pattern + fuzzy matching — Levenshtein-distance keyword matching to catch obfuscated / "typoglycemia" spellings, with a tuned similarity threshold (OWASP).
- Adversarial suffix filtering — detect and strip the gradient-crafted suffix tokens (GCG-style) before they reach the model (arXiv 2505.09602).
- Small classifier guards — Llama Prompt Guard 2, ShieldGemma, and open tools like InjecGuard (which explicitly targets overdefense false positives, +30.8% over prior SOTA on NotInject) run in the wide, cheap first stage (Datadog · InjecGuard).
💡 Spotlighting / data marking is the cheapest layer and you should always ship it. Wrap untrusted content in explicit delimiters and label it
USER_DATA_TO_PROCESS(or an unguessable per-request tag), so both the model and downstream validators treat it as data, not instructions (OWASP). This is exactly Kiya'ssanitize.pyposture for Tier-2 content.
The intent layer — defending on what the user actually asked
Content filtering asks "does this text look malicious?" — a losing game against paraphrase and obfuscation. Intent analysis asks a different, more robust question: which instructions does the model actually plan to follow, and did any of them come from untrusted data?
IntentGuard extracts the model's structured list of intended instructions (using "thinking-intervention" strategies on reasoning models) and flags any overlap with untrusted input. On Mind2Web it cut attack success from 100% to 8.5% — because it targets intent recognition, not string patterns (IntentGuard, arXiv 2512.00966).
IBAC (Intent-Based Access Control) pushes the same idea down to the
authorization layer. An intent parser turns the user request into a
structured Intent; a policy generator compiles it into fine-grained tuples
like tool:read#db:patients[pii=true]; an engine (Cedar / OPA / OpenFGA)
evaluates every tool call against those tuples before it runs
(Ken Huang · IBAC paper).
The security property is the one that matters most: "if the later LLM
logic is hijacked, it doesn't matter — the engine still enforces exactly
what the original intent permitted." It moves authorization from "who can
do what" to "for what purpose, under what conditions," with short TTLs so a
grant doesn't outlive the task (IBAC site · IBAC for coding agents).
The stack-relevant instantiation is agentctl (Apache-2.0): a governance
layer that normalizes the divergent tool vocabularies of Claude Code, Codex,
OpenClaw and Gemini CLI to canonical actions (read/write/delete/deploy/query),
classifies each request into a risk tier (Low→Critical) and evaluates it against
Cedar policies in ~/.agentctl/policies/ — with JWT-signed agent identities,
TTL-based session cleanup, and an immutable JSONL audit trail — before the call
runs (agentctl).
The architecture layer — dual-LLM and information-flow control
The most durable defense predates the current wave. Simon Willison's
dual-LLM pattern splits the system in two: a privileged LLM that
holds tool access but only ever sees trusted user input, and a
quarantined LLM that processes hostile content but has no tools and
cannot emit links or images. A plain-software Controller brokers
between them, passing the privileged model symbolic variables ($VAR1)
instead of raw untrusted text — so contaminated content never reaches the
component that can act on it (Simon Willison).
Willison is honest about the limits: it's clunky, and social-engineering
the user into pasting obfuscated data still breaks it.
Microsoft generalizes this into Information Flow Control (IFC) —
tag data with a trust level, and enforce that untrusted-tagged content can
never influence privileged inference or planning, using quarantined
inference environments (Microsoft).
AEGIS makes IFC concrete as a deployable pre-execution firewall: an SDK
instruments 14 agent frameworks (single line — agentguard.auto()) to intercept
every tool_use, and a server-side Gateway runs a three-stage pipeline —
recursive string extraction from nested args → content risk-scan against 22
patterns across 7 categories (SQLi, path traversal, shell/prompt injection,
sensitive files, exfil, PII) → JSON-Schema policy validation — returning
allow / block / pending-human-review by risk tier, with every decision
written to a tamper-evident log (Ed25519-signed, SHA-256 hash-chained). It blocks
all 48 attacks in its suite at 8.3 ms median overhead (<1% of inference)
and a 1.2% false-positive rate, complementing the I/O firewalls (AEGIS, arXiv 2603.12621).
Runtime resilience — inference-time and multi-agent defenses
Two newer directions spend compute instead of engineering effort:
- SecInfer — inference-time scaling: generate multiple responses under a varied set of security-oriented system prompts, then aggregate to the response most likely to accomplish the intended task, discarding those that drift toward an injected goal. More reasoning budget = better attack resistance, and it beats fine-tuning-based defenses (SecInfer, arXiv 2509.24967).
- Multi-agent defense pipeline — several specialized defense agents coordinate on detection, echoing Microsoft's "critic agents" and "plan drift detection" running alongside the worker (arXiv 2509.14285 · Microsoft).
Trust the boundary only as far as you've tested it
The intent, capability and architecture layers above are all instances of
one 2026 consensus: enforce security out-of-band — a deterministic
policy that mediates the agent's actions — rather than training the model to
refuse. The named systems in that family are CaMeL, FIDES, Progent, RTBAS,
and FORGE (IBAC, dual-LLM/IFC and AEGIS are the same shape), and several
report near-elimination of attacks on AgentDojo (Adaptive Evaluation of
Out-of-Band Defenses, arXiv 2606.26479).
The catch is a measurement one, and it is the same trap in-band defenses
fell into: every one of these was validated on a static benchmark — a
fixed set of injection attempts — the very methodology that made in-band
detectors look strong right up until adaptive, defense-aware attacks broke
twelve of them at >90% success. When the authors actually built a
hand-crafted adaptive attack against Progent on AgentDojo (Qwen2.5-7B,
a setting Progent's own authors never tested), the boundary held — static
25.8% → 4.2% ASR, adaptive 2.6% — early evidence that deterministic
out-of-band enforcement really is more adaptive-robust than in-band
detection. But they are blunt that this is "one small-scale data point" and
stronger white-box attacks are unexplored.
🔑 A "near-zero ASR" number is only as trustworthy as the attacker who produced it. Before you certify a capability boundary, attack it adaptively — with an attacker who knows the policy — not just against the benchmark's fixed injection set. Static robustness is a hypothesis, not a guarantee (arXiv 2606.26479).
The output side — leakage scanning and the defender's own blind spots
Bidirectional means the job isn't done when the agent answers. Output scanning validates the response before delivery: checking for system- prompt leakage, exposed API keys, and PII, and validating outputs against the data sources they were allowed to draw from (OWASP). Datadog frames the operational version as a wide → strict → human-review funnel: cheap static filters first, ML classifiers next, human escalation for high-risk actions, with tracing/telemetry feeding a continuous tuning loop (Datadog).
A practical refinement the redaction tooling converges on: redact with
stable, deterministic placeholders, not blunt deletion. Map each secret to a
consistent token (e.g. a hash-derived tag) so the same value always yields the
same placeholder within a session — the model keeps referential coherence and
can still edit around a redacted value instead of losing the thread. Two
Claude-Code-native tools ship exactly this, in opposite directions:
DataSentry redacts secrets (cloud keys / PEM / JWT / DB creds) from tool
output before the model sees them via a PostToolUse hook, and blocks prompts
that paste live credentials via UserPromptSubmit — a fail-closed, zero-dependency
GPL-3.0 plugin using SHA1-stable placeholders (DataSentry);
Maskit is the egress counterpart — a local privacy gateway that masks keys /
credentials / PII / internal IPs before requests leave the machine to Cursor /
Claude Code / Codex and restores them in the streamed response, reusing
placeholders across multi-turn chats (Maskit).
Between them they cover both directions the deterministic way — keep secrets out
of the model's context, and out of the provider egress.
Two 2026 findings warn that the defender is itself a target:
- Gaslighting the triage model. The in-the-wild
macOS.Gaslightbackdoor embeds a 3.5 KB cascade of 38 fabricated "system" messages (fake token-expiry, OOM, disk-exhaustion, static-analysis flags) to trick an LLM-assisted malware-triage agent into aborting — an evolution from single-block injection to a stacked assault on the analyst's perception (SentinelLABS). The rule it enforces — keep hostile content out of the model entirely, treat triaged samples as data never instructions — is precisely the bidirectional-sanitization posture. - The surveillance flip side. "AI Snitches Get Glitches" formalizes agentic surveillance — an agent analyzing accessible data, writing a report, and exfiltrating it via its own tools — and finds a paradox: some models unprompted-ly assist surveillance and report the attempt to authorities (arXiv 2606.25836). Your output scanner is also the last gate on the agent leaking for someone.
The defense stack at a glance
| Layer | Question it answers | Representative technique | Catches | Blind spot |
|---|---|---|---|---|
| Spotlighting / data marking | "Is this data or instructions?" | Delimiters + USER_DATA tags (OWASP) |
naive injections | model still can obey |
| Input detection | "Does this look malicious?" | PromptArmor, InjecGuard, suffix filtering (PromptArmor) | known + adaptive patterns | novel semantics; latency cost |
| Intent analysis | "Does the model plan to follow it?" | IntentGuard (arXiv 2512.00966) | obfuscated injections | needs reasoning model |
| Capability boundary | "Is this action authorized now?" | IBAC intent tuples (Ken Huang) | hijacked planner | policy must be tight |
| Architecture | "Can untrusted data reach a tool?" | Dual-LLM / IFC (Simon Willison) | data→action flow | UX cost, user social-eng |
| Runtime resilience | "Is this the intended task?" | SecInfer, critic agents (SecInfer) | goal drift | compute cost |
| Output scanning | "Is anything leaking?" | PII/leak filter, wide→strict→human (Datadog) | exfil, secret leak | probabilistic, tunable |
💡 Map it to Kiya. We already run the two most reliable layers:
sanitize.py(spotlighting + quarantine of Tier-2 content) and theredact.pyoutput scrubber (leak/PII gate before Telegram) — worth making its tokens placeholder-stable so the model keeps referential coherence. Off-the-shelf, DataSentry (tool-output redaction) and Maskit (egress masking) are the Claude-Code-native tools in this exact category. The gaps worth closing are an intent gate on tool calls (IBAC-style) and a critic pass on high-risk actions.
📇 Defenses & tools reference
The lesson above is what to learn. This is the catalog behind it — expand when you need a specific method or repo.
Papers by defense family
| Defense | Family | Key result | Source |
|---|---|---|---|
| PromptArmor | LLM-as-detector (input) | <1% ASR/FPR/FNR on AgentDojo | arXiv 2507.15219 · OpenReview |
| IntentGuard | intent analysis | 100% → 8.5% ASR on Mind2Web | arXiv 2512.00966 · OpenReview |
| SecInfer | inference-time scaling | beats SOTA + fine-tuning defenses | arXiv 2509.24967 |
| Adversarial Suffix Filtering | input detection | strips GCG-style suffixes | arXiv 2505.09602 |
| Multi-Agent Defense Pipeline | multi-agent critic | coordinated detection | arXiv 2509.14285 |
| Dual-LLM pattern | architecture | privileged vs quarantined split | Simon Willison |
| AEGIS | pre-exec firewall/audit | 48/48 attacks blocked, 8.3 ms median, 1.2% FPR | arXiv 2603.12621 |
| IBAC | capability boundary | intent tuples, per-call authz, TTL | paper · site |
| Adaptive Eval (CaMeL/FIDES/Progent/RTBAS/FORGE) | eval discipline | out-of-band family; Progent adaptive 2.6% ASR vs static 4.2% — but tested only statically |
arXiv 2606.26479 |
Open-source detectors & enforcers
- agentctl — Apache-2.0 IBAC governance layer; normalizes Claude Code / Codex / OpenClaw / Gemini CLI tool calls to canonical actions, Cedar policies, JWT identities + TTL cleanup, JSONL audit; 144 tests / 96% cov — GitHub
- InjecGuard — prompt-injection guard, +30.8% over SOTA on NotInject; targets overdefense FPs — GitHub
- StackOne Defender — injection detector, 88.7% accuracy, 22MB, CPU-only — GitHub
- zeroleaks — scanner for injection + system-prompt extraction — GitHub
Defender-side threat research
- macOS.Gaslight — 38-message fabricated-system-message cascade gaslights an LLM triage agent — SentinelLABS
- AI Snitches / SurveilBench — agentic surveillance threat model + benchmark — arXiv 2606.25836