Concept. This week the agent flips sides: you use AI as the attacker — autonomous recon, exploit chaining, and payload generation — and you study the 2025-2026 CVEs that prove prompt injection in an agentic framework is a code-execution primitive, not merely a content risk. The through-line is that the same reasoning models that write our tools now write and run our exploits, and the most damaging bugs live in the seam where an LLM's output is trusted by a downstream system that was never designed to receive attacker-controlled text.
🎯 Objectives
By the end of this week you can:
- Explain LLM05 Improper Output Handling and reproduce XSS / SQLi / command-injection through model-generated output — distinct from prompt injection (OWASP LLM05).
- Map an LLM-integrated app's tool/API attack surface (direct + indirect inputs → accessible APIs → classic web exploits through the model) using the PortSwigger excessive-agency methodology (PortSwigger).
- Explain why AI coding agents systematically write broken access control (BOLA/IDOR) — the ownership rule lives in the data model, not the prompt, and authorization has no source→sink shape a scanner can catch (Snyk).
- Chain an AI agent for autonomous recon → exploitation and reason honestly about where it beats and where it loses to a human pentester (ARTEMIS study).
- Walk the prompt-injection → RCE chain in a real agent framework (Semantic Kernel MRO traversal) and name the fix (Microsoft).
- Write a 1-page incident analysis (initial access → exploitation → impact → preventing control) for any CVE in the catalog below.
- Apply the Agents Rule of Two and cost-aware evaluation as design constraints on your own agent stack.
- Recognize the AI gateway / orchestration layer as a control plane — how a LiteLLM / RAGFlow / Kestra compromise converges credential theft, host RCE, and downstream data access — and monitor AI workloads by control-plane role (Microsoft).
The offensive stack: agent as attacker
A modern offensive-AI pipeline wires a reasoning model into standard tooling and lets it plan. Point it at an authorized target and it becomes a recon → enum → exploit → report operator: RAG over vulnerability databases for context-aware exploit selection, workflow engines like N8N chaining stages, and MCP servers such as HexStrike-AI fanning a single prompt across 150+ security tools and a dozen autonomous sub-agents.
The honest capability picture is neither hype nor dismissal. Snyk's framing is that
reasoning-model pentesters find context-dependent flaws with no signature — broken
object-level authorization, business-logic abuse, chained multi-step exploits — that
scanners structurally miss, but they hallucinate findings ~30% of the time without
an independent validator and drop below 26% success on hard system-exploitation
(Snyk). The ARTEMIS field study is the
sober benchmark: on a live ~8,000-host university network, the agent found 9 valid
vulnerabilities at an 82% valid-submission rate, placed second and beat 9 of 10
human professionals — at $18/hr vs $60/hr — yet lost every GUI-driven finding
(80% of humans got RCE through a browser-based console it couldn't drive), and its real
bottleneck was recognizing vulnerability patterns, not technical execution
(ARTEMIS). The same category-specific ceiling
shows up in controlled benchmarks: HackSynth — a planner + summarizer agent looped
inside a firewalled Kali container against 200 CTF challenges (PicoCTF + OverTheWire) —
tops out at 34–40% solved with GPT-4o but scores 0–2 on binary exploitation and
zero on cryptography, and its firewall exists precisely because the agent kept
hallucinating nmap scans against non-existent IPs and, at high temperature, deleted its
own binaries mid-run (HackSynth). Capability is
real but jagged: strong on recon and known patterns, cliff-edged on the classes that
need genuine reasoning under a formal model.
That jagged capability is no longer confined to benchmarks and bug-bounty labs.
UAT-10147 — a Chinese-speaking, financially-motivated crew — is the first publicly
documented actor to wire agentic AI into post-compromise operations as a semi-autonomous
offensive orchestrator, not merely a coding assistant. Cisco Talos recovered the actor's own
artifacts: PentestGPT running on their C2 for dynamic scanning and exploit execution,
DeepAudit for source-code vulnerability scanning, ysoserial with AI-generated Java
deserialization guides, and AI-authored ASP.NET ViewState RCE playbooks with Python
automation (check_paths.py, deploy_implant.py, exfil.py) that document iterative
learning through trial and error — one note concludes "time-based blind testing is entirely
ineffective" for confirming ViewState RCE. Initial access came from one-day RCEs (Zimbra
CVE-2022-27925, Nacos CVE-2021-29441/29442, Telerik CVE-2019-18935, AjaxPro CVE-2021-23758)
across a ~170,000-URL target list split into 17 files of ~10K each for operational
efficiency, followed by a cross-platform SPECTRE implant with a Linux rootkit and EDR
bypass. Talos assesses with moderate-to-high confidence that AI lets actors "scale complex
attacks... while reducing the expertise traditionally required for advanced post-compromise
operations" (Cisco Talos).
The "lower-skilled actor reaches APT-scale post-compromise" ceiling is now an incident, not a
projection.
A white-hat disclosure the same season shows the other face of that capability — and makes the model generation the variable that decides whether an exploit is reachable at all. Hacktron chained a breach of a frontier lab itself: a malicious HEIF image uploaded to OpenAI's public Discourse community forum slipped past FastImage, fell through to ImageMagick's vulnerable libheif 1.19.7 decoder, and triggered a heap overflow → RCE in the forum's image-processing path (CVE-2026-32882 / GHSA-vhm9-85gw-x335, Discourse-side CVSS 8.8). The pivot, though, was not the forum bug — it was an over-privileged SSO token: OpenAI's "Sign in with OpenAI" token, scoped for forum use, was still valid against ChatGPT / Codex APIs, so intercepting the token exchange hijacked an employee's Codex session, which had a GitHub identity connected, and Codex opened a benign proof pull request inside OpenAI's internal monorepo. Two lessons carry: SSO trust boundaries are only as strong as every app wired into them — the OpenAI-side identity flaw (fixed in ~14 h), not the image parser, was the real pivot, so scope tokens least-privilege per service — and a frontier lab fell through a 30-year-old bug class in a boring image decoder, not an AI-specific attack, so secure the whole environment, not just the model. The AI-capability landmark is the sharpest data point in this chapter: Claude Opus 4.8 could not build the ASLR/jemalloc-defeating x86-64 exploit across repeated sessions; Opus 5 produced reliable RCE within hours of its release — expert guidance was still required (not autonomous), the whole two-month project cost under $3K in tokens, and it earned a $6,500 bounty (Hacktron).
🔑 The agent is a fast, cheap, tireless junior with a confident hallucination problem. Every autonomous finding needs an independent validator before it counts — the winning architectures are a reasoning planner + deterministic tools + a separate exploitability checker, never raw autonomy (Snyk).
LLM output exploitation (the COAE class)
The subtlest offensive surface is not the prompt — it is the output. Under LLM05, the model generates an exploit payload (HTML/JS, SQL, a shell string) that a downstream system consumes without sanitization, so the model becomes an untrusted input source to your own backend (OWASP LLM05). The variants share one root cause — trusting model text as safe:
| Sub-class | Mechanism | Downstream sink |
|---|---|---|
| XSS via output | model emits HTML/JS | rendered in a browser |
| SQLi via output | model emits SQL fragment | passed to a DB query |
| Command injection | model output fed to exec/shell |
subprocess |
| Function-calling abuse | manipulated call parameters | privileged tool |
| Hallucination exploit | plausible-but-fake URL/package | attacker pre-registers it |
| Exfiltration | output encodes secrets in URLs/image tags | egress channel |
💡 Sanitize at the boundary the model's output crosses, not just the input it receives. Treat every token the model emits as untrusted until a deterministic layer has validated it against the sink — exactly as you would raw user input.
Building secure code with the same models compounds the risk. The Illusion of Secure LLM Code audited AI-generated authentication across five assistants and four prompting strategies mapped to NIST SP 800-63B: functional or generically "secure" prompts consistently omit brute-force resistance, session management, and robust password handling; single-shot NIST context helps but stays "structurally inadequate," and only iterative reprompting — forcing the model into a self-auditing loop — reaches defense-in-depth (arXiv 2607.23710). The builder lesson: never ship AI-generated auth from one prompt; make security review a re-prompt pass, not a single instruction.
One authorization class is systematically worse, and it is the mirror image of what the agent-as-attacker finds: broken access control (BOLA/IDOR). Snyk's diagnosis is that the rule which would prevent it — "an invoice belongs to an organization, so check the caller's org, not just the invoice id" — lives in the data model and in the heads of the engineers who designed it two years ago, never in the prompt or the surrounding code. Ask an agent to "add an endpoint that returns an invoice by id" and it authenticates correctly, handles the missing-record case and writes clean code — then omits the ownership filter it was never told about, so any authenticated user reads any tenant's invoice by changing one value in the URL. Crucially this is invisible to SAST: unlike SSRF or path traversal, object-level authorization has no "untrusted input → dangerous sink" shape — it is the absence of a comparison only the application's own rules require, and a scanner cannot know which field encodes ownership (CWE-639 / -862 / -863) (Snyk). The five controls that do work are process, not pattern-matching: inventory every endpoint that takes an object id, write the ownership rules down, make "does this verify entitlement, and against which field?" an explicit question in AI-code review, add a cross-tenant test per resource (another tenant must get a 404), and run periodic whole-codebase contextual analysis against an application-context graph. This is the same broken-object-level-authorization the ARTEMIS agent found that scanners missed — one blind spot, now seen from the side that writes the bug.
Finding it on a live target — map the tool surface, then probe it
The theory above tells you what can go wrong; on an authorized engagement you still have to find it. PortSwigger's Web LLM Attacks methodology is the canonical hands-on procedure, and it reduces to three steps: (1) map the inputs — the direct prompt and every indirect channel (scraped web pages, product reviews, emails, documents) the model ingests; (2) map what the LLM can reach — which data and which APIs/tools it is wired to, often by simply asking it ("what APIs can you call?") and, if it stonewalls, re-asking under misleading context or a false pretext until it enumerates its own tools; and (3) probe that surface with classic web exploits — send path-traversal, SQLi and command-injection payloads through the model to each tool it exposes (PortSwigger). The free, Burp-driven lab set drills the exact failure classes this chapter teaches: an excessive-agency assistant that runs privileged SQL with no authorization check when asked to delete a user (the Meta-Instagram confused deputy in miniature), OS command injection through an LLM-exposed internal function, and indirect prompt injection that makes an automated AI scanner leak its own API key or chain into a secondary SSRF → delete-user. The defensive distillation is three rules you can hold a design review to: treat every API handed to an LLM as publicly accessible, never feed the model data it isn't cleared to leak, and never rely on prompting to block an attack — a jailbreak routes around instruction-level guards every time.
🔑 The LLM's tool list is the attack surface — enumerate it before anything else. The model will usually tell you what it can call; once you know the tools you are back to ordinary web-app testing against each one, with the model as a willing proxy that carries no authorization of its own.
When prompts become shells
The 2026 headline is that prompt injection in an agent framework reaches host-level RCE.
In Semantic Kernel (Python), CVE-2026-26030 (CVSS 9.9), unsafe string interpolation
built a filter lambda from AI-controlled input; the exploit walked Python's method
resolution order —
tuple().__class__.__bases__[0].__subclasses__()[…].__init__.__globals__['__builtins__']['eval']
— to reach eval despite an empty __builtins__, defeating the blocklist by using
attribute names it never listed. Microsoft's fix was four layers: an AST node-type
allowlist, a function-call allowlist, a dangerous-attribute blocklist, and name-node
restriction — because "the LLM is not a security boundary"
(Microsoft).
The .NET sibling CVE-2026-25592 accidentally exposed DownloadFileAsync as a kernel
function, letting a chained injection write a payload into the Startup folder for a
sandbox escape. A hands-on CTF reproduces the Python bug end-to-end
(AIAgentCTF). The
coding-agent parallel is gemini-cli's CVSS 10 (headless mode auto-trusts a workspace
.gemini/settings.json, --yolo disables allowlisting), patched in 0.39.1 within two
days (The Register · SecurityWeek).
Amazon Kiro is the same workspace-trust seam widened into a full exfiltration chain,
and it maps directly onto how we run Claude Code: opening a malicious project (via
File → Open Workspace From File) and sending any message — no malicious prompt, no
reference to the attacker's content — lets repo-controlled text steer the agent through
Kiro's POWER.md steering file to read local secrets, write them into a security-relevant
IDE config file, and let a subsequent IDE capability turn that config into network
traffic. The whole chain is legitimate-per-step (repo content → agent → config → egress),
which is exactly why it slips past per-action review; no CVE was assigned, fixed
0.7.45 → 0.8.140
(Mindgard via The Hacker News).
The click layer becomes an injection surface
Retrieval-time injection defenses assume the malicious instruction rides inside the page content the agent scrapes, and scan there. Two verified 2026 findings move the payload into the click and session layer, where those defenses never look — the same seam we already run through with a Gmail-connected assistant.
AI Recommendation Poisoning weaponizes deep-link prefilled prompts. Most assistants honour
a ?q= URL parameter — chatgpt.com/?q=Summarize+this opens your active, logged-in session
and executes the query as if you typed it, no confirmation. Vendors now embed hidden payloads
in "Ask AI" / "Summarize with AI" buttons; a malicious one instructs the assistant to
permanently save the vendor's domain as a "trusted source," silently biasing every future
answer. Because the prompt travels through the click layer rather than scraped content, it
bypasses retrieval-time filtering entirely — a one-click indirect-injection plus persistence
combo (MITRE ATLAS AML.T0080 Memory Poisoning, related AML.T0051 Prompt Injection).
Microsoft catalogued 31 companies across 14 industries, 50+ distinct prompts over 60 days;
clearing memory removes stored payloads, but re-clicking re-poisons
(The Hacker News).
Zenity Labs' Claude-in-Chrome takeover is the same class at maximum severity: a browser agent that reads untrusted content and executes code inside an authenticated session turns inbox access into full account takeover. A malicious email in the victim's Gmail, summarized by Claude in Chrome, hides instructions; direct script execution is blocked, so the payload is smuggled as a benign-looking npm import from a rogue CDN, which runs inside the live session — reading Gmail's Atom feed for Slack / X / Claude.ai magic-link verification codes, triggering password resets, then having Claude relay the OTP to complete the hijack, while silently sharing every Google Drive file to attacker-controlled accounts. SecurityWeek confirms the same shape against ChatGPT Atlas (a planted X comment causes "intent collision" → WhatsApp phishing and Rufus-delegated Amazon purchases). The root cause is architectural, not a patchable bug: agentic browsers intentionally break the Same-Origin Policy, resurrecting CSRF by acting as one entity spanning every authenticated domain at once. Anthropic rated the report "informative"; OpenAI acknowledged "no easy patch" (SecurityWeek).
🔑 Our Gmail-MCP agent has the exact ingredients this class needs — untrusted inbox content + tool-execution authority in an authenticated context. Never let inbox text drive a tool call; scope OTP / magic-link mail out of any "summarize my email" path; and treat any URL-embedded prompt (
?q=…) as untrusted input, not a convenience feature.
Autonomous discovery and the economics of it
Offense scales differently from defense. A cost-aware evaluation across Cybench
(offense) and Splunk BOTS v1 (defense) found red-team CTF performance scales
predictably with test-time compute — scaled open-weight models approach frontier at
competitive cost — while blue-team SOC investigation does not, depending instead on
disciplined tool use and telemetry navigation over raw reasoning budget
(arXiv 2607.15263). On the discovery side, the
open-source OpenAnt shows the winning shape: reachability-filtered decomposition cuts
the analysis surface up to 97%, then adversarial verification plus auto-generated
sandboxed exploit environments kill false positives — finding unknown bugs in OpenSSL,
WordPress, and Flowise (arXiv 2606.19149). A
complementary discovery shape lands the same month: SETYPE reframes the LLM not as a
reachability-decomposing agent but as a semantics-aware type checker — it infers types
directly from the natural-language meaning of symbols and expressions, and a failed type
check is the vulnerability signal (catching semantic bugs that syntactic SAST rules miss).
Its PYSETYPE prototype hit 87% precision / 88% accuracy on Python web apps and surfaced
15 candidate zero-days, 9 developer-confirmed (arXiv 2608.14533).
A new theft surface arrived alongside: Black-Box Skill Stealing shows proprietary agent
skills are extracted cheaply and repeatably, a single successful attempt compromising
the protected skill, with defenses that reduce but don't eliminate leakage
(arXiv 2604.21829).
A sharp measurement caution lands on the reverse-engineering side of this discovery push. LLM decompilers now emit clean, idiomatic C, but they are judged almost entirely on recompilability + re-executability — metrics that reward the wrong path. A decompiled function can build and pass every shipped test yet diverge on other legitimate inputs, and a disclosed vulnerability can silently vanish from the recompiled code with no trace. The paper's Decompile-Diverge oracle needs no fixed tests — it synthesizes a driver, grows a fuzz corpus from the original, and re-executes the decompiled version on identical inputs — and finds 4.9% of "passing" functions diverge (up to 13% on some systems), with ~1 in 10 CVE-grounded functions showing vulnerability erasure; tellingly, one refinement pass lifted Ghidra's build rate 75%→90% while behavioral match fell 74%→62% (arXiv 2609.05370). For any LLM-assisted RE, malware-analysis, or vuln-hunting workflow the lesson is blunt: a decompiler that compiles is not a decompiler that preserves behavior — verify behavior against the original binary, never trust that it built.
🔑 Agents Rule of Two (Microsoft): no AI workflow should simultaneously (1) process untrusted input, (2) hold secrets, and (3) communicate externally. Break any one leg and the exfiltration chain collapses (Microsoft — CI/CD case).
🎯 OSAI exam depth — Reconnaissance for AI Targets (m2)
Everything above assumes you already have a target. On the exam you don't — you get an authorized scope and must discover the AI attack surface yourself, passively first, without tripping the defender's tripwires. AI infrastructure is unusually loud on the wire: unlike generic web apps, model servers leak framework-specific banners, telemetry, default ports and often ship with zero authentication out of the box (Ollama, vLLM, Gradio, MLflow, Qdrant, ChromaDB all bind open by default), which makes fingerprinting cheap (Cisco Talos · Resecurity).
Passive-first, defender-silent discovery. The whole point of "without alerting defenders" is that the first pass touches the target's own logs zero times. Pull from third-party scan indexes that already crawled the internet — Shodan, Censys, ZoomEye, FOFA — so the packets came from their scanners, not you. Only after you've mapped the surface do you send a single low-noise confirmation request. Canonical fingerprints to memorize (ai_osint · Cisco Talos):
| Service | Port | Shodan / Censys signature | Tell |
|---|---|---|---|
| Ollama | 11434 |
port:11434 product:"Ollama"; Server: uvicorn secondary tell |
GET /api/tags lists local models |
| vLLM / OpenAI-compatible | 8000 |
port:8000 "/v1/models" + FastAPI/uvicorn |
mimics OpenAI spec → /v1/models |
| MLflow | 5000 |
http.title:"MLflow" |
unauth experiments/artifacts; path-traversal in errors leaks FS paths |
| Ray dashboard | 8265 |
http.title:"Ray" + /api/jobs |
Jobs API = unauth RCE (ShadowRay) |
| TorchServe | 8080/8081/8082 |
distinctive X-Request-Id/PyTorch header |
mgmt API on 8081 |
| Triton | 8000/8001/8002 |
/v2/health/ready, /v2 KServe routes |
GPU inference server |
| Gradio / Streamlit | 7860/8501 |
http.title:"Gradio" / "Streamlit" |
demo UIs, often internal models |
| Jupyter / Kubeflow | 8888 |
http.title:"Jupyter" -"Login" |
token often disabled → notebook = shell |
| Vector DBs (Qdrant/Weaviate/Milvus) | 6333/8080/19530 |
port:6333 "/collections" |
GET /collections dumps the RAG knowledge base |
Favicon and header hashing collapse the noise: http.favicon.hash: in Shodan matches a
framework's icon across every non-standard port an operator moved it to, so you catch the
20% of Ollama that isn't on 11434 without a single active probe
(Cisco Talos).
Beyond the box, OSINT the dependencies: Google dorks for leaked keys and configs
(filetype:env "OPENAI_API_KEY", site:*.company.com "gradio"), GitHub dorks over the org's
repos for requirements.txt/Dockerfile naming torchserve, mlflow, vllm, and public
grok.com/share / claude.ai/share conversation links that leak internal prompts and
architecture (ai_osint). Subdomain enumeration for
ai., chat., ml., mlflow., notebook. finds the platform even when the app is behind a
CDN.
Purpose-built recon at internet scale. AIMap automates exactly this workflow — 32 preset Shodan-index queries feed a fingerprinting stage that runs Nuclei templates + live HTTP checks to identify protocol, framework, auth state, exposed tools, loaded models, and leaked system prompts across MCP, Ollama, vLLM, LiteLLM, LangServe, Open-WebUI, Gradio, ComfyUI and TGI. The strategic read for a candidate: today's internet has ~25–30K confirmed unauth Ollama hosts and 200,000+ exposed Ray servers, and the exposure population still vastly exceeds active exploitation — the recon-and-catalogue phase is the current frontier, so a defender watching for "AI-aware" scanners often isn't yet (cyberdesserts research · Oligo ShadowRay 2.0).
🧪 Drill: stand up Ollama + MLflow + a Gradio app on non-default ports in a lab, then rediscover them from outside using only banner/favicon fingerprints and
/api/tags,/v1/models,GET /collectionsconfirmation requests. Log every packet that hits the target and prove the discovery pass sent ≤1 request per host.
🎯 OSAI exam depth — AI Infrastructure & Deployment Exploits (m9)
Once recon fingerprints a model server, the exam wants the exploit. The recurring root
cause across the whole class is that these servers load or serve models — and models are
code. Serialized model artifacts (Python .pkl/PyTorch/YAML) run arbitrary code the moment
they are deserialized, and management/inference APIs built for trusted networks ship
internet-reachable. Learn these five as the canonical attack chains:
- Ray — ShadowRay (CVE-2023-48022, CVSS 9.8). The Jobs API on the dashboard (
:8265) has no authorization by design — Anyscale disputes it as intended behavior, so it never got a real patch and doesn't fire in static scanners ("shadow" vuln). Attack:POSTa job to/api/jobs/whoseentrypointis arbitrary shell/Python → unauth RCE. A public PoC and a Metasploit module (exploit/multi/misc/ray_job_rce) exist; ShadowRay 2.0 (late 2025) weaponized it into a self-propagating cryptomining/DDoS botnet across 200K+ exposed clusters, with cron/systemd persistence and worm-like internet scanning (Oligo · PoC). Post-exploit priority: harvest the cluster's model weights and any mounted cloud creds. - TorchServe — ShellTorch (CVE-2023-43654, CVSS 9.8 + CVE-2022-1471 SnakeYAML). Default config
binds the management API (
:8081) to0.0.0.0with no auth and accepts model-config URLs from any domain (SSRF). Chain: hit the mgmt API → register a model whose config URL you control → TorchServe fetches your malicious.mar→ the SnakeYAML deserialization inside it executes → RCE. Tens of thousands were exposed; verify with the open ShellTorchChecker (Oligo). - Ollama — Probllama (CVE-2024-37032).
/api/pulldidn't validate the manifest digest, so a rogue registry serving a manifest with a../path-traversal digest gives arbitrary file write; write a.so, poison/etc/ld.so.preload, then hit/api/chatto spawn a process that loads it → RCE. Worst in Docker, where Ollama runs root on0.0.0.0. Fixed in 0.1.34 but huge exposed tail (Wiz · PoC). - MLflow — model-load RCE (CVE-2024-37052 … -37060).
_load_model_from_local_fileperforms unsafecloudpickledeserialization of a user-controlled model file (CWE-502). Upload a scikit-learn / PyFunc model carrying a hostile serialized object → arbitrary code runs the moment a data scientist opens it (UI:R— needs the victim to load it). Combined with an unauth MLflow tracking server on:5000, an attacker plants the poisoned model for them (GHSA advisory). - Triton — Wiz RCE chain (CVE-2025-23319/-23320/-23334). A shared-memory over-read leaks the Python backend's internal IPC key → register that shared-memory region via the normal API → gain read/write primitives into backend memory → RCE, all through legitimate inference calls so it looks like normal traffic (Wiz).
Pivoting from model server to cloud AI platform. RCE on any of the above is rarely the
objective — it's the foothold. Triton, TorchServe and friends run as custom serving containers
inside SageMaker, Vertex AI, and Azure ML, each with an attached IAM role / service-account
token. From code execution, immediately query the instance metadata service (169.254.169.254
IMDS on AWS, metadata.google.internal on GCP) to steal the workload's temporary credentials,
then pivot into the wider cloud account — the classic container-to-cloud escalation Wiz warns
about (Wiz — Triton).
On multi-tenant AI-as-a-service, the model is the payload: Wiz uploaded a hostile serialized
model to Hugging Face's Inference API, got a reverse shell inside the shared inference
infrastructure, then used a container escape to reach cross-tenant access to other customers'
models and a writable shared container registry (a supply-chain foothold)
(Wiz — Hugging Face).
In-the-wild actors now ship RATs the same way — malicious .pth files whose embedded payload
executes shell on torch.load(), sometimes 7z-wrapped to dodge Picklescan
(Rapid7).
🔑 The universal m9 sink is model deserialization + an unauth management/inference API. Fingerprint the server (m2) → reach its model-load or job-submit path → deliver a poisoned artifact or unauth job → RCE → loot weights + IMDS creds → pivot. Every CVE above is one instantiation of that chain.
🧪 Drill: in an isolated lab, reproduce two chains end-to-end — (1) ShadowRay: submit a job to an unauth Ray dashboard and pop a shell; (2) MLflow/Ollama: craft a hostile serialized model or rogue registry and get code execution on load. Then simulate the pivot by querying a mock IMDS endpoint from the compromised container.
The gateway is the crown jewel — AI infra as a control plane
The model server is not the only prize. The gateway, retrieval, and orchestration layer concentrates every provider credential and holds execution privilege, which makes it the richest single foothold in an AI stack — and it is now actively exploited in the wild. Microsoft's Aug-2026 investigation walks three real intrusions, each treating AI infrastructure as a control plane rather than an isolated app, and each converging on the same trio: credential theft → host RCE → downstream data access (Microsoft):
- LiteLLM gateway — authenticated MCP command-execution (CVE-2026-42271) chained with a
Starlette host-header validation bypass (CVE-2026-48710) yielded unauthenticated RCE; the
gateway process then read
/proc/1/environto harvest every model-provider key, the LiteLLM master key, DB connection strings and UI creds, dropped XMRig, dumped the backend PostgreSQL virtual-keys table, and persisted via SSHauthorized_keys(fix1.83.7). It is the exact/proc/1/environcredential-harvest primitive as the Claude Code /proc case above, one layer up the stack. - RAGFlow — after SSRF recon, the actor planted a hidden Python hook in the LLM-config path that silently captured provider type, model name and API-key material every time a user added a provider — a passive credential harvester that never has to run an exploit again.
- Kestra — an auth-bypass (CVE-2026-49869, CVSS 10) let unauthenticated attackers define workflows running worker-side shell scripts, read the mounted Docker socket to enumerate container-env secrets, and dropped XMRig.
🔑 Monitor AI workloads by their control-plane role, not as isolated apps. A gateway that suddenly spawns a shell, reads
/proc/1/environ, or writes an SSH key is one compromise narrative across three different products — correlate unexpected shells + secret access + persistence + resource hijack, don't chase per-product signals.
📇 CVE & incident reference
The lesson above is what to learn. This is the catalog behind it — expand for the per-incident detail. For each, write the 1-page analysis: initial access → exploitation → impact → preventing control.
Root-cause taxonomy
| Class | Root cause | Preventing control |
|---|---|---|
| Improper output handling | model output trusted by a downstream sink | sanitize at the sink; validate as untrusted |
| Excessive agency / confused deputy | non-deterministic LLM given privileged API access | hard authorization boundary outside the model |
| Agentic-framework RCE | AI-controlled input reaches eval/subprocess/file write |
AST/allowlist gate; LLM is not a boundary |
| Workspace auto-trust | headless/YOLO mode auto-loads attacker config | gate on workspace-trust; no ambient creds |
| Secret exfiltration | untrusted input + secrets + egress in one workflow | Agents Rule of Two |
Copilot / M365 exfiltration lineage
- CVE-2025-32711 — EchoLeak (M365 Copilot): zero-click data exfil via a crafted email; bypassed the XPIA classifier (Wiz).
- CVE-2026-42824 — SearchLeak (M365 Copilot Enterprise, Varonis): one-click theft via Parameter-to-Prompt (P2P) injection — the search
q=param executed as instructions — chained with an HTML streaming-render race and a Bing "Search by Image" SSRF/CSP bypass on an allowlisted*.bing.comdomain. Server-side patched (Varonis). - CVE-2026-24301 — CoSnitch (Copilot Personal, Varonis; companion memory CVE-2026-24299, CVSS 8.8): the P2P lineage taken zero-click — an undocumented URL param auto-executes a
?q=prompt on page-load (runs to completion even if the tab is closed), which queries OAuth-connected Gmail/Drive/Calendar/OneDrive and exfiltrates via Copilot's URL-fetch. Third leg is persistent memory poisoning that survives password change / session revoke / re-enrollment and leaves no process/file/network/log artifact. Disclosed Dec-2025, patched Aug-18-2026 (~8-month lag); pre-patch memory payloads may persist → audit the memory store manually. Enterprise unaffected (Varonis).[Aug-24 pulse]
Browser & click-layer agent takeovers (see "The click layer becomes an injection surface")
- AI Recommendation Poisoning (The Hacker News / Microsoft, Aug 2026): deep-link
?q=prefilled "Ask AI" buttons auto-run a query in the logged-in session and tell the assistant to save the vendor's domain as a "trusted source" — one-click IPI + persistent memory poisoning (ATLAS AML.T0080/T0051); 31 companies / 14 industries / 50+ prompts in 60 days (The Hacker News). - Zenity Labs — Claude in Chrome / ChatGPT Atlas account takeover (SecurityWeek, disclosed Dec-2025/Jan-2026): inbox IPI → payload smuggled as a rogue-CDN npm import runs in the authenticated session → reads Gmail Atom feed for Slack/X/Claude.ai OTPs, resets passwords, exfil Drive files; Atlas variant via X-comment "intent collision." Root cause is architectural — agentic browsers break Same-Origin Policy (CSRF resurrected). Anthropic "informative"; OpenAI "no easy patch" (SecurityWeek).
- BragJack — one malicious extension hijacks the AI agent in five browsers (Gal Weizman / Forever Security, disclosed Sep-16-2026): Chrome Gemini Live, Edge Actions, Opera Neon, Perplexity Comet, and Claude in Chrome. The technique — Weizman calls it "Prompt Forcing," not prompt injection — bypasses the content layer entirely: the attacker writes the whole prompt, picks when it runs, and chains follow-ups, so the model never sees untrusted content and cannot be "on guard." The delivery is DiNneR Serving: two ordinary Chromium permissions millions of extensions already hold — content scripts + the declarativeNetRequest (DNR) API (the ad-blocker permission) — are abused with
modifyHeadersrules to weaken CSP/X-Frame-Options and redirect a JS resource, letting untrusted extension code cross into the privileged AI "body" (the component that reads tabs, files, camera/mic, takes screenshots, acts on sites). Impact ranged local-file/history/camera access → full agentic action on Comet/Neon. Two CVEs —CVE-2026-0628(Chrome, CVSS 8.8, fixed 143.0.7499.192) +CVE-2026-55945(Edge, 4.2); $20K+ across five vendors. Caveat: needs the extension already installed (not a drive-by) — but the abused permissions are the ones nobody hesitates to grant. The lesson pairs sharply against Anthropic's "prompt injection is largely solved" line (Boris Cherny): a neuron-level PI classifier is irrelevant to Prompt Forcing because there is no injected content to detect — the control plane itself is the boundary that failed. (Forever Security · BleepingComputer).
Coding-agent & IDE incidents
- CVE-2025-53773 — GitHub Copilot YOLO-mode RCE (CVSS 9.6): prompt injection via code comments triggering autonomous execution on 100,000+ dev machines (NVD).
- "Comment and Control" (Apr 2026): PR-title injection made Claude Code, Gemini CLI Action, and Copilot Agent leak API keys as GitHub comments (VentureBeat).
- Claude Code GitHub Action /proc exfiltration (Jun 2026): an unsandboxed in-process Read reached
/proc/self/environ; injection in an issue body → read env → trim 7 chars to dodge the secret scanner → exfilANTHROPIC_API_KEY/GITHUB_TOKEN. Origin of the Agents Rule of Two (Microsoft). - AWS AgentCore Harness — /proc heap-read steals plaintext managed credentials (Unit 42 / Sheida Azimi, Sep-18-2026; no CVE, closed "informative" shared-responsibility): the exact
/proc-primitive above, one layer up into a managed agent runtime. AgentCore ships two default-on tools —shell(bash, runs as root) +file_operations— in the same PID-1 memory space where the harness resolves vault credentials into plaintext JWTs. Indirect PI via a support ticket drove the shell tool to read/proc/1/mem, extract a live JWT + MCP-server URL, exfil to a webhook (public egress is the default), then replay the JWT to read customer PII with no AWS creds. A more-restrictive model refused direct attempts; the indirect path + a permissive tool-capable model still succeeded. The red-team lesson: enumerate what the harness can do before the model says a word — the default tool surface (shell/file/broad access) is your attack surface; scope tools, deny public egress (VPC mode), never colocate a shell with credential resolution (Unit 42). - Amazon Kiro workspace-trust exfil (Mindgard, Aug 2026): opening a malicious workspace + sending any message (no malicious prompt) lets repo-controlled content steer the agent via the
POWER.mdsteering file to read local secrets → write them into a security-relevant IDE config → a later IDE capability beacons them out. Legitimate-per-step chain (repo → agent → config → egress); no CVE, fixed0.7.45 → 0.8.140(The Hacker News).
Confused-deputy in production
- Meta Instagram AI account-recovery bot (Jun 1, 2026): prompt injection against a chatbot with direct write access to email-binding and password-reset APIs — attackers changed the bound email, received a reset link, and bypassed 2FA on verified accounts (incl. an Obama-era White House handle and U.S. Space Force). Textbook LLM06 Excessive Agency: natural language became the control plane. Emergency-hotfixed the same day (TechCrunch · Neowin).
Agentic-framework RCE (see "When prompts become shells")
- CVE-2026-26030 — Semantic Kernel Python, CVSS 9.9, MRO traversal → RCE; fixed 1.39.4 (Microsoft · CTF).
- CVE-2026-25592 — Semantic Kernel .NET,
DownloadFileAsyncsandbox escape; fixed 1.71.0. - CVE-2026-90617 / CVE-2026-90618 — GH05TCREW PentestAgent, OS command injection (CVSS 7.3, disclosed Sep-14-2026): an offensive AI agent is itself the target —
run_taskin the MCP HTTP Server (interface/main.py, -90617) andLocalRuntime.execute_command(runtime/runtime.py, -90618) pass model-influenced input to the shell, remotely exploitable. Public exploit; fix PR still pending (rolling release, no fixed version). The canonical "untrusted target data → AI generates command → shell executes" chain, so a hostile HTTP/SSH/DNS response reachable by the agent can reach the shell — interim fix is network-isolate the MCP interface + run as a least-privilege user (VulDB · OffSeq).[Sep-15]
AI-assisted tradecraft
- UAT-10147 — agentic AI as offensive orchestrator (in the wild) (Cisco Talos, Aug 2026): first public actor to run agentic AI in post-compromise ops — PentestGPT on C2, DeepAudit source scanning, AI-authored ViewState-RCE playbooks + Python automation, one-day RCE chains across a ~170K-URL list → SPECTRE implant w/ Linux rootkit + EDR bypass. The "lower-skilled actor reaches APT scale" thesis, operationalized (Cisco Talos).
- Hacktron — reaching OpenAI's internal monorepo (HEIF RCE → over-privileged SSO chain) (Sep 2026): a malicious HEIF upload to OpenAI's public Discourse forum hit an unpatched libheif 1.19.7 heap overflow (CVE-2026-32882 / GHSA-vhm9-85gw-x335, Discourse-side CVSS 8.8, via a FastImage→ImageMagick fallthrough) → forum RCE → intercept the "Sign in with OpenAI" SSO token → hijack an employee ChatGPT/Codex session → connected GitHub identity → benign proof PR in the internal
openaimonorepo. The pivot was the OpenAI-side over-privileged token (fixed ~14 h), not the forum bug. AI-capability landmark: Opus 4.8 failed the exploit across sessions, Opus 5 got reliable x86-64/jemalloc RCE within hours of release — expert-guided, not autonomous; whole ~2-month project <$3K in tokens, $6,500 bounty. Lessons: SSO trust boundaries are only as strong as every wired-in app; secure the whole environment (a boring image parser breached the frontier lab) (Hacktron).[Sep-19] - Patch diffing: compare pre-/post-patch code to locate the exact fix, then weaponize before defenders deploy — the window is now hours (CVE-2026-33626 exploited within 13h of disclosure) (Sysdig).
- AI-generated malware: progressive-stealth PoCs, obfuscation automation, and self-rewriting agentic frameworks (GNAW) that defeat signature detection — AI lowers the skill floor for every evasion class. (hendryadrian)
- CTF/CVE-framing jailbreaks: actors social-engineer their own upstream LLMs by framing exploit-writing as a "CTF challenge," leaving a fingerprint — CVE-templated User-Agents (
ctf-litellm-cve42271-mcp-stdio/1.0) that make good IOCs (Sysdig TRT). - ClickFix via a fake ChatGPT Custom GPT — the AI brand is the lure (Huntress, Sep 2026): no model flaw at all — the trust the AI platform carries is the exploit. A malicious Custom GPT named "Plus 5.6", promoted through Google Ads and hosted on the legitimate
chatgpt.comdomain, funnels victims via a Google-Sites Cloudflare-CAPTCHA ClickFix page into pasting an obfuscated PowerShell one-liner →ISOSimple.msi("Advanced Printer Configuration Reader") → DLL sideload into a Canon-signed binary (COTFileReadApp.exe→ patchedceiinfolog.dll→rdCore.dll) → shellcode carved from a disguised.wav(v2 hides it in a NuGetBuild.dat) → an 806-file per-file-XOR encrypted filesystem (monitor.raw) → a 1.58 MB RAT (remote desktop, cam/mic, 17-browser credential theft, DNS-over-HTTPS C2, AMSI bypass, ntdll unhooking, anti-VM; HKCU-Run + scheduled-task persistence that self-heals every 150–875 s). OpenAI pulled the first GPT Sep-25; a near-identical replacement was live by Sep-27. The lesson: treat AI-platform-hosted content (Custom GPTs, Artifacts,*/sharelinks) and AI-search answers as an untrusted malware-delivery surface — a sponsored search for "chatgpt" lands on the real domain serving the hostile GPT (Huntress).[Sep-30]
Inference-infrastructure (the pipeline, not the model)
- AI gateways as a control plane — LiteLLM / RAGFlow / Kestra (Microsoft, Aug 2026): three in-the-wild intrusions where the credential-concentrating gateway/orchestration layer, not the model, is the target. LiteLLM (CVE-2026-42271 + Starlette host-header bypass CVE-2026-48710 → unauth RCE →
/proc/1/environkey theft + XMRig + PostgreSQL virtual-keys dump; fix1.83.7); RAGFlow (hidden Python hook in the LLM-config path harvests every provider key on setup); Kestra (auth-bypass CVE-2026-49869 CVSS 10 → worker-side shell + Docker-socket secret enumeration). Lesson: monitor by control-plane role, correlate shells + secret access + persistence + resource hijack (Microsoft).[Sep-8] - CVE-2026-20685 — Apple Private Cloud Compute path traversal (CVSS 6.5, info-disclosure): a classic archive-extraction / Zip-Slip bug in
darwin-init(PID 1, root) — its generic tar extractor appended untrusted entry names to output dirs without validation. During boot,darwin-initfetches+extracts cryptexes over HTTP; a privileged-network attacker plants a malicious cryptex → writes attacker-controlled files as root, persisting across reboots, and redirects sealed AI-inference telemetry off-node — defeating PCC's statelessness/attestation/sealed-observability guarantees before the hardened runtime engages. $150K Apple bounty (Drinor Selmanaj / Sentry SARC), found in Apple's Virtual Research Env; fixed in PCC5E290.3. The keystone lesson: a 30-year-old vuln class undermines the inference pipeline even when the model is untouched — secure the whole environment, not just the model (Sentry SARC · NVD).