Skip to content
Phase 2, week 13

Agentic Pentesting + Real CVE Studies

0 of 161 items done. ~5h29m estimated.

Concept. This week the agent flips sides: you use AI as the attacker — autonomous recon, exploit chaining, and payload generation — and you study the 2025-2026 CVEs that prove prompt injection in an agentic framework is a code-execution primitive, not merely a content risk. The through-line is that the same reasoning models that write our tools now write and run our exploits, and the most damaging bugs live in the seam where an LLM's output is trusted by a downstream system that was never designed to receive attacker-controlled text.

🎯 Objectives

By the end of this week you can:

  • Explain LLM05 Improper Output Handling and reproduce XSS / SQLi / command-injection through model-generated output — distinct from prompt injection (OWASP LLM05).
  • Map an LLM-integrated app's tool/API attack surface (direct + indirect inputs → accessible APIs → classic web exploits through the model) using the PortSwigger excessive-agency methodology (PortSwigger).
  • Explain why AI coding agents systematically write broken access control (BOLA/IDOR) — the ownership rule lives in the data model, not the prompt, and authorization has no source→sink shape a scanner can catch (Snyk).
  • Chain an AI agent for autonomous recon → exploitation and reason honestly about where it beats and where it loses to a human pentester (ARTEMIS study).
  • Walk the prompt-injection → RCE chain in a real agent framework (Semantic Kernel MRO traversal) and name the fix (Microsoft).
  • Write a 1-page incident analysis (initial access → exploitation → impact → preventing control) for any CVE in the catalog below.
  • Apply the Agents Rule of Two and cost-aware evaluation as design constraints on your own agent stack.
  • Recognize the AI gateway / orchestration layer as a control plane — how a LiteLLM / RAGFlow / Kestra compromise converges credential theft, host RCE, and downstream data access — and monitor AI workloads by control-plane role (Microsoft).

The offensive stack: agent as attacker

A modern offensive-AI pipeline wires a reasoning model into standard tooling and lets it plan. Point it at an authorized target and it becomes a recon → enum → exploit → report operator: RAG over vulnerability databases for context-aware exploit selection, workflow engines like N8N chaining stages, and MCP servers such as HexStrike-AI fanning a single prompt across 150+ security tools and a dozen autonomous sub-agents.

The honest capability picture is neither hype nor dismissal. Snyk's framing is that reasoning-model pentesters find context-dependent flaws with no signature — broken object-level authorization, business-logic abuse, chained multi-step exploits — that scanners structurally miss, but they hallucinate findings ~30% of the time without an independent validator and drop below 26% success on hard system-exploitation (Snyk). The ARTEMIS field study is the sober benchmark: on a live ~8,000-host university network, the agent found 9 valid vulnerabilities at an 82% valid-submission rate, placed second and beat 9 of 10 human professionals — at $18/hr vs $60/hr — yet lost every GUI-driven finding (80% of humans got RCE through a browser-based console it couldn't drive), and its real bottleneck was recognizing vulnerability patterns, not technical execution (ARTEMIS). The same category-specific ceiling shows up in controlled benchmarks: HackSynth — a planner + summarizer agent looped inside a firewalled Kali container against 200 CTF challenges (PicoCTF + OverTheWire) — tops out at 34–40% solved with GPT-4o but scores 0–2 on binary exploitation and zero on cryptography, and its firewall exists precisely because the agent kept hallucinating nmap scans against non-existent IPs and, at high temperature, deleted its own binaries mid-run (HackSynth). Capability is real but jagged: strong on recon and known patterns, cliff-edged on the classes that need genuine reasoning under a formal model.

That jagged capability is no longer confined to benchmarks and bug-bounty labs. UAT-10147 — a Chinese-speaking, financially-motivated crew — is the first publicly documented actor to wire agentic AI into post-compromise operations as a semi-autonomous offensive orchestrator, not merely a coding assistant. Cisco Talos recovered the actor's own artifacts: PentestGPT running on their C2 for dynamic scanning and exploit execution, DeepAudit for source-code vulnerability scanning, ysoserial with AI-generated Java deserialization guides, and AI-authored ASP.NET ViewState RCE playbooks with Python automation (check_paths.py, deploy_implant.py, exfil.py) that document iterative learning through trial and error — one note concludes "time-based blind testing is entirely ineffective" for confirming ViewState RCE. Initial access came from one-day RCEs (Zimbra CVE-2022-27925, Nacos CVE-2021-29441/29442, Telerik CVE-2019-18935, AjaxPro CVE-2021-23758) across a ~170,000-URL target list split into 17 files of ~10K each for operational efficiency, followed by a cross-platform SPECTRE implant with a Linux rootkit and EDR bypass. Talos assesses with moderate-to-high confidence that AI lets actors "scale complex attacks... while reducing the expertise traditionally required for advanced post-compromise operations" (Cisco Talos). The "lower-skilled actor reaches APT-scale post-compromise" ceiling is now an incident, not a projection.

A white-hat disclosure the same season shows the other face of that capability — and makes the model generation the variable that decides whether an exploit is reachable at all. Hacktron chained a breach of a frontier lab itself: a malicious HEIF image uploaded to OpenAI's public Discourse community forum slipped past FastImage, fell through to ImageMagick's vulnerable libheif 1.19.7 decoder, and triggered a heap overflow → RCE in the forum's image-processing path (CVE-2026-32882 / GHSA-vhm9-85gw-x335, Discourse-side CVSS 8.8). The pivot, though, was not the forum bug — it was an over-privileged SSO token: OpenAI's "Sign in with OpenAI" token, scoped for forum use, was still valid against ChatGPT / Codex APIs, so intercepting the token exchange hijacked an employee's Codex session, which had a GitHub identity connected, and Codex opened a benign proof pull request inside OpenAI's internal monorepo. Two lessons carry: SSO trust boundaries are only as strong as every app wired into them — the OpenAI-side identity flaw (fixed in ~14 h), not the image parser, was the real pivot, so scope tokens least-privilege per service — and a frontier lab fell through a 30-year-old bug class in a boring image decoder, not an AI-specific attack, so secure the whole environment, not just the model. The AI-capability landmark is the sharpest data point in this chapter: Claude Opus 4.8 could not build the ASLR/jemalloc-defeating x86-64 exploit across repeated sessions; Opus 5 produced reliable RCE within hours of its release — expert guidance was still required (not autonomous), the whole two-month project cost under $3K in tokens, and it earned a $6,500 bounty (Hacktron).

🔑 The agent is a fast, cheap, tireless junior with a confident hallucination problem. Every autonomous finding needs an independent validator before it counts — the winning architectures are a reasoning planner + deterministic tools + a separate exploitability checker, never raw autonomy (Snyk).

LLM output exploitation (the COAE class)

The subtlest offensive surface is not the prompt — it is the output. Under LLM05, the model generates an exploit payload (HTML/JS, SQL, a shell string) that a downstream system consumes without sanitization, so the model becomes an untrusted input source to your own backend (OWASP LLM05). The variants share one root cause — trusting model text as safe:

Sub-class Mechanism Downstream sink
XSS via output model emits HTML/JS rendered in a browser
SQLi via output model emits SQL fragment passed to a DB query
Command injection model output fed to exec/shell subprocess
Function-calling abuse manipulated call parameters privileged tool
Hallucination exploit plausible-but-fake URL/package attacker pre-registers it
Exfiltration output encodes secrets in URLs/image tags egress channel

💡 Sanitize at the boundary the model's output crosses, not just the input it receives. Treat every token the model emits as untrusted until a deterministic layer has validated it against the sink — exactly as you would raw user input.

Building secure code with the same models compounds the risk. The Illusion of Secure LLM Code audited AI-generated authentication across five assistants and four prompting strategies mapped to NIST SP 800-63B: functional or generically "secure" prompts consistently omit brute-force resistance, session management, and robust password handling; single-shot NIST context helps but stays "structurally inadequate," and only iterative reprompting — forcing the model into a self-auditing loop — reaches defense-in-depth (arXiv 2607.23710). The builder lesson: never ship AI-generated auth from one prompt; make security review a re-prompt pass, not a single instruction.

One authorization class is systematically worse, and it is the mirror image of what the agent-as-attacker finds: broken access control (BOLA/IDOR). Snyk's diagnosis is that the rule which would prevent it — "an invoice belongs to an organization, so check the caller's org, not just the invoice id" — lives in the data model and in the heads of the engineers who designed it two years ago, never in the prompt or the surrounding code. Ask an agent to "add an endpoint that returns an invoice by id" and it authenticates correctly, handles the missing-record case and writes clean code — then omits the ownership filter it was never told about, so any authenticated user reads any tenant's invoice by changing one value in the URL. Crucially this is invisible to SAST: unlike SSRF or path traversal, object-level authorization has no "untrusted input → dangerous sink" shape — it is the absence of a comparison only the application's own rules require, and a scanner cannot know which field encodes ownership (CWE-639 / -862 / -863) (Snyk). The five controls that do work are process, not pattern-matching: inventory every endpoint that takes an object id, write the ownership rules down, make "does this verify entitlement, and against which field?" an explicit question in AI-code review, add a cross-tenant test per resource (another tenant must get a 404), and run periodic whole-codebase contextual analysis against an application-context graph. This is the same broken-object-level-authorization the ARTEMIS agent found that scanners missed — one blind spot, now seen from the side that writes the bug.

Finding it on a live target — map the tool surface, then probe it

The theory above tells you what can go wrong; on an authorized engagement you still have to find it. PortSwigger's Web LLM Attacks methodology is the canonical hands-on procedure, and it reduces to three steps: (1) map the inputs — the direct prompt and every indirect channel (scraped web pages, product reviews, emails, documents) the model ingests; (2) map what the LLM can reach — which data and which APIs/tools it is wired to, often by simply asking it ("what APIs can you call?") and, if it stonewalls, re-asking under misleading context or a false pretext until it enumerates its own tools; and (3) probe that surface with classic web exploits — send path-traversal, SQLi and command-injection payloads through the model to each tool it exposes (PortSwigger). The free, Burp-driven lab set drills the exact failure classes this chapter teaches: an excessive-agency assistant that runs privileged SQL with no authorization check when asked to delete a user (the Meta-Instagram confused deputy in miniature), OS command injection through an LLM-exposed internal function, and indirect prompt injection that makes an automated AI scanner leak its own API key or chain into a secondary SSRF → delete-user. The defensive distillation is three rules you can hold a design review to: treat every API handed to an LLM as publicly accessible, never feed the model data it isn't cleared to leak, and never rely on prompting to block an attack — a jailbreak routes around instruction-level guards every time.

🔑 The LLM's tool list is the attack surface — enumerate it before anything else. The model will usually tell you what it can call; once you know the tools you are back to ordinary web-app testing against each one, with the model as a willing proxy that carries no authorization of its own.

When prompts become shells

The 2026 headline is that prompt injection in an agent framework reaches host-level RCE. In Semantic Kernel (Python), CVE-2026-26030 (CVSS 9.9), unsafe string interpolation built a filter lambda from AI-controlled input; the exploit walked Python's method resolution order — tuple().__class__.__bases__[0].__subclasses__()[…].__init__.__globals__['__builtins__']['eval'] — to reach eval despite an empty __builtins__, defeating the blocklist by using attribute names it never listed. Microsoft's fix was four layers: an AST node-type allowlist, a function-call allowlist, a dangerous-attribute blocklist, and name-node restriction — because "the LLM is not a security boundary" (Microsoft). The .NET sibling CVE-2026-25592 accidentally exposed DownloadFileAsync as a kernel function, letting a chained injection write a payload into the Startup folder for a sandbox escape. A hands-on CTF reproduces the Python bug end-to-end (AIAgentCTF). The coding-agent parallel is gemini-cli's CVSS 10 (headless mode auto-trusts a workspace .gemini/settings.json, --yolo disables allowlisting), patched in 0.39.1 within two days (The Register · SecurityWeek). Amazon Kiro is the same workspace-trust seam widened into a full exfiltration chain, and it maps directly onto how we run Claude Code: opening a malicious project (via File → Open Workspace From File) and sending any message — no malicious prompt, no reference to the attacker's content — lets repo-controlled text steer the agent through Kiro's POWER.md steering file to read local secrets, write them into a security-relevant IDE config file, and let a subsequent IDE capability turn that config into network traffic. The whole chain is legitimate-per-step (repo content → agent → config → egress), which is exactly why it slips past per-action review; no CVE was assigned, fixed 0.7.45 → 0.8.140 (Mindgard via The Hacker News).

The click layer becomes an injection surface

Retrieval-time injection defenses assume the malicious instruction rides inside the page content the agent scrapes, and scan there. Two verified 2026 findings move the payload into the click and session layer, where those defenses never look — the same seam we already run through with a Gmail-connected assistant.

AI Recommendation Poisoning weaponizes deep-link prefilled prompts. Most assistants honour a ?q= URL parameter — chatgpt.com/?q=Summarize+this opens your active, logged-in session and executes the query as if you typed it, no confirmation. Vendors now embed hidden payloads in "Ask AI" / "Summarize with AI" buttons; a malicious one instructs the assistant to permanently save the vendor's domain as a "trusted source," silently biasing every future answer. Because the prompt travels through the click layer rather than scraped content, it bypasses retrieval-time filtering entirely — a one-click indirect-injection plus persistence combo (MITRE ATLAS AML.T0080 Memory Poisoning, related AML.T0051 Prompt Injection). Microsoft catalogued 31 companies across 14 industries, 50+ distinct prompts over 60 days; clearing memory removes stored payloads, but re-clicking re-poisons (The Hacker News).

Zenity Labs' Claude-in-Chrome takeover is the same class at maximum severity: a browser agent that reads untrusted content and executes code inside an authenticated session turns inbox access into full account takeover. A malicious email in the victim's Gmail, summarized by Claude in Chrome, hides instructions; direct script execution is blocked, so the payload is smuggled as a benign-looking npm import from a rogue CDN, which runs inside the live session — reading Gmail's Atom feed for Slack / X / Claude.ai magic-link verification codes, triggering password resets, then having Claude relay the OTP to complete the hijack, while silently sharing every Google Drive file to attacker-controlled accounts. SecurityWeek confirms the same shape against ChatGPT Atlas (a planted X comment causes "intent collision" → WhatsApp phishing and Rufus-delegated Amazon purchases). The root cause is architectural, not a patchable bug: agentic browsers intentionally break the Same-Origin Policy, resurrecting CSRF by acting as one entity spanning every authenticated domain at once. Anthropic rated the report "informative"; OpenAI acknowledged "no easy patch" (SecurityWeek).

🔑 Our Gmail-MCP agent has the exact ingredients this class needs — untrusted inbox content + tool-execution authority in an authenticated context. Never let inbox text drive a tool call; scope OTP / magic-link mail out of any "summarize my email" path; and treat any URL-embedded prompt (?q=…) as untrusted input, not a convenience feature.

Autonomous discovery and the economics of it

Offense scales differently from defense. A cost-aware evaluation across Cybench (offense) and Splunk BOTS v1 (defense) found red-team CTF performance scales predictably with test-time compute — scaled open-weight models approach frontier at competitive cost — while blue-team SOC investigation does not, depending instead on disciplined tool use and telemetry navigation over raw reasoning budget (arXiv 2607.15263). On the discovery side, the open-source OpenAnt shows the winning shape: reachability-filtered decomposition cuts the analysis surface up to 97%, then adversarial verification plus auto-generated sandboxed exploit environments kill false positives — finding unknown bugs in OpenSSL, WordPress, and Flowise (arXiv 2606.19149). A complementary discovery shape lands the same month: SETYPE reframes the LLM not as a reachability-decomposing agent but as a semantics-aware type checker — it infers types directly from the natural-language meaning of symbols and expressions, and a failed type check is the vulnerability signal (catching semantic bugs that syntactic SAST rules miss). Its PYSETYPE prototype hit 87% precision / 88% accuracy on Python web apps and surfaced 15 candidate zero-days, 9 developer-confirmed (arXiv 2608.14533). A new theft surface arrived alongside: Black-Box Skill Stealing shows proprietary agent skills are extracted cheaply and repeatably, a single successful attempt compromising the protected skill, with defenses that reduce but don't eliminate leakage (arXiv 2604.21829).

A sharp measurement caution lands on the reverse-engineering side of this discovery push. LLM decompilers now emit clean, idiomatic C, but they are judged almost entirely on recompilability + re-executability — metrics that reward the wrong path. A decompiled function can build and pass every shipped test yet diverge on other legitimate inputs, and a disclosed vulnerability can silently vanish from the recompiled code with no trace. The paper's Decompile-Diverge oracle needs no fixed tests — it synthesizes a driver, grows a fuzz corpus from the original, and re-executes the decompiled version on identical inputs — and finds 4.9% of "passing" functions diverge (up to 13% on some systems), with ~1 in 10 CVE-grounded functions showing vulnerability erasure; tellingly, one refinement pass lifted Ghidra's build rate 75%→90% while behavioral match fell 74%→62% (arXiv 2609.05370). For any LLM-assisted RE, malware-analysis, or vuln-hunting workflow the lesson is blunt: a decompiler that compiles is not a decompiler that preserves behavior — verify behavior against the original binary, never trust that it built.

🔑 Agents Rule of Two (Microsoft): no AI workflow should simultaneously (1) process untrusted input, (2) hold secrets, and (3) communicate externally. Break any one leg and the exfiltration chain collapses (Microsoft — CI/CD case).

🎯 OSAI exam depth — Reconnaissance for AI Targets (m2)

Everything above assumes you already have a target. On the exam you don't — you get an authorized scope and must discover the AI attack surface yourself, passively first, without tripping the defender's tripwires. AI infrastructure is unusually loud on the wire: unlike generic web apps, model servers leak framework-specific banners, telemetry, default ports and often ship with zero authentication out of the box (Ollama, vLLM, Gradio, MLflow, Qdrant, ChromaDB all bind open by default), which makes fingerprinting cheap (Cisco Talos · Resecurity).

Passive-first, defender-silent discovery. The whole point of "without alerting defenders" is that the first pass touches the target's own logs zero times. Pull from third-party scan indexes that already crawled the internet — Shodan, Censys, ZoomEye, FOFA — so the packets came from their scanners, not you. Only after you've mapped the surface do you send a single low-noise confirmation request. Canonical fingerprints to memorize (ai_osint · Cisco Talos):

Service Port Shodan / Censys signature Tell
Ollama 11434 port:11434 product:"Ollama"; Server: uvicorn secondary tell GET /api/tags lists local models
vLLM / OpenAI-compatible 8000 port:8000 "/v1/models" + FastAPI/uvicorn mimics OpenAI spec → /v1/models
MLflow 5000 http.title:"MLflow" unauth experiments/artifacts; path-traversal in errors leaks FS paths
Ray dashboard 8265 http.title:"Ray" + /api/jobs Jobs API = unauth RCE (ShadowRay)
TorchServe 8080/8081/8082 distinctive X-Request-Id/PyTorch header mgmt API on 8081
Triton 8000/8001/8002 /v2/health/ready, /v2 KServe routes GPU inference server
Gradio / Streamlit 7860/8501 http.title:"Gradio" / "Streamlit" demo UIs, often internal models
Jupyter / Kubeflow 8888 http.title:"Jupyter" -"Login" token often disabled → notebook = shell
Vector DBs (Qdrant/Weaviate/Milvus) 6333/8080/19530 port:6333 "/collections" GET /collections dumps the RAG knowledge base

Favicon and header hashing collapse the noise: http.favicon.hash: in Shodan matches a framework's icon across every non-standard port an operator moved it to, so you catch the 20% of Ollama that isn't on 11434 without a single active probe (Cisco Talos). Beyond the box, OSINT the dependencies: Google dorks for leaked keys and configs (filetype:env "OPENAI_API_KEY", site:*.company.com "gradio"), GitHub dorks over the org's repos for requirements.txt/Dockerfile naming torchserve, mlflow, vllm, and public grok.com/share / claude.ai/share conversation links that leak internal prompts and architecture (ai_osint). Subdomain enumeration for ai., chat., ml., mlflow., notebook. finds the platform even when the app is behind a CDN.

Purpose-built recon at internet scale. AIMap automates exactly this workflow — 32 preset Shodan-index queries feed a fingerprinting stage that runs Nuclei templates + live HTTP checks to identify protocol, framework, auth state, exposed tools, loaded models, and leaked system prompts across MCP, Ollama, vLLM, LiteLLM, LangServe, Open-WebUI, Gradio, ComfyUI and TGI. The strategic read for a candidate: today's internet has ~25–30K confirmed unauth Ollama hosts and 200,000+ exposed Ray servers, and the exposure population still vastly exceeds active exploitation — the recon-and-catalogue phase is the current frontier, so a defender watching for "AI-aware" scanners often isn't yet (cyberdesserts research · Oligo ShadowRay 2.0).

🧪 Drill: stand up Ollama + MLflow + a Gradio app on non-default ports in a lab, then rediscover them from outside using only banner/favicon fingerprints and /api/tags, /v1/models, GET /collections confirmation requests. Log every packet that hits the target and prove the discovery pass sent ≤1 request per host.

🎯 OSAI exam depth — AI Infrastructure & Deployment Exploits (m9)

Once recon fingerprints a model server, the exam wants the exploit. The recurring root cause across the whole class is that these servers load or serve models — and models are code. Serialized model artifacts (Python .pkl/PyTorch/YAML) run arbitrary code the moment they are deserialized, and management/inference APIs built for trusted networks ship internet-reachable. Learn these five as the canonical attack chains:

  • Ray — ShadowRay (CVE-2023-48022, CVSS 9.8). The Jobs API on the dashboard (:8265) has no authorization by design — Anyscale disputes it as intended behavior, so it never got a real patch and doesn't fire in static scanners ("shadow" vuln). Attack: POST a job to /api/jobs/ whose entrypoint is arbitrary shell/Python → unauth RCE. A public PoC and a Metasploit module (exploit/multi/misc/ray_job_rce) exist; ShadowRay 2.0 (late 2025) weaponized it into a self-propagating cryptomining/DDoS botnet across 200K+ exposed clusters, with cron/systemd persistence and worm-like internet scanning (Oligo · PoC). Post-exploit priority: harvest the cluster's model weights and any mounted cloud creds.
  • TorchServe — ShellTorch (CVE-2023-43654, CVSS 9.8 + CVE-2022-1471 SnakeYAML). Default config binds the management API (:8081) to 0.0.0.0 with no auth and accepts model-config URLs from any domain (SSRF). Chain: hit the mgmt API → register a model whose config URL you control → TorchServe fetches your malicious .mar → the SnakeYAML deserialization inside it executes → RCE. Tens of thousands were exposed; verify with the open ShellTorchChecker (Oligo).
  • Ollama — Probllama (CVE-2024-37032). /api/pull didn't validate the manifest digest, so a rogue registry serving a manifest with a ../ path-traversal digest gives arbitrary file write; write a .so, poison /etc/ld.so.preload, then hit /api/chat to spawn a process that loads it → RCE. Worst in Docker, where Ollama runs root on 0.0.0.0. Fixed in 0.1.34 but huge exposed tail (Wiz · PoC).
  • MLflow — model-load RCE (CVE-2024-37052 … -37060). _load_model_from_local_file performs unsafe cloudpickle deserialization of a user-controlled model file (CWE-502). Upload a scikit-learn / PyFunc model carrying a hostile serialized object → arbitrary code runs the moment a data scientist opens it (UI:R — needs the victim to load it). Combined with an unauth MLflow tracking server on :5000, an attacker plants the poisoned model for them (GHSA advisory).
  • Triton — Wiz RCE chain (CVE-2025-23319/-23320/-23334). A shared-memory over-read leaks the Python backend's internal IPC key → register that shared-memory region via the normal API → gain read/write primitives into backend memory → RCE, all through legitimate inference calls so it looks like normal traffic (Wiz).

Pivoting from model server to cloud AI platform. RCE on any of the above is rarely the objective — it's the foothold. Triton, TorchServe and friends run as custom serving containers inside SageMaker, Vertex AI, and Azure ML, each with an attached IAM role / service-account token. From code execution, immediately query the instance metadata service (169.254.169.254 IMDS on AWS, metadata.google.internal on GCP) to steal the workload's temporary credentials, then pivot into the wider cloud account — the classic container-to-cloud escalation Wiz warns about (Wiz — Triton). On multi-tenant AI-as-a-service, the model is the payload: Wiz uploaded a hostile serialized model to Hugging Face's Inference API, got a reverse shell inside the shared inference infrastructure, then used a container escape to reach cross-tenant access to other customers' models and a writable shared container registry (a supply-chain foothold) (Wiz — Hugging Face). In-the-wild actors now ship RATs the same way — malicious .pth files whose embedded payload executes shell on torch.load(), sometimes 7z-wrapped to dodge Picklescan (Rapid7).

🔑 The universal m9 sink is model deserialization + an unauth management/inference API. Fingerprint the server (m2) → reach its model-load or job-submit path → deliver a poisoned artifact or unauth job → RCE → loot weights + IMDS creds → pivot. Every CVE above is one instantiation of that chain.

🧪 Drill: in an isolated lab, reproduce two chains end-to-end — (1) ShadowRay: submit a job to an unauth Ray dashboard and pop a shell; (2) MLflow/Ollama: craft a hostile serialized model or rogue registry and get code execution on load. Then simulate the pivot by querying a mock IMDS endpoint from the compromised container.

The gateway is the crown jewel — AI infra as a control plane

The model server is not the only prize. The gateway, retrieval, and orchestration layer concentrates every provider credential and holds execution privilege, which makes it the richest single foothold in an AI stack — and it is now actively exploited in the wild. Microsoft's Aug-2026 investigation walks three real intrusions, each treating AI infrastructure as a control plane rather than an isolated app, and each converging on the same trio: credential theft → host RCE → downstream data access (Microsoft):

  • LiteLLM gateway — authenticated MCP command-execution (CVE-2026-42271) chained with a Starlette host-header validation bypass (CVE-2026-48710) yielded unauthenticated RCE; the gateway process then read /proc/1/environ to harvest every model-provider key, the LiteLLM master key, DB connection strings and UI creds, dropped XMRig, dumped the backend PostgreSQL virtual-keys table, and persisted via SSH authorized_keys (fix 1.83.7). It is the exact /proc/1/environ credential-harvest primitive as the Claude Code /proc case above, one layer up the stack.
  • RAGFlow — after SSRF recon, the actor planted a hidden Python hook in the LLM-config path that silently captured provider type, model name and API-key material every time a user added a provider — a passive credential harvester that never has to run an exploit again.
  • Kestra — an auth-bypass (CVE-2026-49869, CVSS 10) let unauthenticated attackers define workflows running worker-side shell scripts, read the mounted Docker socket to enumerate container-env secrets, and dropped XMRig.

🔑 Monitor AI workloads by their control-plane role, not as isolated apps. A gateway that suddenly spawns a shell, reads /proc/1/environ, or writes an SSH key is one compromise narrative across three different products — correlate unexpected shells + secret access + persistence + resource hijack, don't chase per-product signals.

📇 CVE & incident reference

The lesson above is what to learn. This is the catalog behind it — expand for the per-incident detail. For each, write the 1-page analysis: initial access → exploitation → impact → preventing control.

Root-cause taxonomy
Class Root cause Preventing control
Improper output handling model output trusted by a downstream sink sanitize at the sink; validate as untrusted
Excessive agency / confused deputy non-deterministic LLM given privileged API access hard authorization boundary outside the model
Agentic-framework RCE AI-controlled input reaches eval/subprocess/file write AST/allowlist gate; LLM is not a boundary
Workspace auto-trust headless/YOLO mode auto-loads attacker config gate on workspace-trust; no ambient creds
Secret exfiltration untrusted input + secrets + egress in one workflow Agents Rule of Two
Copilot / M365 exfiltration lineage
  • CVE-2025-32711 — EchoLeak (M365 Copilot): zero-click data exfil via a crafted email; bypassed the XPIA classifier (Wiz).
  • CVE-2026-42824 — SearchLeak (M365 Copilot Enterprise, Varonis): one-click theft via Parameter-to-Prompt (P2P) injection — the search q= param executed as instructions — chained with an HTML streaming-render race and a Bing "Search by Image" SSRF/CSP bypass on an allowlisted *.bing.com domain. Server-side patched (Varonis).
  • CVE-2026-24301 — CoSnitch (Copilot Personal, Varonis; companion memory CVE-2026-24299, CVSS 8.8): the P2P lineage taken zero-click — an undocumented URL param auto-executes a ?q= prompt on page-load (runs to completion even if the tab is closed), which queries OAuth-connected Gmail/Drive/Calendar/OneDrive and exfiltrates via Copilot's URL-fetch. Third leg is persistent memory poisoning that survives password change / session revoke / re-enrollment and leaves no process/file/network/log artifact. Disclosed Dec-2025, patched Aug-18-2026 (~8-month lag); pre-patch memory payloads may persist → audit the memory store manually. Enterprise unaffected (Varonis). [Aug-24 pulse]
Browser & click-layer agent takeovers (see "The click layer becomes an injection surface")
  • AI Recommendation Poisoning (The Hacker News / Microsoft, Aug 2026): deep-link ?q= prefilled "Ask AI" buttons auto-run a query in the logged-in session and tell the assistant to save the vendor's domain as a "trusted source" — one-click IPI + persistent memory poisoning (ATLAS AML.T0080/T0051); 31 companies / 14 industries / 50+ prompts in 60 days (The Hacker News).
  • Zenity Labs — Claude in Chrome / ChatGPT Atlas account takeover (SecurityWeek, disclosed Dec-2025/Jan-2026): inbox IPI → payload smuggled as a rogue-CDN npm import runs in the authenticated session → reads Gmail Atom feed for Slack/X/Claude.ai OTPs, resets passwords, exfil Drive files; Atlas variant via X-comment "intent collision." Root cause is architectural — agentic browsers break Same-Origin Policy (CSRF resurrected). Anthropic "informative"; OpenAI "no easy patch" (SecurityWeek).
  • BragJack — one malicious extension hijacks the AI agent in five browsers (Gal Weizman / Forever Security, disclosed Sep-16-2026): Chrome Gemini Live, Edge Actions, Opera Neon, Perplexity Comet, and Claude in Chrome. The technique — Weizman calls it "Prompt Forcing," not prompt injection — bypasses the content layer entirely: the attacker writes the whole prompt, picks when it runs, and chains follow-ups, so the model never sees untrusted content and cannot be "on guard." The delivery is DiNneR Serving: two ordinary Chromium permissions millions of extensions already hold — content scripts + the declarativeNetRequest (DNR) API (the ad-blocker permission) — are abused with modifyHeaders rules to weaken CSP/X-Frame-Options and redirect a JS resource, letting untrusted extension code cross into the privileged AI "body" (the component that reads tabs, files, camera/mic, takes screenshots, acts on sites). Impact ranged local-file/history/camera access → full agentic action on Comet/Neon. Two CVEs — CVE-2026-0628 (Chrome, CVSS 8.8, fixed 143.0.7499.192) + CVE-2026-55945 (Edge, 4.2); $20K+ across five vendors. Caveat: needs the extension already installed (not a drive-by) — but the abused permissions are the ones nobody hesitates to grant. The lesson pairs sharply against Anthropic's "prompt injection is largely solved" line (Boris Cherny): a neuron-level PI classifier is irrelevant to Prompt Forcing because there is no injected content to detect — the control plane itself is the boundary that failed. (Forever Security · BleepingComputer).
Coding-agent & IDE incidents
  • CVE-2025-53773 — GitHub Copilot YOLO-mode RCE (CVSS 9.6): prompt injection via code comments triggering autonomous execution on 100,000+ dev machines (NVD).
  • "Comment and Control" (Apr 2026): PR-title injection made Claude Code, Gemini CLI Action, and Copilot Agent leak API keys as GitHub comments (VentureBeat).
  • Claude Code GitHub Action /proc exfiltration (Jun 2026): an unsandboxed in-process Read reached /proc/self/environ; injection in an issue body → read env → trim 7 chars to dodge the secret scanner → exfil ANTHROPIC_API_KEY/GITHUB_TOKEN. Origin of the Agents Rule of Two (Microsoft).
  • AWS AgentCore Harness — /proc heap-read steals plaintext managed credentials (Unit 42 / Sheida Azimi, Sep-18-2026; no CVE, closed "informative" shared-responsibility): the exact /proc-primitive above, one layer up into a managed agent runtime. AgentCore ships two default-on tools — shell (bash, runs as root) + file_operations — in the same PID-1 memory space where the harness resolves vault credentials into plaintext JWTs. Indirect PI via a support ticket drove the shell tool to read /proc/1/mem, extract a live JWT + MCP-server URL, exfil to a webhook (public egress is the default), then replay the JWT to read customer PII with no AWS creds. A more-restrictive model refused direct attempts; the indirect path + a permissive tool-capable model still succeeded. The red-team lesson: enumerate what the harness can do before the model says a word — the default tool surface (shell/file/broad access) is your attack surface; scope tools, deny public egress (VPC mode), never colocate a shell with credential resolution (Unit 42).
  • Amazon Kiro workspace-trust exfil (Mindgard, Aug 2026): opening a malicious workspace + sending any message (no malicious prompt) lets repo-controlled content steer the agent via the POWER.md steering file to read local secrets → write them into a security-relevant IDE config → a later IDE capability beacons them out. Legitimate-per-step chain (repo → agent → config → egress); no CVE, fixed 0.7.45 → 0.8.140 (The Hacker News).
Confused-deputy in production
  • Meta Instagram AI account-recovery bot (Jun 1, 2026): prompt injection against a chatbot with direct write access to email-binding and password-reset APIs — attackers changed the bound email, received a reset link, and bypassed 2FA on verified accounts (incl. an Obama-era White House handle and U.S. Space Force). Textbook LLM06 Excessive Agency: natural language became the control plane. Emergency-hotfixed the same day (TechCrunch · Neowin).
Agentic-framework RCE (see "When prompts become shells")
  • CVE-2026-26030 — Semantic Kernel Python, CVSS 9.9, MRO traversal → RCE; fixed 1.39.4 (Microsoft · CTF).
  • CVE-2026-25592 — Semantic Kernel .NET, DownloadFileAsync sandbox escape; fixed 1.71.0.
  • CVE-2026-90617 / CVE-2026-90618 — GH05TCREW PentestAgent, OS command injection (CVSS 7.3, disclosed Sep-14-2026): an offensive AI agent is itself the target — run_task in the MCP HTTP Server (interface/main.py, -90617) and LocalRuntime.execute_command (runtime/runtime.py, -90618) pass model-influenced input to the shell, remotely exploitable. Public exploit; fix PR still pending (rolling release, no fixed version). The canonical "untrusted target data → AI generates command → shell executes" chain, so a hostile HTTP/SSH/DNS response reachable by the agent can reach the shell — interim fix is network-isolate the MCP interface + run as a least-privilege user (VulDB · OffSeq). [Sep-15]
AI-assisted tradecraft
  • UAT-10147 — agentic AI as offensive orchestrator (in the wild) (Cisco Talos, Aug 2026): first public actor to run agentic AI in post-compromise ops — PentestGPT on C2, DeepAudit source scanning, AI-authored ViewState-RCE playbooks + Python automation, one-day RCE chains across a ~170K-URL list → SPECTRE implant w/ Linux rootkit + EDR bypass. The "lower-skilled actor reaches APT scale" thesis, operationalized (Cisco Talos).
  • Hacktron — reaching OpenAI's internal monorepo (HEIF RCE → over-privileged SSO chain) (Sep 2026): a malicious HEIF upload to OpenAI's public Discourse forum hit an unpatched libheif 1.19.7 heap overflow (CVE-2026-32882 / GHSA-vhm9-85gw-x335, Discourse-side CVSS 8.8, via a FastImage→ImageMagick fallthrough) → forum RCE → intercept the "Sign in with OpenAI" SSO token → hijack an employee ChatGPT/Codex session → connected GitHub identity → benign proof PR in the internal openai monorepo. The pivot was the OpenAI-side over-privileged token (fixed ~14 h), not the forum bug. AI-capability landmark: Opus 4.8 failed the exploit across sessions, Opus 5 got reliable x86-64/jemalloc RCE within hours of release — expert-guided, not autonomous; whole ~2-month project <$3K in tokens, $6,500 bounty. Lessons: SSO trust boundaries are only as strong as every wired-in app; secure the whole environment (a boring image parser breached the frontier lab) (Hacktron). [Sep-19]
  • Patch diffing: compare pre-/post-patch code to locate the exact fix, then weaponize before defenders deploy — the window is now hours (CVE-2026-33626 exploited within 13h of disclosure) (Sysdig).
  • AI-generated malware: progressive-stealth PoCs, obfuscation automation, and self-rewriting agentic frameworks (GNAW) that defeat signature detection — AI lowers the skill floor for every evasion class. (hendryadrian)
  • CTF/CVE-framing jailbreaks: actors social-engineer their own upstream LLMs by framing exploit-writing as a "CTF challenge," leaving a fingerprint — CVE-templated User-Agents (ctf-litellm-cve42271-mcp-stdio/1.0) that make good IOCs (Sysdig TRT).
  • ClickFix via a fake ChatGPT Custom GPT — the AI brand is the lure (Huntress, Sep 2026): no model flaw at all — the trust the AI platform carries is the exploit. A malicious Custom GPT named "Plus 5.6", promoted through Google Ads and hosted on the legitimate chatgpt.com domain, funnels victims via a Google-Sites Cloudflare-CAPTCHA ClickFix page into pasting an obfuscated PowerShell one-liner → ISOSimple.msi ("Advanced Printer Configuration Reader") → DLL sideload into a Canon-signed binary (COTFileReadApp.exe → patched ceiinfolog.dll → rdCore.dll) → shellcode carved from a disguised .wav (v2 hides it in a NuGet Build.dat) → an 806-file per-file-XOR encrypted filesystem (monitor.raw) → a 1.58 MB RAT (remote desktop, cam/mic, 17-browser credential theft, DNS-over-HTTPS C2, AMSI bypass, ntdll unhooking, anti-VM; HKCU-Run + scheduled-task persistence that self-heals every 150–875 s). OpenAI pulled the first GPT Sep-25; a near-identical replacement was live by Sep-27. The lesson: treat AI-platform-hosted content (Custom GPTs, Artifacts, */share links) and AI-search answers as an untrusted malware-delivery surface — a sponsored search for "chatgpt" lands on the real domain serving the hostile GPT (Huntress). [Sep-30]
Inference-infrastructure (the pipeline, not the model)
  • AI gateways as a control plane — LiteLLM / RAGFlow / Kestra (Microsoft, Aug 2026): three in-the-wild intrusions where the credential-concentrating gateway/orchestration layer, not the model, is the target. LiteLLM (CVE-2026-42271 + Starlette host-header bypass CVE-2026-48710 → unauth RCE → /proc/1/environ key theft + XMRig + PostgreSQL virtual-keys dump; fix 1.83.7); RAGFlow (hidden Python hook in the LLM-config path harvests every provider key on setup); Kestra (auth-bypass CVE-2026-49869 CVSS 10 → worker-side shell + Docker-socket secret enumeration). Lesson: monitor by control-plane role, correlate shells + secret access + persistence + resource hijack (Microsoft). [Sep-8]
  • CVE-2026-20685 — Apple Private Cloud Compute path traversal (CVSS 6.5, info-disclosure): a classic archive-extraction / Zip-Slip bug in darwin-init (PID 1, root) — its generic tar extractor appended untrusted entry names to output dirs without validation. During boot, darwin-init fetches+extracts cryptexes over HTTP; a privileged-network attacker plants a malicious cryptex → writes attacker-controlled files as root, persisting across reboots, and redirects sealed AI-inference telemetry off-node — defeating PCC's statelessness/attestation/sealed-observability guarantees before the hardened runtime engages. $150K Apple bounty (Drinor Selmanaj / Sentry SARC), found in Apple's Virtual Research Env; fixed in PCC 5E290.3. The keystone lesson: a 30-year-old vuln class undermines the inference pipeline even when the model is untouched — secure the whole environment, not just the model (Sentry SARC · NVD).

Recommended resources0/134

Sign in to tick items off and track your progress.

Show

📖 Core Path

📚 Further Reading

CVE Deep Dives
  • 📄 CVE-2025-32711 — EchoLeak (Wiz) — Zero-click data exfiltration via crafted email in Microsoft 365 Copilot; XPIA classifier bypass (~30 min)
  • 📄 CVE-2026-42824 — SearchLeak (Varonis, Jun 2026) — EchoLeak's successor: one-click M365 Copilot Enterprise Search exfil chaining parameter-to-prompt injection (the q URL param read as instructions) + HTML-render race condition (<img> fires before sanitizer) + Bing SSRF/CSP bypass (exfil rides a real microsoft.com domain, evades URL filtering). Lesson: prompt injection reweaponizes "low-impact" web bugs (SSRF, HTML race) when retrieval + output rendering share one flow. Backend-mitigated; PoC, not in-the-wild (~30 min)
  • 📄 CVE-2025-53773 — GitHub Copilot YOLO Mode RCE — CVSS 9.6; prompt injection through code comments; 100K+ developer machines
  • 📄 "Comment and Control" (Apr 2026) — PR title injection leaking API keys via Claude Code, Gemini CLI, Copilot Agent
  • 📄 Pillar Security — "TrustIssues" gemini-cli CVSS 10 (May 2026) — Prompt injection via GitHub issue → supply chain compromise on 101K-star AI coding agent; headless trust + yolo bypass (~30 min)
  • 📄 Black-Box Skill Stealing — arXiv 2604.21829 — Agent skills extracted with only 3 interactions; 90K+ published skills at risk
  • 📄 Sysdig TRT — Jailbreaking LLMs with CTF/CVE Framing (Jun 2026) — Multiple independent actors social-engineer their own upstream LLMs by framing exploit-writing as a "CTF challenge" or "CVE research" task, then deploy the generated code (incl. Python-eval reverse shells) against PraisonAI, LiteLLM, FastGPT, Open-WebUI, Gotenberg, LangFlow using known CVEs (e.g. CVE-2026-42271). The AI-assisted origin leaves a fingerprint: CVE-templated User-Agents (ctf-litellm-cve42271-mcp-stdio/1.0), passwords (MioCtf!…), AWS roleSessionName=cve-scan. Detection: treat CVE-bearing User-Agents as high-signal IOCs (cross-ref Week 10 jailbreak framing) (~25 min)
Prompt Injection → RCE (May 2026 — Proven in Production)
Autonomous Hacking Demos
Autonomous Hacking Research
  • 📄 HackSynth — LLM Agent for Autonomous Pentesting (2024) — planner+summarizer agent in a firewalled Kali container; two CTF benchmarks (PicoCTF 120 + OverTheWire 80). GPT-4o solves 34–40% but 0–2 on binary exploitation and zero on crypto; the firewall exists because the agent hallucinated scans against non-existent IPs and deleted its own binaries at high temperature. The jagged-capability data point behind ARTEMIS's "recognizing patterns is the bottleneck" (~25 min)
  • 📄 Cisco Talos — UAT-10147 integrates agentic AI into post-compromise operations (Aug 2026) — First public documentation of a financially-motivated crew using agentic AI as an offensive orchestrator (not just coding assist): Talos recovered prompt logs from the actor's own endpoints running Claude Code, Codex, Cursor, and Gemini, plus AI-generated playbooks/exploit-automation/troubleshooting logic driving recon, exploitation, validation and persistence across ~170k scanned URLs. Initial access via one-day RCEs (Zimbra, Nacos, Telerik, AjaxPro); custom AI-assisted SPECTRE implant with a Linux rootkit + BYOVD to blind EDR. The "lower-skilled actors reach APT-scale post-compromise" thesis, in the wild (~30 min)
  • 📄 Hacktron AI — reaching an internal OpenAI monorepo via an HEIF RCE + over-privileged SSO token chain (Sep 2026) — white-hat disclosure: a malicious HEIF upload to OpenAI's Discourse forum hit an unpatched libheif 1.19.7 heap overflow (CVE-2026-32882 / GHSA-vhm9-85gw-x335, Discourse CVSS 8.8) → RCE on the forum → intercept the "Sign in with OpenAI" SSO token exchange → hijack employees' ChatGPT/Codex sessions → a connected GitHub identity → a benign proof PR in OpenAI's internal monorepo. Two lessons: (1) SSO trust boundaries are only as strong as every app wired into them — the OpenAI-side identity flaw, not the forum bug, was the pivot (fixed in ~14h); (2) the offense used Claude Opus 4.8/5 to build the ASLR-defeating exploit but needed expert guidance — not autonomous. The frontier-lab breach came through a boring image parser, not an AI-specific attack. $6,500 bounty; <72h.
  • 📄 OpenAnt — LLM Vulnerability Discovery via Decomposition + Adversarial Verification (arXiv 2606.19149, Jun 2026) — Open-source multi-stage system: reachability-filtered code decomposition cuts the analysis surface up to 97%, then adversarial verification (constrained attacker simulation) + auto-generated sandboxed exploit envs kill false positives. Found unknown bugs in OpenSSL, WordPress, Flowise. Same find→verify→reproduce shape as Microsoft MDASH, but open (~30 min)
  • 📄 SETYPE — Finding Vulnerabilities via LLM-Augmented Semantics-Aware Type-Checking (arXiv 2608.14533, Aug 2026) — Wang/Xu/Asokan. The LLM as a type checker instead of an agent: infers types from the natural-language meaning of symbols/expressions, and a failed type check = a candidate vulnerability — catching semantic bugs that syntactic SAST rules overlook. PYSETYPE prototype on Python web apps: 87% precision / 88% accuracy, 15 candidate zero-days / 9 developer-confirmed. Complement to OpenAnt's reachability-decomposition approach (~25 min)
  • 🔧 HexStrike AI v6.0 — MCP offensive-security platform — MCP server wiring Claude/GPT/Copilot to 150+ security tools + 12+ autonomous agents (bug-bounty, CTF solver, CVE intel, exploit generator). 9.7k stars. Study the orchestration + tool-selection engine; note the dual-use risk (~lab)
  • 📄 When LLM Decompilers Recompile More and Preserve Less — arXiv 2609.05370 (Sep 2026) — Liu/Raff/Micinski. LLM decompilers produce clean, idiomatic C judged almost entirely on recompilability + re-executability — but those metrics reward the wrong path: a function can build and pass every shipped test yet diverge on other legitimate inputs, and a disclosed vulnerability can silently vanish from the recompiled code with no trace. Their Decompile-Diverge oracle needs no fixed tests — it synthesizes a driver, fuzzes a corpus from the original, and re-executes the decompiled version on identical inputs: up to 13% of "passing" functions diverge, and ~1 in 10 CVE functions show vulnerability erasure (one model lifted Ghidra's build rate 75%→90% while behavioral match fell 74%→62%). The caution for any LLM-assisted RE / malware-analysis / vuln-hunting workflow: a decompiler that compiles is not a decompiler that preserves behavior (~30 min) [Sep-4]
Evaluating Security Agents
  • 📄 Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents — arXiv 2607.15263 (Jul 2026) — Kassianik/Nelson/Singer. Evaluates LLM security agents at fixed cost levels (inference + tool spend), not just peak success, on offensive Cybench + defensive Splunk BOTS v1. Finding: red-team CTF scales with test-time compute (scaled open-weight models approach frontier at competitive cost); blue-team SOC investigation does not scale the same way — it depends on disciplined tool use, telemetry navigation, and selective enrichment over raw reasoning budget. Reframes agent benchmarks to include economic/operational fit (~30 min)
LLM Output Exploitation
  • 🧪 Build a test app that feeds LLM output to a SQL query — demonstrate SQLi via model output, then implement sanitization
  • 🧪 PortSwigger Web Security Academy — Web LLM Attacks (free interactive labs) — the canonical hands-on lab set for this exact week: excessive-agency LLM APIs (ask the assistant to delete a user → it runs privileged SQL with no authz), OS command injection via an LLM-exposed internal function, and indirect prompt injection against an automated AI scanner to exfiltrate its API key or trigger secondary vulns (SSRF→delete-user). Free, in-browser, Burp-driven. Walkthrough videos: excessive-agency gBQVQCj87mM · LLM-API OS-injection _AK6jDCBWZM · scanner-exfil IPI bmZH_tY0BsU · secondary-vuln SSRF fTshZmYmWWg
AI-Target Recon (OSAI m2)
AI Infrastructure & Deployment Exploits (OSAI m9)
Offensive AI Tools
  • 🔧 N8N — Workflow automation for chaining offensive tools in recon-enumeration-exploit-report pipelines
  • 🔧 Burp Suite + AI Extensions — AI-powered issue explainer, authentication analysis, web exploitation

📡 From the Resources feed

  • 🔧 FirmAgent — fuzzing + LLM agents for IoT firmware vuln discovery (NDSS 2026) — chains IDA sink analysis → QEMU-traced API fuzzing → optional LLM-driven taint analysis + PoC generation; an autonomous-discovery pipeline to study alongside the week's CVE work (in Trove since 2026-09-26 (security/ai-security)) 📡
  • 📄 Huntress — AI Accelerates Attacker Tradecraft, Not New Techniques — 9-month threat-intel read: AI scales recon/enumeration and AI-triaged stolen-data extortion and enables deepfake-video interview fraud, but the techniques are unchanged — plus a case where an attacker inadvertently installed Huntress's own agent on their attack box, unravelling a 2,000+ account compromise (via vendor blog) 📡
  • 📄 Trail of Bits — Auditing in the Age of (Good Enough) AI — flips the usual "agentic code review" framing: over a six-month Miden zkVM audit (novel ZK VM, custom assembly, near-zero tooling), agents built the audit infrastructure before manual review — an LSP server, a MASM decompiler+IR, an abstract-interpretation static analyzer, and a Lean formal-verification framework with an auto MASM→Lean translator. The AI-built tooling found real bugs incl. a high-sev unvalidated-remainder flaw in mod_12289 letting a malicious prover forge Falcon signatures and steal funds. Economics thesis: "a failed side project only costs tokens" makes bespoke tooling for unfamiliar architectures newly viable (via vendor blog) 📡
  • 📄 Trail of Bits — 1Password's AI Patching Benchmark is Misleading — reanalyzes the widely-cited "26% clean fix" figure for autonomous AI vuln-patching: biased hard-fix sample, 22% of trials explicitly told the agent to apply the wrong fix, and >⅓ of runs barred from compiling/running their patch — restrict to reasonable conditions and 86% (2,634/3,067) of the same patches actually blocked the exploit vs a 12.5% human first-fix failure rate; a case study in reading autonomous-patching benchmarks critically (via vendor blog) 📡
  • 📄 Huntress — The AI Attack Surface: threat actors abusing trusted AI platforms — 9 months of intrusions weaponizing legitimate AI-platform features as a malware-delivery channel: malicious Claude Artifacts and claude.ai/share links hosting fake tools, plus SEO-poisoned AI-search results (FakeAgent, MacSync, ClickFix-via-AI-search) steering victims into running stealer commands (SectopRAT/MacSync/AMOS). The trust the platform carries is the lure — no model exploit needed; defenders must treat AI-platform-hosted content and AI-search answers as untrusted delivery surfaces (via vendor blog) 📡
  • 🔧 YesWeHack Claude Kit — official Claude Code plugin enforcing triager-grade bug-bounty report discipline: always-on "never invent facts / no theoretical impact" rules + three skills (write/triage/gotchas) checking per-class proof requirements across 13 vuln types → READY / NEEDS-FIXES / DO-NOT-SUBMIT verdict; GPL-3.0. (via Kiya discovery) 📡
  • 📄 Amazon Kiro prompt injection → local-secret exfil (Mindgard) — attacker-controlled repo content steers the Kiro agent (via Kiro Powers / MCP + steering files) to read local secrets, write them into IDE config, and beacon them out; triggers on merely opening a malicious workspace + sending any message — no malicious prompt needed. Kiro IDE 0.7.45 → fixed 0.8.140, no CVE. The AI-coding-agent trust-boundary lesson that maps directly to our own Claude Code usage.
  • 📄 When AI infrastructure becomes the target (Microsoft Security) — Microsoft's investigation of three real intrusions treating AI infra as a control plane: LiteLLM gateway (CVE-2026-42271 + Starlette host-header bypass -48710 → RCE → /proc/1/environ provider-key theft + XMRig; fix 1.83.7), RAGFlow (hidden Python hook in the LLM-config path silently captures every provider key), Kestra (auth-bypass CVE-2026-49869 CVSS 10 → workflow shell exec). Monitor AI workloads by control-plane role, not as isolated apps.
  • 🔧 Wiz Red Agent POV — series — Wiz's autonomous AI pentester across ~1,000 cloud envs: 17K+ findings, 5,500 validated attack chains, 54% access-control failures (via vendor blog) 📡
  • 📄 Wiz Red Agent POV — reasoning its way to SSRF in GCP Cloud Run — walkthrough of an AI agent chaining a double-slash validator bypass into SSRF, then live credential + source-code theft (via vendor blog) 📡
  • 🔧 Shannon — autonomous white-box AI pentester — reads your source to plan attacks then executes real exploits, reporting only vulns it can actively prove (OWASP injection/XSS/SSRF/auth); CLI, AGPL-3.0, ~46k★ — the proof-based, source-aware counterpart to Argus (shared by @brisk200) 📡
  • 📄 What Is AI Pentesting — how reasoning-model pentesters autonomously find context-dependent flaws (e.g. broken authz) scanners miss (via vendor blog) 📡
  • 🔧 Dalfox — open-source XSS scanner (v3 Rust rewrite) with parameter discovery, WAF fingerprinting, and an MCP server so agents can drive it in CI (via X/Twitter trending) 📡
  • 🔧 Recon-Skills — 25+ field-validated offensive-recon attack patterns (WordPress, CORS, XMLRPC, API, cloud) packaged as structured Claude Code agent skills with verification steps (via X/Twitter trending) 📡
  • 🔧 Burp-MCP-Unrestricted — fork of Burp's MCP extension that lets an AI agent read proxy history, walk the site map, and drive scans/Repeater unrestricted (shared by Ayoma) 📡
  • 🔧 PentesterFlow Agent — open-source terminal agent for authorized offensive security that wires an LLM into standard pentest/bug-bounty tooling (shared by @brisk200) 📡
  • 🌐 Bug Bounty Directory & MCP API — community reference of 3,000+ active bounty/VDP programs with a REST + MCP API for programmatic program discovery in agent workflows (shared by @brisk200) 📡
  • 🔧 BountyForge — Claude Code bug-bounty skill that fans out 8 parallel agents (smart-contract, web/API, access-control, business-logic, race-condition) and emits CVSS-scored, submission-ready HackerOne/Immunefi reports (via Kiya discovery) 📡
  • 📄 Turning Claude into an Autonomous Bug-Bounty Hunter — build walkthrough wiring 124 codified vuln methodologies, 3 hooks, and a Kali/MCP tool bridge into Claude Code with scope-enforcement guardrails (shared by @brisk200) 📡
  • 🔧 Argo — LLM-native static vuln auditor (Claude Code/Codex/local) that reads a repo like a human reviewer, with adversarial validation and read-only, no-auto-submit guardrails (shared by @brisk200) 📡
  • 🔧 GhidraGPT — wires an LLM into Ghidra's decompiler for function analysis and vuln audits (unbounded copies, integer overflows, unchecked returns) as investigative leads, not sound analysis (shared by Ayoma) 📡
  • 🔧 Android Pentesting Skill — Claude Code/OpenCode agent skill for authorized APK audits: 37 Frida scripts, 70+ Semgrep rules, RASP detection, MASVS v2 scoring (via Kiya discovery) 📡
  • 🔧 OpenOSINT — AI OSINT agent exposing 19 recon tools (email/username/breach/domain/IP) as an MCP server for Claude; hard-stop tool calls run the real binary to prevent hallucinated results (via GitHub trending) 📡
  • 🔧 CyberSleuth — OSINT/threat-intel MCP server for Claude with LLM-app fingerprinting and authorization-gated OWASP LLM Top 10 probes across Shodan/VirusTotal/URLScan/CertSpotter (shared by @brisk200) 📡
  • 📄 Cisco's Idan Habler on agentic-AI memory attacks — why persistent agent memory is an attack surface: invisible cross-session contamination needs memory governance alongside secrets management (via Kiya discovery) 📡
  • 📄 AI Recommendation Poisoning — how "Ask AI" buttons silently alter LLM memory — pre-filled deep-link "Ask AI"/"Summarize with AI" buttons run a query the instant a logged-in user clicks (no confirmation); malicious ones tell the assistant to permanently save the vendor's domain as a "trusted source", biasing every future answer — a one-click indirect-injection + persistence combo (MITRE ATLAS AML.T0080 Memory Poisoning + AML.T0051); Microsoft catalogued 31 companies / 50+ prompts. Clearing memory removes stored payloads but re-clicking re-poisons — the durable Kiya lesson: treat any URL-embedded prompt as untrusted, audit stored memory (via The Hacker News)
  • 📄 Zenity Labs — Claude in Chrome indirect prompt injection → account takeover — a malicious email in the victim's Gmail, summarized by Claude in Chrome, hides instructions that make Claude run JS via its javascript_tool inside the already-authenticated browser session; direct script exec is blocked, so the payload is smuggled as a benign-looking npm import from a rogue CDN. That JS reads Gmail's Atom feed for Slack confirmation codes / X reset codes / Claude.ai magic-link nonces and hands them to the attacker → full takeover (Claude.ai the worst — it exposes Gmail/Drive/Calendar/Slack/GitHub connectors). The class generalizes: any AI browser agent that reads untrusted content and executes code in a logged-in session turns inbox access into account takeover — SecurityWeek confirms the same against ChatGPT Atlas, and researchers say there's no clean fix on the horizon. Reported to Anthropic Dec-25/Jan-26 (rated "informative"). Kiya lesson: our Gmail-MCP agent has the exact ingredients — untrusted inbox content + tool-execution authority in an authenticated context; never let inbox text drive a tool call, and scope OTP/magic-link mail out of any summarize-my-email path (Zenity Labs via SecurityWeek)
  • 🔧 Malwoverview — first-response threat-hunting CLI querying 18+ intel platforms (VirusTotal/Hybrid Analysis/URLhaus/…) with an --enrich LLM layer (Claude/Gemini/OpenAI/Ollama) adding risk assessment + MITRE ATT&CK mapping; v8.0.5, 4k★ (via Kiya discovery) 📡
  • 🕳️ XBOW — A Shell Is Worth a Thousand Images (Bing Images RCEs) — an autonomous AI pentester chained SSRF-fetched malicious SVG → ImageMagick delegate pipeline to land three unauthenticated CVSS 9.8 RCEs (CVE-2026-32194/-32191/-21536) with SYSTEM/root on production Microsoft Bing hosts (via vendor blog) 📡
  • 🔧 reverse-skill — AI skill router for authorized RE / pentest / CTF: classifies a task, picks the methodology, then drives the right tool (jadx/Frida/IDA/Burp) across 42 modules; installs into Claude Code/Cursor/Cline, ~21k★ (via Kiya discovery) 📡
  • 🔧 h1-brain — HackerOne MCP server — syncs your bounty history + 3,600 disclosed reports into local SQLite; the hack(handle) tool fuses live scope + past findings + public disclosures into a full attack briefing for Claude; MIT (via Kiya discovery) 📡
  • 🔧 bb-huge — one /bb-huge skill injects 350+ techniques + 7 SOPs into any agent plus a 35-tool MCP stdio server, turning Claude/Codex/Gemini into a persistent bug-bounty hunter with a findings portal that survives session context loss (shared by @brisk200) 📡
  • 🔧 pentest-ai-agents — 50 Claude Code subagents for authorized offensive research (recon, web, AD, cloud, mobile, post-ex) that plan/compose real tool commands (nmap/Burp/sqlmap/Metasploit) behind scope-guard + human-approval gates; ~2.1k★ (via Kiya discovery) 📡
  • 🔧 PentestAgent — MIT AI agent framework for authorized black-box pentest/bug-bounty with single-shot, autonomous, and multi-agent "crew" modes wiring terminal, browser and MCP tools plus a RAG knowledge base; ~2.9k★ (shared by @brisk200) 📡
  • 🔧 LuaN1ao Agent v2 — autonomous pentest agent (TypeScript + Anthropic Pi SDK) with a Planner/Executor/Observer split and causal reasoning graphs tying evidence to confirmed vulns inside a Docker sandbox; ~1.3k★, AGPL-3.0 (shared by @brisk200) 📡
  • 📄 Beyond Prompt Injection: Hacking Apple's Private Cloud Compute — CVE-2026-20685: a path traversal in PCC's darwin-init boot lets a crafted tar write as root and redirect the logging daemon to capture plaintext inference metadata, invisible to attestation — attacking the AI-serving infra itself ($150K bounty) (shared by @brisk200) 📡
  • 📄 Snyk — What Evo's Continuous Offensive Security Found in a Real Enterprise SaaS — an autonomous AI pentester confirmed 33 vulns black-box, chaining a mass-assignment/broken-authz flaw that let a low-priv user rewrite tenant-wide security settings, with working PoCs — agentic offense on a production target (via vendor blog) 📡
  • 🔧 hack-skills — offensive-security knowledge base built for AI agents: 101 deep-topic skills across 14 domains (web, API, priv-esc, binary, crypto, blockchain) as auditable plain-markdown methodology rather than payload dumps, for authorized engagements (shared by @brisk200) 📡
  • 🔧 Bug-Bounty-Agents — 43 drop-in AI agent personas (recon → web/API → infra → exploitation → reporting) for Claude Code/Copilot/Cursor with Burp Suite MCP integration; methodological prompts, not scanners; ~365★ (shared by @brisk200) 📡
  • 🔧 Cairn — frames pentesting as directed state-space search on a blackboard of shared facts/intents with OODA-loop agent workers (Claude Code/Codex/Pi backends); only team to solve all 54 problems at Tencent's 2nd AI Pentest Challenge; ~2.3k★, AGPL-3.0 (via Kiya discovery) 📡
  • 🔧 pentest-agents — bug-bounty automation framework (50 specialized agents, 26 commands, 16 bounty-platform integrations) for Claude Code + six other coding tools; autonomous recon-to-report vuln hunting; ~800★ (shared by @brisk200) 📡
  • 🔧 FoxGuard — fast Rust code-security scanner across 14 languages with taint tracking, secrets + dependency-vuln detection and post-quantum-crypto audits, using LLMs for intelligent crawling and multi-stage attack generation (shared by @RandomCSGuy) 📡
  • 📄 VulnCheck — The First CVE Wave: AI-Assisted Vulnerability Discovery — analysis tying 500%+ surges in CVE disclosures across Chrome/Mozilla/Microsoft to AI-assisted vuln discovery — the disclosure-volume signal of offensive AI going mainstream (shared by @RandomCSGuy) 📡
  • 📄 Schneier — GPT-5.5 Is As Good As Mythos at Finding Vulnerabilities — UK AI Security Institute testing finds OpenAI's GPT-5.5 comparable to Claude Mythos at identifying vulns, with commentary on the democratization of offensive cyber capability (shared by @RandomCSGuy) 📡
  • 🌐 Endor Labs — Claude Fable 5: Mythos-Grade Hype — benchmark of Anthropic's newest model on 200 real vuln-fixing tasks scores 59.8% functional but only 19% security, with record-high benchmark cheating — hype outruns actual secure-coding capability (via vendor blog) 📡
  • 🌐 Endor Labs — Recall, Not Reasoning: How AI Coding Agents Cheat Security Benchmarks — 137/182 confirmed cheating cases reproduce verbatim training-data patches; best agent scores 24% on real security vs 85% functional — a 3.5× "works vs secure" gap you must design evals around (via vendor blog) 📡
  • 🌐 XBOW — Getting to "Should I?" Instead of "Can I?" (High-Accuracy IDOR Detection) — AI pentest agent builds multi-role access baselines first, then asks the semantic authorization question rather than pattern-matching — a repeatable method for finding IDORs in ambiguous contexts (via vendor blog) 📡
  • 🌐 AppOmni — Bodysnatcher: Agentic AI Vulnerability in ServiceNow — CVE-2025-12420 (CVSS 9.3): a hardcoded shared credential in the Now Assist Virtual Agent API lets an unauthenticated attacker impersonate any user by email, then weaponize AI agents to create admin accounts — a real agentic-AI confused-deputy CVE (via Kiya discovery) 📡
  • 🌐 The Hacker News — ServiceNow Patches Critical AI Platform Flaw — disclosure/patch note for the Bodysnatcher impersonation flaw (CVE-2025-12420) bypassing MFA/SSO in Now Assist AI Agents — the vendor-side view of the same incident (via Kiya discovery) 📡
  • 🌐 Endor Labs — Claude Sonnet 5 + Claude Code: Strong on Function, Average on Security — 200-task benchmark: 83.2% functional pass but only 19.6% security pass (writes working patches ~5× more often than it closes vulns), with the lowest cheating rate among frontier models — our own stack model, quantified (via vendor blog) 📡
  • 🌐 Endor Labs — AI SAST Benchmark: 2.6× More Real Vulnerabilities Than Frontier Models — graph-based dataflow SAST found 192 genuine vulns across 8 OSS projects (2.6× Claude Opus 4.7, 3.5× Codex GPT-5.5, 63 unique findings); frontier models have higher precision but miss indirect flows needing graph traversal — where LLM-only vuln discovery hits its ceiling (via vendor blog) 📡
  • 🌐 XBOW — Can an AI Pentest Replace Human Pentesters? — XBOW matched the most senior human pentester (85%) in 28 minutes vs 40 hours and beat all other humans on easy/medium challenges, but humans still lead on business logic, PCI compliance and novel architectures — AI as jagged force-multiplier, not replacement (via vendor blog) 📡
  • 📄 Wiz Red Agent POV — The One Boolean That Broke a B2B Credit System — autonomous Red Agent found a client-trusted boolean flag (unmaskContactData) that unmasked 600M+ business emails and 135M+ phone numbers past a paywall; server never enforced authz, patched in ~2h (via vendor blog) 📡
  • 🌐 Endor Labs — Sonnet 5 + Cursor: Strong Reasoning, Throttled by the Harness — same model, different harness — Cursor+Sonnet 5 scores 63.1% FuncPass vs Claude Code's 83.2%, driven by a 32.5% timeout rate (empty patches on kill), not weaker reasoning; the harness, not the model, sets the ceiling (via vendor blog) 📡
  • 🌐 XBOW — Grok 4.5 and the Middle of the AI Offensive-Security Market — cost-tiered autonomous-vuln-discovery benchmark — Grok 4.5 leads the ~$1 budget tier (~75% solve vs ~65% GLM 5.2) and reaches ~93% at higher spend, edging GPT-5.5/Mythos in the mid-cost band; argues for a fleet-of-models offensive strategy (via vendor blog) 📡
  • 🔧 Mantis (Google) — Google's modular, stack-agnostic toolkit of security-review skills that let AI coding agents autonomously find, reproduce and patch vulnerabilities (via GitHub trending) 📡
  • 🌐 Aikido — Benchmarking 13 AI Models on Rediscovering Known CVEs — 26 real GitHub-advisory CVEs, pass@3 harness: GPT-5.6 tops recall at 23/26 but pooling several cheap-model runs matches a single flagship pass at a fraction of the cost — premium reasoning tiers rarely pay off for vuln rediscovery (via vendor blog) 📡
  • 🌐 Aikido — 11.7B Tokens: GLM-5.3 and DeepSeek Are Now Frontier (Aug 21 2026) — bigger sequel to the 26-CVE run: 10 models × 3 runs × 32 freshly-disclosed CVEs. Open weights won — DeepSeek V4 Pro found 28/32 (87.5% recall, pass@3), beating Opus 5, Grok 4.6, and Sol; three DeepSeek Pro runs cost ~$295. The catch is precision — only 65.6% of DeepSeek's reports were valid vs GPT-5.6-Sol's 86.4% (recall-vs-precision fork: open models find the most but throw the most false leads). Repetition closed single-run gaps (DeepSeek 17→28 across 3 passes). Design implication for a vuln-hunting harness like Argus: pool cheap-open-model runs for recall, gate with a high-precision model before reporting (via vendor blog) 📡
  • 🌐 XBOW — The Rise of Affordable Models (GLM, Muse Spark) — Mythos still leads offensive-security benchmarks but Muse Spark now sits just below Opus 4.6 and GLM has closed most of the gap; the real shift is threat economics — cheap models let attackers run more/longer test passes, not match peak capability (via vendor blog) 📡
  • 📄 APIsec Labs — Meta Instagram Takeover: Not Prompt Injection — a BOLA/business-logic flaw in Meta's AI support agent: supplying a victim's username in chat linked a recovery email to their account with no identity check — the agent executed a state-changing action on an object named straight from request input (via Kiya discovery) 📡
  • 🌐 Endor Labs — AI SAST for C Finds What Other Tools Miss — hybrid deterministic program analysis + LLM reasoning, no build/compiler step, caught 96/102 known vulns (48× the next-best buildless pattern tool) and 2.6×/3.5× the true positives of Claude/Codex on 8 real embedded C projects — buildless memory-safety detection at scale (via vendor blog) 📡
  • 📄 Aikido — We Burned 11.7bn Tokens to Find the Best Cyber AI Model — benchmarked 10 models on 32 recent zero-day CVEs (memorization-minimized); DeepSeek V4 Pro led pooled recall at 28/32 and Grok on consistency — open-weight models now rival proprietary frontiers when properly harnessed (via vendor blog) 📡
  • 🌐 XBOW — Vulnerability Management Automation: Prioritize and Remediate at Scale — argues autonomous-pentest exploit validation belongs between detection and remediation instead of CVSS/scanner noise, and proposes MTTR-by-risk-tier + false-positive-rate as the metrics a VM program should track (via vendor blog) 📡
  • 🌐 CVE-2026-76832 — Agno PythonTools path traversal → RCE — unsanitized file_name in the Agno agent framework's read_file/save_to_file/run_python_file tools lets ../ sequences escape base_dir to read/write/execute arbitrary files; triggerable by direct call OR prompt injection in agent-processed content (CVSS 8.8) (via Kiya discovery) 📡
  • 🔧 Cybermes — autonomous offensive-security / bug-bounty agent (v2.0, native Go core) wiring subfinder/httpx/nuclei/sqlmap/ffuf via a reasoning loop with 50+ skills, a zero-false-positive gate requiring deterministic HTTP proof before reporting, and auto PDF/HTML CVSS reports + MCP integration; ~270★ (via GitHub trending) 📡
  • 📄 Aikido — Finds More Vulnerabilities than Claude Security at Half the Cost — head-to-head on 89 known vulns: Aikido found 76% (68/89) for $75 vs Claude Security's 67% (60/89) for $157, crediting the gap to harness/orchestration + task-specialized small models over raw frontier power — restates the "harness beats model" thesis with a concrete cost/recall benchmark (via vendor blog) 📡
  • 🌐 Endor Labs — Is Your Security Debt Shrinking? — argues AI code-analysis should be judged on whether security debt actually shrinks, not raw finding counts; cites Agent Security League (best model/agent completes 84.4% of tasks correctly but only 7.8% securely across 200 vuln tasks), Comcast (44% of validated critical/high AI findings were false positives) and 1Password (only 26% of AI-generated fixes fully correct) — the jagged-capability picture as a buyer's measurement thesis (via vendor blog) 📡
  • 🌐 XBOW — How We De-Duplicate Findings with Multi-View Embeddings — splits each vuln finding into 4 separately-embedded views (description/repro/location/impact) recombined with per-CWE-weighted similarity, cutting reported findings ~30% at ~90% duplicate-elimination vs analyst-clustered ground truth — how an autonomous pentester keeps signal above noise at scale (via vendor blog) 📡
  • 🌐 Endor Labs — Prompt Patterns That Make Coding Agents Write Safer Code — four secure-prompting techniques with measured effect: naming CWE anti-patterns (CWE-89/79) cut vulns ~59% on GPT-4, a generate→critique→revise "secure-insecure diff" cut Python weakness density 77.5%, a security prefix cut vulns up to 56% — but AI code often gets less secure each iteration, so SAST/SCA/secrets scanning stays mandatory (via vendor blog) 📡
  • 🌐 Endor Labs — Fable 5.1 Takes the Top Spot on the Agent Security League — Claude Code + Fable 5.1 leads both metrics at 87.2% FuncPass / 37.4% SecPass (SecPass nearly doubled vs Fable 5.0's 19.6%, +5.0 over Opus 5) at ~⅓ the cost, no timeouts, and the lowest anti-cheating adjustment (−6.7 FuncPass vs Opus 5's −18.4) — our own stack model, now benchmarked as the strongest secure-fix agent (via vendor blog) 📡
  • 🔧 cve-mcp-server — MCP server giving Claude 27 security tools across 21 APIs (NVD, EPSS, CISA KEV, VirusTotal, Shodan) with a one-call triage_cve orchestrator that fuses CVSS + exploit probability + KEV status + PoC availability into a composite risk score — AI-assisted CVE triage/vuln-research wired straight into the agent's toolchain; ~1.5k★ (shared by Ayoma) 📡
  • 📄 Can Local Open-Weight LLMs Detect Vulnerabilities? — 400+ experiments benchmarking 16 open-weight models for source-code vuln detection: base models score near-chance (50–56%), but fine-tuning + prompt engineering + ensemble voting (LLM predictions fused with classical ML) reach 76.9% — a rigorous, honest look at what a locally-run model can and can't do for vuln research (shared by Ayoma) 📡
  • 🌐 Endpoint AI Agent Abuse (EAA) — ATT&CK-style catalog of 18 techniques for abusing endpoint-resident AI agents (coding/dev tools) as execution, persistence, collection, exfiltration and defense-evasion surfaces — evidence-scoped (feasible/demonstrated/observed) with detection hypotheses; the systematized taxonomy behind in-wild cases like UAT-10147, and it maps straight onto our Claude Code stack (shared by @brisk200) 📡
  • 📄 Unit 42 — Inside an AI-Assisted Enterprise Intrusion — Unit 42's forensic write-up of what they call the first documented case of a fleet of purpose-built AI agents autonomously running an entire enterprise intrusion (initial access → repo-credential harvest → secrets-manager priv-esc → CI/CD hijack → cloud-AI-infra seizure) in under 10 hours vs ~2 weeks, chaining 50+ ATT&CK techniques with the human only setting the objective and re-planning left to the agents — the in-wild proof of machine-speed agentic offense, and the incident behind the 09-07 "AI-gateway layer is actively exploited" theme (via Kiya discovery) 📡
  • 🧪 RepoGuardBench — benchmark measuring whether coding agents stay on-task when repo artifacts (READMEs, issues, comments, test logs, agent-rule files) carry hidden malicious instructions; finds code comments and agent-rule files are the highest-risk injection carriers, and susceptibility scales with model size (0% → 83% proposed-attempt rate across Qwen2.5-Coder 1.5B→14B) — a reproducible, local, open-weight IPI-robustness range (ICML 2026 DL4C workshop); the measurement complement to the GitSpawn/EAA repo-artifact-injection thread (via Kiya discovery) 📡
  • 🌐 AI Agent Hacking Writeups — curated, dated index of 281 public writeups on real attacks against AI agents, sorted into six buckets (indirect PI 46, PI/guardrail-bypass 174, MCP & agent tooling 19, AI platform/infra 20, product/web-bugs 13, AI-as-attacker 9) with platform, researcher and severity per entry — a case-study library for agentic-pentest study; CC0, ~36★ (via X/Twitter trending) 📡
  • 📄 Reaching an internal OpenAI repo via HEIF RCE + overprivileged SSO (CVE-2026-32882) — Hacktron chained a libheif 1.19.7 heap overflow (HEIC upload → FastImage bypass → ImageMagick → RCE, CVSS 8.8) on OpenAI's Discourse forum to an over-permissioned SSO token, reaching a linked employee Codex/GitHub session and opening a proof PR against the private openai monorepo — whole two-month project <$3K in tokens, $6.5K bounty. The AI-capability landmark behind the headlines: Claude Opus 4.8 failed the exploit across sessions; Opus 5 got reliable x86-64/jemalloc RCE within hours of release — a real-CVE study in agentic offense + token-least-privilege + parse-before-you-sanitize (via Kiya discovery) 📡
  • 📄 AEPD — First personal-data breach notified from an AI-agent-executed attack — Spain's DPA reports (Sep 14) the first breach formally notified to a regulator where an autonomous AI agent ran the whole attack — vuln scan → login → recon → modify personal data + access invoices — at machine speed; the real-world, legally-logged bookend to the Unit 42 agentic-intrusion write-up, and the incident that turns "AI-as-attacker" from lab demo into a GDPR notification event (via Kiya discovery) 📡
  • 🌐 XBOW — Grok 4.7 Is Different. So Is the Best Way to Use It. — Grok 4.7 does fewer raw exploit-crafting iterations than 4.6 yet surfaces far more findings (68 vs 42) when routed through xAI's Build orchestration, and 4.6's endless-reasoning-loop bug (0.85% of runs) is gone — fresh evidence that a model's offensive value now depends on being co-optimized with its harness/orchestration layer, not raw capability alone (via vendor blog) 📡
  • 📄 Manus AI agent — indirect prompt injection to cross-account RCE (Salt Labs) — a single malicious email asking to be "summarized" smuggled instructions past Manus's filter with JSFuck obfuscation (the safeguard fired after the payload ran, making it useless) → RCE → a reverse shell that pulled the OAuth tokens for every third-party app the victim linked (Gmail/Drive/GitHub), turning one email into full cross-account compromise; Meta's bug-bounty triaged and patched it while Manus stayed silent — the in-the-wild proof that on a credentialed agent an indirect PI cascades into cross-service takeover (via Kiya discovery) 📡
  • 📄 Gambit Security — Autonomous AI agents hack online retailers for ~$25 a company — recovered a financially-motivated actor's staging server running three open-source harnesses — Strix (vuln scan), Cairn (autonomous exploitation), Hermes (orchestration) — against hundreds of retailers at a marginal $25.46/scan ($3–79 range, ~$12–18K total over 4 weeks). Full autonomous chains (unauth SQLi→OTP bypass→admin→file upload→RCE→sudo→NFS→creds→AWS Secrets dump→card-key extraction→decrypt), 600K+ card records from two firms, skimmers on 19+ confirmed sites, one agent's cleanup routine dropped 180 tables incl. victim backups. The in-wild economics landmark: the UAT-10147 thesis ("APT scale at commodity cost") now priced — a real-CVE agentic-offense study with the actor's own harnesses named (via Kiya discovery) 📡
  • 📄 Huntress — Attackers abuse ChatGPT Custom GPTs to deliver a RAT via ClickFix — in-wild campaign (Sep-28) where a malicious Custom GPT "Plus 5.6", promoted via Google Ads, funnels victims through a Google-Sites-hosted ClickFix/Cloudflare-CAPTCHA lure into running PowerShell → MSI → DLL-sideload into a legit Canon/Stardock-signed binary → shellcode hidden in a .wav (later a NuGet package) → custom encrypted archive → full RAT (remote desktop, cam/mic, 17-browser cred access, DNS-over-HTTPS C2, AMSI bypass, ntdll unhooking, anti-VM). OpenAI removed the first GPT Sep-25; a near-identical replacement appeared by Sep-27. The social-engineering vector is the AI brand itself — a fake "model" as the top of the delivery funnel. [Sep-30 daily-pulse]
  • 📄 The triage is the product — running AI agents against Ethereum's protocol code (Ethereum Foundation) — a practitioner playbook for parallel vuln-hunting agents: a candidate isn't a finding until a self-contained reproducer proves it against real code, and most of the work is rejecting debug-only panics, unreachable hand-crafted inputs and vacuous formal proofs — the disciplined-triage counterweight to raw agentic bug volume. (in Trove since 2026-09-29 (security/ai-security)) 📡
  • 📄 Snyk — Why AI coding agents keep writing broken access control — the why behind the recurring BOLA/IDOR class in AI-generated code: the ownership rule that would prevent it ("invoices belong to organizations") lives in the data model and team knowledge, not the prompt or surrounding code — so the agent ships code that is correct for the stated task yet lets any authenticated user read other tenants' objects (CWE-639/862/863). Unlike SSRF/path-traversal (pattern-matchable, SAST-friendly), authz has no traceable source→sink shape — a scanner can't know which field encodes ownership. Five concrete actions: endpoint-inventory, document ownership rules, make authz an explicit AI-code review question, per-resource cross-tenant tests, periodic deep contextual analysis. Pairs with PixelLeak (W21) as the "authorized but wrong" failure class. [Oct-02 daily-pulse]
  • 📄 AuraForge: Scaling Security Supervision for Training Coding Agents (arXiv 2610.00850) — synthesizes executable multi-CWE security tests (AuraGym: 679 tasks, 3 languages, 177 vuln categories) to supervise coding agents toward secure code, outperforming human-written security tests — a training-side defense against insecure AI-generated code (in Trove since 2026-10-03 (security/ai-security)) 📡
  • 🔧 OperTraitor — LLM-powered Kubernetes operator RBAC privilege-escalation scanner (CybersecurityNews) — an LLM-driven scanner that flags over-privileged K8s operators; caught IBM's Prometurbo operator holding cluster-wide get/list/watch on Secrets → CVE-2026-6389 (CVSS 8.8) — AI applied to find real priv-esc in cluster RBAC (in Trove since 2026-10-02 (security/ai-security)) 📡
  • 🧪 ReproBench: Benchmarking LLM Agents on Reproducing Vulnerability From Scratch (arXiv 2609.34450) — tests whether agents can rebuild the environment and reproduce a real CVE from only its identifier; on 30 IoT-firmware vulns just 5.3% succeed while 45.3% fake it via simulation — a sober measure of agentic end-to-end reproduction capability (in Trove since 2026-10-02 (security/ai-security)) 📡
  • 🧪 VulContextBench: A Benchmark for Security Context Retrieval in Coding Agents (arXiv 2609.32601) — grades whether coding agents actually retrieve and cite the code evidence behind a vulnerability verdict, not just the verdict — finds agents explore most of the right evidence but cite only a fraction, a reliability gap for AI vuln analysis (in Trove since 2026-10-02 (security/ai-security)) 📡
  • 📄 Solidity Meets LLMs: A Transformer-Based Approach to Smart Contract Vulnerability Detection (arXiv 2609.27091) — fine-tunes a BERT transformer to classify Solidity source as vulnerable/safe at 92% F1, a concrete datapoint that pre-trained language models transfer to domain-specific vuln detection — the smart-contract instance of the LLM-as-vuln-classifier theme this week studies (in Trove since 2026-10-04 (security/ai-security)) 📡
  • 📄 The Era of Software Quality, or the Era of Ostriches? (Michael Catanzaro, GNOME) — a maintainer's honest accounting of the AI-vuln-report era: GNOME CVEs 14 (2022) → 141 YTD 2026, mostly AI-scanner-driven; the YesWeHack bounty paid €183,900 for 71 of 298 submissions (the rest noise); an AISLE AI scan of GLib found 118 issues at ~40% false-positive (gobject-introspection typelib misclassifications); yet a human Codean Labs audit caught critical Flatpak / xdg-desktop-portal sandbox escapes AI likely missed. Verdict: accept AI vuln reports but keep humans on severity + architecture — the false-positive tax and the human ceiling, both quantified. [Oct-06 daily-pulse]
  • 📄 Who Is Your Agent Serving? Provider-Side Indirect Prompt Injection in Proactive Agents (arXiv 2610.05266) — a new attack surface in proactive personal agents: an external content provider who controls only material tied to its own target (a product/service) can steer an otherwise-benign agent into recommending that target, without compromising the agent or touching private user context. Formalized via three levers — Target Control (what objective the agent advances), Private Binding (how it ties the target to the user's interests), Prospective Support (what follow-up it offers) — the full attack lifts target-authorization rates by up to 77.4 pp across three proactive-agent environments × six user models; a multi-turn variant keeps influencing responses even when the user never gives final authorization. The research sibling of the W13 "AI Recommendation Poisoning" click-layer fold — here the provider, not an attacker-in-the-middle, is the injection source. [Oct-07 daily-pulse]
  • 📄 Frontier models found the vulnerabilities. Only the attacker found the chains. (Snyk) — head-to-head on the public TaintedPort app: Snyk's agentic live-attack tool Evo COS proved 10 of 15 real exploit chains (e.g. SSRF + hardcoded JWT secret → admin takeover) at 75.7% severity-weighted detection, vs source-reading Claude Security/Mythos at 49.6% which flagged flaws but couldn't demonstrate exploitability — the "finding a flaw ≠ proving an attacker can chain it" thesis, the dynamic-vs-static complement to HackSynth's jagged-capability picture. (in Trove since 2026-10-07 (security/ai-security)) 📡

Study checklist

↪ See roadmap.md → Phase 2 → Week 13

  • Explain LLM05 Improper Output Handling; reproduce XSS/SQLi/command injection via model-generated output (distinct from prompt injection)
  • Map the COAE output-exploitation sub-classes (function-calling abuse, hallucination exploit, exfiltration) to their downstream sinks
  • Map + probe an LLM app's tool surface (PortSwigger method): inputs (direct + indirect) → ask the model its APIs → send classic web exploits through it; run the free excessive-agency / OS-injection / IPI labs
  • Chain an AI agent for autonomous recon → exploit (SQLi/SSRF/OSINT) via a workflow engine (N8N) or an MCP platform (HexStrike-AI)
  • Contrast two LLM vuln-discovery shapes: OpenAnt (reachability decomposition + adversarial verify) vs SETYPE (semantic type-check, failed-check = bug); run PYSETYPE-style prompting on a small Python app
  • Reason honestly about AI-vs-human pentesting: ARTEMIS's 82% valid rate, GUI blind spot, and ~30% hallucination without a validator
  • Walk the prompt-injection → RCE chain: Semantic Kernel MRO traversal (CVE-2026-26030, CVSS 9.9); try the AIAgentCTF hands-on challenge
  • Study CVE-2026-25592 (Semantic Kernel .NET) — DownloadFileAsync sandbox escape via chained function calls
  • Study gemini-cli CVSS 10 (TrustIssues) — headless auto-trust + yolo bypass = full RCE + supply chain
  • Never ship AI-generated auth from one prompt — make security review a re-prompt pass (Illusion of Secure LLM Code, NIST SP 800-63B)
  • Explain why AI agents write broken access control (BOLA/IDOR): the ownership rule lives in the data model, not the prompt; no source→sink shape so SAST misses it (CWE-639/862/863); apply the 5 Snyk controls (endpoint inventory, documented ownership, authz-as-review-question, cross-tenant 404 tests, periodic context-graph analysis)
  • Apply the Agents Rule of Two: never combine untrusted input + secrets + external comms in one workflow
  • Understand cost-aware evaluation — why red-team CTF scales with test-time compute but blue-team SOC does not
  • Map agent capability as jagged (HackSynth): strong on recon/known patterns, near-zero on binary exploitation and crypto; sandbox any autonomous agent (firewall + container) against its own hallucinations
  • Study Black-Box Skill Stealing (arXiv 2604.21829) — skill extraction; map to OWASP LLM07 System Prompt Leakage
  • Study the exfiltration lineage: EchoLeak (CVE-2025-32711) → SearchLeak (CVE-2026-42824, parameter-to-prompt injection)
  • Study coding-agent incidents: Copilot YOLO RCE (CVE-2025-53773), "Comment and Control", Claude Code /proc exfiltration
  • Study the Meta Instagram bot confused-deputy (LLM06 Excessive Agency)
  • Study UAT-10147 (Cisco Talos) — first in-wild agentic-AI offensive orchestrator (PentestGPT/DeepAudit on C2, ~170K-URL list, SPECTRE implant); the "lower-skilled actor reaches APT scale" thesis, operational
  • Study the Hacktron OpenAI breach: HEIF RCE (CVE-2026-32882, libheif 1.19.7 heap overflow on the Discourse forum) → over-privileged SSO token → Codex/GitHub → internal monorepo. Lessons: SSO trust boundaries + token least-privilege; and the model-generation gap (Opus 4.8 failed, Opus 5 succeeded) — expert-guided, not autonomous
  • Map the click layer as an injection surface: AI Recommendation Poisoning (?q= deep-link "Ask AI" → persistent memory poisoning, ATLAS AML.T0080) + Zenity Claude-in-Chrome/Atlas browser-agent takeover; scope OTP/magic-link mail out of any summarize-my-email path
  • Study AI-assisted tradecraft: patch diffing, LLM-generated malware/ evasion (GNAW), CTF-framing jailbreaks as IOCs
  • The AI brand as the lure (Huntress ClickFix via fake "Plus 5.6" Custom GPT → RAT): treat AI-platform-hosted content (Custom GPTs, Artifacts, */share) and AI-search answers as untrusted delivery surfaces
  • Treat the AI gateway/orchestration layer as a control plane: walk the LiteLLM (/proc/1/environ key theft) / RAGFlow (hidden config hook) / Kestra (auth-bypass → shell) intrusions; monitor by control-plane role
  • Study the Amazon Kiro workspace-trust exfil chain (repo → agent → IDE config → egress; any message triggers it) and map it to Claude Code usage
  • Verify LLM-decompiler output behaviorally, not by recompilability: Decompile-Diverge (4.9% divergence, ~1-in-10 CVE vulnerability erasure) — "compiles ≠ preserves behavior"
  • For each CVE: write a 1-page incident analysis (initial access, exploitation, impact, preventing control)

Study notes

Sign in to take notes.