OSAI+ Exam - Guided By RedBlock

Updated 2026-10-05· 102 min read· 121 views
Share:

AI-300: Advanced AI Red Teaming - OffSec AI Red Teamer (OSAI)

OSAI+ Exam - Guided By RedBlock

AI Red Teaming & AI Security Field Guide — OSAI / AI-300 Study Companion

A study companion for advanced AI red teaming (OffSec OSAI / AI-300 topic area) — the security assessment of modern AI systems: generative AI, LLM applications, AI agents, multi-agent systems, RAG pipelines, embeddings, model-context/tool surfaces, the AI/ML supply chain, and AI infrastructure. It maps the attack surface, the published attack taxonomy, reconnaissance and threat-modeling methodology, and — for every attack class — how it works conceptually, how a red teamer tests for it, and how to detect and defend against it.

Scope & framing (read first). AI red teaming is a legitimate discipline whose purpose is to find and fix weaknesses in AI systems so they can be deployed safely. This guide is written for that purpose — assessing and hardening AI deployments you own or are authorized to test. It covers the concepts, methodology, and defenses at the level of the public corpus (OWASP LLM Top 10 & ML Top 10, MITRE ATLAS, NIST AI RMF, and published research). It deliberately does not provide a working jailbreak / prompt-injection cookbook or turnkey payloads to make AI systems behave harmfully — understanding the categories, mechanisms, and defenses is what makes an effective, responsible AI red teamer.
⚠️ Authorized testing only. Assess only AI systems, models, APIs, and infrastructure you own or are explicitly contracted to test. Follow responsible-disclosure and the provider's acceptable-use terms. AI red teaming that improves safety is valuable; attacking third-party AI services without permission is not.


Table of Contents

  1. Introduction — How AI Changes the Attack Surface

  2. The AI Red Team Lifecycle & Engagement Model

  3. Core Frameworks: OWASP LLM Top 10, ML Top 10, MITRE ATLAS, NIST AI RMF

  4. Understanding the AI/ML Stack (what you're assessing)

  5. Reconnaissance for AI Targets

  6. Threat Modeling for AI-Enabled Systems

  7. Prompt Injection — Concept, Testing & Defense

  8. Jailbreaking & Model Guardrail Evaluation

  9. Attacking AI Agents — Memory & Autonomy (Concept & Defense)

  10. Multi-Agent Systems & A2A Trust Boundaries

  11. RAG Pipeline Security & Data-Poisoning Resistance

  12. Embedding & Vector-Database Security

  13. Model Context Protocol (MCP) & Tool-Surface Security

  14. Model Extraction & Adversarial Machine Learning (Concept & Defense)

  15. Sensitive-Data Exposure & Privacy Attacks (Concept & Defense)

  16. AI/ML Supply-Chain Security

  17. AI Infrastructure & Deployment Security

  18. Detection, Monitoring & Guardrail Architecture

  19. Defensive Architecture — Securing AI Systems

  20. MITRE ATLAS — Technique Reference & Examples

  21. AI Infrastructure Hardening Playbook (with examples)

  22. AI/ML Supply-Chain Security Deep-Dive (with examples)

  23. Detection Engineering for AI Systems (with example logic)

  24. AI Security Governance & Program (NIST AI RMF Applied)

  25. Secure AI Development Lifecycle & Design Patterns

  26. Third-Party & Vendor AI Risk Assessment

  27. Expanded Detection-Rule Library (example logic)

  28. AI Security Testing Tools & Techniques (authorized/defensive)

  29. AI Incident Response

  30. Lessons from Real AI Security Incidents (public)

  31. The Capstone Engagement — Methodology

  32. Worked Example — Threat Modeling an Enterprise AI Assistant

  33. Worked Threat Model #2 — Agentic Coding Assistant

  34. Worked Threat Model #3 — Customer-Support Chatbot

  35. Worked Threat Model #4 — Multi-Agent Data Pipeline

  36. Worked Threat Model #5 — AI Embedded in Security Tooling

  37. Worked Example — Capstone Engagement Walkthrough (benign)

  38. Reporting AI Security Findings

  39. Study Plan, Resources & Glossary


1. Introduction — How AI Changes the Attack Surface

AI — especially generative AI and LLM-powered applications — introduces attack surface that traditional application security doesn't cover. Red-teaming AI means understanding what's genuinely new while still applying classic security fundamentals to the infrastructure around the model.

What's new about AI systems:
- Natural language is now an interface — and an attack surface. In a traditional app, code is trusted and data is untrusted. In an LLM app, the instructions and the data arrive through the same channel (text), and the model can't reliably tell them apart. That blurring is the root of prompt injection — the signature AI-specific vulnerability.
- Non-determinism and emergent behavior. The same input can produce different outputs; behavior is probabilistic, not specified. You can't fully enumerate what a model will do, which complicates both attack and defense.
- The model is a learned artifact, not written logic. Its "rules" are statistical patterns in weights, not auditable code. This enables attacks with no classic analog — adversarial examples, model extraction, training-data extraction, embedding inversion.
- Agents act in the world. Modern AI "agents" don't just produce text — they call tools, browse, run code, read/write data, and chain steps autonomously. That turns a text-generation flaw into a path to real actions (data access, API calls, code execution) — the AI equivalent of turning an injection into RCE.
- Trust is distributed. RAG pipelines trust retrieved documents; agents trust tool outputs and each other; the model trusts its context window. Each trust relationship is a boundary an adversary can target.
- The supply chain is data + models + code. Beyond software dependencies, AI systems depend on datasets, pre-trained weights, fine-tuning adapters, and model hubs — new supply-chain vectors.

What stays the same: the infrastructure around AI is still classic security. Model servers are web services (API auth, access control, injection in the surrounding app). ML workloads run in containers and clouds (standard container/cloud hardening). Secrets, networking, IAM, logging — all the traditional discipline still applies, and is often where the most impactful (and most fixable) findings are.

The red teamer's framing. An AI red team assesses the whole system — the model, the application around it, the data pipelines, the agent/tool integrations, and the infrastructure — against a realistic adversary, to find how an attacker could cause harm (data exfiltration, unauthorized actions, safety/policy violations, model theft, service disruption) and to give the defenders a prioritized path to fix it. The goal is safer AI deployments, not exploitation for its own sake.


2. The AI Red Team Lifecycle & Engagement Model

AI red teaming adapts the classic red-team/pentest lifecycle to AI-specific assets while keeping the professional engagement structure.

The lifecycle (mapped to AI):

1. SCOPING & RULES OF ENGAGEMENT
   - which AI systems/models/apps/APIs are in scope; what's off-limits; data-handling rules;
     acceptable-use/provider terms; safety constraints (e.g., don't generate real harmful content —
     use benign proxies to prove a control gap); disclosure process.
2. RECONNAISSANCE
   - discover AI assets: LLM apps, model endpoints, RAG/vector DBs, agents, tool integrations,
     MLOps infra, dependencies, exposed services (§5).
3. THREAT MODELING
   - identify high-value AI assets, trust boundaries, data flows, and plausible attack paths (§6).
4. VULNERABILITY ASSESSMENT / TESTING
   - test each component against its relevant attack classes (prompt injection, RAG poisoning,
     tool abuse, model extraction, infra flaws, supply-chain), prioritizing by threat model.
5. EXPLOITATION / IMPACT VALIDATION
   - demonstrate realistic impact with the minimum necessary proof (a benign canary, a controlled
     data-access demonstration) — enough to prove the control gap without causing real harm.
6. POST-EXPLOITATION / CHAINING
   - show how AI-layer findings combine with infrastructure findings into business impact.
7. REPORTING & REMEDIATION
   - document findings, impact, and prioritized, actionable fixes; support the blue team (§21).

AI-specific engagement considerations:
- Safety-aware testing. The goal is to prove a control gap, not to actually produce harmful outputs or cause damage. Use benign canaries and proxies — e.g., demonstrate that a guardrail can be bypassed by eliciting a harmless-but-policy-restricted token/marker, or that an agent would exfiltrate data by having it move a planted non-sensitive canary file. You document the vulnerability without generating real harm.
- Non-determinism. Test findings multiple times; report reproducibility/likelihood (a flaw that triggers 1-in-5 is still a finding, but characterize it).
- Data sensitivity & privacy. AI tests can surface training data or user data — handle any such exposure per the rules of engagement, minimize collection, and report responsibly.
- Provider terms & legality. Hosted models have acceptable-use policies and often dedicated vulnerability-disclosure/red-team programs — work within them.
- Human + automated. Combine systematic automated probing (where appropriate and authorized) with human creativity and judgment; AI behavior is nuanced.

The professional purpose. Every finding should map to a defensive improvement. An AI red team's deliverable is a safer system: better guardrails, tighter tool permissions, validated data pipelines, hardened infrastructure, and improved detection.


3. Core Frameworks: OWASP LLM Top 10, ML Top 10, MITRE ATLAS, NIST AI RMF

These public frameworks are the shared language and the methodology backbone of AI security. Learn them — they organize the whole field and keep your assessment systematic.

OWASP Top 10 for LLM Applications (the central taxonomy for generative-AI apps). The categories (names evolve across versions, but conceptually):

LLM01  Prompt Injection            — untrusted input manipulating the model's behavior (direct & indirect)
LLM02  Sensitive Information Disclosure — leaking secrets, PII, system prompts, training data
LLM03  Supply Chain                — compromised models, datasets, adapters, plugins, dependencies
LLM04  Data & Model Poisoning      — tampering with training/fine-tuning/RAG data to corrupt behavior
LLM05  Improper Output Handling    — trusting model output unsafely (→ XSS/SSRF/RCE downstream)
LLM06  Excessive Agency            — agents with too many permissions/tools/autonomy
LLM07  System Prompt Leakage       — exposure of the system prompt / instructions
LLM08  Vector & Embedding Weaknesses — RAG/vector-DB attacks, embedding inversion, cross-tenant leakage
LLM09  Misinformation / Overreliance — unsafe trust in incorrect model output
LLM10  Unbounded Consumption       — resource exhaustion, denial-of-wallet, model-extraction via abuse

Use it as your test checklist for any LLM application.

OWASP Machine Learning Security Top 10 (for ML systems broadly): input manipulation (adversarial examples), data poisoning, model inversion, membership inference, model theft/extraction, supply-chain, transfer-learning attacks, model skewing, output integrity, and poisoned/compromised packages.

MITRE ATLAS (Adversarial Threat Landscape for AI Systems) — the ATT&CK-style knowledge base for AI. It provides tactics (reconnaissance, resource development, initial access, ML model access, execution, persistence, defense evasion, discovery, collection, exfiltration, impact) and techniques specific to AI/ML, with real-world case studies. Map every finding to an ATLAS technique for standardized communication — it's the AI equivalent of citing an ATT&CK ID.

NIST AI Risk Management Framework (AI RMF) and the NIST AI 100-2 (Adversarial ML taxonomy) — the governance and risk side: the functions Govern, Map, Measure, Manage, and a formal taxonomy of attack types (evasion, poisoning, privacy, abuse). Useful for framing risk and remediation in the report.

Also relevant: the MITRE ATLAS matrix for techniques, Google's SAIF (Secure AI Framework), Microsoft's AI red team guidance, and Anthropic's / OpenAI's responsible-scaling and red-team publications — all public, all worth reading for how practitioners and labs think about this.

Why frameworks matter: they make your assessment complete (you test every category), standardized (you communicate in shared terms), and credible (you map to recognized references). The combination — OWASP for the "what," ATLAS for the "how/standard naming," NIST for the "risk/governance" — is the professional backbone.


4. Understanding the AI/ML Stack (what you're assessing)

You can't red-team what you don't understand. Know the components of a modern AI application and where each can fail.

The anatomy of a typical LLM application:

USER / CLIENT
   │  (natural-language input — the primary new attack surface)
   ▼
APPLICATION LAYER (the "AI app")
   - system prompt / instructions        - input/output guardrails & filters
   - conversation/session & MEMORY        - orchestration logic
   - prompt templates                      - authentication / authorization / rate limiting
   │
   ├──► RAG PIPELINE  ── retriever ──► VECTOR DATABASE (embeddings of a knowledge base)
   │         (fetches "context" documents to ground the model)
   │
   ├──► TOOLS / FUNCTIONS / PLUGINS (via function-calling / MCP)
   │         - web browsing, code execution, database queries, API calls, file access
   │
   ▼
MODEL / INFERENCE LAYER
   - the LLM (hosted API or self-hosted model server)
   - model weights, adapters (LoRA), tokenizer, inference engine
   │
   ▼
INFRASTRUCTURE / MLOps
   - model servers (e.g., inference frameworks), containers, GPUs, cloud, storage
   - training/fine-tuning pipelines, model registry/hub, datasets, CI/CD

Where each component can fail (the map to later sections):
- Natural-language input → prompt injection, jailbreaking (§7-8).
- System prompt / instructions → system-prompt leakage, instruction override (§7).
- Memory / session → memory poisoning, cross-session leakage (§9).
- RAG / vector DB → data poisoning, retrieval manipulation, cross-tenant/embedding leakage (§11-12).
- Tools / functions / MCP → excessive agency, tool abuse, unsafe output handling → downstream injection/RCE (§13, §9).
- Model / weights / adapters → model extraction, adversarial examples, supply-chain tampering (§14, §16).
- Infrastructure / MLOps → classic infra flaws: exposed model servers, insecure storage, weak IAM, container escape, poisoned CI/CD (§17).

Key concepts to internalize:
- The context window is the "trust zone" — everything in it (system prompt, user input, retrieved docs, tool outputs, memory) influences the model, and the model struggles to distinguish trusted instructions from untrusted data mixed into that window. This single fact underlies most LLM-layer attacks.
- Output handling matters as much as input — model output flowing unsanitized into a browser (→ XSS), a shell (→ RCE), a SQL query (→ SQLi), or another system is "improper output handling," a top cause of real impact.
- Agency = impact — a chatbot that only talks is low-impact; an agent that can act (tools, code, data) turns a prompt-injection into real-world consequences. The permission scope of the tools is the blast radius.

Understanding this stack lets you threat-model (§6) and target your testing where impact is highest.


5. Reconnaissance for AI Targets

Recon maps the AI attack surface before testing — discovering which AI assets exist, how they're exposed, and what they depend on. It blends classic recon with AI-specific discovery.

Goals: identify AI-enabled applications and features, model endpoints/APIs, RAG/vector databases, agent/tool integrations, MLOps infrastructure, and the data/model/dependency supply chain — and understand trust boundaries and exposure.

Discovering AI assets (methodology):
- Application-level signals — chat interfaces, "AI assistant"/"copilot" features, summarization/generation features, semantic search, recommendation engines. Observe the app's behavior and its API traffic (classic web recon, §ewpt/ewptx methodology) to find the AI endpoints behind the UI.
- API & endpoint discovery — look for inference endpoints (/chat, /completions, /embeddings, /v1/..., /predict, /generate), API docs (OpenAPI/Swagger), and the calls the frontend makes to its AI backend. Hidden/undocumented AI endpoints are common.
- Model/service fingerprinting — identify whether a hosted API (OpenAI/Anthropic/etc.) or a self-hosted model server (common inference frameworks) is used; version/behavior fingerprinting via responses, headers, and error messages. Self-hosted model servers sometimes expose management/metrics endpoints.
- RAG / vector-DB indicators — responses that cite internal documents, "based on your knowledge base" behavior, or exposed vector-database services (various vector DBs have default ports/management UIs — treat like any exposed datastore).
- Agent / tool surface — does the AI browse, run code, query databases, send emails, call APIs? Enumerate the tools/functions it can invoke (often discoverable through behavior or the app's function schemas).
- MLOps / infrastructure — model registries/hubs, experiment-tracking UIs, training pipelines, notebook servers (e.g., exposed Jupyter), object storage holding models/datasets, and container/orchestration surfaces. Many are found with standard infra recon (port scans, cloud enumeration, exposed-service discovery).
- Supply-chain intel — which models/datasets/frameworks/libraries are used (from docs, job postings, public repos, SBOMs, dependency manifests). This informs supply-chain testing (§16).

OSINT for AI targets: public repositories (leaked prompts, model names, API endpoints, keys in code), model-hub profiles, research/blog posts describing the system, job listings (reveal the AI/ML stack), and documentation. The same careful, passive-first approach as any recon.

Stealth & footprint. AI endpoints often have logging and anomaly detection (and hosted providers monitor abuse). Reconnaissance should respect scope and minimize noise; prefer passive discovery and documented API use over aggressive probing unless authorized.

Output of recon: an inventory of AI assets with their exposure, dependencies, tools/permissions, and data flows — the input to threat modeling (§6) and your testing plan.


6. Threat Modeling for AI-Enabled Systems

Threat modeling turns the asset inventory into a prioritized plan: what's valuable, who would attack it, how, and where the trust boundaries are. It's the intellectual core of a good AI red team engagement — and a defensive deliverable in its own right.

Identify high-value AI assets:

- the model itself (weights/IP — extraction/theft target)
- training/fine-tuning data & RAG knowledge base (confidentiality, integrity)
- the system prompt / proprietary instructions (IP, and a control surface)
- user data processed by the AI (privacy)
- the tools/actions the AI can take (the "agency" — blast radius)
- API keys / credentials the AI or its infra holds (access to backends, cloud, other systems)
- the AI's decisions/outputs where they drive business actions (integrity)

Map trust boundaries (where trust changes — the places to scrutinize):

user input → application           (untrusted text entering the system)
application → model context        (what gets placed in the context window, from where)
retrieved docs → model             (RAG: is retrieved content trusted as instructions?)
tool output → model                (does the model trust what a tool/API/web page returns?)
agent → agent                      (multi-agent trust: does agent B trust agent A's messages?)
model output → downstream systems  (is output sanitized before it hits a browser/DB/shell?)
model/data → infrastructure        (where models/data are stored, served, and moved)

Each boundary is a candidate attack path: untrusted data crossing into a trusted position is the recurring danger (e.g., a poisoned RAG document being treated as an instruction).

Enumerate attack paths (STRIDE-style, adapted for AI):
- Spoofing — agent impersonation, forged tool/agent messages, identity confusion in multi-agent systems.
- Tampering — data/model poisoning (training, fine-tuning, RAG), prompt/instruction manipulation.
- Repudiation — insufficient logging of AI decisions/actions (can't trace what the AI did and why).
- Information disclosure — system-prompt leakage, training-data/PII extraction, embedding inversion, cross-tenant leakage.
- Denial of service / wallet — resource exhaustion, expensive-query abuse, unbounded consumption.
- Elevation of privilege — excessive agency, tool-permission abuse, orchestration-layer escalation, infra compromise via the AI.

Attack trees & prioritization. For each high-value asset, build an attack tree (goal → sub-goals → techniques) and prioritize branches by likelihood × impact, weighted by the real trust boundaries and the system's agency. A talk-only chatbot with no tools and no sensitive data is low-risk; an agent with database and email access processing untrusted documents is high-risk — your effort should follow the risk.

The dual value. Threat modeling guides your testing and, delivered to the client, is itself a defensive artifact — it tells the blue team where their real risk concentrations and trust-boundary weaknesses are, independent of what you manage to exploit in the time box.


7. Prompt Injection — Concept, Testing & Defense

Prompt injection is the signature AI-application vulnerability (OWASP LLM01). This section covers what it is, why it works, how a red teamer assesses for it, and — most importantly — how to defend against it. (Consistent with the scope note, this is conceptual and defense-oriented, not a working payload cookbook.)

What it is. Prompt injection is when untrusted input causes the model to behave in ways the developer didn't intend — because the model processes developer instructions (the system prompt) and user/third-party data in the same text channel and can't reliably distinguish "instructions I should follow" from "data I should merely process."

Two forms:
- Direct prompt injection — the user's own input tries to override or subvert the system's instructions (e.g., convincing the app to ignore its configured behavior, reveal its system prompt, or act outside its intended role).
- Indirect prompt injection — malicious instructions are planted in content the model will later ingest — a web page the agent browses, a document in a RAG knowledge base, an email the assistant reads, a tool's output. When the model processes that content, the hidden instructions can influence it. This is especially dangerous for agents, because the "attacker" isn't the user at all — it's third-party data, and the impact can be actions taken on the user's behalf.

Why it's hard to fully fix. There's no clean, reliable separation between instructions and data in a single text stream the way there is between code and data in a parameterized SQL query. Models are trained to be helpful and to follow instructions in their context, which is exactly what injection abuses. Research continues, but there is no complete, universal fix today — defense is layered risk reduction, not elimination.

How a red teamer assesses for it (methodology, not payloads):
- Map the trust boundaries (§6) — identify every place untrusted content enters the model's context (user input, RAG docs, tool/web/email content, memory).
- Characterize the system's susceptibility — does untrusted content in the context change the model's behavior, leak its instructions, or influence its tool use? Use benign canaries to demonstrate influence without producing harmful output (e.g., show that planted content can cause the model to emit a harmless marker it otherwise wouldn't, proving the instruction/data boundary is porous).
- Assess impact, not just presence — the severity depends on what the model can do: can injected instructions cause data exfiltration, unauthorized tool calls, or harmful actions (via agency, §9/§13)? A susceptible chatbot with no tools is low impact; a susceptible agent with data/tool access is high.
- Test indirect vectors — this is where real-world impact concentrates: can a document, web page, or email introduced into the system's data flow influence an agent's behavior? Demonstrate with a benign planted canary.

Defenses (the constructive core — what the report should drive):
- Treat all non-system-prompt content as untrusted data, and architect so the model's privileged behavior doesn't depend on the model perfectly resisting injection.
- Least-privilege agency — the single most important mitigation: minimize the tools, permissions, and data the model can access, so even a successful injection has limited blast radius (§13, §19). Require human approval for consequential actions.
- Input/output guardrails — filter and classify inputs and outputs (dedicated guardrail models/classifiers) to catch manipulation attempts and unsafe outputs; strip/neutralize instructions in retrieved content where feasible.
- Strong system-prompt design & separation — clearly delimit system vs. user/retrieved content; use the provider's role/structure features; don't put secrets in the system prompt.
- Output handling — never trust model output as safe: sanitize/encode it before it reaches browsers, shells, databases, or other systems (prevents the injection→XSS/SSRF/RCE chain, LLM05).
- Provenance & isolation for ingested content — validate and tag the source of RAG documents and tool outputs; isolate tenants; don't let one user's content affect another's (§11-12).
- Monitoring & detection — log prompts, tool calls, and outputs; detect anomalous behavior and known injection patterns; rate-limit; alert on sensitive tool use (§18).
- Human-in-the-loop for high-impact actions.

The red teamer's deliverable: a clear picture of where untrusted content reaches the model, what it could influence, the impact given the system's agency, and a prioritized set of these defenses. The point is to make injection low-impact by design, since it can't be fully prevented.


8. Jailbreaking & Model Guardrail Evaluation

"Jailbreaking" refers to getting a model to bypass its safety training / usage policies. For an AI red teamer, the professional objective is evaluating the robustness of a model's/application's guardrails — characterizing how well they hold and where they're weak — so they can be strengthened. (Per the scope note, this section explains the concept, the evaluation methodology, and defenses — not a working jailbreak cookbook.)

What it is & why it matters. Models are trained and configured with safety behaviors and policies. Jailbreaking attempts to circumvent those so the model produces content or takes actions it's meant to refuse. For a deployer, the risk is reputational, legal/compliance, and — when combined with agency — operational (a jailbroken agent taking harmful actions). Guardrail robustness is therefore a legitimate thing to measure.

How guardrail evaluation is approached (methodology):
- Define the policy boundary — what should the system refuse or restrict? You can only evaluate robustness against a defined policy.
- Use benign proxies and canaries — a responsible red team measures whether a control can be bypassed without actually generating real harmful content. For example, configure a harmless but policy-restricted canary behavior and measure whether layered techniques can elicit it; or measure refusal consistency across paraphrases. This demonstrates guardrail weakness without producing dangerous outputs.
- Measure, don't just anecdote — characterize robustness statistically: refusal rate across many benign-proxy variations, consistency under paraphrase/obfuscation/context manipulation, and degradation under multi-turn pressure. Report likelihood and conditions, given non-determinism.
- Assess the whole stack, not just the base model — application-layer guardrails (input/output classifiers), system-prompt constraints, and tool-permission limits all contribute; a weak base model can be well-contained by strong application guardrails, and vice-versa.
- Focus on impact — a guardrail bypass matters most where it connects to agency or sensitive data. A model that can be coaxed into a policy-edge text output but can take no actions and touches no sensitive data is a different risk than one wired to tools.

Why models are susceptible (conceptual). Safety behavior is learned, probabilistic, and operates over an enormous input space that can't be fully covered; models are trained to be helpful and to follow context, which creates tension with refusal; and distribution shift (unusual framings/contexts) can move inputs away from where safety training is strongest. This is an active research area; robustness is improving but imperfect.

Defenses (what to recommend):
- Defense in depth — don't rely on the base model's refusals alone; add input and output guardrail classifiers, system-prompt constraints, and — critically — limit agency so a bypass can't cause real harm.
- Continuous evaluation — guardrail robustness isn't static; new techniques emerge. Recommend ongoing automated safety evaluation (regression testing of refusals), monitoring for bypass attempts, and a process to update guardrails.
- Policy clarity & context control — well-specified policies, tight control of what enters the context, and minimizing untrusted content reduce the attack surface.
- Monitoring & rate limiting — detect and throttle bypass attempts; log and alert; use abuse-detection.
- Use the provider's safety tooling — hosted providers offer moderation/safety endpoints and guidance; self-hosted deployments should add equivalent layers.
- Report honestly — the deliverable is a robustness characterization and a hardening plan, framed around making bypasses low-likelihood and low-impact.

The responsible framing throughout: the aim is to strengthen safety, measured with benign proxies, not to produce a library of working bypasses.


9. Attacking AI Agents — Memory & Autonomy (Concept & Defense)

AI agents go beyond text generation — they maintain memory, call tools, and act autonomously. That autonomy turns LLM-layer weaknesses into real-world impact, so agents are a central AI-red-team target. This section covers the concepts and, centrally, the defenses. (Conceptual and defense-oriented per the scope note.)

What makes agents higher-risk: an agent typically has (1) memory (persistent context across turns/sessions), (2) tools (actions: browse, code, DB, email, APIs, files), and (3) autonomy (it decides and chains steps). The security consequence: a manipulation of the agent's reasoning (via injection, §7) can become an action with real effect — the AI equivalent of escalating from a content flaw to code execution. The blast radius equals the agent's permissions.

Attack-surface concepts (what a red teamer evaluates):
- Memory manipulation / poisoning (concept). If an agent persists information across interactions, untrusted content it ingests could influence future behavior — a planted instruction "remembered" and acted on later. Red teamers assess whether and how long untrusted content persists and influences the agent, using benign canaries. Defense: treat memory contents as untrusted data, scope/isolate memory per user/session, validate what gets written to memory, expire it, and don't let memory grant privilege.
- Tool/action abuse via the agent (concept). The real impact of agent attacks is unauthorized tool use — causing the agent to call a tool in a way the user didn't intend (read data it shouldn't, call an API, move a file). Red teamers assess, with benign canaries (e.g., a planted non-sensitive file the agent is induced to move), whether the agent's tool use can be influenced by untrusted content. Defense: least-privilege tools, per-tool authorization, human approval for consequential actions, and output validation on tool results (§13).
- Goal / workflow manipulation (concept). Influencing the agent's objective or plan so it pursues an attacker-chosen outcome. Defense: constrain the agent's goals and allowed action space; validate plans against policy; checkpoint/approve high-impact steps.
- Excessive agency (OWASP LLM06). The root risk: agents given more tools, permissions, autonomy, or data access than they need. Defense (the #1 mitigation): ruthlessly minimize agency — fewest tools, narrowest permissions, least data, and human-in-the-loop for anything consequential.

Defensive architecture for agents (the constructive core):

- LEAST PRIVILEGE: minimum tools, scoped permissions, narrow data access (blast-radius control)
- HUMAN-IN-THE-LOOP / APPROVAL GATES for consequential or irreversible actions
- TOOL-LEVEL AUTHORIZATION: enforce who/what can call each tool, with its own access control —
  don't let the model's say-so be the only authorization
- INPUT/OUTPUT GUARDRAILS around the agent and around each tool boundary
- OUTPUT & TOOL-RESULT VALIDATION: treat tool outputs and web/doc content as untrusted (indirect injection)
- MEMORY HYGIENE: isolate, scope, validate, and expire memory; never store secrets/privilege in it
- SANDBOXING: run code-execution and risky tools in isolated, least-privilege sandboxes
- MONITORING: log every tool call & decision; alert on sensitive/anomalous actions; rate-limit (§18)
- FAIL-CLOSED: default to refusing/escalating on ambiguity rather than taking action

The key principle: you largely can't guarantee the model won't be manipulated, so you architect so that manipulation can't cause serious harm — least privilege, approval gates, sandboxing, and monitoring. A red team's agent findings should drive exactly these controls. The measure of a secure agent isn't "it can't be injected" (it probably can) but "even if injected, it can't do much damage."


10. Multi-Agent Systems & A2A Trust Boundaries

Multi-agent systems — several AI agents collaborating, sometimes via agent-to-agent (A2A) protocols — add trust relationships between agents, creating new attack surface around those relationships. This section covers the concepts and defenses.

Why multi-agent adds risk: when agents delegate tasks, pass messages, and trust each other's outputs, a compromise or manipulation of one agent can propagate. The system's security now depends on the trust model between agents — which is often implicit and over-trusting ("agent B does whatever agent A's message says").

Trust-boundary concepts a red teamer evaluates:
- Inter-agent message trust. Does a receiving agent treat messages from another agent as trusted instructions? If so, an attacker who can influence one agent (via injection, §7) can influence the others downstream. Defense: agents should treat inbound messages (even from peers) as untrusted data, validate them against policy, and not grant privilege based on message content alone.
- Agent impersonation / spoofing (concept). Can a message be forged to appear from a trusted agent? If identity isn't authenticated, yes. Defense: authenticate agent identities (cryptographic identity/signing of A2A messages), verify provenance, and use mutual authentication in A2A protocols.
- Workflow / orchestration corruption (concept). Manipulating the multi-agent workflow so the collective pursues an attacker-chosen outcome or skips a safety/validation step. Defense: a trusted orchestrator that validates each step, enforces the intended workflow, and doesn't let any single agent redirect the whole process; checkpoints and policy checks between stages.
- Privilege aggregation. Individually low-privilege agents whose combined capabilities (via delegation) exceed any one agent's intended scope. Defense: reason about the aggregate agency of the system; apply least privilege to the composition, not just each agent; constrain delegation.
- Error / hallucination propagation. One agent's mistake or manipulated output cascading through the system as if authoritative. Defense: validation between agents, confidence/consistency checks, and human oversight at critical junctions.

Defensive architecture for multi-agent systems:

- AUTHENTICATE & SIGN inter-agent messages (A2A identity + integrity) — no anonymous trust
- TREAT PEER MESSAGES AS UNTRUSTED DATA — validate against policy; don't execute peer "instructions" blindly
- TRUSTED ORCHESTRATOR that enforces the intended workflow and validates each transition
- LEAST PRIVILEGE per agent AND for the composition (bound aggregate agency)
- CONSTRAIN DELEGATION — explicit, limited, logged; no unbounded task hand-off
- ISOLATION between agents (separate contexts/permissions; a compromise shouldn't be lateral by default)
- END-TO-END MONITORING of the whole workflow, not just individual agents (§18)
- HUMAN OVERSIGHT at high-impact decision points

The principle: multi-agent security is about not extending trust implicitly across agent boundaries. Authenticate, validate, isolate, least-privilege the composition, and keep a trusted orchestrator and monitoring over the whole workflow. As with single agents, assume individual agents can be manipulated and design the system so that doesn't cascade into serious harm.


11. RAG Pipeline Security & Data-Poisoning Resistance

Retrieval-Augmented Generation (RAG) grounds a model's answers in an external knowledge base: a retriever fetches relevant documents (via a vector database) and inserts them into the model's context. Security-wise, RAG means the model now trusts retrieved content — which creates poisoning and manipulation risk (OWASP LLM04/LLM08).

The RAG data flow (and where it's attacked):

knowledge source(s) → ingestion/embedding → VECTOR DATABASE → retriever → model context → answer
       ▲ poisoning at the source           ▲ manipulate what's stored/retrieved   ▲ indirect injection via retrieved docs

Attack-surface concepts a red teamer evaluates:
- Knowledge-source poisoning (concept). If an attacker can influence what goes into the knowledge base (an open wiki, user-submitted content, crawled web pages, shared documents), they can plant content that corrupts answers or carries indirect prompt-injection instructions the model later ingests (§7). This is the central RAG risk. Red teamers assess who can write to the knowledge sources and whether planted (benign canary) content influences outputs.
- Retrieval manipulation (concept). Crafting content so it's preferentially retrieved for target queries (embedding/keyword manipulation to rank the poisoned doc highly), ensuring the malicious content reaches the context. Defense: validate/curate sources; monitor retrieval quality.
- Cross-tenant / access-control leakage. In multi-tenant RAG, does retrieval respect per-user/per-tenant access control, or can one user's query retrieve another's documents? A common, high-impact confidentiality flaw. Defense: enforce document-level authorization at retrieval time, scoped to the requesting user/tenant.
- Sensitive-data exposure via retrieval. The knowledge base may contain secrets/PII that the model then surfaces to unauthorized users. Defense: don't put data in the KB that the user shouldn't see; apply access control and data minimization.

Defenses (the constructive core):

- SOURCE PROVENANCE & VALIDATION: control and vet what enters the knowledge base; trusted,
  curated sources; validate/sanitize ingested content; tag provenance.
- TREAT RETRIEVED CONTENT AS UNTRUSTED DATA (not instructions): clearly delimit it in the context;
  apply guardrails to it; neutralize embedded instructions where feasible (indirect-injection defense, §7).
- ACCESS CONTROL AT RETRIEVAL: enforce per-user/per-tenant document authorization; prevent cross-tenant leaks.
- DATA MINIMIZATION: keep secrets/PII out of the knowledge base unless access-controlled and necessary.
- INTEGRITY MONITORING: detect unexpected changes to the knowledge base; audit ingestion; version data.
- RETRIEVAL QUALITY MONITORING: watch for anomalous retrieval patterns / poisoning indicators.
- VECTOR-DB HARDENING: authenticate and access-control the vector database like any datastore (§12).

The principle: RAG shifts trust onto the knowledge base, so the knowledge base must be governed like a security-sensitive datastore — controlled ingestion, provenance, access control at retrieval, and treating retrieved text as untrusted data rather than as instructions. A red team's RAG findings should drive exactly these controls (who can write to the KB, is retrieval access-controlled, is retrieved content treated as data).


12. Embedding & Vector-Database Security

Embeddings are numerical vector representations of data (text, images) that capture semantic meaning; vector databases store and search them (the backbone of RAG and semantic search). They carry their own security and privacy considerations (OWASP LLM08).

Security & privacy concepts a red teamer evaluates:
- Embedding inversion / information leakage (concept). Embeddings are not a secure one-way hash — research shows that, under some conditions, meaningful information about the original text can be reconstructed or inferred from its embedding. So an exposed embedding store can leak information about the underlying (possibly sensitive) data. Red teamers assess whether embeddings of sensitive data are exposed and what could be inferred. Defense: treat embeddings of sensitive data as sensitive data — access-control and protect them accordingly; don't expose raw embeddings to untrusted parties; consider privacy-preserving techniques.
- Vector-database exposure & access control. Vector DBs are datastores — if deployed without authentication, with default credentials, exposed management UIs/ports, or weak access control, they leak the embedded knowledge base (and enable poisoning). This is often a classic infrastructure finding (exposed/unauth datastore). Defense: authenticate, network-isolate, least-privilege access, no defaults, encrypt at rest/in transit — standard datastore hardening.
- Cross-tenant leakage in shared vector stores. Multiple tenants sharing a vector DB without strict namespace/metadata-based isolation can retrieve each other's data. Defense: enforce tenant isolation at the query layer (metadata filtering tied to the authenticated principal), ideally separate indices/collections per tenant.
- Membership inference (concept). Determining whether a particular item is present in the embedded dataset — a privacy concern for sensitive corpora. Defense: access control, minimization, and monitoring of query patterns.
- Poisoning (ties to §11). Injecting crafted embeddings/documents to manipulate retrieval. Defense: controlled ingestion and integrity monitoring.

Defensive summary:

- Treat embeddings of sensitive data AS sensitive data (inversion risk) — protect & access-control them.
- Harden the vector DB like any datastore: authN, least privilege, network isolation, no defaults,
  encryption, patching, no exposed management endpoints.
- Enforce TENANT ISOLATION at retrieval (authenticated, metadata/namespace-scoped queries).
- Minimize sensitive data in the embedding store; monitor access and query patterns.

The principle: embeddings and vector databases deserve the same data-protection and infrastructure-security rigor as any sensitive datastore — with the extra awareness that embeddings can leak information about their source data (so "it's just vectors" is not a safe assumption). Most real-world vector-DB findings are classic exposure/access-control issues, which are very fixable.


13. Model Context Protocol (MCP) & Tool-Surface Security

Modern AI applications connect models to tools, data, and actions through orchestration layers and integration frameworks — including the Model Context Protocol (MCP) and function-calling/plugin ecosystems. This "tool surface" is where an AI's words become actions, so its security is critical (OWASP LLM05/LLM06).

What the tool surface is: the set of tools/functions/connectors a model can invoke — web browsing, code execution, database queries, file access, API calls, sending messages, interacting with other services — plus the orchestration layer (e.g., an MCP server) that mediates these. MCP and similar frameworks standardize how models discover and call tools and access external context/data.

Security concepts a red teamer evaluates:
- Excessive tool permissions (OWASP LLM06). Tools granted broad access (a DB tool with full read/write, a file tool with wide filesystem access, cloud credentials) mean a manipulated model (§7/§9) can do a lot. Defense (primary): least-privilege every tool — narrowest scope, minimal credentials, read-only where possible.
- Unsafe output handling / injection downstream (LLM05). When model output is passed to a tool (a shell command, a SQL query, an HTTP request, HTML rendering) without validation, classic injection results (command injection, SQLi, SSRF, XSS) — the model becomes an injection source. Defense: validate/parameterize/sanitize everything flowing from model → tool and tool → model; never build a shell command or SQL query by string-concatenating model output.
- Tool/connector authentication & authorization. Does calling a tool enforce its own access control tied to the actual user, or does the model's request alone authorize it? If the latter, a manipulated model acts with the tool's full privilege regardless of the user. Defense: per-tool authorization bound to the authenticated end-user; don't treat the model as an authorization authority.
- Untrusted tool outputs (indirect injection, §7). Content returned by a tool (a web page, an API response, a file) enters the model's context and can carry injection. Defense: treat all tool outputs as untrusted data; guardrail and delimit them.
- MCP-server / orchestration exposure & integrity. The MCP/orchestration layer is itself software/infrastructure: exposed MCP servers, weak auth on the orchestration API, malicious or compromised MCP servers/connectors (supply chain, §16), and over-broad connector scopes. Defense: authenticate and access-control the orchestration layer; vet connectors; least-privilege their scopes; monitor.
- Confused-deputy / privilege escalation. The orchestration layer acting on the model's behalf with its privileges, induced to perform actions the user couldn't. Defense: bind actions to the end-user's authorization; least privilege; approval gates.

Defensive architecture for tool surfaces:

- LEAST PRIVILEGE per tool/connector (scope, credentials, read-only bias) — the master control
- PER-TOOL AUTHORIZATION tied to the authenticated end-user (model ≠ authorization authority)
- SANITIZE/VALIDATE/PARAMETERIZE model→tool inputs (stop injection→RCE/SQLi/SSRF/XSS)
- TREAT TOOL OUTPUTS AS UNTRUSTED (indirect-injection defense) — guardrail & delimit
- HUMAN APPROVAL for consequential/irreversible tool actions
- SANDBOX risky tools (code execution, file ops) in isolated least-privilege environments
- HARDEN the MCP/orchestration layer: authN/authZ, vetted connectors, no exposed admin surfaces
- VET third-party connectors/plugins (supply-chain, §16)
- MONITOR & LOG every tool invocation; alert on sensitive actions; rate-limit (§18)

The principle: the tool surface is the bridge from "AI says" to "system does," so it demands the strongest controls — least privilege, user-bound authorization, injection-safe input/output handling, sandboxing, and monitoring. Most catastrophic AI-agent impact traces back to an over-permissioned, under-validated tool surface — which is exactly what these defenses fix.


14. Model Extraction & Adversarial Machine Learning (Concept & Defense)

Beyond LLM-application attacks, classical adversarial machine learning targets ML models themselves — stealing them, fooling them, or corrupting them. These are well-studied in research and catalogued in MITRE ATLAS and NIST's adversarial-ML taxonomy. (Conceptual and defense-oriented per the scope note.)

Model extraction / model theft (concept). By querying a model many times and observing outputs, an adversary may train a surrogate that approximates the target — stealing the model's functionality/IP, or building a local copy to craft further attacks against. Defense: rate limiting and quotas, query monitoring/anomaly detection (extraction shows distinctive query patterns), output perturbation/rounding, authentication and access control on the inference API, watermarking, and legal/ToS controls. (This overlaps with unbounded consumption / denial-of-wallet, OWASP LLM10 — abusive querying also inflates cost.)

Adversarial examples / evasion (concept). Inputs perturbed (often imperceptibly) to cause a model to misclassify or misbehave — classic in image/ML classifiers, with analogs in other modalities. Defense: adversarial training (train on adversarial examples), input preprocessing/validation, ensemble/robust architectures, detection of anomalous inputs, and not relying solely on an ML model for security-critical decisions.

Data / model poisoning (concept). Corrupting training or fine-tuning data (or the model/adapter artifacts) to implant biases, backdoors, or degraded/triggered behavior (ties to supply chain, §16, and RAG, §11). Defense: data provenance and validation, dataset integrity controls, anomaly detection in training data, trusted sources, model/adapter integrity verification, and evaluation/red-teaming of models before deployment.

Model inversion & membership inference (privacy, §15). Reconstructing training data characteristics, or determining whether a specific record was in the training set — privacy attacks. Defense: differential privacy, data minimization, access control, output restrictions, and monitoring.

How a red teamer engages these (responsibly): characterize the exposure — is the inference API rate-limited and authenticated (extraction/abuse resistance)? are training-data provenance and model-artifact integrity controlled (poisoning resistance)? is the model used for security-critical decisions without robustness considerations (evasion risk)? are privacy protections in place for sensitive training data? The deliverable is a risk assessment and the above defenses — not operational attack tooling.

Defensive summary:

- Inference API: authN, rate limits/quotas, query monitoring, anomaly detection → extraction/abuse resistance
- Training/fine-tuning: data provenance, integrity, validation, trusted sources → poisoning resistance
- Model artifacts: integrity verification (hashes/signatures) for weights & adapters (§16)
- Robustness: adversarial training, input validation, don't rely solely on ML for security decisions
- Privacy: differential privacy, data minimization, access control (§15)
- Pre-deployment: evaluate & red-team models before production; ongoing monitoring

The principle: treat the model as both an asset to protect (IP, integrity, training-data privacy) and a component that can be fooled or stolen — and apply the research-backed defenses (rate limiting + monitoring for extraction, provenance + integrity for poisoning, robustness + validation for evasion, privacy techniques for inversion/inference).


15. Sensitive-Data Exposure & Privacy Attacks (Concept & Defense)

AI systems process and sometimes memorize sensitive data, and can leak it through several channels (OWASP LLM02/LLM07). Protecting confidentiality is a core AI-security objective.

Exposure channels a red teamer evaluates:
- System-prompt / instruction leakage (LLM07). The system prompt may contain proprietary logic, business rules, or (badly) secrets; applications can be induced to reveal it. Defense: never put secrets/credentials in the system prompt; assume it may leak; don't rely on it for security; detect/limit extraction attempts.
- Training-data extraction / memorization. Models can memorize and regurgitate fragments of training data, potentially including PII or secrets present in the training set. Defense: data minimization and scrubbing/PII-removal before training, deduplication, differential privacy, and output filtering; evaluate models for memorization before deployment.
- RAG / knowledge-base leakage (§11). The model surfacing documents or data the user isn't authorized to see. Defense: access control at retrieval, data minimization in the KB, tenant isolation.
- Embedding leakage (§12). Information inferable from exposed embeddings. Defense: protect embeddings of sensitive data.
- Cross-session / cross-user leakage. Memory or context from one user bleeding into another's session. Defense: strict session/tenant isolation; don't share context across users.
- Secrets in the AI's reach. API keys/credentials in the system prompt, in tools, or in accessible data that the model could expose or misuse. Defense: secrets management (never in prompts/code/KB), least privilege, and scoping credentials away from the model's direct output.
- Output disclosure. The model including sensitive data in responses (or in logs/error messages). Defense: output filtering/DLP on model responses, careful logging (don't log secrets/PII), and PII detection.

Defensive summary:

- NO SECRETS in system prompts, code, or knowledge bases — use a secrets manager; scope credentials.
- DATA MINIMIZATION everywhere: train/store/retrieve only what's necessary; scrub PII.
- ACCESS CONTROL & TENANT/SESSION ISOLATION on data, retrieval, memory, and embeddings.
- OUTPUT FILTERING / DLP on model responses; PII detection; careful logging.
- PRIVACY TECHNIQUES for sensitive training data (differential privacy, dedup, memorization testing).
- MONITOR for data-exfiltration patterns and sensitive-data in outputs (§18).

The principle: apply classic data-protection (minimization, access control, isolation, secrets management, DLP) to the new channels AI introduces — the context window, memory, RAG, embeddings, and outputs — and never assume any part of the AI's context is confidential by default.


16. AI/ML Supply-Chain Security

The AI supply chain extends the traditional software supply chain to include datasets, pre-trained model weights, fine-tuning adapters, ML frameworks, and model hubs — each a vector for introducing compromised artifacts (OWASP LLM03). Much of this is standard supply-chain security applied to new artifact types, which makes it concrete and very actionable.

The AI supply-chain attack surface:
- Datasets. Poisoned or tampered training/fine-tuning data (backdoors, bias, degraded behavior). Defense: trusted sources, provenance tracking, integrity verification (hashes), validation and anomaly detection on data, and documentation (datasheets).
- Pre-trained models / weights. Downloading model weights from public hubs introduces risk: tampered weights, backdoored models, or models masquerading as something they're not. Additionally, some model serialization formats can execute code on load (unsafe deserialization) — loading an untrusted model file can be RCE. Defense: obtain models from trusted/verified sources, verify integrity (hashes/signatures), prefer safe serialization formats (e.g., safetensors over pickle-based formats), scan model files, and load untrusted models in sandboxes.
- Fine-tuning adapters (LoRA, etc.). Small add-on weights that modify behavior — a compromised adapter can implant backdoors or unsafe behavior. Defense: verify adapter provenance/integrity; evaluate behavior after applying adapters.
- ML frameworks & libraries. The usual software dependency risk (vulnerable/malicious packages), plus ML-specific tooling. Defense: SBOMs, dependency scanning, pinning, trusted registries, and vulnerability management.
- Model hubs / registries. Typosquatted or malicious model/dataset uploads, compromised accounts, or compromised internal registries. Defense: verify publishers, use signed/verified artifacts, and secure your internal model registry (authN, integrity, access control).
- CI/CD & MLOps pipelines. The build/train/deploy pipeline itself — compromise here injects into everything downstream. Defense: secure the pipeline (least privilege, signed artifacts, provenance/attestation, isolated build/training), as with any software CI/CD.

How a red teamer assesses it: inventory the AI supply chain (which datasets, models, adapters, frameworks, hubs, and pipelines are used) → assess provenance and integrity controls at each stage → check how models/data are obtained, verified, serialized, and loaded → review the MLOps pipeline and registry security. Findings are concrete and fixable (e.g., "model weights loaded from an unverified source using an unsafe serialization format that executes code on load — in a privileged container").

Defensive summary (supply-chain for AI):

- PROVENANCE & INTEGRITY for datasets, weights, and adapters (trusted sources + hashes/signatures)
- SAFE MODEL LOADING: prefer safe serialization (safetensors), scan model files, sandbox untrusted loads
- DEPENDENCY SECURITY: SBOMs, scanning, pinning, trusted registries (ML frameworks & libs)
- SECURE MODEL REGISTRY: authN, access control, signed/verified artifacts
- SECURE MLOps CI/CD: least privilege, signed artifacts, provenance/attestation, isolated training
- EVALUATE artifacts before deployment (behavioral red-teaming of models/adapters)

The principle: the AI supply chain is the classic supply-chain problem with new artifact types (data, weights, adapters) and one sharp new edge — loading a model can execute code — so provenance, integrity verification, safe loading, and pipeline security are the core defenses, and they're highly concrete deliverables.


17. AI Infrastructure & Deployment Security

Behind every AI system is infrastructure — model servers, inference engines, containers, GPUs, cloud services, storage, and MLOps platforms. This layer is largely classic infrastructure security, which means it's concrete, high-impact, and very fixable — and often where the most serious, most certain findings live.

The AI infrastructure attack surface:
- Model / inference servers. The services that serve models over APIs (various inference frameworks and model servers). Risks: exposed/unauthenticated inference or management endpoints, default configurations, version vulnerabilities, and exposed metrics/admin interfaces. Defense: authenticate and access-control all endpoints, don't expose management/metrics publicly, patch, and network-segment. (Treat like any exposed web service — classic recon + hardening.)
- Notebook & experiment platforms. Exposed Jupyter notebooks, experiment-tracking UIs, and MLOps dashboards — frequently found unauthenticated on the internet, granting code execution or data access. Defense: authentication, network isolation, no public exposure.
- Containers & orchestration. ML workloads run in containers/Kubernetes with GPUs. Risks: over-privileged containers, container escape, exposed orchestration APIs, secrets in images/env, and the usual K8s misconfigurations. Defense: standard container/K8s hardening — least-privilege/non-root containers, no privileged mode unless required, network policies, secrets management, image scanning, and RBAC.
- Cloud & storage. Models, datasets, and checkpoints in object storage; GPU instances; managed AI services. Risks: public/misconfigured storage buckets holding model weights or training data, over-permissive IAM, exposed credentials, and insecure managed-service configs. Defense: classic cloud hardening — private storage, least-privilege IAM, encryption, secrets management, logging (ties to the ARTE/AZRTE cloud-security disciplines).
- GPU / compute. Resource exhaustion and multi-tenant isolation concerns on shared GPU infrastructure. Defense: quotas, isolation, monitoring.
- Secrets & credentials. Inference services and pipelines hold API keys and cloud credentials (often with broad access). Defense: secrets managers, least privilege, rotation, never in images/prompts/code.
- Networking. Internal model/vector-DB/tool services exposed more broadly than necessary. Defense: segmentation, zero-trust internal networking, and minimal exposure.

How a red teamer assesses it: this is where AI red teaming overlaps most with traditional pentesting — enumerate and fingerprint the AI infrastructure (model servers, notebooks, registries, vector DBs, orchestration, cloud), find exposed/unauthenticated services, misconfigurations, over-privileged containers/IAM, exposed secrets, and known CVEs, then assess impact (data/model theft, code execution, pivot). The toolkit is your existing infra/cloud/container pentest skill set applied to AI assets.

Defensive summary (AI infrastructure):

- AuthN/authZ on ALL model/inference/management/notebook/registry endpoints — nothing unauth-exposed
- Network segmentation & minimal exposure of internal AI services (model servers, vector DBs, tools)
- Container/K8s hardening: least privilege, non-root, no needless privileged mode, network policies, RBAC
- Cloud hardening: private storage for models/data, least-privilege IAM, encryption, logging
- Secrets management: no secrets in images/env/prompts/code; least privilege; rotation
- Patch & vulnerability management for ML frameworks, servers, and dependencies
- Monitoring & logging across the AI infrastructure (§18)

The principle: the AI layer gets the headlines, but the infrastructure around it is often the easiest and most impactful thing to compromise — and to fix. Apply your full traditional infra/cloud/container security discipline to AI assets; exposed model servers, public buckets of model weights, and unauthenticated notebooks are common, serious, and entirely preventable.


18. Detection, Monitoring & Guardrail Architecture

A secure AI system isn't just hardened — it's observable. Detection and monitoring let defenders catch attacks (injection attempts, abuse, data exfiltration, anomalous agent behavior) that prevention can't fully stop. This is a key deliverable focus for the AI red team.

What to log and monitor:

- PROMPTS & INPUTS: user inputs, retrieved RAG content, tool outputs entering the context
  (watch for injection patterns, anomalies) — mindful of privacy in what you log
- MODEL OUTPUTS: responses (watch for sensitive-data disclosure, policy violations, unsafe content)
- TOOL / ACTION CALLS: every tool invocation, with parameters and the authorizing user
  (the highest-signal telemetry for agents — alert on sensitive/anomalous actions)
- AGENT DECISIONS / WORKFLOW: the reasoning/plan and multi-agent message flow (traceability)
- API USAGE: rate, cost, query patterns (extraction/abuse/denial-of-wallet detection)
- INFRASTRUCTURE: access to model servers, vector DBs, registries, storage (classic security logging)
- GUARDRAIL EVENTS: input/output filter triggers, refusals, blocked actions

Detection opportunities (what to alert on):
- Prompt-injection / jailbreak attempts — known patterns, guardrail-classifier triggers, anomalous instruction-like content in inputs/retrieved docs, sudden behavior changes.
- Sensitive-data in outputs — DLP/PII detection on responses; system-prompt-leakage indicators.
- Anomalous tool use — an agent calling sensitive tools unexpectedly, unusual parameters, or a spike in actions (the agent-attack tell).
- Abuse / extraction / denial-of-wallet — abnormal query volume/patterns, cost spikes.
- RAG/KB integrity — unexpected changes to the knowledge base, anomalous retrievals.
- Infrastructure — access to AI services from unexpected sources, auth failures, config changes.

Guardrail architecture (defense-in-depth layers):

INPUT GUARDRAILS   → classify/filter inputs & retrieved content (injection/abuse detection)
       ▼
SYSTEM PROMPT & CONTEXT CONTROL → clear instruction/data separation; minimal, secret-free context
       ▼
THE MODEL          → provider safety + any fine-tuned safety behavior
       ▼
OUTPUT GUARDRAILS  → filter/validate outputs (DLP, policy, unsafe-content, sanitization for downstream)
       ▼
TOOL AUTHORIZATION & SANDBOXING → per-tool authZ, least privilege, approval gates, isolation
       ▼
MONITORING & RESPONSE → log everything, detect, alert, rate-limit, and have an incident process

The principle: because prevention is imperfect (injection and jailbreaks can't be fully eliminated), detection and layered guardrails are essential — assume some attacks get through and make sure you'll see them and can respond. The AI red team should assess not just "can I get in?" but "would the defenders see it?" — and recommend the logging, detection, and guardrail layers above. Observability of tool calls and outputs is especially high-value.


19. Defensive Architecture — Securing AI Systems

Consolidating the defenses from each section into a reference architecture — the "what good looks like" that your red-team findings should drive toward. This is the constructive heart of responsible AI red teaming.

The layered defense model for an AI application:

1. GOVERNANCE & THREAT MODEL
   - know your AI assets, trust boundaries, and risks (§6); define policies; NIST AI RMF alignment.
2. MINIMIZE ATTACK SURFACE & AGENCY (the master principle)
   - least privilege everywhere: fewest tools, narrowest permissions, least data, minimal exposure.
   - the single most effective control — it caps the blast radius of every AI-layer attack.
3. INPUT & CONTEXT CONTROL
   - treat ALL non-system content (user input, RAG docs, tool outputs, memory) as untrusted data.
   - input guardrails/classifiers; clear instruction/data separation; no secrets in context.
4. MODEL & GUARDRAIL LAYER
   - provider/base safety + application guardrails (input/output classifiers); continuous eval.
5. OUTPUT HANDLING
   - never trust output: sanitize/validate/encode before it hits browsers, shells, DBs, other systems.
   - output DLP/PII filtering; policy checks.
6. TOOL / ACTION SECURITY
   - per-tool authorization bound to the end-user; injection-safe tool I/O; human approval for
     consequential actions; sandboxing for risky tools (§13).
7. DATA & RETRIEVAL SECURITY
   - provenance & validation of training/RAG data; access control at retrieval; tenant/session isolation;
     protect embeddings; data minimization; no secrets in data (§11-12, §15).
8. SUPPLY CHAIN
   - provenance/integrity for data, weights, adapters; safe model loading; secure MLOps pipeline (§16).
9. INFRASTRUCTURE
   - authN/authZ on all endpoints; network segmentation; container/cloud hardening; secrets management;
     patching (§17).
10. MONITORING, DETECTION & RESPONSE
   - log prompts/outputs/tool-calls/infra; detect injection/abuse/exfiltration/anomalies; rate-limit;
     incident response; continuous red-teaming (§18).
11. HUMAN OVERSIGHT & FAIL-CLOSED
   - humans in the loop for high-impact actions; default to safe/refuse/escalate on ambiguity.

The guiding principles:
- Assume the model can be manipulated (injection/jailbreak aren't fully solvable) → architect so manipulation can't cause serious harm (least privilege, approval gates, sandboxing, output handling).
- Agency = blast radius → minimizing what the AI can do is the highest-leverage control.
- Treat everything in the context as untrusted data except the system prompt → and don't put secrets even there.
- Defense in depth → layered guardrails, because any single layer can fail.
- Observe everything → detection compensates for imperfect prevention.
- Classic security still applies → the infrastructure, data, and supply chain around the model are where much of the real, fixable risk lives.

A red-team report's value is measured by how well it moves the system toward this architecture.


20. MITRE ATLAS — Technique Reference & Examples

MITRE ATLAS (Adversarial Threat Landscape for Artificial-Intelligence Systems) is the ATT&CK-style matrix for AI. Mapping findings to ATLAS gives your assessment a standard vocabulary. Below is a tactic-by-tactic walkthrough with concrete example scenarios (framed defensively — what the behavior looks like and how to detect/defend it).

Reconnaissance — the adversary learns about the AI system.
- Example: an attacker reads your public docs, job posts, and GitHub to learn you run a self-hosted LLM behind a RAG pipeline over a specific vector DB. Detect/defend: minimize public disclosure of the AI stack; monitor for scraping/enumeration of AI endpoints.

Resource Development — the adversary builds capabilities.
- Example: an attacker obtains a copy of the same open-weights base model you use, to probe it offline before attacking your deployment. Defend: assume the base model is known; don't rely on model secrecy; put security in the application/infra layers.

Initial Access — getting into the AI system.
- Example A (via the app): untrusted content reaches the model's context (user input, or an indirect vector like a web page the agent browses). Example B (via infra): an exposed, unauthenticated model server or notebook. Detect/defend: authenticate endpoints; treat context content as untrusted (§7); network-segment model servers (§17).

ML Model Access — gaining query or deeper access to the model.
- Example: an attacker with API access queries the model at high volume to characterize it (precursor to extraction). Detect/defend: authenticate the inference API, rate-limit, monitor query patterns (§14).

Execution — causing unintended behavior/actions.
- Example: a manipulated agent invokes a tool the user didn't intend (e.g., a database query). Detect/defend: least-privilege tools, per-tool authZ, output validation, approval gates (§13); alert on anomalous tool calls (§18).

Persistence — maintaining influence over time.
- Example: poisoned content in a RAG knowledge base that keeps influencing answers; or a planted instruction written into agent memory. Detect/defend: KB integrity monitoring & provenance (§11); memory hygiene/expiry (§9).

Defense Evasion — avoiding guardrails/detection.
- Example: obfuscated/encoded content that slips past an input filter (the AI analog of WAF evasion). Detect/defend: layered guardrails (don't rely on one classifier), canonicalize inputs, monitor for evasion patterns.

Discovery — exploring what the AI/system can do.
- Example: probing to enumerate the agent's available tools or the system prompt's constraints. Detect/defend: don't expose tool schemas unnecessarily; detect enumeration behavior; assume the system prompt may leak (§15).

Collection — gathering target data.
- Example: inducing the system to surface documents/PII from the RAG store or memory. Detect/defend: retrieval access control + tenant isolation (§11-12); output DLP (§15).

Exfiltration — getting data out.
- Example: an injected instruction causing an agent with a web/email tool to send collected data outward. Detect/defend: egress controls on tools, least privilege, approval for outbound actions, monitoring of tool outputs (§13, §18).

Impact — the adversary's end goal.
- Examples: data breach (confidentiality), corrupted outputs driving bad business decisions (integrity), model theft (IP), service disruption or denial-of-wallet via resource abuse (availability/cost). Defend: the full defensive architecture (§19) prioritized by which impact matters most to the business.

How to use ATLAS in practice: for each finding, tag the tactic (attacker goal) and technique (method), cite it in the report (e.g., "indirect prompt injection via a poisoned RAG document → unauthorized tool use → data exfiltration"), and pair each with the detection/defense. ATLAS also publishes real-world case studies — read them; they're the AI equivalent of ATT&CK's threat-group reports and show how these techniques chain in actual incidents.


21. AI Infrastructure Hardening Playbook (with examples)

The infrastructure around the model is where the most certain, highest-impact, and most fixable findings usually live — and it's classic security territory. This playbook gives concrete example findings and their fixes.

Model / inference servers.

EXAMPLE FINDING:  A self-hosted inference server is reachable on the internet with no authentication;
                  its /metrics and a management endpoint are exposed. Anyone can query the model
                  (cost/abuse/extraction) and read internal telemetry.
IMPACT:           Model abuse & extraction, information disclosure, potential DoS/denial-of-wallet.
FIX:              Require authN on all endpoints; bind to internal interfaces / put behind an
                  authenticated gateway; never expose /metrics or admin surfaces publicly; rate-limit;
                  patch the server; network-segment so only the app tier can reach it.

Notebook & MLOps platforms.

EXAMPLE FINDING:  A Jupyter notebook server is exposed to the internet without a token/password,
                  granting arbitrary code execution on a GPU host with access to datasets and cloud creds.
IMPACT:           Full code execution, data theft, cloud pivot — critical.
FIX:              Never expose notebooks publicly; require authN; run in isolated, least-privilege
                  environments; no long-lived cloud credentials on the host; network isolation.
# Same pattern applies to exposed experiment-tracking UIs, pipeline dashboards, and model registries.

Object storage (model weights & datasets).

EXAMPLE FINDING:  A cloud storage bucket holding model weights and a training dataset is public-read.
IMPACT:           Model/IP theft; training-data (possibly PII) breach; enables offline attack prep.
FIX:              Private buckets; least-privilege IAM; block public access; encryption at rest;
                  access logging; verify no weights/datasets are world-readable (classic cloud hygiene).

Containers & Kubernetes (ML workloads).

EXAMPLE FINDING:  ML inference pods run as root in privileged containers, mount broad host paths, and
                  hold a cloud credential with wide permissions; no network policies restrict pod egress.
IMPACT:           Container escape → node/cluster compromise; credential abuse → cloud takeover; free
                  lateral movement.
FIX:              Non-root, least-privilege containers; drop privileged mode/capabilities; minimal
                  mounts; scoped, short-lived credentials (workload identity); network policies limiting
                  egress; image scanning; RBAC on the orchestration API; no exposed kubelet/dashboard.

Secrets & credentials.

EXAMPLE FINDING:  An API key for the LLM provider and a database credential are baked into a container
                  image / sitting in an env var / embedded in the system prompt.
IMPACT:           Credential theft → model/account abuse, data access, cost.
FIX:              Use a secrets manager; inject at runtime; least privilege & rotation; NEVER put
                  secrets in images, code, prompts, or the RAG knowledge base.

Vector database & internal AI services.

EXAMPLE FINDING:  The vector DB backing RAG is reachable on the internal network with default/no auth;
                  any internal host can read or write the embedded knowledge base.
IMPACT:           Knowledge-base disclosure and poisoning (§11-12).
FIX:              AuthN + least-privilege access; network isolation; no defaults; encryption;
                  tenant isolation; monitor access.

GPU / compute & availability.

EXAMPLE FINDING:  No quotas or rate limits on an expensive inference endpoint.
IMPACT:           Denial-of-wallet / resource exhaustion (OWASP LLM10).
FIX:              Quotas, rate limits, cost alerts, autoscaling guards, and abuse detection.

The hardening checklist (apply to every AI deployment):

[ ] AuthN/authZ on ALL model, inference, management, notebook, registry, and vector-DB endpoints
[ ] Nothing AI-related exposed publicly that doesn't need to be; network segmentation of internal services
[ ] Private storage for weights & datasets; least-privilege IAM; encryption; access logging
[ ] Containers: non-root, unprivileged, minimal mounts, scoped creds, network policies, image scanning
[ ] Secrets in a manager — never in images/env/prompts/code/KB; rotation
[ ] Rate limits + quotas + cost alerts on inference endpoints (abuse/extraction/denial-of-wallet)
[ ] Patch & vuln-manage ML frameworks, servers, and dependencies
[ ] Centralized logging & monitoring across the AI infrastructure (§23)

The principle: bring your full traditional infra/cloud/container pentest-and-hardening discipline to AI assets. The glamorous AI-layer attacks get attention, but an exposed model server, a public bucket of weights, or an unauthenticated notebook is more common, more certain, and just as damaging — and entirely preventable.


22. AI/ML Supply-Chain Security Deep-Dive (with examples)

The AI supply chain (datasets, weights, adapters, frameworks, hubs, pipelines) is classic supply-chain security plus one sharp new edge — loading a model can execute code. Concrete examples and fixes:

Unsafe model deserialization (the sharp edge).

EXAMPLE FINDING:  The app downloads model weights in a pickle-based serialization format from an
                  unverified source and loads them in a privileged service. The format can execute
                  arbitrary code on load.
IMPACT:           Remote code execution the moment the model is loaded — critical.
FIX:              Prefer safe serialization (e.g., safetensors) that doesn't execute code; obtain
                  models only from verified sources; verify integrity (hash/signature) before loading;
                  scan model files; load untrusted/new models in a sandbox with no privileges/credentials.

Tampered / backdoored weights or adapters.

EXAMPLE FINDING:  A fine-tuning adapter (LoRA) is pulled from a public hub with no integrity check; it
                  subtly alters behavior under a specific trigger (a backdoor).
IMPACT:           Hidden malicious behavior in production; integrity compromise.
FIX:              Provenance tracking + integrity verification (hashes/signatures) for weights AND
                  adapters; evaluate/red-team behavior after applying adapters; prefer trusted publishers.

Poisoned datasets.

EXAMPLE FINDING:  A training/fine-tuning dataset is assembled partly from an open, writable web source
                  that an attacker contributed poisoned samples to.
IMPACT:           Degraded/biased/triggered model behavior; backdoors (OWASP LLM04).
FIX:              Trusted/curated sources; provenance (datasheets); integrity verification; anomaly
                  detection & validation on training data; deduplication; document data lineage.

Typosquatted / malicious hub artifacts.

EXAMPLE FINDING:  A developer installs a model or package whose name closely mimics a popular one
                  (typosquat) from a public registry; it's malicious.
IMPACT:           Compromised model/dependency entering the pipeline.
FIX:              Verify publishers/namespaces; pin exact versions/digests; use trusted/mirrored
                  registries; dependency scanning; SBOMs.

Vulnerable ML framework / dependency.

EXAMPLE FINDING:  An outdated ML framework with a known RCE CVE is used in the inference service.
IMPACT:           Exploitable server compromise.
FIX:              SBOM + dependency scanning + patch management + version pinning (standard SCA
                  applied to the ML stack).

Compromised MLOps pipeline / registry.

EXAMPLE FINDING:  The model-build CI/CD pipeline has weak access control and unsigned artifacts; a
                  compromise there would inject into every downstream deployment.
IMPACT:           Supply-chain compromise of all models built by the pipeline.
FIX:              Least privilege on the pipeline; signed artifacts + provenance/attestation (e.g.,
                  SLSA-style); isolated build/training; secure the model registry (authN, access
                  control, verified/signed artifacts); audit.

The AI-SBOM concept. Just as software has an SBOM, maintain an inventory of your AI artifacts: which models (source, version, hash, license), datasets (source, provenance), adapters, frameworks/libraries, and hubs/registries are used — so you can verify integrity, respond to vulnerabilities, and prove provenance. Some ecosystems now support model/dataset cards and signing — use them.

Red-team assessment of the AI supply chain:

[ ] Inventory all datasets, models, adapters, frameworks, hubs, and pipelines (AI-SBOM)
[ ] How are models/datasets obtained? From trusted/verified sources? Integrity verified?
[ ] What serialization format are models loaded from? Is loading sandboxed? (the RCE edge)
[ ] Are adapters and datasets provenance-tracked and integrity-checked?
[ ] Dependency security on the ML stack (scanning, pinning, SBOM)
[ ] MLOps pipeline & registry security (least privilege, signed artifacts, isolation)
[ ] Are models evaluated/red-teamed before deployment?

The principle: apply mature software-supply-chain practices — provenance, integrity verification, signing, SBOMs, dependency scanning, secure CI/CD — to the new artifact types (data, weights, adapters), and treat model loading as a code-execution event that demands safe formats and sandboxing. These are concrete, high-certainty, very actionable findings.


23. Detection Engineering for AI Systems (with example logic)

Because prevention is imperfect (injection/jailbreaks can't be fully eliminated), detection is essential. This section gives concrete, example telemetry and detection logic patterns (illustrative pseudo-rules — adapt to your SIEM/platform).

Telemetry to collect (the AI observability stack):

- input events:   user prompt, retrieved RAG docs, tool outputs entering context (+ source/provenance)
- model events:   response text, refusals, guardrail-classifier scores/triggers, token usage
- tool events:    tool name, parameters, authorizing user, result, duration (HIGHEST-signal for agents)
- agent events:   plan/decision trace, multi-agent messages (sender/receiver/intent)
- api events:     per-user request rate, cost, query patterns
- infra events:   access to model servers, vector DBs, registries, storage (classic security logs)
# privacy note: logging prompts/outputs may capture sensitive data — apply access control, retention
#   limits, and redaction to the logs themselves.

Example detection patterns (illustrative logic):

# 1) Possible prompt-injection / instruction content in UNTRUSTED sources (RAG doc / tool output / web)
WHEN  retrieved_doc OR tool_output  CONTAINS instruction-like patterns
      (e.g., imperative "ignore/override/disregard" directed at the assistant, embedded role markers,
       hidden/zero-width or heavily-encoded text)
THEN  flag for review; elevate if the agent subsequently changes behavior or calls a sensitive tool.

# 2) Anomalous / sensitive tool use by an agent  (the core agent-attack tell)
WHEN  tool_call IN (sensitive_tools: db_write, email_send, file_move, code_exec, external_http)
      AND (unusual_for_this_user OR not_in_expected_workflow OR spike_in_tool_calls)
THEN  alert; (ideally require human approval for these tools by design, §13).

# 3) Sensitive data in model OUTPUT  (DLP on responses)
WHEN  model_output MATCHES (PII patterns, secret/key patterns, system-prompt fingerprint)
THEN  block/redact + alert (system-prompt-leakage / data-exfiltration indicator, §15).

# 4) Guardrail-bypass / evasion attempts
WHEN  input guardrail_score HIGH repeatedly from one user/session
      OR heavy obfuscation/encoding in inputs
      OR many paraphrased retries after refusals
THEN  rate-limit + alert (evasion behavior).

# 5) Model extraction / abuse / denial-of-wallet
WHEN  one principal's query_volume OR cost >> baseline
      OR systematic, diverse querying consistent with surrogate-model building
THEN  throttle + alert (OWASP LLM10 / extraction, §14).

# 6) RAG / knowledge-base integrity
WHEN  unexpected writes/changes to the knowledge base
      OR anomalous retrieval (a rarely-retrieved doc suddenly dominating answers)
THEN  alert; verify provenance (poisoning indicator, §11).

# 7) Cross-tenant / access-control anomaly
WHEN  a retrieval or tool action touches data outside the requesting user's authorized scope
THEN  block + alert (isolation failure, §11-12).

# 8) Infrastructure
WHEN  access to a model server / vector DB / registry from an unexpected source,
      auth failures, or config changes
THEN  alert (classic security monitoring applied to AI assets, §17).

Metrics worth tracking: refusal/guardrail-trigger rates (and drift), sensitive-tool-call frequency, blocked-output rate, per-user cost/volume outliers, KB change frequency, and time-to-detect for simulated attacks. A red team should test detection — run a benign canary attack and verify the blue team would see it; "silent success" is itself a finding.

The principle: build AI observability (especially around tool calls and outputs), write detections for the attack behaviors (not just signatures), and continuously validate that defenders can see attacks — because you're compensating for prevention you can't make perfect.


24. AI Security Governance & Program (NIST AI RMF Applied)

Technical testing sits inside an organizational program. The NIST AI Risk Management Framework (AI RMF) gives the governance structure; a red team's findings should feed it. This matters for the "risk management" learning objectives.

NIST AI RMF — the four functions (applied):

GOVERN   — culture, policies, roles, accountability for AI risk.
           Examples: an AI acceptable-use & security policy; an AI inventory with owners; a
           model-approval/red-team gate before production; incident-response plans that cover AI.
MAP      — understand context & identify risks.
           Examples: catalogue AI systems, their purpose, data, trust boundaries, and impact
           (this is threat modeling, §6, at program scale).
MEASURE  — assess, analyze, and track risks.
           Examples: red-team assessments (this guide), automated safety/robustness evals, guardrail
           metrics, monitoring KPIs (§23), reproducibility/likelihood of findings.
MANAGE   — prioritize and act on risks.
           Examples: remediation roadmap toward the defensive architecture (§19), risk acceptance
           decisions, continuous monitoring, and re-assessment cadence.

An AI security program checklist (what mature looks like):

[ ] AI inventory: every AI system/model/app with an owner, purpose, data, and risk rating
[ ] Policies: AI acceptable-use, data-handling, model-approval, and secure-development standards
[ ] Pre-deployment gate: threat model + security review + red-team/eval before production
[ ] Secure-by-design standards: least-privilege agency, untrusted-context handling, output handling,
    secrets management, infra & supply-chain controls (the §19 architecture as policy)
[ ] Continuous evaluation: automated safety/robustness regression tests; guardrail monitoring
[ ] Monitoring & detection: AI observability + detections + alerting (§23)
[ ] Incident response: playbooks that cover AI-specific incidents (injection abuse, data leakage,
    model compromise, agent misbehavior) with roles and escalation
[ ] Third-party/vendor AI risk: assess the security of AI services, models, and data you consume
[ ] Supply-chain controls: AI-SBOM, provenance, integrity, signing (§22)
[ ] Training & awareness: developers understand AI-specific risks and secure patterns
[ ] Periodic red-teaming & re-assessment: AI changes fast; test on a cadence, not once

Risk framing for leadership (translate technical → business):

Confidentiality:  data/model/IP exposure (RAG leakage, extraction, embedding/inversion, prompt leakage)
Integrity:        poisoned data/models, manipulated outputs driving bad decisions, backdoors
Availability/cost:denial-of-service / denial-of-wallet, resource exhaustion
Safety/compliance:policy/regulatory violations from unsafe outputs or actions
Operational:      unauthorized actions by over-permissioned agents (the "excessive agency" risk)

The red team's role in the program. An assessment is one MEASURE activity; its value is realized when findings flow into MANAGE (prioritized remediation) and back into GOVERN (policy/standard updates) and MAP (updated risk picture). Recommend not just fixes but process: a pre-production red-team gate, continuous evals, monitoring, and a re-assessment cadence — because a one-time test of a fast-changing AI system has a short shelf life.

The principle: secure AI is a program, not a project. Ground your work in NIST AI RMF, deliver findings that update the organization's risk picture and controls, and push for the governance that keeps AI secure over time — pre-deployment gates, continuous evaluation, monitoring, and recurring red-teaming.


25. Secure AI Development Lifecycle & Design Patterns

Red-team findings are most valuable when they drive secure-by-design development. This section gives the secure-AI development lifecycle (AI-SDLC) and concrete, reusable secure design patterns that prevent whole classes of findings — the "build it right" counterpart to the attacks.

The AI Secure Development Lifecycle (AI-SDLC):

REQUIREMENTS   define the AI system's purpose, data sensitivity, required agency, and risk tolerance.
               Example: "a support bot that answers from public docs, takes NO actions, touches NO PII"
               has a very different security bar than "an agent that edits records and sends email."
DESIGN         threat-model (§6); choose an architecture that minimizes agency and isolates trust
               (apply the §19 defensive architecture up front). Decide the trust boundaries deliberately.
BUILD          implement secure patterns (below); least-privilege tools; output handling; secrets mgmt.
TEST/EVAL      automated safety & robustness evals; security review; red-team/eval gate BEFORE prod.
DEPLOY         hardened infra (§21); monitoring & detection wired in (§23); rate limits & cost controls.
OPERATE        continuous monitoring, guardrail-drift checks, incident response (§28), periodic re-test.

Secure design patterns (reusable, each prevents a class of attacks):

Pattern 1 — Minimize agency (dual control for actions). Default an AI to read/advise, not act. For any action, require either (a) narrow, scoped, read-biased tools, or (b) human approval. Prevents: the injection→harmful-action chain (§7/§9).

BAD:   agent has a general "run_shell" or "send_email to anyone" tool, auto-executed.
GOOD:  agent proposes an action → a scoped tool with per-user authZ executes it → human approves
       anything consequential/irreversible.

Pattern 2 — Untrusted-by-default context. Treat everything that isn't the system prompt (user input, RAG docs, tool outputs, memory, peer-agent messages) as untrusted data; delimit it clearly; never let its content grant privilege. Prevents: direct & indirect injection impact.

GOOD:  structure the context so retrieved/tool content is in a clearly-marked "data" region the model
       is instructed to treat as information, not instructions; apply guardrails to it; strip/neutralize
       instruction-like artifacts where feasible.

Pattern 3 — Injection-safe output handling. Never pass model output into a sink (shell, SQL, HTML, HTTP, file path) without the same validation/parameterization/encoding you'd apply to user input. Prevents: LLM output → XSS/SQLi/SSRF/RCE (OWASP LLM05).

BAD:   db.query("SELECT ... WHERE x='" + model_output + "'")   |  element.innerHTML = model_output
GOOD:  parameterized queries; context-aware output encoding; allow-list/validate tool arguments derived
       from model output; a dedicated "model output is untrusted" boundary.

Pattern 4 — Per-user authorization at every boundary. The model/agent is never the authorization authority. Each tool call, retrieval, and data access re-checks the authenticated end-user's permissions. Prevents: BOLA-style and cross-tenant leakage, confused-deputy escalation.

GOOD:  kb_search and directory_lookup filter results to what THIS user may see; send_email checks the
       user is allowed to email that recipient — enforced server-side, not by the model's say-so.

Pattern 5 — Isolation & least privilege everywhere. Isolate tenants/sessions/agents; scope every credential narrowly; sandbox risky tools (code execution). Prevents: lateral impact, credential abuse, blast-radius growth.

Pattern 6 — No secrets in the context. Secrets live in a secrets manager and are used by the application layer with scoped permissions — never in the system prompt, code, logs, or knowledge base. Prevents: prompt-leakage credential exposure (§15).

Pattern 7 — Fail closed & bound the loop. On ambiguity or guardrail uncertainty, refuse/escalate rather than act. Bound autonomous loops (max steps, max tool calls, cost/time budgets). Prevents: runaway agents, denial-of-wallet, and "keep trying until it works" evasion.

Pattern 8 — Observability by design. Log prompts, outputs, tool calls, and decisions from day one; wire detections and alerts (§23). Prevents: silent compromise; enables response.

A secure-AI design checklist (for the build phase):

[ ] Agency minimized; actions gated (scoped tools / human approval)
[ ] All non-system context treated as untrusted data; guardrailed
[ ] Model output validated/encoded before every sink
[ ] Per-user authZ enforced at every tool/retrieval/data boundary
[ ] Tenants/sessions/agents isolated; credentials scoped; risky tools sandboxed
[ ] No secrets in prompts/code/logs/KB (secrets manager)
[ ] Fail-closed defaults; bounded autonomous loops & budgets
[ ] Logging/monitoring/detection built in; infra hardened (§21)
[ ] Threat-modeled in design; red-team/eval gate before prod

The principle: most AI-security findings are preventable by architecture. These patterns — above all minimize agency and treat context as untrusted — turn "the model can be manipulated" (unavoidable) into "manipulation can't cause serious harm" (achievable). A red team's highest-value output is pushing the organization to adopt these patterns as standards.


26. Third-Party & Vendor AI Risk Assessment

Most organizations consume AI — hosted model APIs, AI SaaS products, third-party agents/plugins/connectors, and open models/datasets — rather than building everything. Assessing the security of AI you don't control is a core, practical part of AI risk management.

The categories of third-party AI risk:
- Hosted model APIs (you send data to a provider's model). Risks: data handling/retention by the provider, where your prompts/outputs go, training-on-your-data concerns, availability/SLA, provider-side security, and the provider's own safety/guardrail posture. Assess: the provider's data-use & retention terms, certifications (SOC 2, ISO), data-residency, zero-retention/enterprise options, and incident history. Control: data minimization before sending; use enterprise/no-train tiers; avoid sending secrets/regulated data; contractual terms.
- AI SaaS products (a vendor's AI feature embedded in a product). Risks: the vendor inherits all the OWASP LLM risks on your behalf; you have limited visibility. Assess: vendor security questionnaire extended for AI — how do they handle prompt injection, tool permissions, data isolation, and model supply chain? Control: due diligence, contractual security requirements, data-handling limits.
- Third-party agents / plugins / MCP connectors (external capabilities your AI calls). Risks: a malicious or vulnerable connector = supply-chain + tool-surface risk (§13/§22); over-broad connector scopes; the connector seeing your data. Assess: vet the connector's publisher, permissions/scopes requested, and data access; review its code if open. Control: least-privilege scopes, allow-listed connectors, sandboxing, monitoring.
- Open models / datasets / adapters (consumed from hubs). Risks: the full supply-chain set (§22) — tampering, backdoors, unsafe loading. Assess & control: provenance, integrity verification, safe loading, evaluation before use.

An AI vendor-assessment questionnaire (example — extend your standard vendor review):

DATA:        What data do we send? Where is it stored/processed? Retention? Is it used for training?
             Can we opt out / use a zero-retention tier? Data residency? Deletion on request?
SECURITY:    Certifications (SOC 2 / ISO 27001)? Pen-test / AI red-team cadence? Vulnerability disclosure?
             Encryption in transit/at rest? Tenant isolation? Access controls?
AI-SPECIFIC: How do they mitigate prompt injection & jailbreaks? Guardrails in place? If agentic —
             tool permissions, authorization model, human-in-the-loop for actions? Model supply-chain
             controls? Monitoring/abuse detection? Safety evaluations?
OPERATIONS:  Availability/SLA? Incident-response & notification process? Sub-processors? Change notice?
             Who is liable for AI errors/harms? Logging/audit availability to us?

The "data boundary" question (the most important one): for any third-party AI, trace exactly what data leaves your control, where it goes, who can see it, how long it's kept, and whether it trains a model. This single analysis drives most of the real risk and most of the controls (minimize, anonymize, use no-train/zero-retention tiers, or keep sensitive data out entirely).

Shared-responsibility for AI (like cloud). With hosted/SaaS AI, security is split: the provider secures their model/infrastructure; you remain responsible for what data you send, how you configure and integrate it, the permissions you grant, and how you handle its outputs. Know the split for each vendor — gaps live at the boundary.

Red-team angle on third-party AI: even when you can't test the vendor's internals, you can assess your integration — what data you send, what permissions you grant their connector/agent, how you handle their outputs (injection-safe?), and whether your use respects the data boundary. Those integration-side findings are within your control and often the real risk. The deliverable: a vendor risk rating + required controls + data-handling limits, feeding the governance program (§24).


27. Expanded Detection-Rule Library (example logic)

Building on §23, a broader library of illustrative detection patterns across the AI attack surface. These are conceptual rule shapes to adapt to your SIEM/guardrail platform — detections for attacker behaviors, not signatures.

Input / prompt-layer detections:

D-IN-1  High guardrail-classifier score on input, repeated from one user/session   → rate-limit + alert
D-IN-2  Heavy obfuscation/encoding, zero-width/hidden chars, or unusual token ratios in input
D-IN-3  Instruction-like imperatives in UNTRUSTED sources (RAG doc / tool output / web / email) directed
        at the assistant (override/ignore/exfiltrate semantics)                     → flag + correlate
D-IN-4  Many paraphrased retries immediately after refusals (persistence/evasion)   → throttle + alert
D-IN-5  Input attempting to elicit the system prompt / internal config              → alert (LLM07)

Output-layer detections (DLP for AI):

D-OUT-1 PII patterns in model output (emails, SSNs, card numbers, phone)            → redact/block + alert
D-OUT-2 Secret/key/credential patterns in output                                    → block + alert (critical)
D-OUT-3 System-prompt fingerprint appearing in output (leakage)                     → block + alert
D-OUT-4 Policy/safety-classifier trigger on output                                  → block + alert
D-OUT-5 Output containing executable/markup that will hit a sink unsanitized         → sanitize + alert (LLM05)

Agent / tool-layer detections (highest signal):

D-TOOL-1 Sensitive tool call (db_write, email_send, file_move, code_exec, external_http) that is
         unusual for the user / outside the expected workflow                       → require approval/alert
D-TOOL-2 Spike in tool-call volume or an unexpected tool-call sequence (loop/runaway) → throttle + alert
D-TOOL-3 Tool action touching data outside the authenticated user's scope           → BLOCK + alert (authZ fail)
D-TOOL-4 Outbound action (email/HTTP) shortly after the agent ingested untrusted content (injection→exfil
         correlation)                                                               → approval/alert
D-TOOL-5 Tool arguments derived from model output that look like injection (shell/SQL/URL metacharacters)
         → validate/block (LLM05)

Multi-agent detections:

D-MA-1  Inter-agent message failing authentication/integrity check                  → drop + alert (spoof)
D-MA-2  An agent deviating from the orchestrated workflow / skipping a checkpoint    → halt + alert
D-MA-3  Aggregate action across agents exceeding any single agent's intended scope   → alert

Data / RAG / embedding detections:

D-RAG-1 Unexpected write/change to the knowledge base                               → alert (poisoning, §11)
D-RAG-2 Anomalous retrieval (rarely-used doc suddenly dominating; retrieval for a query that shouldn't
        match it)                                                                   → alert
D-RAG-3 Cross-tenant/namespace retrieval                                            → BLOCK + alert (§12)
D-RAG-4 Bulk / systematic querying of the vector store (extraction/scraping)        → throttle + alert

Model / API detections:

D-MOD-1 One principal's query volume or cost >> baseline (extraction / denial-of-wallet, LLM10)
D-MOD-2 Systematic, diverse querying consistent with surrogate-model building        → throttle + alert
D-MOD-3 Repeated adversarial-looking inputs against an ML classifier (evasion probing) → alert

Infrastructure detections (classic security, applied to AI assets):

D-INF-1 Access to a model server / vector DB / registry / notebook from an unexpected source/network
D-INF-2 Authentication failures or new admin access on AI infrastructure
D-INF-3 Config changes to model-serving / guardrail components                      → alert + review
D-INF-4 A model artifact loaded from an unverified source / unexpected hash          → block + alert (§22)

Correlation = the real wins. The highest-fidelity detections chain signals: untrusted content ingested (D-IN-3) → behavior change → sensitive tool call (D-TOOL-1) → outbound action (D-TOOL-4) is a far stronger "injection→exfiltration" detection than any single rule. Build correlation across the input→model→tool→output flow.

Operationalizing: route AI telemetry to your SIEM; implement input/output guardrails inline (block in real time) and behavioral detections in the SIEM (alert/correlate); tune to reduce false positives (benign tool use, legitimate bulk queries); and validate with benign red-team canaries — if your simulated injection→tool-call→exfil chain doesn't light up these rules, that gap is a finding. Track metrics: detection coverage per ATLAS technique, time-to-detect on canary attacks, and false-positive rates.


28. AI Security Testing Tools & Techniques (authorized/defensive)

A reference to the categories of tooling used in authorized AI security assessment and continuous evaluation — the defensive/assessment toolkit, described by what each category does (not as a weaponization cookbook).

Automated LLM vulnerability scanning / eval frameworks. Open-source and commercial frameworks probe an LLM application for known weakness categories (prompt-injection susceptibility, data-leakage, toxicity/safety regressions, robustness) and produce a report — the AI analog of a web vulnerability scanner. Organizations use them for continuous safety/robustness regression testing in CI, and red teams use them for breadth before manual depth. (Examples in the ecosystem include open frameworks such as garak, Microsoft PyRIT, Giskard, and others.) Use them to triage; verify findings manually and assess impact in context.

Guardrail / moderation frameworks. Libraries and services that implement input/output guardrails — content classification, PII detection, topic/role constraints, and output validation — e.g., provider moderation endpoints and open-source guardrail toolkits. In an assessment, you evaluate whether these are deployed and how robust they are; in defense, you deploy and tune them (§18).

Red-team / adversarial-ML evaluation libraries. Research toolkits for evaluating model robustness to adversarial examples, membership inference, and extraction (classic ML security). Used to measure robustness and drive adversarial training and other defenses (§14).

RAG / data-pipeline testing. Tooling and techniques to assess retrieval access control, tenant isolation, and knowledge-base integrity — often a mix of custom scripts and standard data-store security testing against the vector DB (§11-12).

AI-SBOM / supply-chain tooling. Model/dataset scanning, integrity verification, signing, and dependency/SBOM tools applied to the AI stack — to verify provenance and catch tampered or unsafe artifacts (§22).

Infrastructure & cloud security tooling (classic, applied to AI). Your existing arsenal — network/service discovery, cloud-config and container scanners, secrets scanners, vulnerability scanners — pointed at AI assets (model servers, notebooks, vector DBs, registries, buckets of weights). This is where much of the highest-certainty testing happens (§17/§21).

Observability & detection tooling. SIEM/logging platforms plus AI-specific observability (tracing of prompts, tool calls, and agent decisions) to implement the detections in §23/§31.

Techniques (methodology, not payloads):

- Benign-canary testing: prove a control gap with harmless markers (no real harmful output).
- Differential/behavioral testing: compare behavior across inputs/paraphrases to characterize robustness.
- Boundary mapping: systematically enumerate where untrusted content reaches the model and what it can
  influence (the trust-boundary method, §6).
- Impact-driven assessment: weight every finding by the system's agency and data access.
- Detection validation: run benign canary chains and verify the blue team would see them.
- Continuous evaluation: regression-test safety/robustness as the model/app changes (AI drifts).

Responsible-use note. These tools are for systems you own or are authorized to test, used to find and fix weaknesses. Hosted providers have acceptable-use policies and often dedicated AI red-team/disclosure programs — work within them. The goal of the toolkit is a measurably safer, better-evaluated AI deployment — not a library of working attacks.

Build vs. buy for continuous assurance. Mature programs wire automated evals + guardrail monitoring + infra scanning + detection validation into CI/CD and operations, so AI security is continuous (because AI systems change fast), with periodic human red-teaming for depth. That combination — automation for breadth and regression, humans for creativity and impact — is the practical state of the art.


29. AI Incident Response

Even well-defended AI systems can be attacked or misbehave. An AI-aware incident-response (IR) capability — adapting classic IR to AI-specific incidents — closes the loop and is a key governance recommendation (§24).

AI-specific incident types to prepare for:

- Prompt-injection abuse / jailbreak leading to policy-violating output or unauthorized action
- Data leakage via the AI (PII, secrets, system prompt, cross-tenant/RAG leakage)
- Agent misbehavior / tool abuse (unauthorized or harmful actions taken by an agent)
- Data/model poisoning (corrupted RAG KB, tampered training data, backdoored model/adapter)
- Model/IP theft (extraction) or inference-API abuse (denial-of-wallet)
- AI-infrastructure compromise (exposed model server, poisoned pipeline, credential theft via the AI)
- Supply-chain compromise (malicious model/dataset/dependency introduced)

The IR lifecycle, adapted for AI (NIST 800-61 shape):

PREPARE   - AI asset inventory, logging/telemetry in place (§23), playbooks for the incident types above,
            roles, and the ability to disable/roll back AI features & revoke tool permissions quickly.
DETECT &  - alerts from the detections (§31); user/analyst reports; anomalous cost/behavior; poisoning
ANALYZE     indicators. Analyze: what was the vector (direct/indirect injection? infra? supply chain?),
            scope (which users/data/tenants/actions), and impact.
CONTAIN   - AI-specific containment: disable or constrain the affected AI feature/tool; revoke/scope-down
            tool permissions; isolate the affected agent/service; roll back a poisoned KB/model to a known-
            good version; block abusive principals; rotate exposed credentials.
ERADICATE - remove poisoned data / malicious artifacts; fix the root cause (the missing guardrail, the
            over-permissioned tool, the exposed infra, the unvetted model); restore clean data/models.
RECOVER   - restore the AI feature with the fix in place; re-validate behavior (eval/red-team) before
            re-enabling; monitor closely for recurrence.
LESSONS   - post-incident review feeding governance (§24): update guardrails, permissions, data controls,
            detections, and standards; re-threat-model; add a regression test for the specific attack.

AI-specific IR considerations:
- Reproducibility/non-determinism. An AI incident may be hard to reproduce exactly; capture the full context (prompt, retrieved content, tool calls, outputs) at the time — hence the importance of logging (§23).
- Rollback matters. For poisoning (KB/model/data), the key capability is versioning + rollback to a known-good state — build this in advance.
- Fast "kill switch." The ability to quickly disable an AI feature or revoke an agent's tools is a critical containment control — design for it.
- Data-exposure handling. If the AI leaked PII/regulated data, invoke the data-breach process (legal/compliance/notification) per policy.
- Preserve evidence. Logs of prompts, retrieved content, tool calls, and outputs are the forensic record — protect and analyze them (mindful they may contain sensitive data).
- Cross-functional. AI incidents can span app, data, infra, legal/compliance, and comms — coordinate like any major incident.

Readiness checklist:

[ ] AI telemetry & detections live (so you'll SEE incidents) (§23/§31)
[ ] Playbooks for the AI incident types above
[ ] Kill switch: can disable AI features & revoke tool permissions fast
[ ] Versioning + rollback for KB, models, datasets (poisoning recovery)
[ ] Credential rotation process for AI-held secrets
[ ] Data-breach process wired in for AI data leakage
[ ] Post-incident → governance feedback loop (fix + regression test + re-threat-model)

The principle: assume incidents will happen, and prepare the AI-specific capabilities that classic IR doesn't cover — telemetry to detect, a kill switch and permission-revocation to contain, versioning/rollback to recover from poisoning, and a feedback loop into governance. A red team should assess IR readiness too: "if this attack succeeded, could you detect, contain, and recover?" — gaps there are findings.


30. Lessons from Real AI Security Incidents (public)

AI security isn't theoretical — there's a growing body of publicly-documented incidents, research demonstrations, and MITRE ATLAS case studies. Rather than reproduce specifics, here are the recurring patterns and lessons that show up again and again — study the public cases (ATLAS, OWASP, vendor advisories, research papers) for detail.

Recurring pattern 1 — Indirect prompt injection via ingested content is real and impactful. Numerous public demonstrations show AI assistants and agents being influenced by instructions hidden in web pages, documents, emails, and other content they process — not by the user, but by third-party data. Lesson: this is the signature real-world AI risk; treat all ingested content as untrusted and minimize agency so influence can't become harm (§7, §19).

Recurring pattern 2 — Agents + tools turn text flaws into real actions. The most serious demonstrations involve agents that could act (send data, call APIs, execute code) being manipulated into doing so. Lesson: agency is the amplifier — least-privilege tools, human approval, and sandboxing are the controls that matter most (§9, §13).

Recurring pattern 3 — Exposed AI infrastructure is common. Researchers repeatedly find unauthenticated model servers, notebooks, vector databases, and public buckets of model weights/datasets on the internet. Lesson: classic exposure/misconfiguration is widespread in AI deployments and is the highest-certainty, most-preventable risk (§17, §21).

Recurring pattern 4 — Unsafe model loading = code execution. The ecosystem has had real cases of malicious models on public hubs exploiting unsafe serialization to run code on load. Lesson: verify provenance/integrity and use safe serialization + sandboxed loading (§22).

Recurring pattern 5 — Data leakage through AI channels. Public incidents include system-prompt leakage, training-data regurgitation, and cross-user/cross-tenant data exposure. Lesson: no part of the AI's context is confidential by default; apply data minimization, access control, isolation, and output DLP (§15).

Recurring pattern 6 — Supply-chain and dependency issues. Compromised or typosquatted models/datasets/packages, and vulnerable ML frameworks, have all appeared. Lesson: AI-SBOM, provenance, integrity, and dependency security (§16, §22).

Recurring pattern 7 — Guardrails are bypassable; defense in depth is essential. Every published guardrail/safety measure has been shown to be imperfect under determined probing. Lesson: never rely on a single guardrail or on the base model's refusals alone; layer input/output guardrails, limit agency, and monitor (§8, §18, §19).

Recurring pattern 8 — Overreliance causes harm. Incidents where humans trusted incorrect or manipulated AI output (bad decisions, fabricated information acted upon). Lesson: keep humans in the loop for consequential decisions; present evidence, not just conclusions; verify AI output (§9, §30 — the "AI is an assistant, not an oracle" principle).

How to learn from public cases: MITRE ATLAS publishes real-world case studies mapped to its techniques — the single best structured source. OWASP's GenAI/LLM project, NIST publications, vendor security blogs/advisories, academic papers (the research community publishes extensively on injection, extraction, poisoning, and privacy attacks), and the AI labs' own red-team and safety publications round it out. Reading these builds the pattern-recognition that lets you anticipate how a given AI system will fail.

The meta-lesson: across all the public incidents, the same handful of root causes recur — untrusted content treated as trusted, excessive agency, exposed infrastructure, weak supply-chain integrity, and overreliance. Which is exactly why this guide keeps returning to the same defenses: treat context as untrusted, minimize agency, harden infrastructure and the supply chain, and keep humans and monitoring in the loop. Master those, and you address the majority of real-world AI risk.


31. The Capstone Engagement — Methodology

A full AI red-team engagement ties everything together: a structured, authorized assessment of a realistic enterprise AI environment that produces a prioritized path to a safer system. Here's the methodology (not a target-specific attack script).

Phase-by-phase:

1. SCOPING & RULES OF ENGAGEMENT (§2)
   - define in-scope AI systems/models/apps/APIs/infra; off-limits items; data-handling & safety rules
     (benign-canary approach — prove gaps without real harm); provider-ToS/legal; disclosure process;
     success criteria (what the client wants assurance on).
2. RECONNAISSANCE (§5)
   - inventory AI assets: LLM apps/features, model endpoints, RAG/vector DBs, agents & tools,
     MLOps/registries/notebooks, cloud/containers, and the data/model/dependency supply chain.
   - map exposure, dependencies, tools/permissions, and data flows.
3. THREAT MODELING (§6)
   - identify high-value assets, trust boundaries, and prioritized attack paths (likelihood × impact).
   - produces the testing plan AND a standalone defensive artifact for the client.
4. TESTING (per component, prioritized by the threat model)
   - LLM-app layer: prompt-injection susceptibility & impact; guardrail robustness; system-prompt/
     sensitive-data exposure (concept + benign canaries, §7-8, §15).
   - agents & multi-agent: memory/tool/workflow manipulation surface; excessive agency;
     A2A trust boundaries (§9-10).
   - RAG/embeddings: KB write access & poisoning resistance; retrieval access control; vector-DB
     exposure; tenant isolation (§11-12).
   - tool/MCP surface: tool permissions, output handling (injection downstream), orchestration
     exposure, connector supply chain (§13).
   - model/ML: extraction/abuse resistance (rate-limit/auth), robustness, privacy posture (§14-15).
   - supply chain: provenance/integrity of data/weights/adapters; safe loading; MLOps/registry
     security (§16).
   - infrastructure: exposed/unauth model servers, notebooks, vector DBs, registries; container/cloud
     misconfig; secrets; CVEs (§17) — often the highest-certainty findings.
5. IMPACT VALIDATION & CHAINING
   - demonstrate realistic impact with minimal, benign proof (canaries); show how AI-layer + infra
     findings combine into business impact (e.g., exposed vector DB → KB poisoning → agent influence,
     or exposed model server → model theft; injection + over-permissioned tool → data access).
6. DETECTION ASSESSMENT (§18)
   - would the blue team have SEEN this? assess logging/detection coverage and gaps.
7. REPORTING & REMEDIATION (§21)
   - prioritized findings with impact and concrete fixes; map to OWASP LLM Top 10 / MITRE ATLAS /
     NIST; a remediation roadmap toward the defensive architecture (§19); debrief the blue team.

The mindset for the capstone: think like a realistic adversary across the whole system (model + app + data + agents/tools + infra + supply chain), prioritize by real risk, validate impact responsibly with benign canaries, and — always — convert findings into a clear, prioritized path to a more secure, better-monitored AI deployment. The engagement succeeds when the client can act on it.


32. Worked Example — Threat Modeling an Enterprise AI Assistant

A concrete, benign threat-modeling walkthrough for a realistic system, showing how §6 is applied in practice.

The system (example): "CorpAssist," an internal AI assistant. It's an LLM app with:
- a chat UI for employees; a system prompt defining its role;
- a RAG pipeline over an internal knowledge base (HR docs, policies, wiki) in a vector DB;
- tools: search the KB, look up an employee directory, create IT tickets, and send notification emails;
- memory of the conversation; hosted model via an API; running in the company cloud/K8s.

Step 1 — High-value assets:

- the HR/policy knowledge base (confidential employee data)       [confidentiality]
- the employee directory data                                      [confidentiality/privacy]
- the email + ticket tools (actions taken on employees' behalf)    [integrity/operational]
- the system prompt (internal logic)                               [IP, minor]
- cloud credentials the app/infra holds                            [access to backends]

Step 2 — Trust boundaries:

employee input → app                 (untrusted text)
KB documents → model context         (are HR docs trusted as instructions? → indirect-injection risk)
directory/tool outputs → model       (untrusted tool outputs)
model output → email/ticket tools    (does the AI's say-so authorize sending email/creating tickets?)
app → cloud/K8s/vector DB            (infra & data access)

Step 3 — Attack paths (prioritized by likelihood × impact):

HIGH:  indirect injection via a poisoned KB doc (anyone who can edit the wiki) → influences the agent
       → induces an unintended email to all staff OR exfiltration of directory data via email tool.
       (combines §7 injection + §9 agency + §11 RAG poisoning — the top risk)
HIGH:  cross-tenant/authorization gap in retrieval → an employee retrieves HR data about others.
MED:   over-permissioned email tool (can email anyone, no approval) → large blast radius if misused.
MED:   infra: vector DB or model endpoint exposed/unauth on the internal network (§17).
LOW:   system-prompt leakage (minor IP; ensure no secrets in it).
LOW:   denial-of-wallet via unbounded querying.

Step 4 — Prioritized defensive recommendations (the deliverable):

1. LEAST-PRIVILEGE AGENCY: scope the email tool (templates only / limited recipients) and require
   HUMAN APPROVAL for mass or external emails — caps the blast radius of any injection.
2. TREAT KB CONTENT AS UNTRUSTED DATA: delimit retrieved docs; guardrail them; control who can write
   to the KB; integrity-monitor it (§11).
3. RETRIEVAL ACCESS CONTROL: enforce per-employee authorization on HR/directory data; tenant/role
   isolation so one employee can't retrieve others' sensitive records (§11-12).
4. OUTPUT HANDLING & DLP: filter outputs for PII before display/sending (§15).
5. INFRA: authN + network isolation for the vector DB and model endpoint; secrets in a manager (§17).
6. MONITORING: log tool calls (esp. email/ticket) and alert on anomalies; detect KB changes (§23).

The takeaway: without touching a single exploit, the threat model already tells the client their top risk is an over-permissioned email tool combined with a writable, trusted knowledge base — and the single highest-leverage fix is least-privilege agency + human approval on the email tool, because it caps the damage of the whole injection→exfiltration chain. That prioritized, architecture-level insight is the core value of AI red teaming.


33. Worked Threat Model #2 — Agentic Coding Assistant

Different AI system types have different risk profiles. Here's a threat model for an agentic coding assistant — an AI that reads a codebase, runs commands, edits files, and executes code. High agency = high stakes.

The system (example): "DevAgent" — assists developers by reading repos, writing/editing code, running build/test commands, executing code in a workspace, searching the web for docs, and installing packages. It runs with access to a developer's workspace and sometimes CI.

High-value assets:

- source code & secrets in the repo (API keys, .env, credentials)        [confidentiality]
- the ability to EXECUTE code & shell commands                            [integrity/RCE-equivalent]
- package installation (supply-chain entry point)                          [integrity/supply chain]
- CI/CD access (if wired in)                                               [critical — pipeline compromise]
- the developer's workstation / cloud credentials                          [access/pivot]

Trust boundaries & the key risk:

web content / docs the agent reads → agent     (INDIRECT INJECTION via a malicious web page or
                                                 README/dependency doc — the top risk here)
repo content / issues / PRs → agent             (untrusted content in the project itself)
agent decision → code execution / shell         (the agent's say-so → real command execution)
agent → package install                          (pulls untrusted packages → supply chain)
agent → CI/CD                                    (highest blast radius)

Attack paths (prioritized):

CRITICAL: indirect injection via web/doc/repo content → agent runs attacker-chosen code or shell
          commands in the workspace → exfiltrate repo secrets, or pivot using workstation/cloud creds.
          (the "agent with code execution reading untrusted content" is the canonical high-risk pattern)
HIGH:     agent induced to install a malicious/typosquatted package → supply-chain compromise.
HIGH:     agent with CI/CD access manipulated → poisoned build/deploy (affects everything downstream).
MED:      agent reads & leaks repo secrets into its output / web search / logs.
MED:      over-broad file access → modifies files outside the intended project.

Defensive recommendations (prioritized):

1. SANDBOX EXECUTION: run all code/shell execution in an isolated, ephemeral sandbox with NO access to
   real secrets, cloud credentials, or the broader network — the single most important control.
   A compromised agent then executes only in a throwaway box.
2. LEAST PRIVILEGE: no standing cloud/CI credentials in the agent's environment; scope file access to
   the project; read-only where possible; human approval for pushes/deploys/installs.
3. TREAT WEB/REPO/DOC CONTENT AS UNTRUSTED (indirect injection) — guardrail it; don't let it drive
   unbounded execution without checkpoints.
4. PACKAGE/SUPPLY-CHAIN GUARDRAILS: allow-listed or reviewed dependencies; pin versions; scan installs.
5. SECRETS HYGIENE: keep secrets out of the agent's reach; scan for secret exposure in output/logs.
6. BOUND THE LOOP: max steps/commands; cost/time budgets; human-in-the-loop for consequential actions.
7. MONITOR: log every command/file-edit/install; alert on sensitive operations (§23).

The takeaway: a coding agent is essentially remote code execution as a feature, so the dominant control is sandboxing + credential isolation — assume it can be manipulated via the untrusted content it reads, and ensure that when it is, it's executing in a disposable box with nothing valuable to steal or break. This is the clearest illustration of the core principle: you can't prevent manipulation, so you cap the blast radius by design.


34. Worked Threat Model #3 — Customer-Support Chatbot

A customer-facing support chatbot has a different profile: lower agency, but public exposure, brand/safety risk, and potential access to customer data.

The system (example): "SupportBot" — public-facing, answers questions from a product KB, looks up a customer's order/account (after auth), can issue refunds up to a limit, and escalates to humans. Exposed to anyone on the internet.

High-value assets:

- customer account/order data (PII)                         [confidentiality/privacy]
- the refund/account-action capability                      [integrity/financial]
- brand reputation & safety (public outputs)                [reputational/compliance]
- the KB (product info)                                     [integrity]

Trust boundaries & key risks:

public user input → bot        (fully untrusted, high-volume, adversarial — direct injection/jailbreak)
KB content → bot context       (if KB is editable/crawled, indirect injection)
bot → account-lookup tool      (does it enforce the authenticated customer's scope?)
bot → refund tool              (financial action — authorization & limits critical)
bot output → public user       (brand/safety: harmful/off-policy/embarrassing outputs)

Attack paths (prioritized):

HIGH:  BOLA/cross-customer access — a user induces the bot to look up ANOTHER customer's order/account
       (if the lookup isn't strictly scoped to the authenticated user).
HIGH:  refund abuse — manipulating the bot into issuing refunds beyond policy / to the wrong account
       (financial loss), if refund authorization is weak or model-authorized.
MED:   jailbreak → harmful/off-brand public output → reputational/compliance harm.
MED:   data leakage — bot surfaces internal KB content or PII it shouldn't.
MED:   denial-of-wallet — high-volume adversarial querying inflates cost.
LOW:   system-prompt leakage.

Defensive recommendations (prioritized):

1. STRICT PER-CUSTOMER AUTHORIZATION on all account/order lookups — scoped to the AUTHENTICATED user,
   enforced server-side (not by the model). Prevents cross-customer data access.
2. FINANCIAL ACTION CONTROLS: refunds require server-side policy enforcement (limits, eligibility),
   bound to the authenticated account; human approval above a threshold; full audit log.
3. OUTPUT GUARDRAILS: filter outputs for safety/brand/PII; moderation layer on public responses.
4. INPUT GUARDRAILS + RATE LIMITING: handle adversarial public input; throttle abuse; cost controls.
5. KB HYGIENE: control KB content; keep internal-only/PII data out of the public bot's KB.
6. ESCALATION & FAIL-CLOSED: on uncertainty, escalate to a human rather than act.
7. MONITORING: log lookups, refunds, guardrail triggers; alert on anomalies (§23).

The takeaway: for a public support bot, the agency is narrower but the exposure is total, so the priorities are strict per-customer authorization (to stop cross-customer data access — the classic BOLA risk) and tight financial-action controls (to stop refund abuse), backed by output moderation for brand/safety. Low agency doesn't mean low risk when real customer data and money are one tool-call away.


35. Worked Threat Model #4 — Multi-Agent Data Pipeline

A multi-agent system where specialized agents collaborate on a workflow illustrates the inter-agent trust risks of §10.

The system (example): "DataCrew" — a pipeline of agents: a Researcher agent gathers data (web + internal docs), an Analyst agent processes it, a Writer agent drafts a report, and an Orchestrator coordinates them. Agents pass messages and intermediate results to each other.

High-value assets:

- the internal data the Researcher accesses                     [confidentiality]
- the integrity of the final report (drives decisions)          [integrity]
- any tools the agents hold (web, DB, file, publish)            [agency/blast radius]
- the orchestration workflow                                     [integrity]

Trust boundaries & key risks:

external/internal content → Researcher     (indirect injection enters the pipeline HERE)
Researcher → Analyst → Writer messages     (do downstream agents TRUST upstream output as instructions?)
any agent → its tools                       (aggregate agency across the crew)
Orchestrator ↔ agents                        (can an agent redirect the whole workflow?)
Writer → publish/output                     (where the corrupted result exits)

Attack paths (prioritized):

HIGH:  indirect injection at the Researcher (a poisoned web page/doc) propagates DOWNSTREAM — the
       Analyst and Writer trust the Researcher's output and carry the injected instruction/content
       through to the final report or a tool action. (injection + inter-agent trust = cascade)
HIGH:  privilege aggregation — individually modest agents whose COMBINED tools (web read + DB + publish)
       enable exfiltration or harmful action no single agent was meant to do.
MED:   workflow corruption — manipulating the Orchestrator/messages so the crew skips a validation step
       or pursues an attacker-chosen goal.
MED:   agent impersonation — if inter-agent messages aren't authenticated, a forged message steers an agent.
MED:   error/hallucination propagation — one agent's bad output treated as authoritative downstream.

Defensive recommendations (prioritized):

1. TREAT INTER-AGENT MESSAGES AS UNTRUSTED DATA — each agent validates inbound messages against policy;
   no agent executes another's "instructions" blindly. (breaks the downstream-cascade, §10)
2. BOUND AGGREGATE AGENCY — reason about the crew's COMBINED capabilities; least-privilege the composition;
   the Researcher that reads untrusted web content should NOT also hold publish/DB-write power.
3. TRUSTED ORCHESTRATOR enforces the intended workflow, validates each transition, and is the only
   component that can advance/branch the pipeline — individual agents can't redirect it.
4. AUTHENTICATE inter-agent messages (identity + integrity) — no anonymous trust between agents.
5. VALIDATION CHECKPOINTS between stages (consistency/policy checks); human review before publish.
6. ISOLATE agents (separate contexts/permissions) so a compromise isn't automatically lateral.
7. END-TO-END MONITORING of the whole workflow, not just per agent (§23).

The takeaway: multi-agent risk is about trust cascading across agent boundaries — the fix is to not extend trust implicitly: validate inter-agent messages as untrusted, bound the aggregate agency (especially keeping untrusted-content-reading agents away from powerful tools), and keep a trusted orchestrator with validation checkpoints over the whole workflow. Design so that injecting one agent can't quietly steer the entire crew.


36. Worked Threat Model #5 — AI Embedded in Security Tooling

AI is increasingly embedded in security tools (SOC copilots, automated triage, AI-assisted response). These are high-trust systems where manipulation has outsized consequences — a worthwhile threat model and a nice tie-in to defensive operations.

The system (example): "SOCpilot" — an AI assistant in the SOC that reads alerts and logs, summarizes incidents, enriches indicators, suggests (or takes) response actions like isolating a host or blocking an IP, and drafts tickets. It ingests security telemetry and has response tooling.

High-value assets:

- the integrity of its security analysis & recommendations   [integrity — analysts trust it]
- response actions (isolate host, block IP, disable account) [operational — can disrupt the business]
- access to sensitive security telemetry                      [confidentiality]
- the SOC's trust in the tool                                 [operational]

Trust boundaries & the key (ironic) risk:

attacker-influenced log/alert content → AI     (the AI reads data an ATTACKER can shape — logs,
                                                 alert fields, file names, user-agents — INDIRECT
                                                 INJECTION from the very adversary being investigated)
AI analysis → analyst decisions                 (poisoned analysis → wrong response)
AI → response tooling                           (manipulated AI → harmful/disruptive actions)

Attack paths (prioritized):

HIGH:  indirect injection via telemetry — an attacker plants injection content in a field the AI will
       read (a crafted log entry, filename, user-agent, alert note) to manipulate the AI's analysis
       (hide their activity, misdirect the analyst) or its response actions. (the attacker poisoning
       the tool investigating them)
HIGH:  manipulated AUTO-RESPONSE — if the AI can take actions, inducing it to isolate the wrong host /
       block legitimate IPs / disable accounts = self-inflicted denial of service.
MED:   AI misses or downplays a real incident (integrity/availability of detection).
MED:   sensitive-telemetry leakage via the AI.

Defensive recommendations (prioritized):

1. HUMAN-IN-THE-LOOP FOR RESPONSE ACTIONS — the AI RECOMMENDS; a human approves isolate/block/disable.
   (an AI that auto-responds based on attacker-influenceable data is dangerous) — the top control.
2. TREAT TELEMETRY AS UNTRUSTED (it literally contains attacker-controlled content) — guardrail inputs;
   don't let log/alert content act as instructions; sanitize/delimit fields fed to the AI.
3. LEAST-PRIVILEGE RESPONSE TOOLING — scoped actions, reversible where possible, with approval gates.
4. ANALYST SKEPTICISM / EXPLAINABILITY — present evidence, not just conclusions; analysts verify (the AI
   is an assistant, not an oracle — the eSOC "verify AI output" principle).
5. MONITOR THE AI's actions & recommendations; log and alert; detect manipulation attempts.
6. FAIL-CLOSED — on uncertainty, escalate to a human rather than auto-act.

The takeaway: AI in security tooling is a high-trust, high-irony target — it reads data the adversary can influence and may take disruptive actions — so the controls are human-in-the-loop for response, treating telemetry as untrusted input, and analyst verification. It's a vivid reminder of the universal principle (don't let untrusted content drive consequential action) in the one place where getting it wrong hands the attacker your own defenses.


37. Worked Example — Capstone Engagement Walkthrough (benign)

A benign, end-to-end example of assessing a realistic enterprise AI environment — showing how the whole methodology (§20-method, §5-6, §19) comes together into a report. All findings are illustrative and proven with benign canaries.

Scope (example): assess "CorpAssist" (§26) plus its supporting AI infrastructure; benign-canary rules; no real harmful output; internal network access provided; 2-week box.

Phase 1 — Recon (§5). Inventory the AI assets:

- the chat app + its backend /chat and /embeddings API calls (observed in traffic)
- a self-hosted vector DB on the internal network
- a hosted LLM provider API (fingerprinted from traffic)
- tools exposed to the agent: kb_search, directory_lookup, create_ticket, send_email
- MLOps: a model registry and an experiment-tracking UI; cloud storage with an embeddings export
- supply chain: the embedding model + framework versions (from docs/SBOM)

Phase 2 — Threat model (§6/§26). Produces the asset/boundary/path map; top risk = KB poisoning → agent → email exfiltration; secondary = retrieval authorization and infra exposure.

Phase 3 — Testing (prioritized by the threat model), example findings:

FINDING 1 (High) — Over-permissioned agent tool (excessive agency, OWASP LLM06 / ATLAS Execution+Exfil)
  The send_email tool can email ANY address with ANY content, authorized solely by the model's request,
  with no human approval. Benign proof: a planted canary instruction in a test KB doc caused the agent
  to attempt emailing a (test) canary token to a test address.
  Impact: a single injected instruction could exfiltrate data or spam staff.
  Fix: scope the tool (internal recipients / templates), require human approval for external/bulk email,
       per-tool authZ bound to the requesting user, egress monitoring (§13).

FINDING 2 (High) — Knowledge base is writable by any employee + treated as trusted (LLM04/01, ATLAS Persistence)
  Any employee can edit wiki pages ingested into RAG; retrieved content is injected into context without
  being treated as untrusted. Benign proof: a canary marker placed in a wiki page later appeared in the
  agent's behavior, showing retrieved content influences it (indirect-injection path).
  Fix: control/curate KB writes; provenance + integrity monitoring; delimit & guardrail retrieved content;
       least-privilege agency so influence can't become harmful action (§11, §7).

FINDING 3 (High) — Retrieval lacks per-user authorization (LLM08, ATLAS Collection)
  HR documents are retrievable regardless of the requesting employee's role. Benign proof: a standard
  test account retrieved a restricted (test) HR document via a normal query.
  Fix: enforce document-level authorization at retrieval, scoped to the authenticated user; tenant/role
       isolation (§11-12).

FINDING 4 (Critical) — Exposed, unauthenticated vector DB (infra, ATLAS Initial Access)
  The vector DB is reachable on the internal network with no authentication; any internal host can read
  or write the embedded knowledge base.
  Impact: KB disclosure AND poisoning at the source.
  Fix: authN + least-privilege access + network isolation + encryption (§17).

FINDING 5 (High) — Public cloud storage of an embeddings export (infra/privacy, ATLAS Exfiltration)
  An embeddings export of the HR KB sits in a public-read bucket. Given embedding-inversion risk, this
  may leak information about the underlying confidential documents.
  Fix: make the bucket private; least-privilege IAM; treat embeddings of sensitive data as sensitive (§12,§17).

FINDING 6 (Medium) — Model registry / experiment UI exposed without auth (infra, ATLAS ML Model Access)
  Fix: authN + network isolation; no public exposure (§17).

FINDING 7 (Medium) — No rate limits / cost controls on the inference path (LLM10)
  Fix: rate limits, quotas, cost alerts, abuse detection (§14).

FINDING 8 (Low) — System prompt retrievable; contains no secrets (good) but discloses internal logic (LLM07)
  Fix: assume it may leak; keep secrets out (already the case); minor.

Phase 4 — Chaining & impact (§20). The capstone insight: Findings 4+2+1 chain — an attacker with internal access writes a poisoned doc directly into the exposed vector DB (F4), which is trusted content (F2), influencing the agent to misuse the over-permissioned email tool (F1) to exfiltrate HR data — a realistic, high-impact path combining infra + data + agency weaknesses. Finding 3 independently enables broad HR-data access.

Phase 5 — Detection assessment (§23). Tool calls (incl. email) were not logged; KB writes were not monitored; no alerting on anomalous retrieval. The blue team would not have seen the chain — itself a key finding.

Phase 6 — Reporting & roadmap (§21/§19). Prioritized remediation:

1 (now):   least-privilege agency + human approval on send_email (F1) — caps the worst chain cheaply
2 (now):   authN + isolation on the vector DB (F4) and private the embeddings bucket (F5)
3 (next):  retrieval authorization & tenant isolation (F3); control KB writes + integrity monitoring (F2)
4 (next):  log & alert on tool calls, KB changes, anomalous retrieval (detection gaps, §23)
5 (ongoing): rate limits/cost controls (F7); registry/UI authN (F6); pre-deployment red-team gate (§24)

The takeaway: the engagement's value isn't the individual findings — it's the prioritized, architecture-level roadmap showing the client that a few high-leverage fixes (least-privilege email tool, vector-DB auth, retrieval authorization, and tool-call logging) break the dangerous chain and dramatically reduce risk. That is what a great AI red-team delivers: realistic adversary insight converted into a clear path to a safer, better-monitored AI deployment.


38. Reporting AI Security Findings

The report is the deliverable — where your assessment becomes safer systems. AI security reporting follows professional pentest-reporting discipline, adapted for AI-specific nuances.

Report structure:

1. Executive Summary — overall AI risk posture, key findings, and business impact, in plain language.
2. Scope & Methodology — systems/models/apps/infra tested; approach; standards (OWASP LLM Top 10,
   MITRE ATLAS, NIST AI RMF); safety/benign-canary approach used.
3. Threat Model — high-value AI assets, trust boundaries, and the attack paths assessed (a valuable
   standalone artifact for the client).
4. Findings — per finding:
      • Title & severity (impact × likelihood, noting non-determinism/reproducibility)
      • Affected AI component & trust boundary
      • Description — what it is and WHY it exists (the mechanism)
      • Evidence — benign-canary proof, logs, configs, requests/responses (reproducible, no real harm)
      • Impact — concrete consequence given the system's agency/data (data exposure, unauthorized
        action, model theft, policy violation, service/cost impact)
      • Mapping — OWASP LLM Top 10 ID, MITRE ATLAS technique, NIST category
      • Remediation — specific, actionable, prioritized fix (toward the §19 architecture)
5. Detection & Monitoring Gaps — what the blue team couldn't see; logging/detection recommendations.
6. Remediation Roadmap — prioritized plan toward a secure AI architecture.
7. Appendices — methodology detail, references, reproduction steps (benign).

AI-specific reporting principles:
- Benign evidence. Prove the control gap with canaries/proxies, not real harmful output — the report should demonstrate risk responsibly and be safe to share.
- Severity with nuance. Weight by the system's agency and data access (a susceptible talk-only bot ≠ a susceptible agent with DB/email access) and note reproducibility/likelihood given non-determinism.
- Impact framing. Translate AI-layer findings into business terms — data breach, unauthorized actions, IP/model theft, compliance/safety violations, cost/denial-of-wallet.
- Actionable, prioritized remediation. Map each finding to a specific control from §19; lead with the highest-leverage fixes (usually least-privilege agency, access control/isolation on data, and infrastructure exposure).
- Standards mapping. OWASP LLM Top 10 + MITRE ATLAS + NIST gives credibility and a shared language for the defenders.
- Support the blue team. Include detection/monitoring recommendations and offer a debrief — the goal is a measurably safer, better-observed AI system.

The bottom line: a great AI red-team report doesn't just list how the AI could be attacked — it gives the organization a clear, prioritized, standards-aligned plan to make its AI deployment safe, and the monitoring to keep it that way.


39. Study Plan, Resources & Glossary

Study approach. AI red teaming sits at the intersection of (a) AI/ML understanding, (b) classic application and infrastructure security, and (c) the new AI-specific attack taxonomy and defenses. Build all three, and ground everything in the public frameworks and in defensive thinking.

Study plan (flexible):

Phase 1  Foundations: how AI changes the attack surface (§1), the AI/ML stack (§4), and the frameworks
         — OWASP LLM Top 10 & ML Top 10, MITRE ATLAS, NIST AI RMF (§3). Read these primary sources.
Phase 2  Methodology: AI red-team lifecycle (§2), reconnaissance (§5), threat modeling (§6).
Phase 3  LLM-app security: prompt injection concept & defense (§7), guardrail evaluation (§8),
         sensitive-data exposure (§15) — focus on mechanisms and defenses.
Phase 4  Agents & orchestration: AI agents (§9), multi-agent/A2A (§10), MCP/tool surfaces (§13).
Phase 5  Data & model layer: RAG security (§11), embeddings/vector DBs (§12), model extraction &
         adversarial ML (§14).
Phase 6  Supply chain (§16) & infrastructure (§17) — apply your classic security skills to AI assets.
Phase 7  Defense & detection: guardrail/detection architecture (§18), defensive architecture (§19).
Phase 8  Capstone methodology (§20) & reporting (§21) — practice assessing & writing up a whole AI system.
Hands-on: practice on INTENTIONALLY-vulnerable, authorized AI labs and your own deployments (below).

Authorized practice & learning resources (public):

Frameworks/standards:  OWASP Top 10 for LLM Applications; OWASP ML Security Top 10; OWASP GenAI
                       Security Project; MITRE ATLAS; NIST AI RMF & NIST AI 100-2 (adversarial ML);
                       Google SAIF; cloud/vendor AI security guidance; lab red-team publications.
Intentionally-vulnerable / practice (authorized):  OWASP tooling and deliberately-vulnerable LLM apps
                       for learning prompt-injection/guardrail concepts; AI/ML security CTFs and ranges;
                       your own test deployments. (Practice only where explicitly authorized.)
Classic security base: your existing web/API, cloud, container, and supply-chain security skills apply
                       directly to AI infrastructure and the application around the model.

Glossary:
- LLM / Generative AI — large language model / AI that generates content.
- Prompt injection (direct/indirect) — untrusted input/content manipulating model behavior (OWASP LLM01).
- Jailbreaking / guardrail evaluation — bypassing (or measuring the robustness of) safety controls.
- System prompt — the developer's instructions configuring the model's behavior.
- AI agent — an AI that uses memory, tools, and autonomy to take actions.
- Multi-agent / A2A — multiple agents collaborating; agent-to-agent communication.
- RAG (Retrieval-Augmented Generation) — grounding model answers in retrieved documents.
- Vector database / embeddings — stores of semantic vector representations; the RAG/search backbone.
- Embedding inversion — inferring/reconstructing source information from an embedding.
- MCP (Model Context Protocol) — a standard for connecting models to tools/context; part of the "tool surface."
- Excessive agency — an AI with more tools/permissions/autonomy than needed (OWASP LLM06).
- Improper output handling — trusting model output unsafely → downstream injection (OWASP LLM05).
- Model extraction / theft — approximating/stealing a model via querying.
- Adversarial example / evasion — crafted input causing model misbehavior.
- Data/model poisoning — corrupting training/fine-tuning/RAG data or model artifacts.
- Model inversion / membership inference — privacy attacks recovering training-data info / membership.
- AI supply chain — datasets, weights, adapters, frameworks, hubs, and pipelines (OWASP LLM03).
- MITRE ATLAS — the ATT&CK-style knowledge base of AI/ML adversary tactics & techniques.
- OWASP LLM Top 10 / ML Top 10 — the standard AI vulnerability taxonomies.
- NIST AI RMF — the AI risk-management framework (Govern/Map/Measure/Manage).
- Guardrails — input/output filters and controls constraining AI behavior.
- Least-privilege agency — minimizing an AI's tools/permissions/data — the master mitigation.


End of guide. Responsible AI red teaming assesses the whole AI system — model, application, data pipelines, agents/tools, supply chain, and infrastructure — against realistic adversaries, in order to find and fix weaknesses and make deployments safer. Ground your work in the public frameworks (OWASP LLM Top 10, MITRE ATLAS, NIST AI RMF), map the attack surface and trust boundaries, assess each component's exposure and impact (using benign canaries, never real harm), and above all convert every finding into concrete, prioritized defenses — with least-privilege agency, untrusted-by-default context handling, injection-safe tool surfaces, data/supply-chain/infrastructure hardening, and monitoring as the pillars. The goal, always, is safer AI.

Leave a heart if you found this helpful

Comments

Sign in to leave a comment