Education hub

Learn AI security testing

Every attack module explained, what it is, how it works, the real-world incidents behind it, and how to defend against it. Covers the full attack chain from the model layer through the API layer to agentic pipelines.

173 red team attacks25 attack modulesOWASP LLM Top 109 compliance frameworksModel IntelligenceISO 42001 · EU AI ActEchoLeak · Copilot CVEs
New · Interactive tool

Map your AI build to the major AI-security frameworks

Pick a risk in your stack, whether it lives in the input, output, model, infrastructure, or agentic layer, and see which controls apply across OWASP, MITRE, and the major compliance frameworks. These are the same risks Nemesis tests for.

OWASP LLM Top 10MITRE ATLASMITRE ATT&CKOWASP APIOWASP WebNIST AI RMFISO 42001EU AI ActNIST 800-53SOC 2HIPAAHITRUSTNIST AI 600-1
Open the crosswalk →
Input
Prompts, RAG context, files
Output
Responses, downstream handling
Model
Model, training, embeddings
Infrastructure
API, keys, hosting, limits
Agentic
Tools, autonomy, actions

Model Intelligence phase, new in 2026

Before any attacks fire, Nemesis runs 8 lightweight probes to fingerprint the target model. This phase runs in about 30 seconds and produces a Model Intelligence Report that appears in your results and compliance report.

Identity detection
Detects Claude, GPT, Gemini, Llama, Mistral from responses
Capability scan
Vision, function calling, knowledge cutoff
Mismatch detection
Alerts when detected model differs from specified
Guardrail detection
Confirms whether safety controls are active

How Nemesis compares to other tools

vs Garak

Matches or exceeds on all LLM model-layer tests. Adds API security, injection probing, agentic chain attacks, and embedding leakage that Garak does not cover.

vs Promptfoo

Matches on OWASP LLM Top 10. Adds model identity fingerprinting, agentic chain, EchoLeak and Copilot CVE tests, and full injection probing.

vs Akto / PyRIT

Adds full LLM model-layer and agentic testing Akto lacks, plus API-layer tests PyRIT lacks. Includes LLM-as-judge scoring for accuracy.

Real-world incidents Nemesis now tests for

CVE-2025-32711 · CVSS 9.3
EchoLeak (Microsoft Copilot)

Hidden instructions in a shared document caused Copilot to return a user’s private recent emails when asked for a summary. No click, no download, just a question to an AI assistant. Module: Embedding & RAG Leakage.

CVE-2025-53773 · CVSS 9.6
GitHub Copilot source code injection

A markdown image tag hidden in a source code file caused Copilot to send sensitive data to an attacker-controlled URL. Over 10M developers in scope. Module: Embedding & RAG Leakage.

McKinsey breach · March 2026
SQL injection via AI chatbot

SQL injection delivered through an AI chatbot interface reached a backend database. $20 and 2 hours to full breach. Module: Injection Probing.

Multi-agent pipelines · 2025-2026
Cross-agent instruction propagation

Malicious instructions injected into agent A propagate to agent B in automated pipelines, bypassing per-agent safety checks. Module: Agentic Chain Attacks.

OWASP LLM Top 10 (2025)

All 10 categories covered across the attack modules. Click any category to expand.

Extended coverage

Additional modules beyond OWASP LLM Top 10: systemic bias, multimodal injection, conversation poisoning, firewall bypass, API security, injection attacks, toxicity, model identity, agentic pipelines, and embedding leakage. Click any module to expand.

25
Attack modules
173
Red team attacks
10/10
OWASP LLM covered
9
Compliance frameworks
Free · No account

Run a free passive scan

13 checks on your API surface, run instantly from a URL. No account, nothing stored.

Run free scan →
Managed · Full assessment

Request a Red Team assessment

173 adversarial attacks across 25 modules with LLM-as-judge scoring, run by our team, delivered as a compliance report across 9 frameworks.

Request assessment →