# Awesome AI Security Tools [![Awesome](https://awesome.re/badge.svg)](https://awesome.re) > A curated list of **public-source, research, and commercial** tools for AI security and AI-assisted cybersecurity โ€” autotriage, agent security, AI/ML supply chain, pentest agents, AI SAST, LLM-driven fuzzing, threat intelligence, SOC/SIEM triage, reverse engineering, LLM red-teaming, and more. [![License: CC0-1.0](https://img.shields.io/badge/license-CC0--1.0-lightgrey.svg)](https://creativecommons.org/publicdomain/zero/1.0/) [![PRs Welcome](https://img.shields.io/badge/PRs-welcome-brightgreen.svg)](#contributing) **Type legend:** ๐ŸŸข public source / open-source ยท ๐Ÿ”ฌ research (paper / benchmark / dataset / framework) ยท ๐ŸŸ  commercial with open components ยท โš ๏ธ restrictive, non-commercial, or unclear/no license โ€” check before use. GitHub-hosted entries show static **โ˜… stars** and **last-commit** snapshots; refresh them with `python3 scripts/update_github_metrics.py` before release. Most recently refreshed entry: 2026-10-08. Hugging Face model entries show license, access, and artifact metadata. Ordering within a section favors flagship and actively maintained projects. --- ## Contents - [Autotriage of Security Findings](#autotriage-of-security-findings) - [AI Agent & Coding-Agent Security](#ai-agent--coding-agent-security) - [Scanners & Auditors](#scanners--auditors) - [Frameworks, Rule Standards & Benchmarks](#frameworks-rule-standards--benchmarks) - [Runtime Protection & Enforcement](#runtime-protection--enforcement) - [AI/ML Supply Chain & Model Security](#aiml-supply-chain--model-security) - [Pentest & Red-Team Agents](#pentest--red-team-agents) - [AI-Powered Recon & Narrow ML Tools](#ai-powered-recon--narrow-ml-tools) - [Subdomain & DNS Prediction](#subdomain--dns-prediction) - [Recon Screenshot Triage](#recon-screenshot-triage) - [Software / Tech Fingerprinting](#software--tech-fingerprinting) - [AI-Assisted Fuzzing](#ai-assisted-fuzzing) - [Password / Credential ML](#password--credential-ml) - [Phishing Detection (Visual / URL)](#phishing-detection-visual--url) - [AI/ML-Assisted Detection Rules & Engines](#aiml-assisted-detection-rules--engines) - [Defensive Trained-Model Detectors](#defensive-trained-model-detectors) - [AI-Powered SAST & Secure Code Review](#ai-powered-sast--secure-code-review) - [AI-Powered Threat Modeling](#ai-powered-threat-modeling) - [LLM-Driven Fuzzing](#llm-driven-fuzzing) - [Harness / target generation](#harness--target-generation) - [Fuzzing the LLM](#fuzzing-the-llm) - [Threat Intelligence](#threat-intelligence) - [Log Analysis / SIEM / SOC Triage](#log-analysis--siem--soc-triage) - [Reverse Engineering](#reverse-engineering) - [LLM Red-Teaming & Guardrails](#llm-red-teaming--guardrails) - [Scanners, Evals & Guardrails](#scanners-evals--guardrails) - [Prompt-Injection Classifier Models](#prompt-injection-classifier-models) - [Specialty Security LLMs](#specialty-security-llms) - [LLM Honeypots & Deception](#llm-honeypots--deception) - [CTF / Exploit / Bug-Bounty Agents & Benchmarks](#ctf--exploit--bug-bounty-agents--benchmarks) - [Cloud / IaC / DFIR / OSINT / Phishing](#cloud--iac--dfir--osint--phishing) - [Related Awesome Lists](#related-awesome-lists) - [Contributing](#contributing) - [Contact](#contact) - [License](#license) --- ## Autotriage of Security Findings AI/LLM tools that triage, deduplicate, prioritize, or validate the output of scanners and finding sources. - **[Metis](https://github.com/arm/metis)** ๐ŸŸข โ€” Security-review framework that triages external SARIF findings using source-navigation evidence, optional C/C++ code graphs, and structured LLM decisions. *(Arm)* **Caveat:** hosted providers receive selected source and evidence and require credentials; local providers need separately provisioned weights, and indexing may use a separate embedding provider. Debug traces can contain sensitive content. Model verdicts are not proof of exploitability or grounds for automatic suppression. *(โ˜… 872 ยท updated 2026-09-30)* - **[nuclei-autotriage](https://github.com/cyberok-org/nuclei-autotriage)** ๐ŸŸขโš ๏ธ โ€” Two-stage LLM triage (falsifier + red-team pass) of Nuclei JSONL findings via OpenAI-compatible endpoints (vLLM/Ollama). *(CyberOK)* โ€” **note:** restrictive personal/non-commercial EULA, not a permissive OSS license. *(โ˜… 1 ยท updated 2026-05-25)* - **Related:** [agent-audit](https://github.com/scadastrangelove/agent-audit) ยท [asamm](https://github.com/scadastrangelove/asamm) - **[seclab-taskflow-agent](https://github.com/GitHubSecurityLab/seclab-taskflow-agent)** ๐ŸŸข โ€” YAML-driven taskflow agent framework for triaging CodeQL/SAST alerts and filtering false positives. *(GitHub Security Lab)* *(โ˜… 221 ยท updated 2026-08-17)* - **Related:** [SigmaOptimizer](https://github.com/YusukeJustinNakajima/SigmaOptimizer) - **[honeyslop](https://github.com/gadievron/honeyslop)** ๐ŸŸข โ€” Code-canary decoys to triage AI-hallucinated ("slop") vulnerability reports flooding bug-bounty programs. *(โ˜… 97 ยท updated 2026-05-20)* - **[nano-analyzer](https://github.com/weareaisle/nano-analyzer)** ๐ŸŸข๐Ÿ”ฌ โ€” Minimal three-stage LLM pipeline (context โ†’ scan โ†’ skeptical triage) for zero-day discovery in C/C++. *(AISLE)* *(โ˜… 307 ยท updated 2026-04-14)* - **[SigmaOptimizer](https://github.com/YusukeJustinNakajima/SigmaOptimizer)** ๐ŸŸข โ€” Generates, tests, and refines Sigma rules from real logs with false-positive checking. *(โ˜… 11 ยท updated 2025-08-01)* - **Related:** [soctalk](https://github.com/soctalk/soctalk) ยท [seclab-taskflow-agent](https://github.com/GitHubSecurityLab/seclab-taskflow-agent) - **[ai-soc-triage-assistant](https://github.com/pranavibunny/ai-soc-triage-assistant)** ๐ŸŸขโš ๏ธ โ€” SOC alert triage assistant with prompt-injection guardrails, output validation, and MITRE ATT&CK mapping. *(โ˜… 0 ยท updated 2026-02-23)* > See also: OpenAI's *Aardvark* research preview โ€” public references exist, but there is no standalone installable repository to badge here. --- ## AI Agent & Coding-Agent Security Securing the AI agents themselves โ€” auditing coding agents (Claude Code, Codex, OpenClaw), scanning skills / plugins / MCP manifests, and governance for agentic development. A fast-moving 2026 category, split below by role. ### Scanners & Auditors - **[Geiger](https://github.com/Atomburstofficial/geiger)** ๐ŸŸข โ€” Local CLI that inventories AI-agent installations, MCP configurations, hooks, and extensions, and reports credential-shaped values and configuration drift. **Caveat:** configuration inventory, not runtime monitoring or enforcement; exposure labels are heuristics, and coverage depends on known locations and readable files. Detector errors can leave gaps even when the strict-mode exit code is clean; explicit report options write local files. *(โ˜… 158 ยท updated 2026-09-21)* - **Sources:** [Reviewed scan engine](https://github.com/Atomburstofficial/geiger/blob/f8b8edbdbd18c3714022a2f5f84707f78a03006f/src/engine.js) - **[Guardana](https://github.com/guardana/guardana)** ๐ŸŸข โ€” Verifies AI artifacts, live-system checks, and recorded agent traces, including evidence of unapproved side effects, outside the request path. **Caveat:** beta; trace checks depend on complete, trustworthy evidence and do not enforce tool calls. Built-in artifact checks are local, while active probes contact configured targets and optional collectors receive reports. Not a general SAST scanner or compliance certification. *(โ˜… 145 ยท updated 2026-10-05)* - **[agent-audit](https://github.com/scadastrangelove/agent-audit)** ๐ŸŸข โ€” Forensic auditor for local AI coding agents (Claude Code, Codex CLI, OpenClaw) **and** project-surface scanner for repos shipping skills, plugins, and MCP manifests; 296 bundled rules across native + imported detector families, with optional LLM cross-verification. *(CyberOK / S. Gordeychik)* *(โ˜… 15 ยท updated 2026-07-15)* - **Sources:** [asamm](https://github.com/scadastrangelove/asamm) ยท [ATR โ€“ Agent Threat Rules](https://github.com/Agent-Threat-Rule/agent-threat-rules) ยท [aguara](https://github.com/garagon/aguara) ยท [Cisco AI Defense โ€“ skill-scanner](https://github.com/cisco-ai-defense/skill-scanner) - **Related:** [asamm](https://github.com/scadastrangelove/asamm) ยท [aguara](https://github.com/garagon/aguara) ยท [agentguard](https://github.com/GoPlusSecurity/agentguard) ยท [agentic-radar](https://github.com/splx-ai/agentic-radar) ยท [nuclei-autotriage](https://github.com/cyberok-org/nuclei-autotriage) - **[AI-Infra-Guard](https://github.com/Tencent/AI-Infra-Guard)** ๐ŸŸข โ€” Full-stack AI red-teaming platform covering OpenClaw security scan, agent scan, skills scan, MCP scan, AI-infra vulnerability scan, and LLM jailbreak evaluation. *(Tencent Zhuque Lab)* *(โ˜… 6,789 ยท updated 2026-10-08)* - **Related:** [agent-audit](https://github.com/scadastrangelove/agent-audit) ยท [aguara](https://github.com/garagon/aguara) ยท [Cisco AI Defense โ€“ skill-scanner](https://github.com/cisco-ai-defense/skill-scanner) ยท [Cisco AI Defense โ€“ mcp-scanner](https://github.com/cisco-ai-defense/mcp-scanner) - **[SkillSpector](https://github.com/NVIDIA/SkillSpector)** ๐ŸŸข โ€” Security scanner for AI-agent skills used by Claude Code, Codex CLI, Gemini CLI, and similar ecosystems; combines static analysis, AST/YARA/taint checks, optional LLM semantic review, MCP least-privilege/tool-poisoning checks, risk scoring, and SARIF/JSON/Markdown output. *(NVIDIA)* *(โ˜… 14,696 ยท updated 2026-08-15)* - **Related:** [Cisco AI Defense โ€“ skill-scanner](https://github.com/cisco-ai-defense/skill-scanner) ยท [skilltotal](https://github.com/pezhik/skilltotal) ยท [Snyk Agent Scan](https://github.com/snyk/agent-scan) ยท [Cisco AI Defense โ€“ mcp-scanner](https://github.com/cisco-ai-defense/mcp-scanner) - **[Ramparts](https://github.com/highflame-ai/ramparts)** ๐ŸŸข โ€” Rust scanner for MCP servers and agent-skill bundles with YARA rules, optional LLM analysis, OSV/CVE lookups, OWASP MCP Top 10 mapping, and SARIF/JSON/Markdown reports. *(โ˜… 96 ยท updated 2026-08-07)* - **Related:** [SkillSpector](https://github.com/NVIDIA/SkillSpector) ยท [Cisco AI Defense โ€“ skill-scanner](https://github.com/cisco-ai-defense/skill-scanner) ยท [Cisco AI Defense โ€“ mcp-scanner](https://github.com/cisco-ai-defense/mcp-scanner) - **[lintlang](https://github.com/hermes-labs-ai/lintlang)** ๐ŸŸข โ€” Local deterministic linter for AI-agent instructions, tool definitions, and system prompts that flags ambiguous descriptions, missing limits, conflicting directives, and schema gaps, with CLI and SARIF output. โ€” **note:** heuristic static instruction/config review only; not runtime enforcement, prompt-injection detection, semantic-correctness proof, or a security guarantee. *(โ˜… 134 ยท updated 2026-10-02)* - **Related:** [little-canary](https://github.com/hermes-labs-ai/little-canary) ยท [rule-audit](https://github.com/roli-lpci/rule-audit) - **[mcp-armor](https://github.com/aira-security/mcp-armor)** ๐ŸŸข โ€” Local MCP security scanner with auto-discovery for agentic IDE configs, tool/resource/prompt inventory, prompt-injection checks, rug-pull and tool-poisoning detection, baseline drift monitoring, and JSON/Markdown reports. *(Aira Security)* *(โ˜… 120 ยท updated 2026-03-27)* - **Related:** [SkillSpector](https://github.com/NVIDIA/SkillSpector) ยท [Ramparts](https://github.com/highflame-ai/ramparts) ยท [Cisco AI Defense โ€“ mcp-scanner](https://github.com/cisco-ai-defense/mcp-scanner) - **[aguara](https://github.com/garagon/aguara)** ๐ŸŸข โ€” Single-binary static scanner (Go, no LLM) for AI-agent skills and MCP servers; multi-layer engine (pattern + NLP + taint tracking + rug-pull detection). Companion **[aguara-mcp](https://github.com/garagon/mcp-aguara)** exposes scanning as an MCP tool. *(โ˜… 86 ยท updated 2026-08-12)* - **Related:** [aguara-mcp](https://github.com/garagon/mcp-aguara) ยท [agent-audit](https://github.com/scadastrangelove/agent-audit) ยท [Snyk Agent Scan](https://github.com/snyk/agent-scan) ยท [Cisco AI Defense โ€“ skill-scanner](https://github.com/cisco-ai-defense/skill-scanner) - **[agent-scan](https://github.com/snyk/agent-scan)** ๐ŸŸข โ€” Security scanner for AI agents, MCP servers, and agent skills; the successor path for the original Invariant Labs mcp-scan work. *(Snyk)* *(โ˜… 2,913 ยท updated 2026-08-13)* - **Related:** [aguara](https://github.com/garagon/aguara) ยท [Cisco AI Defense โ€“ mcp-scanner](https://github.com/cisco-ai-defense/mcp-scanner) ยท [Cisco AI Defense โ€“ skill-scanner](https://github.com/cisco-ai-defense/skill-scanner) - **[inkog](https://github.com/inkog-io/inkog)** ๐ŸŸ  โ€” Commercial-backed static security scanner for AI agents across LangChain, LangGraph, CrewAI, AutoGen, and no-code workflows; Apache-2.0 CLI with proprietary deep-scan engine. *(Inkog)* *(โ˜… 28 ยท updated 2026-06-07)* - **Related:** [Snyk Agent Scan](https://github.com/snyk/agent-scan) ยท [agentic-radar](https://github.com/splx-ai/agentic-radar) - **[AgentShield](https://github.com/affaan-m/agentshield)** ๐ŸŸข โ€” Security scanner for AI-agent configurations, MCP servers, hooks, and tool permissions with CLI, GitHub Action, and app workflows. *(โ˜… 1,070 ยท updated 2026-07-22)* - **Related:** [agent-audit](https://github.com/scadastrangelove/agent-audit) ยท [Snyk Agent Scan](https://github.com/snyk/agent-scan) - **[repo-forensics](https://github.com/alexgreensh/repo-forensics)** ๐ŸŸขโš ๏ธ โ€” Offline scanner for AI-agent repos, skills, plugins, and MCP servers; license is PolyForm Noncommercial. *(โ˜… 155 ยท updated 2026-08-08)* - **Related:** [agent-audit](https://github.com/scadastrangelove/agent-audit) ยท [aguara](https://github.com/garagon/aguara) - **[skill-scanner](https://github.com/cisco-ai-defense/skill-scanner)** ๐ŸŸ  โ€” Scanner for agent skills combining YAML + YARA patterns, LLM-as-a-judge, and behavioral dataflow analysis (Codex / Cursor skill formats). *(Cisco AI Defense)* *(โ˜… 2,440 ยท updated 2026-08-04)* - **Related:** [defenseclaw](https://github.com/cisco-ai-defense/defenseclaw) ยท [aguara](https://github.com/garagon/aguara) ยท [Cisco AI Defense โ€“ mcp-scanner](https://github.com/cisco-ai-defense/mcp-scanner) - **[mcp-scanner](https://github.com/cisco-ai-defense/mcp-scanner)** ๐ŸŸขโš ๏ธ โ€” Scanner for MCP servers and agentic tool surfaces, covering tools, prompts, resources, package risk, malware indicators, and deployment readiness. *(Cisco AI Defense)* *(โ˜… 1,037 ยท updated 2026-08-07)* - **Related:** [Cisco AI Defense โ€“ skill-scanner](https://github.com/cisco-ai-defense/skill-scanner) ยท [Snyk Agent Scan](https://github.com/snyk/agent-scan) ยท [aguara](https://github.com/garagon/aguara) - **[mcp-guardian](https://github.com/alexandriashai/mcp-guardian)** ๐ŸŸข โ€” JS/TS library and CLI for detecting prompt injection in MCP tool descriptions and pinning tool definitions. *(โ˜… 6 ยท updated 2026-07-29)* - **Related:** [Cisco AI Defense โ€“ mcp-scanner](https://github.com/cisco-ai-defense/mcp-scanner) ยท [Snyk Agent Scan](https://github.com/snyk/agent-scan) - **[MCP Observatory](https://github.com/KryptosAI/mcp-observatory)** ๐ŸŸข๐ŸŸ  โ€” CI-native MCP-server testing tool for schema drift, safe attack simulation, record/replay verification, health scoring, and SARIF evidence before agents depend on a server. *(KryptosAI)* โ€” **note:** the local evidence engine is open source; hosted telemetry intelligence, fleet workflows, and commercial ranking remain outside the package. *(โ˜… 176 ยท updated 2026-08-15)* - **Related:** [Cisco AI Defense โ€“ mcp-scanner](https://github.com/cisco-ai-defense/mcp-scanner) ยท [skilltotal](https://github.com/pezhik/skilltotal) - **[agentic-radar](https://github.com/splx-ai/agentic-radar)** ๐ŸŸ  โ€” CLI security scanner for agentic workflows (LangGraph, CrewAI, n8n, etc.) โ€” maps tools/data flows and flags risks. *(SplxAI)* *(โ˜… 1,036 ยท updated 2025-11-27)* - **[skilltotal](https://github.com/pezhik/skilltotal)** ๐ŸŸข โ€” Offline deterministic static scanner (regex + AST, no LLM, no account) for AI components โ€” agent skills/plugins, MCP servers, npm & PyPI packages, and git repos; flags supply-chain risk, dangerous capabilities, prompt-injection surfaces, MCP tool poisoning/shadowing, and data-exfiltration paths, maps to the OWASP Agentic Skills Top 10, and emits JSON + SARIF 2.1.0. *(skilltotal.ai)* *(โ˜… 1 ยท updated 2026-08-17)* - **Related:** [aguara](https://github.com/garagon/aguara) ยท [Snyk Agent Scan](https://github.com/snyk/agent-scan) ยท [Cisco AI Defense โ€“ skill-scanner](https://github.com/cisco-ai-defense/skill-scanner) ยท [Cisco AI Defense โ€“ mcp-scanner](https://github.com/cisco-ai-defense/mcp-scanner) - **[Sunglasses](https://github.com/sunglasses-dev/sunglasses)** ๐ŸŸข โ€” Local input/content scanner for AI agents that checks prompts, files, media metadata, skills, and tool descriptions against pattern and mechanism-based prompt-injection, exfiltration, command-injection, and agent-threat rules. โ€” **note:** early-stage project; published precision/recall benchmark is self-reported by the project. *(โ˜… 4 ยท updated 2026-08-14)* - **Related:** [skilltotal](https://github.com/pezhik/skilltotal) ยท [Armorer Guard](https://github.com/ArmorerLabs/Armorer-Guard) - **[trentclaw](https://github.com/trnt-ai/trent-openclaw-security-assessment)** ๐ŸŸ  โ€” Client-side security auditor for OpenClaw deployments: applies pattern-based secret redaction locally, then uploads config/skill metadata and confirm-gated skill archives to Trent AI's API, which identifies misconfigurations, risky skills (prompt injection, permission escalation, data exfiltration), and chained attack paths. *(Trent AI)* โ€” **note:** core detection runs server-side via the Trent AI API (requires an API key); the Apache-2.0 client collects OpenClaw config/skill metadata, applies pattern-based secret redaction locally, and uploads skill archives only after an explicit in-terminal confirmation. *(โ˜… 23 ยท updated 2026-07-27)* - **[A2A Security Scanner](https://github.com/cisco-ai-defense/a2a-scanner)** ๐ŸŸข โ€” CLI and PyPI scanner for Agent-to-Agent (A2A) agent cards, source code, registries, and live endpoints using specification validation, YARA rules, heuristics, endpoint testing, and an optional LLM analyzer. *(Cisco AI Defense)* *(โ˜… 163 ยท updated 2026-04-16)* - **Related:** [Cisco AI Defense โ€“ mcp-scanner](https://github.com/cisco-ai-defense/mcp-scanner) ยท [Agent Threat Rules](https://github.com/Agent-Threat-Rule/agent-threat-rules) - **[Ship Safe](https://github.com/asamassekou10/ship-safe)** ๐ŸŸข๐ŸŸ  โ€” Local security CLI for application code, AI-agent and MCP configuration, secrets, dependencies, CI, and cloud/IaC surfaces, with deterministic checks, optional AI-assisted analysis, and SARIF output. โ€” **note:** the MIT CLI works without an account for its core checks; optional AI modes can send selected context to the configured provider, while hosted dashboards and organization workflows are separate commercial features. Coverage and benchmark figures are project-reported. *(โ˜… 827 ยท updated 2026-08-24)* - **Related:** [SkillTotal](https://github.com/pezhik/skilltotal) ยท [agent-audit](https://github.com/scadastrangelove/agent-audit) - **[AgentSeal](https://github.com/getagentseal/agentseal)** ๐ŸŸ โš ๏ธ โ€” Agent-security CLI and runtime library for scanning MCP servers, skills, prompts, and configuration, plus local guard and monitoring workflows with optional model-assisted red teaming. โ€” **note:** source-available under FSL-1.1-Apache-2.0 rather than an OSI-approved open-source license; offline guard and MCP/skill scanning are available locally, while red-team modes require local Ollama or configured provider credentials. Catalog and coverage figures are project-reported. *(โ˜… 343 ยท updated 2026-06-11)* - **Related:** [agent-scan](https://github.com/snyk/agent-scan) ยท [AgentLock](https://github.com/webpro255/agentlock) - **[SlowMist Agent Security](https://github.com/slowmist/slowmist-agent-security)** ๐ŸŸข โ€” Security-review skill and workflow for auditing agent skills, MCP servers, repositories, URLs, and documents before installation or use. *(SlowMist)* โ€” **note:** Markdown-based workflow skill executed by a compatible host agent, not a deterministic standalone scanner; findings depend on the selected model and its available tools. *(โ˜… 502 ยท updated 2026-04-17)* - **Related:** [MCP-Security-Checklist](https://github.com/slowmist/MCP-Security-Checklist) ยท [sast-skills](https://github.com/utkusen/sast-skills) - **[Sandbox Probe](https://github.com/controlplaneio/sandbox-probe)** ๐ŸŸข โ€” Static Go probe that measures the effective filesystem, network, process, credential, and runtime capabilities exposed inside an AI-agent sandbox, then compares sandbox and host baselines. *(ControlPlane)* โ€” **note:** boundary-measurement auditor, not an enforcement layer; some integration scripts can ask real agents to execute the probe, while deterministic model-free stubs are available for CI. *(โ˜… 25 ยท updated 2026-08-25)* - **Related:** [Sandlock](https://github.com/multikernel/sandlock) ยท [AIO Sandbox](https://github.com/agent-infra/sandbox) ### Frameworks, Rule Standards & Benchmarks - **[RepoGuardBench](https://github.com/DaoyuanLi2816/RepoGuardBench)** ๐ŸŸข๐Ÿ”ฌ โ€” Repository prompt-injection benchmark with coding tasks, attack carriers, defenses, and separate scoring of attack attempts, observed outcomes, and task utility. **Caveat:** its subprocess workspace is not an OS sandbox: allowing Python or pytest does not prevent host-file access or networking. Use a separate disposable environment without secrets; marker-based outcomes and model behavior are not proof of production isolation. *(โ˜… 70 ยท updated 2026-10-03)* - **[Project CodeGuard](https://github.com/cosai-oasis/project-codeguard)** ๐ŸŸขโš ๏ธ โ€” Model-agnostic secure-coding rules and skills framework with translators for popular coding agents, validators, release artifacts, and an MCP server for centrally distributing the rules. *(CoSAI / OASIS)* โ€” **note:** framework and ruleset, not a deterministic scanner or runtime enforcement boundary; repository content uses CC BY 4.0 rather than a conventional software license. *(โ˜… 338 ยท updated 2026-09-18)* - **Related:** [OWASP review skill for Claude Code (model-dependent; evaluations may require paid API)](https://github.com/agamm/claude-code-owasp) ยท [Cisco Foundry evaluation specification (CC BY 4.0; no implementation code)](https://github.com/CiscoDevNet/foundry-security-spec) - **[asamm](https://github.com/scadastrangelove/asamm)** ๐Ÿ”ฌ โ€” *Agentic SAMM* โ€” an OWASP SAMM extension for AI-driven development: an entry-point-based threat taxonomy plus 17 controls across 5 SAMM functions (Governance, Design, Implementation, Verification, Operations) with L1/L2/L3 maturity. License: CC BY-SA 4.0. *(CyberOK / S. Gordeychik)* *(โ˜… 17 ยท updated 2026-07-26)* - **Sources:** [OWASP SAMM](https://owaspsamm.org/) ยท [NIST AI RMF](https://www.nist.gov/itl/ai-risk-management-framework) ยท [NCSC Secure AI Guidelines](https://www.ncsc.gov.uk/collection/guidelines-secure-ai-system-development) ยท [MCP Security Best Practices](https://modelcontextprotocol.io/) - **Related:** [agent-audit](https://github.com/scadastrangelove/agent-audit) ยท [AI governance crosswalk (expert review required; not certification)](https://github.com/jeanmalaquias/ai-governance-mapping) ยท [AATMF threat-modeling reference (CC BY-SA 4.0; not an automated verifier)](https://github.com/SnailSploit/AATMF-Adversarial-AI-Threat-Modeling-Framework) - **[agent-threat-rules (ATR)](https://github.com/Agent-Threat-Rule/agent-threat-rules)** ๐ŸŸข โ€” Open, versioned, machine-readable detection rules for AI-agent threats (prompt injection, tool poisoning, MCP attacks, and skill compromise) โ€” "Sigma for agents"; 768 rules across 10 categories with integrations for Microsoft AGT, Cisco AI Defense, MISP, OWASP, FINOS, and SigmaHQ. *(โ˜… 371 ยท updated 2026-08-17)* - **Related:** [agent-audit](https://github.com/scadastrangelove/agent-audit) ยท [aguara](https://github.com/garagon/aguara) ยท [Sigma-AI detection content (temporal/similarity extensions require a compatible engine)](https://github.com/agentshield-ai/sigma-ai) - **[Agent Governance Toolkit](https://github.com/microsoft/agent-governance-toolkit)** ๐ŸŸข โ€” Multi-language toolkit for policy-enforced agent tool calls and audit records, with optional identity, MCP-gateway, sandboxing, reliability, and compliance components. *(Microsoft)* โ€” **note:** official public preview; APIs and deployment patterns may change before general availability. *(โ˜… 5,962 ยท updated 2026-08-12)* - **Related:** [ATR โ€“ Agent Threat Rules](https://github.com/Agent-Threat-Rule/agent-threat-rules) ยท [ToolHive](https://github.com/stacklok/toolhive) - **[MCP-Security-Checklist](https://github.com/slowmist/MCP-Security-Checklist)** ๐ŸŸข โ€” Security checklist for MCP clients, servers, multi-MCP deployments, lifecycle controls, authz/authn, isolation, and crypto-specific MCP integrations. *(SlowMist)* *(โ˜… 835 ยท updated 2025-04-28)* - **Related:** [Cisco AI Defense โ€“ mcp-scanner](https://github.com/cisco-ai-defense/mcp-scanner) ยท [ATR โ€“ Agent Threat Rules](https://github.com/Agent-Threat-Rule/agent-threat-rules) ยท [MCP security hardening and operations guidance (CC0; not deployment certification)](https://github.com/ModelContextProtocol-Security/modelcontextprotocol-security.io) - **[Anthropic-Cybersecurity-Skills](https://github.com/mukul975/Anthropic-Cybersecurity-Skills)** ๐ŸŸข โ€” Large community cybersecurity skill library for AI agents, mapped to MITRE ATT&CK, NIST CSF, MITRE ATLAS, D3FEND, and NIST AI RMF. โ€” **note:** independent community project, not affiliated with Anthropic. *(โ˜… 28,060 ยท updated 2026-08-08)* - **Related:** [sast-skills](https://github.com/utkusen/sast-skills) ยท [Cisco AI Defense โ€“ skill-scanner](https://github.com/cisco-ai-defense/skill-scanner) - **[Claude-BugHunter](https://github.com/elementalsouls/Claude-BugHunter)** ๐ŸŸข โ€” Claude Code / agent-skill bundle for authorized bug hunting and external red-team workflows across web, API, identity, cloud, recon, reporting, Burp MCP, slash commands, and the cbh CLI. โ€” **note:** skill bundle and workflow knowledge base, not a standalone scanner. *(โ˜… 3,630 ยท updated 2026-08-17)* - **Related:** [sast-skills](https://github.com/utkusen/sast-skills) ยท [Anthropic-Cybersecurity-Skills](https://github.com/mukul975/Anthropic-Cybersecurity-Skills) - **[AgentDojo](https://github.com/ethz-spylab/agentdojo)** ๐ŸŸข๐Ÿ”ฌ โ€” Benchmark environment for prompt-injection attacks and defenses in tool-using LLM agents. *(โ˜… 753 ยท updated 2026-06-02)* - **Related:** [agent-audit](https://github.com/scadastrangelove/agent-audit) ยท [ATR โ€“ Agent Threat Rules](https://github.com/Agent-Threat-Rule/agent-threat-rules) - **[Agent3Sigma-Canary](https://github.com/antgroup/Agent3Sigma-Canary)** ๐ŸŸข๐Ÿ”ฌ โ€” Sandboxed research framework for evaluating AI-agent security over complete execution trajectories, covering direct/indirect injection, skill and memory poisoning, and practical risk outcomes. *(Ant Group)* โ€” **note:** research framework that requires Docker plus target and auxiliary LLM configuration; use only in controlled environments. *(โ˜… 35 ยท updated 2026-08-09)* - **Related:** [AgentDojo](https://github.com/ethz-spylab/agentdojo) ยท [HarmBench](https://github.com/centerforaisafety/HarmBench) - **[Agent Security Bench (ASB)](https://github.com/agiresearch/ASB)** ๐ŸŸข๐Ÿ”ฌ โ€” Official ICLR 2025 benchmark for evaluating attacks and defenses in LLM-based agents across ten scenarios, including direct and indirect prompt injection, memory poisoning, and defensive strategies. โ€” **note:** research benchmark rather than a production control; reproducing evaluations requires configured target and evaluator models. *(โ˜… 285 ยท updated 2026-04-16)* - **Related:** [AgentDojo](https://github.com/ethz-spylab/agentdojo) ยท [Agent3Sigma-Canary](https://github.com/antgroup/Agent3Sigma-Canary) - **[Skill-Inject](https://github.com/aisa-group/skill-inject)** ๐ŸŸข๐Ÿ”ฌ โ€” Benchmark for measuring prompt-injection vulnerabilities carried by agent skill files across Claude Code, Codex CLI, and Gemini CLI under multiple safety-policy conditions. โ€” **note:** benchmark artifact that executes controlled malicious skill scenarios; run only in an isolated test environment with synthetic data and accounts. *(โ˜… 91 ยท updated 2026-07-01)* - **Related:** [SkillSpector](https://github.com/NVIDIA/SkillSpector) ยท [AgentDojo](https://github.com/ethz-spylab/agentdojo) - **[AI Security Verification Standard (AISVS)](https://github.com/OWASP/AISVS)** ๐Ÿ”ฌโš ๏ธ โ€” Stable verification standard defining testable security requirements for AI applications across model lifecycle, supply chain, data handling, agentic systems, and MCP integrations. *(OWASP)* โ€” **note:** security standard and checklist, not an executable scanner; share-alike terms apply to adapted material. *(โ˜… 429 ยท updated 2026-07-30)* - **Related:** [asamm](https://github.com/scadastrangelove/asamm) ยท [Agent Threat Rules](https://github.com/Agent-Threat-Rule/agent-threat-rules) ยท [GenAI data-security guidance (CC BY-SA 4.0; separate scanner not included)](https://github.com/GenAI-Security-Project/GenAI-Data-Security-Initiative) ยท [OWASP LLMSVS (distinct documentation project; CC BY-SA 4.0)](https://github.com/OWASP/www-project-llm-verification-standard) - **[OWASP Agent Security Regression Harness](https://github.com/OWASP/Agent-Security-Regression-Harness)** ๐ŸŸข โ€” Vendor-neutral harness for running repeatable agent and MCP abuse scenarios, evaluating policy assertions over execution traces, and emitting machine-readable regression results for local development and CI. *(OWASP)* โ€” **note:** early OWASP Incubator project; it is a regression harness for known abuse cases, not a scanner, leaderboard, general benchmark, or guarantee of agent security. *(โ˜… 49 ยท updated 2026-07-27)* - **Related:** [Agent Security Bench (ASB)](https://github.com/agiresearch/ASB) ยท [AgentDojo](https://github.com/ethz-spylab/agentdojo) ### Runtime Protection & Enforcement - **[OpenAPPA](https://github.com/archestra-ai/OpenAPPA)** ๐ŸŸข โ€” Rust information-flow policy engine and agent-tool runtime that tracks trajectory labels and effects and binds approvals to actions using recorded events. *(Archestra)* **Caveat:** preview/RFC with evolving interfaces. Policy evaluation is deterministic for fixed policy and recorded evidence, not for whole-agent behavior; protection requires correct tool contracts, trusted labels, and complete mediation. Configured annotation authorities may use LLMs, and project-run benchmarks do not establish universal prompt-injection or exfiltration prevention. *(โ˜… 1,427 ยท updated 2026-10-05)* - **Sources:** [Reviewed policy checks](https://github.com/archestra-ai/OpenAPPA/blob/85465feb8214f662dd21d40df2c01964e5b099fa/appa-engine/src/check.rs) ยท [Recorded-state projection](https://github.com/archestra-ai/OpenAPPA/blob/85465feb8214f662dd21d40df2c01964e5b099fa/appa-engine/src/projection.rs) - **[Vetto](https://github.com/shleder/vetto)** ๐ŸŸข โ€” Rust command wrapper that applies native OS containment to coding-agent subprocesses through Linux Landlock/namespaces and platform-specific sandbox backends. **Caveat:** guarantees vary by OS and selected tier: Linux FULL requires kernel support, while FS-ONLY lacks mount/PID/network namespaces and has descendant-cleanup limits. Allowed paths, environment values, and network destinations remain agent capabilities; this is not semantic prompt-injection protection. Optional telemetry is opt-in, and release binaries and escape resistance were not independently verified. *(โ˜… 45 ยท updated 2026-10-05)* - **Sources:** [Reviewed threat model](https://github.com/shleder/vetto/blob/0335a49d5b90507d4ce90f71e82a133b4d946d8a/docs/threat-model.md) ยท [Telemetry configuration](https://github.com/shleder/vetto/blob/0335a49d5b90507d4ce90f71e82a133b4d946d8a/docs/telemetry.md) - **[Spring MCP Security](https://github.com/spring-ai-community/mcp-security)** ๐ŸŸข โ€” Spring security components for MCP authentication, JWT audience validation, principal-bound sessions, and transport Origin/Host checks. *(Spring AI Community)* **Caveat:** match library and Spring AI versions and the supported transport/stack combination. Operators must configure issuers, scopes, and application authorization; these controls do not detect prompt injection. *(โ˜… 117 ยท updated 2026-09-20)* - **Related:** [.NET MCP OAuth/Entra/DPoP samples (configured identity provider and compatible clients required)](https://github.com/damienbod/McpSecurity) - **[yoloAI](https://github.com/kstenerud/yoloai)** ๐ŸŸข โ€” Coding-agent sandbox runner with multiple backends, credential-brokering integrations, and a copy/diff/apply workflow for reviewing workspace changes. **Caveat:** public beta; isolation depends on the backend and configuration. Select network isolation explicitly, avoid privileged modes, and review mounts; agents without broker support receive mounted credentials, and hosted agents still send context to their providers. *(โ˜… 216 ยท updated 2026-08-21)* - **[Jailoc](https://github.com/seznam/jailoc)** ๐ŸŸข โ€” Container launcher for OpenCode with configurable workspace mounts, network/DNS policy, resource limits, and secret references. *(Seznam)* **Caveat:** Docker and the agent image remain trusted dependencies; extra mounts, SSH-agent forwarding, optional Docker access, and supplied credentials expand the boundary. Container configuration is not a guarantee against escape or data disclosure to model providers. *(โ˜… 32 ยท updated 2026-09-24)* - **[Agent Sandbox (agentbox)](https://github.com/mattolson/agent-sandbox)** ๐ŸŸข โ€” Container-based coding-agent environment with proxy-mediated egress policies and server-side credential injection. **Caveat:** the reviewed firewall permits the detected Docker host subnet, not only the proxy, and has no IPv6 rules; proxy enforcement is inactive outside enforce mode. Review IPv6, neighboring services, writable mounts, and the IDE/control plane rather than assuming complete network or host isolation. *(โ˜… 208 ยท updated 2026-09-26)* - **[Crust](https://github.com/BakeLens/crust)** ๐ŸŸขโš ๏ธ โ€” Local Go gateway that extracts tool calls from supported model API responses and applies rules before forwarding them to an agent. **Caveat:** source-available under Elastic License 2.0, including managed-service restrictions, not a permissive OSS license. Only intercepted supported paths are covered; parse failures can pass the original body through. The sample forwards to OpenRouter, disables local telemetry, and leaves the storage encryption key empty; review provider disclosure and any enabled local records. *(โ˜… 440 ยท updated 2026-03-30)* - **[Agent Safehouse](https://github.com/eugene1g/agent-safehouse)** ๐ŸŸข โ€” macOS launcher that assembles Seatbelt sandbox-exec profiles to scope coding agents' filesystem and integration access. **Caveat:** macOS-only hardening, not a VM or prompt-injection detector. The default profile permits inbound and outbound networking; exfiltration prevention is explicitly out of scope. Granted workspaces and integrations remain accessible, and launched agents retain their own credential, billing, and network behavior. *(โ˜… 2,092 ยท updated 2026-09-30)* - **[OpenShell](https://github.com/NVIDIA/OpenShell)** ๐ŸŸข โ€” Policy-governed runtime for autonomous and coding agents with container or microVM-backed sandboxes, filesystem/process/network controls, endpoint-bound credential injection, and audit logs. *(NVIDIA)* โ€” **note:** pre-release runtime whose effective isolation depends on the selected compute driver and policy; Kubernetes and GPU paths are experimental. Anonymous operational telemetry is enabled by default but can be disabled at deployment time or compiled out. *(โ˜… 8,715 ยท updated 2026-09-21)* - **Sources:** [NVIDIA agent-stack security model](https://developer.nvidia.com/blog/where-security-fits-in-an-ai-agent-stack/) - **[Numbat](https://github.com/perplexityai/numbat)** ๐ŸŸข โ€” Endpoint-local visibility and detection for AI-agent activity across hooks, plugins, OTLP, and on-disk artifacts, with CEL rules, multi-step sequence detections, and forensic reconstruction. *(Perplexity AI)* โ€” **note:** monitoring is the default posture; blocking is opt-in, limited to supported synchronous pre-action hooks, and the shipped rules are monitor-only until operators explicitly configure enforcement. *(โ˜… 1,069 ยท updated 2026-09-15)* - **[agentsh](https://github.com/canyonroad/agentsh)** ๐ŸŸข โ€” Execution-layer policy shell for AI agents that intercepts file, network, process, signal, and selected database activity and emits structured audit events. *(Canyon Road)* โ€” **note:** the shell shim bypasses policy for non-TTY stdin unless `--force` is used, which is critical for headless agents and CI. Linux is the production target; native macOS enforcement is alpha and native Windows drivers are not yet production-ready. *(โ˜… 385 ยท updated 2026-09-15)* - **[brood-box](https://github.com/stacklok/brood-box)** ๐ŸŸข โ€” Experimental runner for coding agents in hardware-isolated microVMs with copy-on-write workspace snapshots, egress profiles, selective secret forwarding, and file-by-file review before applying changes. *(Stacklok)* โ€” **note:** APIs and behavior are explicitly experimental; `workspace-mode=direct` bypasses snapshot isolation and writes directly to the workspace, so it is suitable only for already-trusted interactive work. *(โ˜… 72 ยท updated 2026-09-18)* - **[Greywall](https://github.com/GreyhavenHQ/greywall)** ๐ŸŸข โ€” Kernel-enforced filesystem, network, syscall, and command-policy wrapper for coding agents on Linux and macOS, with a separate traffic-observability mode and least-privilege profile generation. *(Greyhaven)* โ€” **note:** deny-by-default applies to `greywall`; `greywatch` is intentionally permissive and records rather than blocks activity. Verify the effective platform backend and selected mode before treating it as an enforcement boundary. *(โ˜… 298 ยท updated 2026-08-13)* - **[nono](https://github.com/nolabs-ai/nono)** ๐ŸŸข โ€” Least-privilege sandbox for AI coding agents that isolates the agent and delegated tools with composable filesystem, network, credential-proxy, and command policies. *(NoLabs)* โ€” **note:** APIs are still stabilizing ahead of the 1.0 release; review every pulled profile before use. *(โ˜… 3,687 ยท updated 2026-08-17)* - **Related:** [microsandbox](https://github.com/superradcompany/microsandbox) ยท [ToolHive](https://github.com/stacklok/toolhive) - **[Coi](https://github.com/coipond/coi)** ๐ŸŸข โ€” Incus-based session manager for coding agents with protected workspace mounts, configurable network policy, optional host-side monitoring, and persistent or ephemeral Linux containers. โ€” **note:** shared-kernel containers are not a microVM boundary; workspace writes reach the host and configured credentials can be copied into the container. Default `restricted` networking allows public-internet egress and monitoring is off. The hardened profile enables monitoring and reduces kernel attack surface, but is not a general exfiltration or hostile-agent containment guarantee. *(โ˜… 733 ยท updated 2026-10-02)* - **Sources:** [Containment threat model](https://github.com/coipond/coi/wiki/Threat-Model-Containment-Limits) ยท [Reviewed defaults](https://github.com/coipond/coi/blob/20ba7981e7a5b00728d671431a024a9666115b0e/internal/config/embedded/default_config.toml) - **[cplt](https://github.com/navikt/cplt)** ๐ŸŸข โ€” Kernel-backed sandbox wrapper for AI coding agents that applies Seatbelt on macOS or Landlock and seccomp on Linux, content-pins approvals for repository policy, filters environment and resource access, and gates selected git and GitHub commands. *(NAV (Norwegian Labour and Welfare Administration))* โ€” **note:** no native Windows backend; the standard posture permits outbound HTTPS on port 443 and warns rather than blocks `git push`, while stricter egress and push blocking require explicit configuration. Linux has documented limitations around Git metadata, localhost, and SSH-agent isolation, so review the threat model and effective policy for the target platform. *(โ˜… 114 ยท updated 2026-09-06)* - **Related:** [nono](https://github.com/nolabs-ai/nono) ยท [microsandbox](https://github.com/superradcompany/microsandbox) - **[Arcjet Guard](https://github.com/arcjet/arcjet-js)** ๐ŸŸข๐ŸŸ  โ€” JavaScript runtime guard for AI-agent tool calls and MCP handlers, with prompt-injection detection, sensitive-data detection/redaction, and custom local policy rules. *(Arcjet)* โ€” **note:** open SDK packages integrate with Arcjet's hosted platform; assess the service, account, and data-processing requirements for the protections you enable. *(โ˜… 681 ยท updated 2026-08-15)* - **Related:** [LLM Guard](https://github.com/protectai/llm-guard) ยท [AgentLock](https://github.com/webpro255/agentlock) ยท [Arcjet Python SDK (hosted rules send inputs, require key; core fail-open, helpers fail-closed; configure tool coverage)](https://github.com/arcjet/arcjet-py) - **[ToolHive](https://github.com/stacklok/toolhive)** ๐ŸŸข โ€” Platform for running MCP servers in isolated containers with per-request identity/access policy, registry and gateway workflows, audit logs, Kubernetes operator support, and observability hooks. *(Stacklok)* *(โ˜… 2,019 ยท updated 2026-08-14)* - **Related:** [microsandbox](https://github.com/superradcompany/microsandbox) ยท [defenseclaw](https://github.com/cisco-ai-defense/defenseclaw) - **[Pipelock](https://github.com/luckyPipewrench/pipelock)** ๐ŸŸข โ€” AI-agent firewall and verifiable egress-control layer mediating HTTP, WebSocket, CONNECT, MCP, and A2A traffic to detect prompt injection, secret exfiltration, SSRF, and suspicious outbound actions. โ€” **note:** open-source core is Apache-2.0; commercial reporting/features are also advertised. *(โ˜… 796 ยท updated 2026-08-17)* - **Related:** [ToolHive](https://github.com/stacklok/toolhive) ยท [mcp-context-protector](https://github.com/trailofbits/mcp-context-protector) - **[node9](https://github.com/node9-ai/node9-proxy)** ๐ŸŸข โ€” Local policy and human-approval gate for supported coding-agent tool calls routed through agent hooks or an MCP wrapper, with credential-path checks, secret detection, and audit logging. *(Node9)* โ€” **note:** cooperative hooks/MCP enforcement, not process isolation or a complete credential boundary; scripts are judged by their command line rather than their contents, and the egress allowlist is opt-in. Local use needs no account; optional hosted team features can transmit operational data. `node9 init` asks before sending an install ping, with Yes preselected; the payload includes a persistent machine ID, detected agents, OS, and version, so it is pseudonymous rather than anonymous. *(โ˜… 216 ยท updated 2026-09-27)* - **Related:** [Pipelock](https://github.com/luckyPipewrench/pipelock) ยท [agentsh](https://github.com/canyonroad/agentsh) - **[SourceryKit](https://github.com/ProvablyAI/sourcerykit)** ๐ŸŸ โš ๏ธ โ€” Python SDK for agent guardrails that intercepts outbound HTTP calls, enforces trusted-endpoint policies, logs requests, and checks agent handoff claims against a Provably backend before propagation. *(ProvablyAI)* โ€” **note:** BSL-1.1 licensed; requires Provably backend/API credentials and database setup. *(โ˜… 18 ยท updated 2026-08-12)* - **Related:** [Pipelock](https://github.com/luckyPipewrench/pipelock) ยท [mcp-context-protector](https://github.com/trailofbits/mcp-context-protector) - **[emisar](https://github.com/AndrewDryga/emisar)** ๐ŸŸ โš ๏ธ โ€” Agent-infrastructure control plane that exposes declared, typed actions through MCP, applies policy and approval gates before dispatch, revalidates calls on an outbound-only host runner, and records separate control-plane and host audit trails. โ€” **note:** runner, MCP bridge, and packs are Apache-2.0; the hosted portal/control plane is BSL-1.1 and converts to Apache-2.0 on 2029-07-26. Connecting a runner requires an emisar account and outbound HTTPS to the control plane. *(โ˜… 416 ยท updated 2026-08-17)* - **[mcp-context-protector](https://github.com/trailofbits/mcp-context-protector)** ๐ŸŸข โ€” MCP security wrapper that sits in front of downstream MCP servers, scans tool responses with guardrail providers, and supports quarantine/review workflows for desktop and coding-agent MCP configs. *(Trail of Bits)* *(โ˜… 222 ยท updated 2026-02-13)* - **Related:** [Cisco AI Defense โ€“ mcp-scanner](https://github.com/cisco-ai-defense/mcp-scanner) ยท [ToolHive](https://github.com/stacklok/toolhive) ยท [Pipelock](https://github.com/luckyPipewrench/pipelock) - **[MCP Defender](https://github.com/MCP-Defender/MCP-Defender)** ๐ŸŸขโš ๏ธ โ€” Desktop app that proxies MCP tool-call requests and responses for Cursor, Claude, VS Code, and Windsurf, checks intercepted traffic against signatures, and prompts users to allow or block suspicious calls. โ€” **note:** AGPL-3.0 licensed; project has been acquired by Docker. *(โ˜… 255 ยท updated 2026-06-05)* - **Related:** [ToolHive](https://github.com/stacklok/toolhive) ยท [mcp-context-protector](https://github.com/trailofbits/mcp-context-protector) - **[MCP Gateway](https://github.com/lasso-security/mcp-gateway)** ๐ŸŸข โ€” Plugin-based MCP gateway that proxies configured MCP servers, sanitizes sensitive request/response data, supports guardrail plugins such as basic masking and Presidio, and runs a server reputation/risk check before loading MCP servers. *(Lasso Security)* *(โ˜… 385 ยท updated 2026-01-22)* - **Related:** [ToolHive](https://github.com/stacklok/toolhive) ยท [Pipelock](https://github.com/luckyPipewrench/pipelock) ยท [mcp-context-protector](https://github.com/trailofbits/mcp-context-protector) ยท [Claude PostToolUse hook sample (pattern warnings; not isolation or guaranteed context cleanup)](https://github.com/lasso-security/claude-hooks) - **[Parallax](https://github.com/agent-defense/parallax)** ๐ŸŸข โ€” Rust runtime policy engine for AI agents: evaluates lifecycle events with regex, keyword, Sigma, CEL, and SQL rules to block or redact prompt injection, data exfiltration, dangerous tool calls, and secret leakage. โ€” **note:** early-stage project with limited adoption signal. *(โ˜… 35 ยท updated 2026-06-05)* - **Related:** [Pipelock](https://github.com/luckyPipewrench/pipelock) ยท [Armorer Guard](https://github.com/ArmorerLabs/Armorer-Guard) - **[Armorer Guard](https://github.com/ArmorerLabs/Armorer-Guard)** ๐ŸŸข โ€” Local Rust scanner and MCP proxy for AI-agent prompt injection, credential leakage, exfiltration, and risky tool-call arguments, with structured reasons and no scanner network calls. โ€” **note:** young project with limited independent adoption signal. *(โ˜… 42 ยท updated 2026-08-09)* - **Related:** [agentguard](https://github.com/GoPlusSecurity/agentguard) ยท [ATR โ€“ Agent Threat Rules](https://github.com/Agent-Threat-Rule/agent-threat-rules) - **[onecli](https://github.com/onecli/onecli)** ๐ŸŸข โ€” Credential gateway and encrypted vault for AI agents; injects real API credentials at the gateway so agents only see placeholder keys. *(โ˜… 3,104 ยท updated 2026-07-31)* - **Related:** [agentguard](https://github.com/GoPlusSecurity/agentguard) ยท [defenseclaw](https://github.com/cisco-ai-defense/defenseclaw) - **[microsandbox](https://github.com/superradcompany/microsandbox)** ๐ŸŸข โ€” Local-first, microVM-backed programmable sandboxes for AI agents with SDKs, CLI, MCP support, and rootless hardware isolation. *(โ˜… 7,569 ยท updated 2026-08-17)* - **Related:** [agentguard](https://github.com/GoPlusSecurity/agentguard) ยท [defenseclaw](https://github.com/cisco-ai-defense/defenseclaw) ยท [agents-sandbox launcher for microsandbox (GPL-3.0)](https://github.com/inoio/agents-sandbox) - **[agentguard](https://github.com/GoPlusSecurity/agentguard)** ๐ŸŸข โ€” Real-time security layer for coding agents: hooks scan every new skill, block dangerous actions before execution, run daily posture patrols, and track which skill triggered each action (incl. Web3-specific checks). *(โ˜… 456 ยท updated 2026-06-25)* - **Related:** [agent-audit](https://github.com/scadastrangelove/agent-audit) ยท [defenseclaw](https://github.com/cisco-ai-defense/defenseclaw) - **[AgentAegis](https://github.com/antgroup/agent-aegis)** ๐ŸŸข โ€” OpenClaw security plugin that observes or blocks selected prompt, tool-call, memory, exfiltration, and protected-path risks through configurable agent lifecycle hooks. *(Ant Group)* โ€” **note:** the code defaults blocking defenses to `enforce`; the README's `observe` configuration is an optional staged-rollout example. Coverage depends on enabled hooks, modes, and protected assets; this is OpenClaw-specific defense in depth, not a general sandbox or injection-proof boundary, and published effectiveness demonstrations are project-operated. *(โ˜… 199 ยท updated 2026-06-26)* - **Sources:** [Configuration defaults](https://github.com/antgroup/agent-aegis/blob/23d59b8986b86de448d24ff973e662295768d533/src/config.ts#L394) - **[defenseclaw](https://github.com/cisco-ai-defense/defenseclaw)** ๐ŸŸ  โ€” Enforcement and evidence layer for agentic deployments: static CodeGuard checks, sandboxing, registry ingestion with SSRF guards, and audit/observability. *(Cisco AI Defense)* *(โ˜… 819 ยท updated 2026-08-17)* - **Related:** [Cisco AI Defense โ€“ skill-scanner](https://github.com/cisco-ai-defense/skill-scanner) ยท [agentguard](https://github.com/GoPlusSecurity/agentguard) - **[clawsec](https://github.com/prompt-security/clawsec)** ๐ŸŸขโš ๏ธ โ€” Security skill suite for OpenClaw-family agents; AGPL-3.0 licensed. *(Prompt Security)* *(โ˜… 1,083 ยท updated 2026-08-05)* - **Related:** [agentguard](https://github.com/GoPlusSecurity/agentguard) ยท [Cisco AI Defense โ€“ skill-scanner](https://github.com/cisco-ai-defense/skill-scanner) - **[AgentLock](https://github.com/webpro255/agentlock)** ๐ŸŸข๐Ÿ”ฌโš ๏ธ โ€” Pre-action authorization gate for LLM agent tool calls that decides from session provenance rather than content, with deny-by-default tool permissions, parameter lineage, Ed25519 signed receipts, and a hash-chained audit log; AGPL-3.0 licensed with commercial options. โ€” **note:** evaluated on AgentDojo with predictions pre-registered before the runs; the published results include a suite where the defense costs more utility than the attack it prevents. *(โ˜… 19 ยท updated 2026-08-12)* - **[h5i](https://github.com/h5i-dev/h5i)** ๐ŸŸข โ€” Local Rust CLI for auditable coding-agent workspaces: per-agent worktrees with sandbox policies, provenance capture, peer review, neutral verification, secret/prompt-injection audit signals, and refs/h5i/* run metadata. โ€” **note:** security-adjacent agent-workspace governance tool, not a vulnerability scanner or VM-equivalent sandbox. *(โ˜… 531 ยท updated 2026-08-16)* - **Related:** [microsandbox](https://github.com/superradcompany/microsandbox) ยท [defenseclaw](https://github.com/cisco-ai-defense/defenseclaw) - **[DvalinCode](https://github.com/arthurpanhku/dvalincode)** ๐ŸŸข โ€” Local-first AI coding agent with runtime governance controls: org/repo policy gates for tools, models, MCP servers, paths, and commands, plus provider/shell/MCP egress controls and hash-chained audit logs. โ€” **note:** young project with limited independent adoption signal. *(โ˜… 112 ยท updated 2026-08-17)* - **Related:** [h5i](https://github.com/h5i-dev/h5i) ยท [Pipelock](https://github.com/luckyPipewrench/pipelock) ยท [Armorer Guard](https://github.com/ArmorerLabs/Armorer-Guard) - **[TAP](https://github.com/holonym-foundation/tap-oss)** ๐ŸŸข๐ŸŸ  โ€” Credential-isolation proxy and MCP server for AI agents: agents send placeholder credentials, TAP injects real secrets server-side after per-action policy checks, with optional human approval on sensitive calls. *(human.tech)* โ€” **note:** Apache-2.0 runtime is self-hostable, but the hosted dashboard and managed-service deployment glue are proprietary; self-hosting puts credential/key isolation and policy-engine hardening on the operator. *(โ˜… 12 ยท updated 2026-07-23)* - **[Agent Memory Guard](https://github.com/OWASP/www-project-agent-memory-guard)** ๐ŸŸข โ€” Runtime middleware for AI-agent memory reads and writes, screening prompt injection, memory poisoning, secret/PII leakage, protected-key tampering, and size anomalies before persisted memory is reused. *(OWASP)* โ€” **note:** OWASP Incubator project; published benchmark numbers are project-reported and should be independently reproduced before production enforcement. *(โ˜… 125 ยท updated 2026-08-16)* - **[AIO Sandbox](https://github.com/agent-infra/sandbox)** ๐ŸŸขโš ๏ธ โ€” All-in-one Docker workspace for AI agents with browser, shell, file, code-execution, MCP, and VSCode interfaces, plus API-key/JWT controls and private-deployment guidance. โ€” **note:** the public repository ships SDKs, integrations, and docs rather than the core runtime service; official Chromium-enabled deployments use `seccomp=unconfined`. Treat it as a trusted execution environment, not a hardened isolation boundary; use separate VMs or a hardened runtime plus network policy for hostile workloads. *(โ˜… 5,727 ยท updated 2026-08-17)* - **Related:** [Kubernetes Agent Sandbox](https://github.com/kubernetes-sigs/agent-sandbox) ยท [microsandbox](https://github.com/superradcompany/microsandbox) ยท [EdgeBox desktop sandbox (GPL-3.0; README recommends a successor; escape resistance unverified)](https://github.com/BIGPPWONG/EdgeBox) ยท [Ephemeral cooperative workspace isolation (shared sandbox/provider trust; not hostile-tenant containment)](https://github.com/Ephemeral-AI-Lab/ephemeral-sandbox) - **[Agentgateway](https://github.com/agentgateway/agentgateway)** ๐ŸŸข โ€” Agent-native proxy and gateway for MCP and A2A traffic with OAuth/JWT/API-key authentication, CEL-based RBAC policies, TLS, rate limiting, and OpenTelemetry observability. *(โ˜… 4,385 ยท updated 2026-08-14)* - **Related:** [MCP Gateway](https://github.com/lasso-security/mcp-gateway) ยท [Pipelock](https://github.com/luckyPipewrench/pipelock) - **[Kubernetes Agent Sandbox](https://github.com/kubernetes-sigs/agent-sandbox)** ๐ŸŸข โ€” Kubernetes CRDs and controllers for isolated, stateful singleton agent workloads, delegating low-level isolation to configured runtimes such as gVisor or Kata Containers. *(Kubernetes SIG Apps)* โ€” **note:** sandbox orchestrator, not an isolation runtime itself; security depends on the selected RuntimeClass, network policy, and workload configuration. *(โ˜… 3,544 ยท updated 2026-08-16)* - **Sources:** [Kubernetes announcement](https://kubernetes.io/blog/2026/03/20/running-agents-on-kubernetes-with-agent-sandbox/) - **Related:** [AIO Sandbox](https://github.com/agent-infra/sandbox) - **[Prismor](https://github.com/PrismorSec/prismor)** ๐ŸŸข โ€” Self-hosted runtime control plane for coding agents with pre-tool-call hooks, policy-driven observe/approve/block decisions, an MCP gateway, secret and egress controls, and tamper-evident audit evidence. *(โ˜… 291 ยท updated 2026-08-16)* - **Related:** [Armorer Guard](https://github.com/ArmorerLabs/Armorer-Guard) ยท [AgentLock](https://github.com/webpro255/agentlock) - **[Gram](https://github.com/speakeasy-api/gram)** ๐ŸŸข๐ŸŸ โš ๏ธ โ€” Open-source stack behind Speakeasy's AI control plane that centrally manages MCPs, Skills, and Assistants with granular permissions, policy enforcement, threat detection, and observability. *(Speakeasy)* โ€” **note:** AGPL-3.0 copyleft license; Gram is also available as Speakeasy's hosted commercial control plane, while local development uses a full Docker/Mise stack. *(โ˜… 266 ยท updated 2026-09-09)* - **[tirith](https://github.com/sheeki03/tirith)** ๐ŸŸขโš ๏ธ โ€” Terminal guard for developers and AI coding agents that intercepts homograph and terminal-injection tricks, obfuscated execution chains, pipe-to-shell patterns, credential exfiltration, and malicious skill/config files. โ€” **note:** AGPL-3.0 with a separate commercial license; shell interception is a host-side guard, not a substitute for sandboxing or least-privilege tool access. *(โ˜… 2,664 ยท updated 2026-08-13)* - **Related:** [nono](https://github.com/nolabs-ai/nono) ยท [gate.cat](https://github.com/BGMLAI/gate.cat) - **[ADR](https://github.com/uber/ADR)** ๐ŸŸข๐Ÿ”ฌ โ€” Agentic AI Detection and Response system combining cross-client agent telemetry, ADR-Bench security scenarios, and a dual-agent detector for suspicious intent, tool use, and execution traces. *(Uber)* โ€” **note:** deployed at Uber and published with an MLSys 2026 paper; the open release includes the Sensor, benchmark, and Detector, but not ADR Prevention or the offline ADR Explorer. Default detector configurations require model-provider credentials. *(โ˜… 1,442 ยท updated 2026-08-16)* - **Related:** [Pipelock](https://github.com/luckyPipewrench/pipelock) ยท [AgentLock](https://github.com/webpro255/agentlock) - **[xaidr](https://github.com/delphisecurity/xaidr)** ๐ŸŸข โ€” In-process runtime security sensor for AI agents that inspects input, tool calls, output, and agent-to-agent envelopes inside the agent process, with structured shell-command classification, YAML policy, privilege tiers, an opt-in circuit breaker, and OpenTelemetry export. *(Delphi Security)* โ€” **note:** early-stage in-process sensor, not an isolation boundary; monitor mode is the default and unexpected scan failures fail open while emitting degraded telemetry. Published detection and false-positive figures are project-reported on its committed corpus. *(โ˜… 22 ยท updated 2026-08-17)* - **Related:** [agentguard](https://github.com/GoPlusSecurity/agentguard) ยท [defenseclaw](https://github.com/cisco-ai-defense/defenseclaw) - **[Agentmetry](https://github.com/blitzcrieg1/agentmetry)** ๐ŸŸข โ€” Local-first flight recorder for AI coding agents and MCP servers that writes a hash-chained JSONL trail with RFC 6962 Merkle roots, applies MITRE-mapped sequence detection across a session, attests every 300s which agent surfaces are covered, uncovered, absent or unknown, and forwards to Splunk HEC, Elastic ECS, Google SecOps UDM or CloudEvents. โ€” **note:** young project with limited independent adoption signal; it records and detects rather than blocks, and its README states that hooks are cooperative and that tamper-evident is not the same as attributable. *(โ˜… 12 ยท updated 2026-08-24)* - **Related:** [HOL Guard](https://github.com/hashgraph-online/hol-guard) ยท [ADR](https://github.com/uber/ADR) - **[piighost](https://github.com/Athroniaeth/piighost)** ๐ŸŸข โ€” Local runtime pseudonymization layer that replaces detected PII with stable placeholders before model calls and restores the original values in responses and selected tool arguments, with integrations for LangChain, Pydantic AI, LlamaIndex, and OpenAI-compatible clients. โ€” **note:** reversible de-identification, not anonymization; the placeholder mapping remains sensitive, restored tool arguments can contain the original PII, and optional LLM-backed detectors may send content to the configured model provider. *(โ˜… 11 ยท updated 2026-08-23)* - **Related:** [Kiji](https://github.com/Dataiku/kiji-proxy) ยท [onecli](https://github.com/onecli/onecli) - **[HOL Guard](https://github.com/hashgraph-online/hol-guard)** ๐ŸŸข๐ŸŸ  โ€” Local-first runtime security layer for coding agents that evaluates commands, package installs, skills, MCP configuration, and sensitive actions through policy, approval, evidence, and audit workflows. *(Hashgraph Online)* โ€” **note:** integrations are cooperative hooks rather than a hard isolation boundary, and documented hook failures can fail open; synchronized evidence, team policy, and fleet workflows use the optional hosted service. *(โ˜… 466 ยท updated 2026-08-24)* - **Related:** [emisar](https://github.com/MatteoColelli/emisar) ยท [Prismor](https://github.com/PrismorSec/prismor) - **[StackOne Defender](https://github.com/StackOneHQ/defender)** ๐ŸŸข โ€” Offline TypeScript runtime guard for indirect prompt injection in tool results, using bundled ONNX classifiers, deterministic checks, sanitization, and allow/block verdicts. *(StackOne)* โ€” **note:** blocking high-risk results is opt-in by default; missing Tier-2 dependencies can fall back to simpler Tier-1 detection unless strict mode is enabled, so configure fail behavior explicitly for enforcement use. *(โ˜… 117 ยท updated 2026-08-19)* - **Related:** [LLM Guard](https://github.com/protectai/llm-guard) ยท [little-canary](https://github.com/hermes-labs-ai/little-canary) - **[Earl](https://github.com/mathematic-inc/earl)** ๐ŸŸข โ€” Capability proxy for AI agents that exposes approved operation names while keeping request templates and credentials outside the model, with HCL policy, audit, and egress controls. *(Mathematic)* โ€” **note:** operation templates and policy configuration form part of the trusted boundary and require review; the proxy reduces credential and request-construction exposure but is not a sandbox for the agent process. *(โ˜… 113 ยท updated 2026-08-23)* - **Related:** [onecli](https://github.com/onecli/onecli) ยท [agentgateway](https://github.com/agentgateway/agentgateway) - **[Bifrost](https://github.com/maximhq/bifrost)** ๐ŸŸข๐ŸŸ  โ€” Apache-licensed AI and MCP gateway with multi-provider routing, virtual-key access controls, budgets, rate limits, MCP aggregation, OAuth, automatic fallbacks, and load balancing. *(Maxim)* โ€” **note:** model-provider credentials and proxied request/response data are sensitive; restrict and authenticate the gateway and admin interface, use TLS, and load only trusted plugins. Content guardrails, RBAC/SSO, clustering, and other advanced governance controls require Bifrost Enterprise. *(โ˜… 7,548 ยท updated 2026-08-25)* - **Related:** [Portkey AI Gateway](https://github.com/Portkey-AI/gateway) ยท [Agentgateway](https://github.com/agentgateway/agentgateway) - **[Portkey AI Gateway](https://github.com/Portkey-AI/gateway)** ๐ŸŸข๐ŸŸ  โ€” Open AI gateway with provider routing, fallback and retry controls, guardrail integrations, observability, and MCP traffic support for model and agent applications. *(Portkey)* โ€” **note:** the gateway is general infrastructure rather than a standalone security scanner; model-provider credentials are required, and some RBAC, analytics, and managed control-plane capabilities belong to Portkey's hosted or enterprise products. *(โ˜… 12,815 ยท updated 2026-05-25)* - **Related:** [agentgateway](https://github.com/agentgateway/agentgateway) ยท [Pipelock](https://github.com/luckyPipewrench/pipelock) - **[Casbin AI Gateway](https://github.com/apache/casbin-gateway)** ๐ŸŸข โ€” Local gateway and policy layer for model-provider and MCP traffic, combining provider-key mediation, access control, prompt records, and configurable request handling. *(Apache Casbin)* โ€” **note:** binds to localhost by default because its local UI can exercise administrative operations and the relay can access stored provider keys; do not expose it beyond a trusted host without adding authentication and reviewing prompt-recording settings. *(โ˜… 570 ยท updated 2026-08-24)* - **Related:** [agentgateway](https://github.com/agentgateway/agentgateway) ยท [Portkey AI Gateway](https://github.com/Portkey-AI/gateway) - **[Adrian](https://github.com/secureagentics/Adrian)** ๐ŸŸข๐ŸŸ  โ€” Runtime monitoring and intervention layer that correlates agent actions with available reasoning traces through Python and TypeScript SDKs, a Claude Code integration, and managed or self-hosted deployments. *(Secure Agentics)* โ€” **note:** the bundled self-hosted stack requires Docker, substantial disk space, and an NVIDIA GPU for the recommended local classifier; the README's `+35%` and `4x` headline extrapolates cited third-party research rather than an Adrian benchmark. *(โ˜… 552 ยท updated 2026-08-20)* - **Related:** [Agentic Radar](https://github.com/splx-ai/agentic-radar) ยท [Armorer Guard](https://github.com/ArmorerLabs/Armorer-Guard) - **[Sandlock](https://github.com/multikernel/sandlock)** ๐ŸŸข โ€” Unprivileged Linux process sandbox using Landlock, seccomp-BPF, and seccomp user notification to apply per-process filesystem, network, syscall, and execution policies without a container or VM. *(Multikernel)* โ€” **note:** Linux-kernel process isolation rather than prompt-injection detection; protection depends on available Landlock and seccomp features and is not a multi-tenant VM boundary. *(โ˜… 381 ยท updated 2026-08-23)* - **Related:** [Sandbox Probe](https://github.com/controlplaneio/sandbox-probe) ยท [AIO Sandbox](https://github.com/agent-infra/sandbox) - **[Lunar](https://github.com/TheLunarCompany/lunar)** ๐ŸŸข๐ŸŸ  โ€” API and MCP gateway combining outbound-traffic visibility, policy enforcement, rate limits, retries, circuit breakers, and centralized MCP server aggregation for agent workloads. *(Lunar.dev)* โ€” **note:** the repository is active but its latest formal GitHub release is from 2024; the README positions the open core for non-production or personal use and directs production deployments to commercial platform tiers. *(โ˜… 485 ยท updated 2026-08-28)* - **Related:** [agentgateway](https://github.com/agentgateway/agentgateway) ยท [Bifrost](https://github.com/maximhq/bifrost) - **[Norviq](https://github.com/norviq-dev/norviq)** ๐ŸŸข โ€” Kubernetes policy enforcement point for LLM agent tool calls that content-hash pins each tool definition at discovery and evaluates every tools/call against OPA/Rego policy before the upstream MCP server receives it. โ€” **note:** ships in audit mode โ€” the installed baseline records what it would have refused and lets the call proceed, and all shipped policy presets default to allow, so enforcement starts with the first rule an operator writes. The discovery-time description scanner is a documented heuristic and the project publishes the payloads that defeat it. Young project with limited independent adoption signal. *(โ˜… 24 ยท updated 2026-09-01)* - **Related:** [Prismor](https://github.com/PrismorSec/prismor) ยท [mcp-context-protector](https://github.com/trailofbits/mcp-context-protector) - **[sofagent](https://github.com/KongFangXun/sofagent)** ๐ŸŸข โ€” Commit-time audit and governance suite for AI coding agents that scans git diffs against deterministic rules, records local audit history, and exposes MCP tools for governance aggregation. โ€” **note:** HMAC signing is optional, while local hooks, configuration, and key material remain accessible to same-user agents; the default setup is not fail-closed and Git hooks can be bypassed, so treat the history as local audit evidence rather than a hardened tamper-proof boundary. *(โ˜… 42 ยท updated 2026-09-03)* - **Related:** [Pipelock](https://github.com/luckyPipewrench/pipelock) - **[Tenuo](https://github.com/tenuo-ai/tenuo)** ๐ŸŸข๐ŸŸ  โ€” Capability-based authorization for agent tool calls using signed, holder-bound warrants that constrain tools and arguments, verify locally, and narrow authority across delegation. โ€” **note:** the Rust core, SDKs, and authorizer sidecar are Apache-2.0; the managed control plane is commercial. Enforcement requires trusted verification on every effecting tool path, outside agent control when bypass is possible; this is not a sandbox. The stateless core permits identical-call replay within its proof-of-possession window, so state-changing operations need application-level nonce or idempotency controls. *(โ˜… 100 ยท updated 2026-10-03)* - **Related:** [Agentgateway](https://github.com/agentgateway/agentgateway) ยท [AgentLock](https://github.com/webpro255/agentlock) --- ## AI/ML Supply Chain & Model Security Tools for securing model artifacts, serialized ML files, AI/ML supply-chain surfaces, and malicious-package detection datasets/benchmarks. - **[MindArmour](https://github.com/mindspore-ai/mindarmour)** ๐ŸŸข๐Ÿ”ฌ โ€” MindSpore toolkit for adversarial robustness and model-privacy research, including gradient-based targeted and untargeted attacks. *(MindSpore)* **Caveat:** requires a compatible MindSpore, hardware, model, and dataset stack; package metadata points to Gitee for source and issue tracking. Model-specific evaluations are not robustness guarantees. *(โ˜… 92 ยท updated 2026-02-25)* - **[FedREDefense](https://github.com/xyq7/FedREDefense)** ๐ŸŸข๐Ÿ”ฌ โ€” ICML 2024 research implementation of federated-learning model-poisoning defense using reconstruction error of client updates. **Caveat:** an experimental training/evaluation pipeline, not a production federation security layer; GPU, datasets, and attack configuration affect results. Consult the published erratum for the FLTrust baseline rather than copying the original comparison figures. *(โ˜… 31 ยท updated 2026-09-16)* - **Related:** [FLDetector poisoning-detection research (historical; no root license, legacy MXNet 1.9.1)](https://github.com/zaixizhang/FLDetector) ยท [FL-WBC client-side poisoning defense (historical; no root license, legacy PyTorch 1.2)](https://github.com/jeremy313/FL-WBC) ยท [NDSS21 model-poisoning research (historical; no root license, legacy datasets/frameworks)](https://github.com/vrt1shjwlkr/NDSS21-Model-Poisoning) ยท [ModelPoisoning FL research (historical; no root license, legacy TensorFlow 1.8)](https://github.com/inspire-group/ModelPoisoning) - **[ModelAudit](https://github.com/promptfoo/modelaudit)** ๐ŸŸข โ€” Model-artifact scanner with format-specific static parsers, a Rust-backed pickle-analysis companion, metadata checks, and JSON/SARIF reporting. *(Promptfoo)* **Caveat:** installed usage enables telemetry by default; disable with `PROMPTFOO_DISABLE_TELEMETRY=1` or `NO_ANALYTICS=1`. Local scans need no hosted LLM, while remote sources need network access and sometimes credentials. Reports can retain raw secrets; optional `metadata --trust-loaders` can deserialize artifacts. A clean static scan does not establish model safety. *(โ˜… 76 ยท updated 2026-10-04)* - **[AI BOM](https://github.com/cisco-ai-defense/aibom)** ๐ŸŸข โ€” Inventories models, agents, tools, MCP clients and servers, datasets, prompts, guardrails, secrets, and cloud AI resources, with CycloneDX 1.6 output and policy workflows. *(Cisco AI Defense)* โ€” **note:** core inventory is distinct from the narrower model-file AIsbom already listed; the extended `analyze` pipeline requires an external LLM provider and can send analysis context off-host. Optional sanitized Galileo telemetry is separately configurable. *(โ˜… 110 ยท updated 2026-09-17)* - **Sources:** [Cisco announcement](https://blogs.cisco.com/ai/know-your-ai-stack-introducing-ai-bom-in-cisco-ai-defense) - **Related:** [AIsbom](https://github.com/Lab700xOrg/aisbom) - **[Fraim](https://github.com/fraim-dev/fraim)** ๐ŸŸข โ€” Framework for AI-powered security workflows including LLM SAST and IaC analysis with SARIF/HTML output. *(โ˜… 160 ยท updated 2025-12-01)* - **Related:** [sast-skills](https://github.com/utkusen/sast-skills) - **[Adversarial Robustness Toolbox (ART)](https://github.com/Trusted-AI/adversarial-robustness-toolbox)** ๐ŸŸข โ€” Flagship machine-learning security library for evaluating and defending models against evasion, poisoning, extraction, and inference attacks across major ML frameworks. *(LF AI & Data / IBM)* *(โ˜… 6,179 ยท updated 2025-11-13)* - **Related:** [Foolbox](https://github.com/bethgelab/foolbox) ยท [PrivacyRaven](https://github.com/trailofbits/PrivacyRaven) ยท [XLab adversarial-ML course (incomplete sections; no root license, external models/cloud Colab)](https://github.com/zroe1/xlab-ai-security) ยท [Machine Learning Security course (no root license; separate dependencies/datasets)](https://github.com/unica-mlsec/mlsec) ยท [OWASP ML Security Top 10 (CC BY-SA 4.0 risk reference, not certification)](https://github.com/OWASP/www-project-machine-learning-security-top-10) - **[Foolbox](https://github.com/bethgelab/foolbox)** ๐ŸŸข โ€” Classic Python toolbox for generating adversarial examples and benchmarking robustness of PyTorch, TensorFlow, and JAX models. *(โ˜… 2,972 ยท updated 2024-03-04)* - **Related:** [Adversarial Robustness Toolbox](https://github.com/Trusted-AI/adversarial-robustness-toolbox) - **[modelscan](https://github.com/protectai/modelscan)** ๐ŸŸข โ€” Scans ML model files for unsafe serialization patterns and embedded code, with a focus on model serialization attacks. *(Protect AI)* *(โ˜… 762 ยท updated 2026-02-18)* - **Related:** [Fickling](https://github.com/trailofbits/fickling) ยท [picklescan](https://github.com/mmaitre314/picklescan) ยท [ai-exploits](https://github.com/protectai/ai-exploits) - **[Fickling](https://github.com/trailofbits/fickling)** ๐ŸŸข โ€” Python pickle decompiler, rewriter, and static analyzer for inspecting and detecting malicious pickle/PyTorch payloads. *(Trail of Bits)* *(โ˜… 662 ยท updated 2026-08-13)* - **Related:** [modelscan](https://github.com/protectai/modelscan) ยท [picklescan](https://github.com/mmaitre314/picklescan) - **[picklescan](https://github.com/mmaitre314/picklescan)** ๐ŸŸข โ€” Lightweight CLI/library for detecting suspicious Python pickle operations in ML and model artifacts. *(โ˜… 418 ยท updated 2026-07-01)* - **Related:** [modelscan](https://github.com/protectai/modelscan) ยท [Fickling](https://github.com/trailofbits/fickling) - **[AIsbom](https://github.com/Lab700xOrg/aisbom)** ๐ŸŸข โ€” AI software bill of materials tooling for AI/ML supply-chain inventory and provenance metadata. *(โ˜… 76 ยท updated 2026-08-17)* - **Related:** [modelscan](https://github.com/protectai/modelscan) ยท [model-provenance-kit](https://github.com/cisco-ai-defense/model-provenance-kit) - **[model-provenance-kit](https://github.com/cisco-ai-defense/model-provenance-kit)** ๐ŸŸข โ€” Toolkit for model-family provenance and fingerprinting across model weights, tokenizers, and architecture signals. *(Cisco AI Defense)* *(โ˜… 101 ยท updated 2026-08-12)* - **Related:** [AIsbom](https://github.com/Lab700xOrg/aisbom) - **[pickle-fuzzer](https://github.com/cisco-ai-defense/pickle-fuzzer)** ๐ŸŸข โ€” Structure-aware fuzzer for pickle scanners, useful for hardening tools such as modelscan, Fickling, and picklescan. *(Cisco AI Defense)* *(โ˜… 17 ยท updated 2026-08-03)* - **Related:** [modelscan](https://github.com/protectai/modelscan) ยท [Fickling](https://github.com/trailofbits/fickling) ยท [picklescan](https://github.com/mmaitre314/picklescan) - **[Medusa](https://github.com/Pantheon-Security/medusa)** ๐ŸŸขโš ๏ธ โ€” AI-first security scanner for AI/ML repos, agents, and MCP surfaces; AGPL-3.0 licensed. *(Pantheon Security)* *(โ˜… 964 ยท updated 2026-06-24)* - **Related:** [agent-audit](https://github.com/scadastrangelove/agent-audit) ยท [modelscan](https://github.com/protectai/modelscan) - **[PrivacyRaven](https://github.com/trailofbits/PrivacyRaven)** ๐ŸŸข๐Ÿ”ฌ โ€” Privacy-testing library for deep-learning systems, covering model extraction and membership-inference style attacks. *(Trail of Bits)* โ€” **note:** archived/hiatus project, but still a useful reference implementation. *(โ˜… 214 ยท updated 2025-09-05)* - **Related:** [Adversarial Robustness Toolbox](https://github.com/Trusted-AI/adversarial-robustness-toolbox) - **[gym-malware](https://github.com/endgameinc/gym-malware)** ๐ŸŸข๐Ÿ”ฌ โ€” OpenAI Gym environment for reinforcement-learning agents that mutate PE malware to evade static ML malware detectors. *(โ˜… 636 ยท updated 2018-06-15)* - **[deep-pwning](https://github.com/cchio/deep-pwning)** ๐ŸŸข๐Ÿ”ฌ โ€” Historical "Metasploit for machine learning" framework for experimenting with adversarial robustness of ML models. *(โ˜… 571 ยท updated 2022-05-17)* - **[open-malicious-code-benchmark](https://github.com/False-Positive-Community/open-malicious-code-benchmark)** ๐ŸŸข๐Ÿ”ฌ โ€” OMCBench benchmark suite for malicious-code/package detection: labeled Python and JavaScript package archives, common runners, and published precision/recall/F1 metrics. *(False Positive Community)* โ€” **note:** evaluates an unreleased commercial ML detector (MOLOT / PT Application Inspector) alongside open-source baselines. *(โ˜… 17 ยท updated 2026-06-09)* - **Related:** [GuardDog](https://github.com/DataDog/guarddog) ยท [OSSGadget](https://github.com/microsoft/OSSGadget) ยท [malicious-code-ruleset](https://github.com/apiiro/malicious-code-ruleset) ยท [bandit4mal](https://github.com/lyvd/bandit4mal) - **[malicious-software-packages-dataset](https://github.com/DataDog/malicious-software-packages-dataset)** ๐ŸŸข๐Ÿ”ฌ โ€” Human-vetted dataset of malicious software packages across npm, PyPI, IDE extensions, and AI Skills, useful for detector training and evaluation. *(Datadog Security Labs)* โ€” **note:** contains real malware samples; Datadog notes selection bias because many samples were identified by GuardDog. *(โ˜… 372 ยท updated 2026-08-17)* - **Related:** [GuardDog](https://github.com/DataDog/guarddog) ยท [pypi_malregistry](https://github.com/lxyeternal/pypi_malregistry) - **[GuardDog](https://github.com/DataDog/guarddog)** ๐ŸŸข โ€” CLI for detecting malicious PyPI, npm, Go, RubyGems, GitHub Actions, and VSCode extension packages using Semgrep rules and package-metadata heuristics. *(Datadog)* *(โ˜… 1,184 ยท updated 2026-08-14)* - **Related:** [malicious-software-packages-dataset](https://github.com/DataDog/malicious-software-packages-dataset) ยท [Packj](https://github.com/ossillate-inc/packj) - **[package-analysis](https://github.com/ossf/package-analysis)** ๐ŸŸข๐Ÿ”ฌ โ€” Sandboxed static/dynamic analysis pipeline for open-source packages, capturing filesystem, process, and network behavior and publishing data for malicious-package research. *(OpenSSF)* *(โ˜… 903 ยท updated 2026-07-21)* - **Related:** [malicious-packages](https://github.com/ossf/malicious-packages) ยท [package-feeds](https://github.com/ossf/package-feeds) - **[malicious-code-ruleset](https://github.com/apiiro/malicious-code-ruleset)** ๐ŸŸข โ€” Focused Semgrep ruleset for malicious-code patterns such as dynamic execution and obfuscation, used as an OMCBench baseline. *(Apiiro)* *(โ˜… 149 ยท updated 2025-02-24)* - **Related:** [open-malicious-code-benchmark](https://github.com/False-Positive-Community/open-malicious-code-benchmark) - **[pypi_malregistry](https://github.com/lxyeternal/pypi_malregistry)** ๐Ÿ”ฌโš ๏ธ โ€” ASE'23 / USENIX Security'26 malicious-PyPI dataset with more than 10k malicious package versions. โ€” **note:** no LICENSE file found and the repository contains malware samples; handle in an isolated environment. *(โ˜… 129 ยท updated 2026-07-21)* - **Related:** [malicious-software-packages-dataset](https://github.com/DataDog/malicious-software-packages-dataset) - **[Activation-based Model Scanner (AMS)](https://github.com/GoogleCloudPlatform/activation-model-scanner)** ๐ŸŸข๐Ÿ”ฌ โ€” PyPI scanner that uses safety-related activation fingerprints to detect degraded or removed safety training and compare an open-weight model with a known baseline. *(Google Cloud Platform)* โ€” **note:** unofficial and unsupported Google research-derived project; GPU execution is recommended, calibration covers a limited set of model families, and the documented method can miss modifications that preserve measured safety directions. *(โ˜… 32 ยท updated 2026-08-27)* - **Related:** [AASE research](https://research.google/pubs/aase-activation-based-ai-safety-enforcement-via-lightweight-probes/) ยท [model-provenance-kit](https://github.com/cisco-ai-defense/model-provenance-kit) --- ## Pentest & Red-Team Agents Autonomous and semi-autonomous AI agents for penetration testing, exploitation, and attack simulation. - **[Cairn](https://github.com/oritera/Cairn)** ๐ŸŸข๐Ÿ”ฌโš ๏ธ โ€” Fact-intent graph coordinator that schedules Claude Code, Codex, or Pi workers for penetration-testing and CTF exploration through a shared persistent blackboard. **Caveat:** workers can execute arbitrary tools; Docker mode exposes the host Docker socket and local mode has no OS sandbox. Configured providers receive target context; use a dedicated isolated host. Root license is AGPL-3.0; README describes personal/educational use and offers a separate commercial license for use without AGPL obligations. *(โ˜… 3,198 ยท updated 2026-09-07)* - **[PentestGPT](https://github.com/GreyDGL/PentestGPT)** ๐ŸŸข๐Ÿ”ฌ โ€” The original USENIX'24 LLM pentest agent; re-released as an autonomous pipeline with strong benchmark results. *(โ˜… 14,906 ยท updated 2026-07-14)* - **[PentAGI](https://github.com/vxcontrol/pentagi)** ๐ŸŸข โ€” Fully autonomous multi-agent pentest framework with Docker sandboxing. *(VXControl)* *(โ˜… 21,860 ยท updated 2026-08-06)* - **[CAI โ€“ Cybersecurity AI](https://github.com/aliasrobotics/cai)** ๐ŸŸข๐ŸŸ  โ€” Modular, bug-bounty-ready agent framework supporting 300+ LLM models. MIT for research; separate commercial license for production/on-prem. *(Alias Robotics)* *(โ˜… 9,745 ยท updated 2026-07-14)* - **[Strix](https://github.com/usestrix/strix)** ๐ŸŸข โ€” Autonomous "AI hackers" that dynamically run code and validate vulnerabilities with PoCs (Apache-2.0). *(โ˜… 53,569 ยท updated 2026-08-17)* - **[hackingBuddyGPT](https://github.com/ipa-lab/hackingBuddyGPT)** ๐ŸŸข๐Ÿ”ฌ โ€” Minimal (~50 LOC) research framework for LLM-driven Linux priv-esc and web pentesting (FSE'23). *(โ˜… 1,209 ยท updated 2026-08-10)* - **[Nebula](https://github.com/berylliumsec/nebula)** ๐ŸŸข๐ŸŸ  โ€” AI pentesting CLI assistant with local-LLM support (Llama-3.1, Mistral, DeepSeek). *(โ˜… 1,087 ยท updated 2026-07-26)* - **[HexStrike-AI](https://github.com/0x4m4/hexstrike-ai)** ๐ŸŸข โ€” MCP server exposing 150+ security tools (nmap, gobuster, nuclei, โ€ฆ) to AI agents (MIT). *(โ˜… 11,116 ยท updated 2026-08-03)* - **[Deep Eye](https://github.com/zakirkun/deep-eye)** ๐ŸŸข โ€” AI-assisted penetration-testing scanner that orchestrates multiple LLM providers for payload generation, 45+ vulnerability checks, CVE/RAG-assisted testing, AI triage, scan diffing, browser automation, proxying, and multi-format reports. โ€” **note:** MIT-licensed; authorized use only. Heavy runtime surface: optional browser automation/proxying, plaintext API-key config, plugins with full OS access, and pickle model files called out in SECURITY.md. *(โ˜… 1,972 ยท updated 2026-08-14)* - **Related:** [HexStrike-AI](https://github.com/0x4m4/hexstrike-ai) ยท [pentest-ai](https://github.com/0xSteph/pentest-ai) - **[Burp Suite MCP Server](https://github.com/PortSwigger/mcp-server)** ๐ŸŸขโš ๏ธ โ€” Official Burp Suite extension exposing Burp to AI clients through MCP. *(PortSwigger)* โ€” **note:** GPL-3.0 licensed. *(โ˜… 1,071 ยท updated 2026-08-12)* - **Related:** [HexStrike-AI](https://github.com/0x4m4/hexstrike-ai) ยท [Burp MCP workflow helpers (Burp/Claude required; payloads and tokens may enter model context)](https://github.com/SnailSploit/Burp-MCP-Security-Analysis-Toolkit) - **[pentest-ai](https://github.com/0xSteph/pentest-ai)** ๐ŸŸข โ€” Offensive-security MCP server with 200+ wrapped tools, specialist agents, and OWASP-oriented probes for authorized testing. *(โ˜… 1,598 ยท updated 2026-08-17)* - **Related:** [pentest-ai-agents](https://github.com/0xSteph/pentest-ai-agents) - **[Transilience Community Tools](https://github.com/transilienceai/communitytools)** ๐ŸŸข โ€” Claude Code skill and coordination-role bundle with supporting tools for authorized penetration testing, reconnaissance, source review, and AI threat testing. *(Transilience AI)* โ€” **note:** agent-interpreted workflows, not a deterministic scanner or enforcement boundary; results depend on the configured coding agent/model, and published CTF results are maintainer-run after iterative benchmark tuning. The Docker helper uses host networking and `--dangerously-skip-permissions`, while `env-reader.py` prints requested secrets. Review and pin content, use disposable environments and scoped test credentials, and treat transcripts and context sent to model providers as sensitive. *(โ˜… 555 ยท updated 2026-07-29)* - **Sources:** [Container helper](https://github.com/transilienceai/communitytools/blob/95fdc128af4ca1ae16b3226f9430f1bad97b0656/scripts/kali-claude-setup.sh) ยท [Credential helper](https://github.com/transilienceai/communitytools/blob/95fdc128af4ca1ae16b3226f9430f1bad97b0656/tools/env-reader.py) - **[pentest-ai-agents](https://github.com/0xSteph/pentest-ai-agents)** ๐ŸŸข โ€” Collection of Claude Code offensive-security subagents for authorized penetration-testing research. *(โ˜… 2,132 ยท updated 2026-08-16)* - **Related:** [pentest-ai](https://github.com/0xSteph/pentest-ai) - **[DarkMoon](https://github.com/ASCIT31/Dark-Moon)** ๐ŸŸขโš ๏ธ โ€” Autonomous AI penetration-testing platform that orchestrates specialized web, AD, Kubernetes, CMS, and framework agents through an MCP-controlled Docker toolbox with local privacy-tokenization for sensitive target data. โ€” **note:** GPL-3.0 licensed; heavy Docker/LLM stack, use only for authorized testing. *(โ˜… 843 ยท updated 2026-08-06)* - **Related:** [PentAGI](https://github.com/vxcontrol/pentagi) ยท [HexStrike-AI](https://github.com/0x4m4/hexstrike-ai) ยท [pentest-ai](https://github.com/0xSteph/pentest-ai) - **[T3MP3ST](https://github.com/elder-plinius/T3MP3ST)** ๐ŸŸขโš ๏ธ โ€” Autonomous offensive-security meta-harness that wraps local or API-backed coding agents into a multi-agent recon-to-exploit workflow with MCP/API, War Room UI, tool arsenal, and committed benchmark artifacts. โ€” **note:** very new AGPL-3.0 project with bold benchmark claims; use only for authorized testing and verify independently before operational use. *(โ˜… 5,596 ยท updated 2026-08-12)* - **Related:** [PentAGI](https://github.com/vxcontrol/pentagi) ยท [HexStrike-AI](https://github.com/0x4m4/hexstrike-ai) ยท [pentest-ai](https://github.com/0xSteph/pentest-ai) - **[Shannon](https://github.com/KeygraphHQ/shannon)** ๐ŸŸข๐ŸŸ โš ๏ธ โ€” White-box autonomous AI pentester with strong XBOW-benchmark results. *Shannon Lite is AGPL-3.0; Shannon Pro is commercial.* *(โ˜… 46,882 ยท updated 2026-08-12)* - **[AIDA](https://github.com/Vasco0x4/AIDA)** ๐ŸŸขโš ๏ธ โ€” Model-agnostic autonomous pentest agent running inside an isolated Docker environment; AGPL-3.0 licensed. *(โ˜… 471 ยท updated 2026-07-19)* - **[HackSynth](https://github.com/aielte-research/HackSynth)** ๐ŸŸข๐Ÿ”ฌโš ๏ธ โ€” Planner/summarizer LLM-agent framework for autonomous penetration testing and benchmark evaluation; AGPL-3.0 licensed. *(โ˜… 314 ยท updated 2025-06-24)* - **[VulnBot](https://github.com/KHenryAegis/VulnBot)** ๐ŸŸข๐Ÿ”ฌ โ€” Multi-agent collaborative penetration-testing framework with RAG support. *(โ˜… 185 ยท updated 2025-04-07)* - **[PentestAgent](https://github.com/GH05TCREW/pentestagent)** ๐ŸŸข โ€” Black-box AI pentest framework with MCP, multi-agent spawning, and persistent sessions. *(โ˜… 2,958 ยท updated 2026-08-04)* - **[cyber-security-llm-agents](https://github.com/NVISOsecurity/cyber-security-llm-agents)** ๐ŸŸขโš ๏ธ โ€” AutoGen-based agents for cybersecurity tasks (shown at RSAC 2024). *(NVISO)* *(โ˜… 388 ยท updated 2024-05-07)* - **[Pentest-Swarm-AI](https://github.com/Armur-Ai/Pentest-Swarm-AI)** ๐ŸŸข โ€” Swarm-intelligence multi-agent pentest with stigmergic blackboard coordination (Go). *(โ˜… 2,205 ยท updated 2026-08-04)* - **[hackGPT](https://github.com/NoDataFound/hackGPT)** ๐ŸŸขโš ๏ธ โ€” LLM offensive-security toolkit. *(โ˜… 1,199 ยท updated 2026-08-12)* - **[ShiftGrid](https://github.com/BuFuuu/shiftgrid)** ๐ŸŸข โ€” Prompt engine that turns Claude Code into a transparent, human-in-the-loop pentester, structuring engagements through checklists, observations, and notes exposed via an agent-facing API. โ€” **note:** early-stage local Docker application with no built-in authentication; keep its API and UI ports bound to localhost. *(โ˜… 39 ยท updated 2026-08-02)* - **[BugTraceAI](https://github.com/BugTraceAI/BugTraceAI-CLI)** ๐ŸŸขโš ๏ธ โ€” Self-hosted autonomous web-application security scanner that combines reconnaissance, specialist exploit agents, Go fuzzers, and Playwright validation to produce evidence-backed findings. *(BugTraceAI)* โ€” **note:** AGPL-3.0 licensed and beta; use only for authorized testing. Requires an LLM provider or local Ollama endpoint and a substantial Docker/Playwright/Go runtime. *(โ˜… 174 ยท updated 2026-07-30)* - **Related:** [Project overview](https://github.com/BugTraceAI/BugTraceAI) ยท [Web dashboard](https://github.com/BugTraceAI/BugTraceAI-WEB) ยท [Docker launcher](https://github.com/BugTraceAI/BugTraceAI-Launcher) - **[HunterX](https://github.com/nullc0d30/HunterX)** ๐ŸŸข โ€” AI-assisted offensive security engine that orchestrates reconnaissance, security-tool coordination, vulnerability detection and validation, evidence collection, and report-ready findings in one workflow. *(NullC0d3)* โ€” **note:** early-stage and tool-orchestration-heavy; run only in an isolated, authorized assessment environment. *(โ˜… 12 ยท updated 2026-08-16)* - **[MCP Security Hub](https://github.com/FuzzingLabs/mcp-security-hub)** ๐ŸŸข โ€” Collection of Dockerized MCP servers that expose offensive-security tools such as Nmap, Nuclei, SQLMap, Ghidra, Hashcat, and related assessment utilities to MCP-capable assistants. *(FuzzingLabs)* โ€” **note:** orchestration and wrapper collection rather than a security boundary; its containers invoke offensive tools and must be used only in isolated, explicitly authorized environments. *(โ˜… 761 ยท updated 2026-04-08)* - **Related:** [pentest-ai](https://github.com/0xSteph/pentest-ai) ยท [Burp Suite MCP Server](https://github.com/PortSwigger/mcp-server) - **[Forefy .context](https://github.com/forefy/.context)** ๐ŸŸข โ€” MIT-licensed collection of AI-agent Skills, Goals, and Dynamic Workflows for security auditing, authorized penetration testing, and research across web, cloud, blockchain, and defensive workflows. *(Forefy)* โ€” **note:** agent-interpreted skill and workflow bundle rather than a deterministic scanner; includes active offensive procedures, so review and pin content before use and run it only in isolated, authorized assessments. The hosted AI Security Registry is a separate SaaS-backed catalog that publishes commit provenance and project-generated scan/audit metadata. *(โ˜… 133 ยท updated 2026-08-31)* - **Related:** [AI Security Registry](https://forefy.com/asr) ยท [Review methodology](https://forefy.com/asr/docs/supply-chain-defense) ยท [OpenAPI schema](https://forefy.com/docs/openapi.json) - **[RedAmon](https://github.com/samugit83/redamon)** ๐ŸŸขโš ๏ธ โ€” Self-hosted AI pentest framework combining reconnaissance, a Neo4j attack-surface graph, Kali-based tooling, configurable human approvals, and agent-assisted remediation pull requests, with MCP client and server interfaces. โ€” **note:** authorized testing only, on a dedicated isolated host. The heavy Docker stack includes seccomp-unconfined components and services with Docker-socket access; approval gates are configurable, not an unavoidable boundary. Cloud LLMs receive target context; local inference avoids that provider transfer but does not disable external reconnaissance or other configured services. Bundled tools retain separate licenses, including copyleft and WPScan restrictions. *(โ˜… 2,906 ยท updated 2026-10-02)* --- ## AI-Powered Recon & Narrow ML Tools Hyper-specific AI/ML tools for a single offensive-security, recon, or detection step โ€” the subwiz/eyeballer pattern rather than broad autonomous agents. ๐Ÿ… = self-contained trained model or learned model/pattern engine; ๐Ÿ…‘ = LLM wrapper that calls an external API. ### Subdomain & DNS Prediction - **[subwiz](https://github.com/hadriansecurity/subwiz)** ๐ŸŸข โ€” ๐Ÿ… Lightweight nanoGPT model that predicts resolvable subdomains via beam search; model weights are published on Hugging Face. *(Hadrian Security)* *(โ˜… 388 ยท updated 2025-12-18)* - **Related:** [HadrianSecurity/subwiz model](https://huggingface.co/HadrianSecurity/subwiz) - **[regulator](https://github.com/cramppet/regulator)** ๐ŸŸขโš ๏ธ โ€” ๐Ÿ… Learns and ranks regex-like naming patterns from known subdomains to generate likely new candidates. โ€” **note:** no LICENSE file found; treat as source-available until clarified. *(โ˜… 392 ยท updated 2023-02-18)* - **Related:** [subwiz](https://github.com/hadriansecurity/subwiz) ### Recon Screenshot Triage - **[eyeballer](https://github.com/BishopFox/eyeballer)** ๐ŸŸขโš ๏ธ โ€” ๐Ÿ… Convolutional neural network that classifies pentest/recon screenshots (login pages, webapps, old-looking sites, parked domains, and custom 404s) for attack-surface triage. *(Bishop Fox)* โ€” **note:** GPL-3.0 licensed. *(โ˜… 1,290 ยท updated 2024-02-19)* ### Software / Tech Fingerprinting - **[GyoiThon](https://github.com/gyoisamurai/GyoiThon)** ๐ŸŸข๐Ÿ”ฌ โ€” ๐Ÿ… Machine-learning-assisted web intelligence tool that fingerprints products, versions, CVEs, login pages, debug messages, and related web-server signals from HTTP responses. โ€” **note:** historical research reference; Apache-2.0 licensed, but maintenance is low. *(โ˜… 826 ยท updated 2021-06-29)* ### AI-Assisted Fuzzing - **[ffufai](https://github.com/jthack/ffufai)** ๐ŸŸขโš ๏ธ โ€” ๐Ÿ…‘ AI wrapper around the ffuf web fuzzer that suggests file extensions and paths from the target URL and headers using OpenAI or Anthropic models. *(Joseph Thacker)* โ€” **note:** requires an LLM API key; README states MIT but no LICENSE file was found. *(โ˜… 801 ยท updated 2025-12-04)* ### Password / Credential ML - **[PassGPT](https://huggingface.co/javirandor/passgpt-10characters)** ๐Ÿ”ฌโš ๏ธ โ€” ๐Ÿ… GPT-style password model trained on leaked passwords for research on password generation and strength estimation. *(Rando et al.)* *license: CC BY-NC-4.0 ยท access: open 10-char model; 16-char variant gated ยท artifacts: PyTorch/Safetensors.* Research-only / non-commercial use; related code: [javirandor/passgpt](https://github.com/javirandor/passgpt). - **[PassGAN](https://github.com/brannondorsey/PassGAN)** ๐Ÿ”ฌ โ€” ๐Ÿ… WGAN that learns password distributions from leaks to generate guesses; historical reference implementation of the PassGAN paper (MIT). โ€” **note:** historical research reference; not an actively maintained password-auditing product. *(โ˜… 2,009 ยท updated 2018-09-30)* - **[neural_network_cracking](https://github.com/cupslab/neural_network_cracking)** ๐Ÿ”ฌ โ€” ๐Ÿ… RNN password-guessing model from *Fast, Lean, and Accurate: Modeling Password Guessability Using Neural Networks* (USENIX Security 2016); Apache-2.0 licensed. *(CMU CUPS Lab)* โ€” **note:** historical USENIX research implementation, not a maintained password-auditing product. *(โ˜… 243 ยท updated 2018-11-30)* - **Related:** [PassGPT](https://huggingface.co/javirandor/passgpt-10characters) ยท [PassGAN](https://github.com/brannondorsey/PassGAN) ### Phishing Detection (Visual / URL) - **[phishing-url-detection](https://huggingface.co/pirocheto/phishing-url-detection)** ๐ŸŸข โ€” ๐Ÿ… Packaged URL phishing classifier with ONNX and pickle artifacts. *license: MIT ยท access: open ยท artifacts: ONNX, pickle.* Model card recommends ONNX over pickle for safer inference. - **[Phishing Email Detection DistilBERT v2.4.1](https://huggingface.co/cybersectony/phishing-email-detection-distilbert_v2.4.1)** ๐ŸŸข โ€” DistilBERT text-classification model for email and URL phishing detection, trained on a public Hugging Face phishing-email dataset. *license: Apache-2.0 ยท access: open ยท artifacts: Safetensors.* โ€” **note:** strong download signal, but independently verify the very high published metrics before production use. - **[PhishIntention](https://github.com/lindsey98/PhishIntention)** ๐Ÿ”ฌ โ€” ๐Ÿ… Deep-vision phishing detector that infers both brand intention and credential-taking intention from webpage appearance and dynamics (USENIX Security 2022). โ€” **note:** CC0-1.0 licensed. *(โ˜… 263 ยท updated 2026-06-04)* - **[VisualPhishNet](https://github.com/S-Abdelnabi/VisualPhishNet)** ๐Ÿ”ฌโš ๏ธ โ€” ๐Ÿ… Triplet CNN for zero-day phishing detection by visual similarity to trusted websites (ACM CCS 2020). *(CISPA)* โ€” **note:** no LICENSE file found; dataset access is research-request based. *(โ˜… 30 ยท updated 2022-02-09)* ### AI/ML-Assisted Detection Rules & Engines - **[SYARA](https://github.com/nabeelxy/syara)** ๐ŸŸข๐Ÿ”ฌ โ€” ๐Ÿ… Semantic YARA-like rule engine for text and multimodal signals, adding embedding similarity, classifier-backed rules, LLM evaluators, and pHash matching to familiar YARA-style syntax. โ€” **note:** early-stage engine; useful for LLM-era intent signals such as phishing, prompt injection, jailbreaks, hallucination, and disinformation rather than classic binary-only YARA matching. *(โ˜… 18 ยท updated 2026-03-05)* - **[AutoYara](https://github.com/FutureComputing4AI/AutoYara)** ๐ŸŸข๐Ÿ”ฌ โ€” ๐Ÿ… Research implementation of automatic YARA rule generation via biclustering over byte n-grams for malware-family samples. โ€” **note:** Apache-2.0 research code from the ACM AISec 2020 paper; README explicitly says it comes with no warranty or support. *(โ˜… 79 ยท updated 2025-10-08)* - **Related:** [Automatic Yara Rule Generation Using Biclustering](https://arxiv.org/abs/2009.03779) - **[yaraml_rules](https://github.com/sophos/yaraml_rules)** ๐ŸŸข๐Ÿ”ฌ โ€” Research code that trains scikit-learn classifiers on malware and benign corpora, then compiles the learned model into deployable YARA rules. *(Sophos)* โ€” **note:** historical research reference; the maintained value is the ML-to-YARA technique, not a current detection product. *(โ˜… 215 ยท updated 2020-12-18)* - **[RuleLLM](https://github.com/zhang-xr/RuleLLM)** ๐ŸŸข๐Ÿ”ฌ โ€” ๐Ÿ…‘ LLM-assisted malware-rule generator that clusters malicious code samples and produces/refines/validates YARA and Semgrep rules. โ€” **note:** MIT-licensed research prototype; requires OpenAI-compatible API access plus YARA/Semgrep validators. *(โ˜… 12 ยท updated 2025-04-25)* ### Defensive Trained-Model Detectors - **[open-appsec](https://github.com/openappsec/openappsec)** ๐ŸŸข๐ŸŸ โš ๏ธ โ€” ML-based web application and API firewall combining an offline-trained model with environment-specific traffic learning and block/log decisions. **Caveat:** the bundled Apache-2.0 basic model is recommended for monitor/test use; the production-oriented advanced model is a separate portal download with its own terms. Review standalone versus SaaS management, telemetry/data flows, and deployment privileges such as host IPC; detection claims are not a security guarantee. *(โ˜… 1,718 ยท updated 2026-09-01)* - **Related:** [Open-appsec HTTP attachment (requires separate agent/models/configuration; review their terms)](https://github.com/openappsec/attachment) - **[DeepSQLi](https://github.com/gatewayd-io/DeepSQLi)** ๐ŸŸขโš ๏ธ โ€” ๐Ÿ… Deep-learning SQL-injection detector with dataset, trained models, and a Flask Prediction API for GatewayD IDS/IPS integration. *(GatewayD)* โ€” **note:** AGPL-3.0 licensed; defensive detector rather than offensive generator. *(โ˜… 7 ยท updated 2026-02-21)* - **[deepsecrets](https://github.com/ntoskernel/deepsecrets)** ๐ŸŸข โ€” Semantic secrets scanner using lexing/parsing, entropy checks, and hashed-known-secret matching across 500+ languages. โ€” **note:** useful narrow detector, but not a trained ML model. *(โ˜… 172 ยท updated 2026-06-04)* - **[VLAI Vulnerability Severity Classifier](https://huggingface.co/CIRCL/vulnerability-severity-classification-roberta-base)** ๐ŸŸข๐Ÿ”ฌ โ€” RoBERTa-based vulnerability-severity classifier trained on CIRCL vulnerability scores to assist triage before manual CVSS scoring. *(CIRCL)* *license: CC-BY-4.0 ยท access: open ยท artifacts: Safetensors.* --- ## AI-Powered SAST & Secure Code Review Static analysis and secure code review enhanced with LLMs. - **[Symfony Security Auditor](https://github.com/vinceAmstoutz/symfony-security-auditor)** ๐ŸŸข โ€” Symfony-aware source-audit package with attacker/reviewer model passes, confidence filtering, baseline deduplication, and budget-bounded partial results. **Caveat:** reviewer-validated findings are model judgments, not confirmed exploitation. Offline-only mode defaults to false, so hosted providers receive selected source; local inference needs separately provisioned models. Requires PHP 8.3+ and a compatible Symfony version. *(โ˜… 95 ยท updated 2026-10-04)* - **[Clearwing](https://github.com/Lazarus-AI/clearwing)** ๐ŸŸข โ€” Agent-driven source-security research pipeline that ranks files, coordinates vulnerability hunting and validation, optionally generates patches, and exports SARIF, JSON, and Markdown findings. **Caveat:** can compile untrusted targets and execute generated proofs or exploits; use isolated authorized environments and explicit spending limits. Hosted models receive source context; local providers are supported. Optional OTLP/Phoenix tracing can export run metadata. Findings need independent verification, and released packages can lag main. *(โ˜… 1,072 ยท updated 2026-10-02)* - **[Cloudflare Security Audit Skill](https://github.com/cloudflare/security-audit-skill)** ๐ŸŸข โ€” Coding-agent skill for a multi-phase source-security audit with reconnaissance, coverage-led hunting, independent verification, structured findings, and a separate record-validation pass. *(Cloudflare)* โ€” **note:** workflow skill, not a standalone scanner; it requires a capable coding agent with parallel subagents and an OS-enforced sandbox. Results are nondeterministic, and Cloudflare reports that one run finds only about half of the vulnerabilities found across repeated runs. *(โ˜… 18,424 ยท updated 2026-09-14)* - **Sources:** [Cloudflare vulnerability-harness write-up](https://blog.cloudflare.com/build-your-own-vulnerability-harness/) - **[Mantis](https://github.com/google/mantis)** ๐ŸŸข๐Ÿ”ฌ โ€” Security-review skills and an ADK reference harness for vulnerability discovery, triage, reproduction, patching, and deterministic verification gates. *(Google)* โ€” **note:** can generate and execute code while reproducing findings; use only on authorized source in an isolated restricted environment. Model output and proposed findings still require expert verification. *(โ˜… 1,678 ยท updated 2026-09-18)* - **Sources:** [Google Cloud announcement](https://cloud.google.com/blog/products/identity-security/getting-started-with-the-mantis-harness-to-find-and-fix-bugs) - **[VulnHunter (Capital One)](https://github.com/capitalone/VulnHunter)** ๐ŸŸข โ€” Attacker-oriented source-review workflow with forward analysis from exposed entry points, a finding-falsification pass, evidence-backed remediation, and a separate fix-verification flow. *(Capital One)* โ€” **note:** built and optimized for Claude Code with an Opus-class model, so source context is processed under the selected provider's terms and results remain model-dependent. Use only on code you are authorized to assess. *(โ˜… 1,016 ยท updated 2026-08-15)* - **Sources:** [Capital One announcement](https://www.capitalone.com/tech/open-source/) - **[Vulnhuntr](https://github.com/protectai/vulnhuntr)** ๐ŸŸข โ€” Zero-shot vulnerability discovery in Python repos via LLM call-chain analysis; credited with a 0-day RCE in Ragflow. *(Protect AI)* *(โ˜… 2,741 ยท updated 2025-02-06)* - **Related:** [IRIS](https://github.com/iris-sast/iris) - **[deepsec](https://github.com/vercel-labs/deepsec)** ๐ŸŸข โ€” Agent-powered security harness for scanning large codebases with coding agents, resumable parallel runs, custom matchers, and optional revalidation. *(Vercel Labs)* *(โ˜… 7,695 ยท updated 2026-08-13)* - **Related:** [claude-code-security-review](https://github.com/anthropics/claude-code-security-review) ยท [sast-skills](https://github.com/utkusen/sast-skills) - **[Codex Security](https://github.com/openai/codex-security)** ๐ŸŸ  โ€” CLI and TypeScript SDK that use Codex Security to find, validate, and help fix vulnerabilities in a codebase, with scan comparison and containerized bulk-scan support. *(OpenAI)* โ€” **note:** the CLI/SDK are open source, but scans require Codex Security access and, for best results, OpenAI Trusted Access. *(โ˜… 9,911 ยท updated 2026-08-16)* - **Related:** [deepsec](https://github.com/vercel-labs/deepsec) ยท [defending-code-reference-harness](https://github.com/anthropics/defending-code-reference-harness) - **[openยทkritt](https://github.com/Kritt-ai/open-kritt)** ๐ŸŸขโš ๏ธ โ€” Self-hosted platform that orchestrates Codex or Claude Code across focused vulnerability-research workflows, then validates, de-duplicates, ranks, and reports resulting findings. *(Kritt AI)* โ€” **note:** jobs run as root in disposable Docker containers with writable target copies and direct internet access; the stack has no application auth by default and sends scanned code to the configured model provider. Deploy only on a dedicated, access-controlled host and scan authorized targets. *(โ˜… 1,877 ยท updated 2026-08-16)* - **Related:** [deepsec](https://github.com/vercel-labs/deepsec) ยท [Visa Vulnerability Agentic Harness](https://github.com/visa/visa-vulnerability-agentic-harness) - **[Visa Vulnerability Agentic Harness](https://github.com/visa/visa-vulnerability-agentic-harness)** ๐ŸŸข โ€” Agentic SAST pipeline for autonomous vulnerability discovery, exploitability verification, SARIF/Markdown reporting, remediation, and validation using frontier AI models. *(Visa)* โ€” **note:** Apache-2.0; authorized use only. The default `scan` profile can continue into remediation and edit target source files; use `--stop-after s9` for detection-only runs. *(โ˜… 2,556 ยท updated 2026-08-04)* - **Related:** [deepsec](https://github.com/vercel-labs/deepsec) ยท [defending-code-reference-harness](https://github.com/anthropics/defending-code-reference-harness) - **[defending-code-reference-harness](https://github.com/anthropics/defending-code-reference-harness)** ๐ŸŸข โ€” Reference Claude Code skills and autonomous vulnerability-discovery pipeline for threat modeling, static scanning, triage, execution-verified C/C++ memory-bug discovery, reporting, and patch generation. *(Anthropic)* โ€” **note:** official reference implementation, not maintained as a product; the autonomous pipeline executes target code and should be run only inside the documented gVisor sandbox. *(โ˜… 7,270 ยท updated 2026-08-06)* - **Related:** [deepsec](https://github.com/vercel-labs/deepsec) ยท [claude-code-security-review](https://github.com/anthropics/claude-code-security-review) - **[rust-in-peace](https://github.com/scadastrangelove/rust-in-peace)** ๐ŸŸข โ€” Rust-security fork of Anthropic's defending-code reference harness, adding a Rust profile for agentic review of unsafe/FFI memory bugs, panic-DoS, deserialization-trust issues, and Miri/ASan/panic/hang-verified findings. *(Sergey Gordeychik)* โ€” **note:** very new Apache-2.0 fork; autonomous runs execute target code and should use the documented sandbox. *(โ˜… 14 ยท updated 2026-08-11)* - **Related:** [defending-code-reference-harness](https://github.com/anthropics/defending-code-reference-harness) ยท [deepsec](https://github.com/vercel-labs/deepsec) - **[Reproof](https://github.com/scadastrangelove/reproof)** ๐ŸŸข โ€” Multi-language Kimi Code port of rust-in-peace, orchestrating agentic vulnerability discovery, PoC replay in fresh containers, comparison with known bugs, and patch re-testing. *(Sergey Gordeychik)* **Caveat:** newly published, maintainer-authored project; benchmark results are project-reported, not independently reproduced here. The autonomous pipeline executes target code and generated PoCs; use the documented Linux/Docker/gVisor setup on authorized targets. The configured model provider receives source and execution context and may incur costs; review generated patches and sensitive findings before use or disclosure. *(โ˜… 1 ยท updated 2026-10-05)* - **Related:** [rust-in-peace](https://github.com/scadastrangelove/rust-in-peace) ยท [defending-code-reference-harness](https://github.com/anthropics/defending-code-reference-harness) - **[claude-code-security-review](https://github.com/anthropics/claude-code-security-review)** ๐ŸŸ  โ€” Official Claude-based semantic SAST GitHub Action that reviews PR diffs. *(Anthropic)* *(โ˜… 5,866 ยท updated 2026-02-11)* - **[IRIS](https://github.com/iris-sast/iris)** ๐ŸŸข๐Ÿ”ฌ โ€” Neurosymbolic SAST combining LLMs with CodeQL for Java vulnerability detection (MIT). *(โ˜… 413 ยท updated 2026-07-02)* - **[sast-skills](https://github.com/utkusen/sast-skills)** ๐ŸŸข โ€” Agent skills that turn AI coding assistants into a multi-agent SAST scanner. *(โ˜… 1,276 ยท updated 2026-04-08)* - **Related:** [Fraim](https://github.com/fraim-dev/fraim) ยท [llm-sast-scanner](https://github.com/SunWeb3Sec/llm-sast-scanner) ยท [AWS MCP code-scanning sample (external scanners; not MCP protocol analysis)](https://github.com/aws-samples/sample-mcp-security-scanner) - **[llm-sast-scanner](https://github.com/SunWeb3Sec/llm-sast-scanner)** ๐ŸŸข โ€” SAST skill for AI coding agents with structured source-to-sink analysis across 34 vulnerability classes. *License: MIT stated in README.* *(โ˜… 274 ยท updated 2026-04-07)* - **Related:** [sast-skills](https://github.com/utkusen/sast-skills) - **[sast-ai-workflow](https://github.com/RHEcosystemAppEng/sast-ai-workflow)** ๐ŸŸข โ€” LangGraph workflow for reviewing static-analysis findings, reducing false positives, and producing vulnerability review output. *(Red Hat Ecosystem AppEng)* *(โ˜… 20 ยท updated 2026-06-29)* - **Related:** [seclab-taskflow-agent](https://github.com/GitHubSecurityLab/seclab-taskflow-agent) ยท [Fraim](https://github.com/fraim-dev/fraim) - **[llm-security-scanner](https://github.com/iknowjason/llm-security-scanner)** ๐ŸŸขโš ๏ธ โ€” LLM-powered code scanner that opens GitHub issues for findings. *(โ˜… 22 ยท updated 2025-04-02)* - **[Trail of Bits Skills](https://github.com/trailofbits/skills)** ๐ŸŸขโš ๏ธ โ€” Claude Code- and Codex-compatible security workflow skills for code review, differential review, false-positive analysis, supply-chain checks, GitHub Actions auditing, Semgrep rule generation, and vulnerability research. *(Trail of Bits)* โ€” **note:** reusable agent workflow instructions rather than a standalone deterministic scanner; share-alike terms apply to adapted material. *(โ˜… 6,624 ยท updated 2026-08-14)* - **Related:** [sast-skills](https://github.com/utkusen/sast-skills) ยท [Claude-BugHunter](https://github.com/elementalsouls/Claude-BugHunter) ยท [Claude Code cybersecurity skills (host permissions govern execution)](https://github.com/Masriyan/Claude-Code-CyberSecurity-Skill) ยท [Defensive Claude Code playbooks (host/model-dependent; production claims unverified)](https://github.com/GoldenWing-360/claude-security-skills) - **[OpenHack](https://github.com/hadriansecurity/OpenHack)** ๐ŸŸข โ€” File-based source-guided white-box security-review workspace that orchestrates agents through reconnaissance, vulnerability hunting, validation, evidence capture, and reporting. *(Hadrian Security)* โ€” **note:** requires an external coding harness/model and can consume substantial model tokens; OpenHack provides workflow state and review artifacts, not sandboxing or an execution-security boundary. Use only on authorized targets. *(โ˜… 730 ยท updated 2026-06-01)* - **Related:** [deepsec](https://github.com/vercel-labs/deepsec) ยท [rust-in-peace](https://github.com/scadastrangelove/rust-in-peace) - **[Buttercup](https://github.com/trailofbits/buttercup)** ๐ŸŸข๐Ÿ”ฌโš ๏ธ โ€” Multi-component cyber reasoning system for finding, validating, and patching software vulnerabilities with coordinated agent workflows. *(Trail of Bits)* โ€” **note:** AGPL-3.0 research/competition system rather than a lightweight scanner; deployment uses multiple services, Docker, and configured model providers. *(โ˜… 1,676 ยท updated 2026-08-10)* - **Related:** [OpenHack](https://github.com/hadriansecurity/OpenHack) ยท [Visa Vulnerability Agentic Harness](https://github.com/visa/visa-vulnerability-agentic-harness) --- ## AI-Powered Threat Modeling Architecture-level threat model generation and design-phase risk analysis driven by LLM reasoning. - **[ThreatForest](https://github.com/aws-samples/sample-agentic-attack-tree-generator)** ๐ŸŸข๐Ÿ”ฌ โ€” Strands-based threat-modeling pipeline that analyzes repositories, generates attack trees and mitigations, and maps attack steps to ATT&CK and other TTP frameworks through an embedding service. *(AWS Samples)* **Caveat:** sends project context to configured model providers and requires separately licensed embedding weights; optional Langfuse export defaults off. Keep its unauthenticated API on loopback. Generated paths and embedding-based TTP mappings require expert review, not automatic acceptance as demonstrated attack paths. *(โ˜… 80 ยท updated 2026-08-01)* - **[tachi](https://github.com/davidmatousek/tachi)** ๐ŸŸข โ€” Threat modeling and AI-reasoning vulnerability detection harness for Claude Code that dispatches 14 specialized threat agents (6 STRIDE, 5 LLM, 3 agentic) against an architecture description in Mermaid, C4, PlantUML, ASCII, or free text, producing SARIF 2.1.0 for code scanning, MAESTRO seven-layer classification, attack trees, CVSS-aligned composite risk scores, compensating-controls analysis of the target codebase, and a PDF report. *(David Matousek)* โ€” **note:** runs inside Claude Code; architecture descriptions are processed by the configured Claude model. *(โ˜… 89 ยท updated 2026-08-13)* - **Related:** [STRIDE GPT](https://github.com/mrwadams/stride-gpt) ยท [defending-code-reference-harness](https://github.com/anthropics/defending-code-reference-harness) - **[STRIDE GPT](https://github.com/mrwadams/stride-gpt)** ๐ŸŸข โ€” LLM-powered threat modeling tool that generates STRIDE threat models, attack trees, data flow diagrams, DREAD risk scores, mitigations, and Gherkin test cases from application descriptions, architecture diagrams, or codebases (agentic analysis mode), with OWASP LLM Top 10 and Agentic (ASI) coverage, MITRE ATT&CK/ATLAS mapping, Markdown/JSON/SARIF/HTML output, and broad LLM provider support via LiteLLM including local hosting. *(Matthew Adams)* *(โ˜… 1,101 ยท updated 2026-08-12)* - **Related:** [tachi](https://github.com/davidmatousek/tachi) --- ## LLM-Driven Fuzzing Two families: (a) LLMs generating harnesses/targets for traditional fuzzing, and (b) fuzzing the LLM itself. ### Harness / target generation - **[seclab-taskflows-fuzzing](https://github.com/GitHubSecurityLab/seclab-taskflows-fuzzing)** ๐ŸŸข โ€” Taskflow package for LLM-assisted C/C++ fuzz-harness generation and repair, AFL++ campaigns, coverage feedback, and crash triage. *(GitHub Security Lab)* **Caveat:** requires a configured model provider and a Linux build/fuzzing toolchain; provider-backed analysis can disclose source code. Host-shell and file-write tools are not a sandbox, and setup can install system tools. Use a disposable isolated environment for authorized targets. *(โ˜… 20 ยท updated 2026-09-21)* - **Related:** [seclab-taskflow-agent](https://github.com/GitHubSecurityLab/seclab-taskflow-agent) - **[Ultrafuzz](https://github.com/monad-developers/ultrafuzz)** ๐ŸŸข โ€” Agentic smart-contract fuzzing orchestrator that generates tests, coordinates analysis campaigns, and collects findings and reports. *(Monad Foundation)* **Caveat:** agents deliberately bypass permission and sandbox approvals and can access host files and the network. Run only on an ephemeral isolated VM without unrelated credentials; review a published release rather than the default unstable branch. Model/provider configuration determines cost and data disclosure; the harness is not an isolation boundary. *(โ˜… 98 ยท updated 2026-10-02)* - **[PromeFuzz](https://github.com/pvz122/PromeFuzz)** ๐ŸŸข๐Ÿ”ฌ โ€” Research framework for C/C++ fuzz-harness generation using code metadata, documentation, API relationships, iterative compilation repair, and crash analysis (CCS 2025). **Caveat:** requires target build metadata, a native LLVM/Clang toolchain, and configured generation/embedding models. Hosted providers receive source and documentation context; generated harnesses and target code execute during campaigns. Use isolated authorized environments and do not generalize the paper's coverage or bug-finding results to arbitrary targets. *(โ˜… 61 ยท updated 2026-07-30)* - **[oss-fuzz-gen](https://github.com/google/oss-fuzz-gen)** ๐ŸŸข โ€” LLM-driven fuzz-harness generation for OSS-Fuzz; reported 26 real vulnerabilities (incl. CVE-2024-9143 in OpenSSL). *(Google)* *(โ˜… 1,430 ยท updated 2026-03-02)* - **[PromptFuzz](https://github.com/PromptFuzz/PromptFuzz)** ๐ŸŸข๐Ÿ”ฌโš ๏ธ โ€” LLM-mutated prompts to generate fuzz drivers for C/C++ libraries (Rust). *(โ˜… 342 ยท updated 2026-05-15)* - **[Fuzz4All](https://github.com/fuzz4all/fuzz4all)** ๐ŸŸข๐Ÿ”ฌ โ€” "Universal" LLM-based fuzzer across compilers/languages (ICSE 2024). *(โ˜… 336 ยท updated 2025-08-11)* - **[ChatAFL](https://github.com/ChatAFLndss/ChatAFL)** ๐ŸŸข๐Ÿ”ฌ โ€” LLM-guided protocol fuzzing extending AFLNet (NDSS'24). *(โ˜… 392 ยท updated 2025-06-20)* - **[TitanFuzz](https://github.com/ise-uiuc/TitanFuzz)** ๐ŸŸข๐Ÿ”ฌโš ๏ธ โ€” First LLM-based fuzzer for PyTorch/TensorFlow (ISSTA'23). *(โ˜… 94 ยท updated 2023-09-10)* ### Fuzzing the LLM - **[LLMFuzzer](https://github.com/mnns/LLMFuzzer)** ๐ŸŸข๐Ÿ”ฌ โ€” First open-source fuzzing framework for LLM API integrations. โ€” **note:** historical research reference; maintenance appears low compared with current LLM security scanners. *(โ˜… 377 ยท updated 2024-02-12)* - **[ps-fuzz](https://github.com/prompt-security/ps-fuzz)** ๐ŸŸ  โ€” System-prompt hardening fuzzer; 16 attacks ร— 16 providers. *(Prompt Security)* *(โ˜… 703 ยท updated 2026-02-16)* - **[FuzzyAI](https://github.com/cyberark/FuzzyAI)** ๐ŸŸ  โ€” Automated LLM fuzzer for jailbreaks/prompt injection. *(CyberArk)* *(โ˜… 1,563 ยท updated 2026-02-06)* - **[spikee](https://github.com/ReversecLabs/spikee)** ๐ŸŸข โ€” Prompt-injection evaluation and exploitation kit with dataset generation, Burp integration, and pluggable judges. *(ReversecLabs / WithSecure)* *(โ˜… 263 ยท updated 2026-09-11)* - **Related:** [promptmap](https://github.com/utkusen/promptmap) - **[promptmap](https://github.com/utkusen/promptmap)** ๐ŸŸขโš ๏ธ โ€” Prompt-injection scanner for custom LLM applications in white-box and black-box modes; GPL-3.0 licensed. *(โ˜… 1,250 ยท updated 2025-12-01)* - **Related:** [spikee](https://github.com/ReversecLabs/spikee) - **[ai-prompt-fuzzer](https://github.com/PortSwigger/ai-prompt-fuzzer)** ๐ŸŸข โ€” Burp Suite extension fuzzing GenAI/LLM prompts. *(PortSwigger)* *(โ˜… 36 ยท updated 2025-09-04)* --- ## Threat Intelligence AI/LLM tooling for CTI gathering, IOC/TTP extraction, and analysis. - **[ZettelForge](https://github.com/ThreatRecall/zettelforge)** ๐ŸŸข โ€” CTI investigation memory with entity and IOC extraction, graph and vector retrieval, local storage, and an MCP interface. **Caveat:** core storage and retrieval are local, but embedding models require an initial download and optional remote LLM providers receive analysis content. Rich extraction and synthesis depend on the configured model; synthesis without an LLM is a placeholder, not a completed analysis. Generated links and attribution need analyst validation. *(โ˜… 64 ยท updated 2026-07-10)* - **[VoidAccess](https://github.com/KatrielMoses/voidaccess)** ๐ŸŸข โ€” OSINT and CTI pipeline combining source collection, IOC/entity extraction, enrichment, relationship analysis, and intelligence exports. **Caveat:** collection contacts external and potentially hostile sources; Tor is not an anonymity guarantee. Cloud model adapters disclose submitted content, while local models are configurable. Neural embeddings are optional and the SHA-256 fallback is not semantic ML; generated attribution and detection rules require analyst review. *(โ˜… 770 ยท updated 2026-08-04)* - **[cti-skills](https://github.com/Liberty91LTD/cti-skills)** ๐ŸŸข โ€” CTI workflow skills and integration clients for source assessment, IOC investigation, ATT&CK mapping, detection-rule drafting, and MISP/OpenCTI workflows. *(Liberty91)* **Caveat:** skills guide a host agent rather than provide an independent detection engine. Live integrations need provider credentials and can disclose intelligence; some clients can create, modify, or delete records. Confirmation instructions are not enforced authorization, so scope credentials and review writes separately. *(โ˜… 25 ยท updated 2026-09-30)* - **[trs](https://github.com/deadbits/trs)** ๐ŸŸข โ€” LLM + ChromaDB tool to summarize threat reports and extract MITRE TTPs and IOCs. *(โ˜… 10 ยท updated 2023-11-15)* - **[TI-Mindmap-GPT](https://github.com/format81/TI-Mindmap-GPT)** ๐ŸŸข โ€” Streamlit app: AI summaries, mindmaps, IOC/TTP extraction, and ATT&CK Navigator layers. *(โ˜… 111 ยท updated 2026-02-16)* - **[aiocrioc](https://github.com/referefref/aiocrioc)** ๐ŸŸข โ€” LLM + OCR IOC extraction (pulls IOCs from images/PDFs). *(โ˜… 38 ยท updated 2024-12-04)* - **[ThreatIngestor](https://github.com/InQuest/ThreatIngestor)** ๐ŸŸข โ€” Extracts/aggregates IOCs from feeds; integrates with MISP/ThreatKB (pairs well with LLM post-processing). *(โ˜… 922 ยท updated 2026-05-26)* - **[IATelligence](https://github.com/fr0gger/IATelligence)** ๐ŸŸข โ€” Explains imported Windows APIs in PE files via GPT and maps to MITRE ATT&CK. *(โ˜… 384 ยท updated 2022-12-09)* - **Related:** [MCP_Security](https://github.com/fr0gger/MCP_Security) - **[MCP_Security](https://github.com/fr0gger/MCP_Security)** ๐ŸŸขโš ๏ธ โ€” MCP server (ORKL) for querying the ORKL threat-intel API. *(โ˜… 51 ยท updated 2025-01-22)* - **Related:** [IATelligence](https://github.com/fr0gger/IATelligence) - **[threat-intelligence-cti-analysis](https://github.com/AnandBinuArjun/threat-intelligence-cti-analysis)** ๐ŸŸข โ€” NLP/LLM pipeline for IOC extraction, MITRE ATT&CK mapping, and knowledge-graph generation from unstructured CTI. *(โ˜… 4 ยท updated 2025-11-03)* - **Related:** [soctalk](https://github.com/soctalk/soctalk) - **[CTINexus](https://github.com/peng-gao-lab/ctinexus)** ๐ŸŸข๐Ÿ”ฌ โ€” LLM-assisted framework for data-efficient extraction of cyber-threat intelligence and construction of structured cybersecurity knowledge graphs from unstructured reports. โ€” **note:** ships tests, releases, Docker configuration, and a Python package; extraction workflows require a configured supported LLM provider. *(โ˜… 85 ยท updated 2026-02-25)* - **Related:** [threat-intelligence-cti-analysis](https://github.com/AnandBinuArjun/threat-intelligence-cti-analysis) ยท [CTIBench](https://github.com/maveryn/cti-bench) - **[CTIBench](https://github.com/maveryn/cti-bench)** ๐Ÿ”ฌโš ๏ธ โ€” NeurIPS 2024 Spotlight benchmark with 4,610 examples across CTI knowledge, CWE root-cause mapping, CVSS prediction, ATT&CK technique extraction, and threat-actor attribution. โ€” **note:** non-commercial research benchmark; the repository publishes data, evaluation notebooks, model outputs, and raw logs rather than a production CTI service. *(โ˜… 92 ยท updated 2026-05-07)* - **Related:** [CTINexus](https://github.com/peng-gao-lab/ctinexus) ยท [CTI-BERT](https://huggingface.co/ibm-research/CTI-BERT) - **[CTI-BERT](https://huggingface.co/ibm-research/CTI-BERT)** ๐ŸŸข๐Ÿ”ฌ โ€” BERT model pretrained from scratch on a large cybersecurity text corpus for downstream CTI extraction, classification, and question-answering tasks. *(IBM Research)* *license: Apache-2.0 ยท access: open ยท artifacts: PyTorch.* --- ## Log Analysis / SIEM / SOC Triage AI agents for SOC alert triage, investigation, and incident response. - **[AI-SOC-Agent](https://github.com/M507/ai-soc-agent)** ๐ŸŸข โ€” Black Hat 2025 MCP server exposing security-investigation tools (ELK, IRIS). *(โ˜… 47 ยท updated 2025-12-28)* - **[soctalk](https://github.com/soctalk/soctalk)** ๐ŸŸข โ€” LangGraph SOC automation agent with MCP integrations for Wazuh, Cortex, TheHive, and MISP plus mock-agent test lab. *(โ˜… 78 ยท updated 2026-08-12)* - **Related:** [SigmaOptimizer](https://github.com/YusukeJustinNakajima/SigmaOptimizer) - **[Vigil SOC](https://github.com/Vigil-SOC/vigil)** ๐ŸŸข โ€” Open-source AI SOC with readable Python agents, Markdown playbooks, and MCP integrations for triage, investigation, hunting, response, reporting, and forensics. *(Vigil SOC)* *(โ˜… 247 ยท updated 2026-08-17)* - **Related:** [soctalk](https://github.com/soctalk/soctalk) - **[agentic-soc-platform](https://github.com/FunnyWolf/agentic-soc-platform)** ๐ŸŸข โ€” Agentic SOC platform (LangGraph/Dify) with local-LLM support. *(โ˜… 1,144 ยท updated 2026-08-05)* - **[SigmAIQ](https://github.com/AttackIQ/SigmAIQ)** ๐ŸŸขโš ๏ธ โ€” pySigma wrapper and LangChain toolkit for automatic Sigma rule creation and translation; LGPL-2.1 licensed. *(AttackIQ)* *(โ˜… 97 ยท updated 2025-11-03)* - **Related:** [SigmaOptimizer](https://github.com/YusukeJustinNakajima/SigmaOptimizer) - **[RulePilot](https://github.com/LLM4SOC-Topic/RulePilot)** ๐ŸŸข๐Ÿ”ฌ โ€” LLM-powered security-rule generation agent for Splunk, Microsoft Sentinel, and Elastic, with field detection from log samples, multi-stage refinement, and cross-platform rule conversion. โ€” **note:** MIT-licensed ICSE 2026 research prototype; requires an OpenAI API key. *(โ˜… 15 ยท updated 2025-10-20)* - **Related:** [SigmAIQ](https://github.com/AttackIQ/SigmAIQ) ยท [SigmaOptimizer](https://github.com/YusukeJustinNakajima/SigmaOptimizer) - **[SOCGPT](https://github.com/Ninadjos/SOCGPT-AI-Powered-SOC-Assistant)** ๐ŸŸข โ€” LLM log summarization, severity triage, MITRE mapping, and Q&A. *(โ˜… 7 ยท updated 2025-06-11)* - **[AttackGen](https://github.com/mrwadams/attackgen)** ๐ŸŸข โ€” LLM-driven incident-response scenario generator using MITRE ATT&CK + ATLAS. *(โ˜… 1,234 ยท updated 2026-08-13)* - **[Google Security Operations and Threat Intelligence MCP Server](https://github.com/google/mcp-security)** ๐ŸŸข โ€” MCP servers and packages that let MCP clients access Google Security Operations, SOAR, Google Threat Intelligence, and Security Command Center for investigation, hunting, and security automation workflows. *(Google Cloud)* โ€” **note:** integration layer rather than an MCP-defense tool; requires Google credentials and access to the connected security products/services. *(โ˜… 517 ยท updated 2026-04-29)* - **Related:** [MCP_Security](https://github.com/fr0gger/MCP_Security) ยท [Vigil SOC](https://github.com/Vigil-SOC/vigil) - **[ExCyTIn-Bench (SecRL)](https://github.com/microsoft/SecRL)** ๐ŸŸข๐Ÿ”ฌ โ€” ICML 2026 benchmark for evaluating LLM agents on cyber-threat investigation and threat hunting through security question-answering over eight anonymized incident databases. *(Microsoft)* โ€” **note:** evaluation requires model-provider credentials, Dockerized MySQL incident databases, and roughly 10 GB for the standard eight-container setup (up to 33 GB for the combined database). *(โ˜… 143 ยท updated 2026-08-03)* - **Related:** [CTIBench](https://github.com/maveryn/cti-bench) ยท [Google Security Operations MCP](https://github.com/google/mcp-security) --- ## Reverse Engineering LLM-assisted binary analysis and traffic inspection. - **[IDAssist](https://github.com/symgraph/IDAssist)** ๐ŸŸข โ€” Native IDA Pro plugin for LLM-assisted function explanation, symbol renaming, binary knowledge graphs, document retrieval, and tool-driven analysis through local or hosted model providers. **Caveat:** requires separately licensed IDA Pro; hosted models receive analysis context, optional SymGraph sharing exports symbols and graph data, and MCP tools can modify the IDB or start local processes. Findings require analyst verification. *(โ˜… 729 ยท updated 2026-09-20)* - **[DeepZero](https://github.com/416rehman/DeepZero)** ๐ŸŸข๐Ÿ”ฌ โ€” Resumable Windows driver research framework combining Ghidra decompilation, static analysis, and optional LLM assessment through YAML pipelines. โ€” **note:** the Windows-driver workflow requires an operator-provided Ghidra installation and optional LLM-provider credentials; analyze only drivers you are authorized to research. *(โ˜… 712 ยท updated 2026-09-10)* - **[Gepetto](https://github.com/JusticeRage/Gepetto)** ๐ŸŸข โ€” IDA Pro plugin: GPT adds comments and meaningful variable names. *(โ˜… 3,457 ยท updated 2026-08-15)* - **[ida-pro-mcp](https://github.com/mrexodia/ida-pro-mcp)** ๐ŸŸข โ€” MCP bridge for IDA Pro exposing decompile, disassemble, xref, rename, and debugging workflows to LLM clients. *(โ˜… 11,393 ยท updated 2026-08-17)* - **[GhidraMCP](https://github.com/LaurieWired/GhidraMCP)** ๐ŸŸข โ€” MCP server exposing Ghidra reverse-engineering ops to any MCP-capable LLM. *(โ˜… 9,805 ยท updated 2025-06-23)* - **Related:** [GhidrOllama](https://github.com/lr-m/GhidrOllama) ยท [OGhidra](https://github.com/llnl/OGhidra) - **[ReVa](https://github.com/cyberkaida/reverse-engineering-assistant)** ๐ŸŸข โ€” Ghidra-focused reverse-engineering assistant with MCP support, Claude Skills integration, and long-form analysis workflows. *(โ˜… 805 ยท updated 2026-07-28)* - **Related:** [GhidraMCP](https://github.com/LaurieWired/GhidraMCP) ยท [GhidrAssistMCP](https://github.com/symgraph/GhidrAssistMCP) - **[GhidrAssistMCP](https://github.com/symgraph/GhidrAssistMCP)** ๐ŸŸข โ€” Native Ghidra MCP extension with broad tool coverage, headless support, and security-sensitive tool gating. *(โ˜… 719 ยท updated 2026-08-02)* - **Related:** [ReVa](https://github.com/cyberkaida/reverse-engineering-assistant) ยท [ghidra-mcp](https://github.com/bethington/ghidra-mcp) - **[ghidra-mcp](https://github.com/bethington/ghidra-mcp)** ๐ŸŸข โ€” Ghidra MCP server with large tool coverage, GUI plugin, headless server, and lazy tool loading. *(โ˜… 3,333 ยท updated 2026-08-12)* - **Related:** [GhidrAssistMCP](https://github.com/symgraph/GhidrAssistMCP) - **[GhidrOllama](https://github.com/lr-m/GhidrOllama)** ๐ŸŸขโš ๏ธ โ€” Ghidra script using the Ollama API for function analysis/renaming. *(โ˜… 154 ยท updated 2024-11-29)* - **Related:** [OGhidra](https://github.com/llnl/OGhidra) ยท [GhidraMCP](https://github.com/LaurieWired/GhidraMCP) - **[GhidraGPT](https://github.com/weirdmachine64/GhidraGPT)** ๐ŸŸข โ€” Ghidra plugin that integrates LLMs for automated code refactoring and analysis. *(โ˜… 658 ยท updated 2026-07-22)* - **Related:** [GhidraMCP](https://github.com/LaurieWired/GhidraMCP) ยท [ReVa](https://github.com/cyberkaida/reverse-engineering-assistant) - **[LLM4Decompile](https://github.com/albertan017/LLM4Decompile)** ๐ŸŸข๐Ÿ”ฌโš ๏ธ โ€” Research project for binary-to-C decompilation with LLMs; code is MIT, but model weights use a more restrictive license. *(โ˜… 6,965 ยท updated 2026-02-12)* - **[x64dbg_mcp](https://github.com/bromoket/x64dbg_mcp)** ๐ŸŸข โ€” MCP server exposing x64dbg debugging and reverse-engineering operations to AI clients. *(โ˜… 104 ยท updated 2026-06-08)* - **[binaryninja-mcp](https://github.com/MCPPhalanx/binaryninja-mcp)** ๐ŸŸข โ€” MCP server for Binary Ninja-assisted reverse engineering. *(โ˜… 47 ยท updated 2025-05-13)* - **[OGhidra](https://github.com/llnl/OGhidra)** ๐ŸŸข โ€” Natural-language Ghidra analysis via Ollama. *(Lawrence Livermore National Lab)* *(โ˜… 410 ยท updated 2026-08-14)* - **Related:** [GhidrOllama](https://github.com/lr-m/GhidrOllama) ยท [GhidraMCP](https://github.com/LaurieWired/GhidraMCP) - **[ghidra_tools (G-3PO)](https://github.com/tenable/ghidra_tools)** ๐ŸŸข โ€” Ghidra plugin for AI-assisted decompiled-code analysis. *(Tenable)* *(โ˜… 312 ยท updated 2023-05-10)* - **[gpt-wpre](https://github.com/moyix/gpt-wpre)** ๐Ÿ”ฌ โ€” Whole-program reverse engineering with GPT-3. *(โ˜… 407 ยท updated 2022-12-31)* - **[burpgpt](https://github.com/aress31/burpgpt)** ๐ŸŸข โ€” Burp Suite extension integrating GPT for passive scanning. *(โ˜… 2,347 ยท updated 2024-06-09)* - **Related:** [Burp-extension-for-GPT](https://github.com/tenable/Burp-extension-for-GPT) - **[Burp-extension-for-GPT](https://github.com/tenable/Burp-extension-for-GPT)** ๐ŸŸข โ€” Burp extension to analyze HTTP traffic with GPT. *(Tenable)* *(โ˜… 116 ยท updated 2023-05-01)* - **Related:** [burpgpt](https://github.com/aress31/burpgpt) - **[REA](https://github.com/morluto/rea)** ๐ŸŸข โ€” Local CLI and MCP toolkit for agent-assisted reverse engineering of native binaries, managed PE/CLI files, JavaScript/Electron apps, and browser runtimes, using Hopper or an operator-provided Ghidra installation. โ€” **note:** deep native analysis uses separately licensed Hopper or operator-installed Ghidra. Setup can modify agent MCP registrations after interactive approval, and dynamic providers run with the current user's permissions; analyze only authorized artifacts. *(โ˜… 346 ยท updated 2026-08-14)* - **Related:** [GhidraMCP](https://github.com/LaurieWired/GhidraMCP) ยท [ReVa](https://github.com/cyberkaida/reverse-engineering-assistant) - **[Open-ReverseLab](https://github.com/LING71671/open-reverselab)** ๐ŸŸขโš ๏ธ โ€” Agent-native reverse-engineering workspace and MCP server combining Ghidra headless analysis with Frida, x64dbg, Rizin, YARA, and a runnable CTF/APK/PE knowledge base. โ€” **note:** the repository additionally asserts a mandatory disclaimer with use, redistribution, derivative-work, and indemnification terms beyond GPL-3.0. Optional installers download and execute numerous third-party offensive tools, sometimes from latest or moving upstream refs without checksums; review them before use and run only in an isolated, authorized lab. *(โ˜… 1,101 ยท updated 2026-09-08)* - **Related:** [GhidraMCP](https://github.com/LaurieWired/GhidraMCP) ยท [ReVa](https://github.com/cyberkaida/reverse-engineering-assistant) --- ## LLM Red-Teaming & Guardrails Tools for attacking and defending LLM applications themselves. ### Scanners, Evals & Guardrails - **[DrAttack](https://github.com/xirui-li/DrAttack)** ๐ŸŸข๐Ÿ”ฌ โ€” Research jailbreak implementation that decomposes requests, searches prompt substitutions, and evaluates target-model responses. **Caveat:** historical defaults use the old OpenAI ChatCompletion API and gpt-3.5-turbo-0613, requiring compatible dependencies or adaptation. Refusal-prefix and model-judge heuristics do not prove attack success; configured generation and embedding APIs receive research inputs. *(โ˜… 69 ยท updated 2026-09-12)* - **[RL-Hammer (rl-injector)](https://github.com/facebookresearch/rl-injector)** ๐Ÿ”ฌโš ๏ธ โ€” Research training pipeline for prompt-injection attacker models, with InjecAgent tool-response substitution and target-based reward computation. *(Meta)* **Caveat:** noncommercial code; GPU resources, models, and data are separate dependencies. Hosted targets receive attack contexts, and the reviewed tokenizer path enables trust_remote_code, so model repositories require separate review. Published outcomes are not independently reproduced here. *(โ˜… 57 ยท updated 2025-11-24)* - **[Prompt Injection as Role Confusion](https://github.com/role-confusion/prompt-injection-as-role-confusion)** ๐ŸŸข๐Ÿ”ฌโš ๏ธ โ€” Research code for probing role representations and studying prompt injection through confusion between instruction and data roles. **Caveat:** the root license contains the MIT permission text without a copyright header. Experiments require compatible model weights and GPU resources; some configurations are marked untested, and API-backed paths disclose inputs to providers. This is not a production guardrail. *(โ˜… 134 ยท updated 2026-05-31)* - **[GuidedBench](https://github.com/SproutNan/GuidedBench)** ๐ŸŸข๐Ÿ”ฌ โ€” Guideline-grounded jailbreak evaluator that scores saved model responses against case-specific criteria using configurable judge backends. **Caveat:** the bundled dataset integration is gated and needs approval/HF_TOKEN; dataset and model terms are separate from the code license. Judge models can disagree or fail, and hosted backends receive evaluated content; this is an evaluator, not an attack generator or ground-truth oracle. *(โ˜… 32 ยท updated 2026-08-31)* - **[OpenART](https://github.com/AI45Lab/OpenART)** ๐ŸŸข๐Ÿ”ฌโš ๏ธ โ€” Research benchmark and Docker-based runtime for red-teaming tool-using agents through adversarial changes to stateful, long-running task environments. **Caveat:** bundled runnable tasks are a subset of the paper's broader evaluation, not the entire headline corpus. Model-backed attackers, targets, and judges require configured providers and can disclose evaluation content. Executable evaluators, container mounts, and network access require disposable authorized environments; this is not a production guardrail or independently verified containment layer. *(โ˜… 229 ยท updated 2026-10-03)* - **[ZeroLeaks](https://github.com/ZeroLeaks/zeroleaks)** ๐ŸŸ โš ๏ธ โ€” Source-available TypeScript library and CLI for model-based system-prompt extraction and injection testing with adaptive attacks and configurable providers. **Caveat:** FSL restricts competing commercial use before the future Apache conversion. The public package tests a model plus system prompt, not a deployed application's actual tools, memory, or side effects; hosted deployed-agent testing is separate and paid. Hosted providers receive prompt material and require credentials, while local compatible endpoints are supported. Model-graded results are not proof of exploitation. *(โ˜… 733 ยท updated 2026-09-26)* - **[RAMPART](https://github.com/microsoft/RAMPART)** ๐ŸŸข โ€” Pytest-native framework for repeatable adversarial and benign safety regression tests against AI agents, with statistical trials and evaluators for responses, tool calls, and external side effects. *(Microsoft)* โ€” **note:** test framework built on PyRIT, not a runtime protection layer; users must supply target adapters and model credentials, and passing its scenarios does not establish production safety beyond the tested behaviors. *(โ˜… 414 ยท updated 2026-09-18)* - **Sources:** [Microsoft Security announcement](https://www.microsoft.com/en-us/security/blog/2026/05/20/introducing-rampart-and-clarity-open-source-tools-to-bring-safety-into-agent-development-workflow/) - **Related:** [PyRIT](https://github.com/microsoft/PyRIT) - **[Agent OPFOR](https://github.com/KeyValueSoftwareSystems/agent-opfor)** ๐ŸŸข โ€” Adversary-emulation toolkit for AI agents, LLM applications, and MCP servers with multi-turn attack templates, tool and memory tests, trace-aware evaluation, CLI, SDK, browser, MCP, and skill interfaces. *(KeyValue Software Systems)* โ€” **note:** attacker and judge workflows depend on configured LLM providers and add inference cost; trace integrations can expose internal agent data to the selected observability service. Use only against systems you own or are authorized to test. *(โ˜… 580 ยท updated 2026-09-13)* - **[NuGuard](https://github.com/NuGuardAI/nuguard)** ๐ŸŸข โ€” Generates an AI-SBOM, statically analyzes agentic applications, red-teams live targets for prompt injection/tool misuse/data exfiltration, and validates behavioral policy compliance with SARIF, JSON, and Markdown reports. *(NuGuard AI)* โ€” **note:** beta project; live red-team scans actively probe the target and require authorization. LLM-assisted features need provider credentials, and the optional NuGuard.ai hosted offering adds commercial features. *(โ˜… 23 ยท updated 2026-08-14)* - **Related:** [garak](https://github.com/NVIDIA/garak) ยท [Medusa](https://github.com/Pantheon-Security/medusa) - **[garak](https://github.com/NVIDIA/garak)** ๐ŸŸข โ€” The LLM vulnerability scanner โ€” probes for prompt injection, jailbreaks, data leakage, and more. *(NVIDIA)* *(โ˜… 8,834 ยท updated 2026-08-14)* - **Related:** [PyRIT](https://github.com/microsoft/PyRIT) ยท [promptfoo](https://github.com/promptfoo/promptfoo) - **[PyRIT](https://github.com/microsoft/PyRIT)** ๐ŸŸข โ€” Python Risk Identification Tool; battle-tested across 100+ GenAI red-team operations. *(Microsoft)* *(โ˜… 4,316 ยท updated 2026-08-14)* - **Related:** [Arcanum prompt-injection taxonomy (CC BY 4.0 reference, not a detector)](https://github.com/Arcanum-Sec/arc_pi_taxonomy) - **[promptfoo](https://github.com/promptfoo/promptfoo)** ๐ŸŸข โ€” LLM eval + red-teaming/pentesting CLI with 50+ attack plugins (MIT). *Note: OpenAI announced an acquisition agreement in March 2026; remains MIT-licensed โ€” track governance.* *(โ˜… 24,302 ยท updated 2026-08-17)* - **Related:** [Promptfoo CI Action (checkout evaluation, not base/head diff; defaults to latest, sharing possible; review permissions/data policy)](https://github.com/promptfoo/promptfoo-action) - **[Augustus](https://github.com/praetorian-inc/augustus)** ๐ŸŸข โ€” Single-binary LLM security testing framework for prompt injection, jailbreaks, and adversarial attacks across many providers. *(Praetorian)* *(โ˜… 280 ยท updated 2026-08-17)* - **Related:** [garak](https://github.com/NVIDIA/garak) ยท [PyRIT](https://github.com/microsoft/PyRIT) - **[agentic_security](https://github.com/msoedov/agentic_security)** ๐ŸŸข โ€” Agentic LLM vulnerability scanner and AI red-team kit for jailbreaks, prompt injection, fuzzing, and API stress testing. *(โ˜… 1,966 ยท updated 2026-07-31)* - **Related:** [garak](https://github.com/NVIDIA/garak) ยท [spikee](https://github.com/ReversecLabs/spikee) - **[HackAgent](https://github.com/AISecurityLab/hackagent)** ๐ŸŸข โ€” Python SDK and CLI for red-teaming AI agents with research-backed attacks such as AdvPrefix, AutoDAN-Turbo, PAIR, TAP, FlipAttack, BoN, and static templates across agent frameworks. โ€” **note:** works locally without an API key; optional cloud reporting is available. *(โ˜… 358 ยท updated 2026-08-15)* - **Related:** [PyRIT](https://github.com/microsoft/PyRIT) ยท [garak](https://github.com/NVIDIA/garak) ยท [agentic_security](https://github.com/msoedov/agentic_security) - **[wallbreaker](https://github.com/JailbrokenAI/wallbreaker)** ๐ŸŸข๐Ÿ”ฌโš ๏ธ โ€” Claude-Code-style terminal and red-team harness for authorized LLM safety testing, with HarmBench, PAIR/TAP, Crescendo, GCG-style workflows, Parseltongue transforms, MCP tooling, LLM judges, and reproducible run artifacts. โ€” **note:** offensive jailbreak toolkit for authorized testing only; AGPL-3.0 licensed and NOTICE flags external jailbreak corpora with their own or missing upstream licenses. *(โ˜… 1,208 ยท updated 2026-08-11)* - **Related:** [HarmBench](https://github.com/centerforaisafety/HarmBench) ยท [PyRIT](https://github.com/microsoft/PyRIT) ยท [garak](https://github.com/NVIDIA/garak) - **[HiveTrace Red](https://github.com/HiveTrace/HiveTraceRed)** ๐ŸŸข โ€” Early-stage LLM red-teaming framework with 80+ attack templates, async evaluation pipelines, WildGuard evaluators, multi-provider support, and HTML reporting. โ€” **note:** young project with limited independent adoption signal. *(โ˜… 29 ยท updated 2026-08-13)* - **Related:** [garak](https://github.com/NVIDIA/garak) ยท [PyRIT](https://github.com/microsoft/PyRIT) ยท [promptfoo](https://github.com/promptfoo/promptfoo) - **[DeepTeam](https://github.com/confident-ai/deepteam)** ๐ŸŸข โ€” Open-source framework for red-teaming LLMs and LLM systems across jailbreaks, prompt injection, data leakage, and safety risks. *(โ˜… 2,451 ยท updated 2026-08-12)* - **[Moonshot](https://github.com/aiverify-foundation/moonshot)** ๐ŸŸข โ€” Modular tool for benchmarking, red-teaming, and evaluating LLM applications with custom connectors and recipes. *(AI Verify Foundation)* *(โ˜… 347 ยท updated 2026-02-05)* - **[Guardrails AI](https://github.com/guardrails-ai/guardrails)** ๐ŸŸข โ€” Python framework for adding input/output guards, validators, structured-output controls, and Guardrails Hub checks to LLM applications. *(Guardrails AI)* *(โ˜… 7,294 ยท updated 2026-08-14)* - **Related:** [NeMo Guardrails](https://github.com/NVIDIA-NeMo/Guardrails) ยท [LLM Guard](https://github.com/protectai/llm-guard) ยท [Pydantic AI Shields integration (README recommends a successor)](https://github.com/vstorm-co/pydantic-ai-shields) - **[Giskard](https://github.com/Giskard-AI/giskard-oss)** ๐ŸŸข โ€” Open-source evaluation, testing, and red-teaming framework for LLM agents, including agent vulnerability scanning and RAG evaluation workflows. *(Giskard AI)* *(โ˜… 5,755 ยท updated 2026-08-17)* - **Related:** [Moonshot](https://github.com/aiverify-foundation/moonshot) ยท [promptfoo](https://github.com/promptfoo/promptfoo) - **[LangKit](https://github.com/whylabs/langkit)** ๐ŸŸข โ€” LLM monitoring toolkit extracting safety/security signals such as jailbreak similarity, prompt-injection similarity, hallucination checks, PII patterns, toxicity, and refusal metrics. *(WhyLabs)* *(โ˜… 994 ยท updated 2024-11-22)* - **[LLM Guard](https://github.com/protectai/llm-guard)** ๐ŸŸข โ€” Suite of input/output scanners (PII, prompt injection, etc.). *(Protect AI)* โ€” **note:** archived by the maintainer; retained as a historical reference for local input/output guardrails. *(โ˜… 3,201 ยท updated 2026-07-08)* - **Related:** [Rebuff](https://github.com/protectai/rebuff) ยท [Pytector prompt-injection classifier integration](https://github.com/MaxMLang/pytector) - **[Rebuff](https://github.com/protectai/rebuff)** ๐ŸŸข โ€” Archived prompt-injection detector (heuristics + LLM + vector DB + canary tokens). *(Protect AI)* โ€” **note:** archived by the maintainer; retained as a historical prompt-injection defense reference. *(โ˜… 1,520 ยท updated 2024-01-25)* - **Related:** [LLM Guard](https://github.com/protectai/llm-guard) - **[NeMo Guardrails](https://github.com/NVIDIA-NeMo/Guardrails)** ๐ŸŸข โ€” Programmable guardrails (input/output/dialog/retrieval rails) for LLM apps. *(NVIDIA)* *(โ˜… 6,965 ยท updated 2026-08-17)* - **Related:** [Backtranslation-defense reproduction (historical; legacy FastChat/GPU and model judges, separate bundled-code terms)](https://github.com/YihanWang617/LLM-Jailbreaking-Defense-Backtranslation) ยท [Jailbreak-defense research library (historical; local models or OpenAI, not an authority boundary)](https://github.com/YihanWang617/llm-jailbreaking-defense) - **[PurpleLlama](https://github.com/meta-llama/PurpleLlama)** ๐ŸŸข โ€” Llama Guard classifiers, CodeShield, and CyberSecEval. *(Meta)* *(โ˜… 4,355 ยท updated 2026-08-14)* - **[LLAMATOR](https://github.com/LLAMATOR-Core/llamator)** ๐ŸŸขโš ๏ธ โ€” Red-teaming framework for chatbots and GenAI systems; CC BY-NC-SA 4.0 licensed. *(โ˜… 215 ยท updated 2026-01-15)* - **[Vigil](https://github.com/deadbits/vigil-llm)** ๐ŸŸข๐Ÿ”ฌ โ€” Library/REST API to scan prompts and responses for prompt injection. *(โ˜… 496 ยท updated 2024-01-31)* - **[Counterfit](https://github.com/Azure/counterfit)** ๐ŸŸข โ€” ML/AI penetration-testing automation tool. *(Microsoft)* *(โ˜… 935 ยท updated 2025-07-18)* - **[AI-Red-Teaming-Playground-Labs](https://github.com/microsoft/AI-Red-Teaming-Playground-Labs)** ๐ŸŸข โ€” CTFd-based AI red-team training challenges. *(Microsoft)* *(โ˜… 2,046 ยท updated 2025-10-07)* - **[EasyJailbreak](https://github.com/EasyJailbreak/EasyJailbreak)** ๐ŸŸข๐Ÿ”ฌ โ€” Framework for building and testing adversarial jailbreak prompts. *(โ˜… 886 ยท updated 2026-03-30)* - **[TextAttack](https://github.com/QData/TextAttack)** ๐ŸŸข๐Ÿ”ฌ โ€” Python framework for adversarial attacks, data augmentation, and training for NLP models; useful for robustness testing beyond chat-only LLM scanners. *(โ˜… 3,467 ยท updated 2026-08-15)* - **[GPTFuzz](https://github.com/sherdencooper/GPTFuzz)** ๐ŸŸข๐Ÿ”ฌ โ€” Research framework for red-teaming LLMs with auto-generated jailbreak prompts. *(โ˜… 604 ยท updated 2026-02-27)* - **[HarmBench](https://github.com/centerforaisafety/HarmBench)** ๐ŸŸข๐Ÿ”ฌ โ€” ICML 2024 standardized evaluation framework for automated red-teaming and robust-refusal benchmarking. *(Center for AI Safety)* *(โ˜… 1,029 ยท updated 2024-08-05)* - **Related:** [JailbreakBench](https://github.com/JailbreakBench/jailbreakbench) ยท [JailBreakV-28K (historical benchmark; no root license, education/research scope; separate image/model terms)](https://github.com/SaFo-Lab/JailBreakV_28K) - **[llm-attacks (GCG)](https://github.com/llm-attacks/llm-attacks)** ๐ŸŸข๐Ÿ”ฌ โ€” Canonical Greedy Coordinate Gradient adversarial-suffix attack implementation for transferable attacks on aligned language models. *(โ˜… 4,761 ยท updated 2024-08-02)* - **Related:** [nanoGCG](https://github.com/GraySwanAI/nanoGCG) ยท [LLM embedding-attack research (white-box embeddings, not an ordinary text API)](https://github.com/SchwinnL/LLM_Embedding_Attack) - **[nanoGCG](https://github.com/GraySwanAI/nanoGCG)** ๐ŸŸข โ€” Fast, lightweight PyTorch implementation of the GCG adversarial-suffix algorithm. *(โ˜… 346 ยท updated 2025-05-13)* - **Related:** [llm-attacks (GCG)](https://github.com/llm-attacks/llm-attacks) - **[JailbreakBench](https://github.com/JailbreakBench/jailbreakbench)** ๐ŸŸข๐Ÿ”ฌ โ€” NeurIPS 2024 open robustness benchmark and leaderboard for generating and defending against LLM jailbreaks. *(โ˜… 654 ยท updated 2025-03-31)* - **Related:** [MEGA security evaluation artifacts (limited vendor-run experiments, not independent rankings)](https://github.com/mega-edo/mega-security-leaderboard) ยท [Chasing Shadows evaluation-methodology artifacts (external data/models; some API/GPU experiments)](https://github.com/Dormant-Neurons/llm-pitfalls) - **[Open-Prompt-Injection](https://github.com/liu00222/Open-Prompt-Injection)** ๐ŸŸข๐Ÿ”ฌ โ€” Open-source toolkit and benchmark for implementing and evaluating prompt-injection attacks, defenses, and LLM-integrated applications. *(โ˜… 479 ยท updated 2025-10-29)* - **Related:** [PromptInject (historical prompt-injection research)](https://github.com/agencyenterprise/PromptInject) - **[PINT Benchmark](https://github.com/lakeraai/pint-benchmark)** ๐ŸŸข๐Ÿ”ฌ โ€” Prompt-injection test benchmark for evaluating detectors and guardrails across multilingual prompt injection, jailbreak, benign, and hard-negative inputs. *(Lakera)* โ€” **note:** archived benchmark retained as a historical research reference; use newer maintained corpora for current detector comparisons. *(โ˜… 198 ยท updated 2026-04-02)* - **Related:** [Open-Prompt-Injection](https://github.com/liu00222/Open-Prompt-Injection) ยท [Prompt Guard 86M](https://huggingface.co/meta-llama/Prompt-Guard-86M) ยท [Tensor Trust game/data tooling (historical; separate data, dependencies and deployment need review)](https://github.com/HumanCompatibleAI/tensor-trust) - **[PIArena](https://github.com/sleeepeer/PIArena)** ๐ŸŸข๐Ÿ”ฌ โ€” ACL 2026 toolbox and benchmark for prompt-injection attacks and defenses, with ready-to-use attacks/defenses, evaluation pipelines, agent benchmarks, a Hugging Face dataset, and leaderboard. *(โ˜… 47 ยท updated 2026-04-20)* - **Related:** [PINT Benchmark](https://github.com/lakeraai/pint-benchmark) ยท [Open-Prompt-Injection](https://github.com/liu00222/Open-Prompt-Injection) - **[Whistleblower](https://github.com/Repello-AI/whistleblower)** ๐ŸŸขโš ๏ธ โ€” Offensive testing tool for inferring system prompts and discovering capabilities of LLM applications exposed through APIs. *(Repello AI)* โ€” **note:** no LICENSE file found. *(โ˜… 177 ยท updated 2025-10-27)* - **[LLMmap](https://github.com/pasquini-dario/LLMmap)** ๐ŸŸข๐Ÿ”ฌ โ€” Minimal-query fingerprinting tool for identifying LLMs from behavioral traces, with a pretrained open-set inference model. *(โ˜… 424 ยท updated 2025-07-24)* - **[llm-security](https://github.com/greshake/llm-security)** ๐Ÿ”ฌ โ€” Original PoC for indirect prompt-injection attacks. *(โ˜… 2,128 ยท updated 2025-07-17)* - **Related:** [Dropbox LLM extraction research (historical code and results)](https://github.com/dropbox/llm-security) ยท [PIPE threat-model primer (2023 reference, not current API guidance; no root license)](https://github.com/jthack/PIPE) ยท [HouYi (historical prompt-injection research)](https://github.com/LLMSecurity/HouYi) ยท [RAG poisoning teaching PoC (AGPL-3.0; embedding weights/model endpoint required; not a benchmark)](https://github.com/prompt-security/RAG_Poisoning_POC) - **[JailbreakLLMs](https://github.com/TrustAIRLab/JailbreakLLMs)** ๐Ÿ”ฌโš ๏ธ โ€” Research dataset of 6,387 ChatGPT prompts, including in-the-wild jailbreak prompts from Reddit, Discord, websites, and open datasets. *(โ˜… 23 ยท updated 2024-02-21)* - **[Do-Not-Answer](https://github.com/Libr-AI/do-not-answer)** ๐ŸŸข๐Ÿ”ฌ โ€” Dataset for evaluating LLM safeguards on unsafe or policy-sensitive prompts. *(โ˜… 341 ยท updated 2024-06-07)* - **[prompt-injection-defenses](https://github.com/tldrsec/prompt-injection-defenses)** ๐ŸŸขโš ๏ธ โ€” Curated catalog of practical defenses against prompt injection. *(โ˜… 724 ยท updated 2025-02-22)* - **[little-canary](https://github.com/hermes-labs-ai/little-canary)** ๐ŸŸข๐Ÿ”ฌ โ€” Prompt-injection preflight sensor that probes untrusted input with a powerless canary model and returns PASS, FLAG, or BLOCK with explicit coverage state before primary-agent action; includes opt-in Claude Code prompt-submission and OpenAI Agents SDK input-boundary adapters. โ€” **note:** experimental sensing layer, not a security guarantee or runtime containment. A failed canary can return availability-first routing with explicit degraded coverage; remote/OpenAI-compatible canary or judge endpoints receive raw input. *(โ˜… 30 ยท updated 2026-09-14)* - **Related:** [Rebuff](https://github.com/protectai/rebuff) ยท [prompt-injection-defenses](https://github.com/tldrsec/prompt-injection-defenses) - **[Kiji Privacy Proxy](https://github.com/Dataiku/kiji-proxy)** ๐ŸŸข โ€” Local privacy proxy for OpenAI-compatible AI API traffic that detects and masks 26 PII types with an ONNX model before forwarding requests, then restores mappings in responses. *(Dataiku 575 Lab)* โ€” **note:** protects configured proxied traffic, not every path by which an application or agent can disclose data; operators retain responsibility for proxy routing and local mapping storage. *(โ˜… 422 ยท updated 2026-07-27)* - **Related:** [LLM Guard](https://github.com/protectai/llm-guard) ยท [Kiji PII model](https://huggingface.co/DataikuNLP/kiji-pii-model-onnx) - **[Anamorpher](https://github.com/trailofbits/anamorpher)** ๐ŸŸข๐Ÿ”ฌ โ€” Research tool with a frontend and Python API for crafting and visualizing image-scaling attacks that reveal hidden prompt injections to multimodal AI systems after downscaling. *(Trail of Bits)* โ€” **note:** active beta research tool for authorized testing; generated payloads are sensitive to the target's image-resampling implementation and preprocessing pipeline. *(โ˜… 1,076 ยท updated 2026-02-19)* - **Sources:** [Trail of Bits research announcement](https://blog.trailofbits.com/2025/08/21/weaponizing-image-scaling-against-production-ai-systems/) - **Related:** [PromptFuzz](https://github.com/PromptFuzz/PromptFuzz) ยท [promptfoo](https://github.com/promptfoo/promptfoo) ยท [Invisible prompt-injection fixtures (hidden-content samples)](https://github.com/bountyyfi/invisible-prompt-injection) - **[Prompt SIREN](https://github.com/facebookresearch/prompt-siren)** ๐ŸŸข๐Ÿ”ฌ โ€” Research workbench for developing and evaluating prompt-injection attacks and defenses with state-machine agent control, AgentDojo/SWE-bench integrations, configuration sweeps, and reproducible result aggregation. *(Meta AI)* โ€” **note:** experiment harness; running target evaluations requires model-provider credentials and may need Docker or browser extras depending on the selected environment. *(โ˜… 61 ยท updated 2026-05-18)* - **Related:** [AgentDojo](https://github.com/ethz-spylab/agentdojo) ยท [PyRIT](https://github.com/microsoft/PyRIT) - **[Argus](https://github.com/gy15901580825/Argus)** ๐ŸŸข๐Ÿ”ฌ โ€” Black-box red-team framework for LLM applications and agents with target adapters, attack probes, deterministic and model-assisted judging, and SARIF, JUnit, HTML, and JSON reporting. โ€” **note:** pre-1.0 project; full deployments use an orchestrator and PostgreSQL, while model-assisted judges require provider credentials. Attack-success and guardrail-improvement figures are project-reported; probe only authorized targets. *(โ˜… 204 ยท updated 2026-08-14)* - **Related:** [PyRIT](https://github.com/microsoft/PyRIT) ยท [promptfoo](https://github.com/promptfoo/promptfoo) - **[Cryptex OSS](https://github.com/m4xx101/cryptex-oss)** ๐ŸŸข๐ŸŸ  โ€” Browser-based and self-hostable adversarial prompt workbench with transformation pipelines, campaign workflows, TAP/PAIR/Crescendo methods, and HarmBench/JailbreakBench-style corpora. โ€” **note:** authorized-testing toolkit; provider keys are supplied by the operator and retained in browser storage. Included benchmark scoring is explicitly heuristic rather than paper-accurate, and the separate production product is closed-source. *(โ˜… 339 ยท updated 2026-06-09)* - **Related:** [wallbreaker](https://github.com/JailbrokenAI/wallbreaker) ยท [HarmBench](https://github.com/centerforaisafety/HarmBench) - **[API Relay Audit](https://github.com/toby-bridges/api-relay-audit)** ๐ŸŸขโš ๏ธ โ€” Local audit CLI for third-party LLM relays and proxies, testing prompt injection, model substitution, tool-call rewriting, streaming behavior, and relay-specific trust assumptions. โ€” **note:** sends the supplied API credential to the relay being tested, so use a scoped disposable key and only assess services you are authorized to probe. Results identify suspicious behavior but do not certify a relay as safe. *(โ˜… 803 ยท updated 2026-08-15)* - **Related:** [garak](https://github.com/NVIDIA/garak) ยท [LLMmap](https://github.com/pasquini-dario/LLMmap) - **[AIDR Bastion](https://github.com/socprime/AIDR-Bastion)** ๐ŸŸข๐ŸŸ โš ๏ธ โ€” Runtime input-protection service combining detection rules, similarity search, classifiers, optional LLM analysis, and code-oriented checks to allow, block, or notify on suspicious GenAI traffic. *(SOC Prime)* โ€” **note:** multi-service deployment can require OpenSearch/Elasticsearch, Qdrant, Kafka, and an optional local or hosted model; some rule and integration workflows connect to SOC Prime services. Treat it as a layered signal source, not a complete isolation boundary. *(โ˜… 107 ยท updated 2026-08-20)* - **Related:** [LLM Guard](https://github.com/protectai/llm-guard) ยท [NeMo Guardrails](https://github.com/NVIDIA-NeMo/Guardrails) - **[CaMeL](https://github.com/google-research/camel-prompt-injection)** ๐ŸŸข๐Ÿ”ฌ โ€” Research implementation of the capability-based CaMeL interpreter architecture for separating trusted control flow from untrusted data while evaluating prompt-injection defenses on AgentDojo. *(Google Research / Google DeepMind / ETH Zurich)* โ€” **note:** unsupported paper-reproduction artifact whose maintainers explicitly warn that the interpreter may contain bugs and may not be fully secure; do not treat it as a production guardrail. *(โ˜… 379 ยท updated 2025-06-20)* - **Related:** [Defeating Prompt Injections by Design](https://arxiv.org/abs/2503.18813) ยท [AgentDojo](https://github.com/ethz-spylab/agentdojo) ยท [Agent-security lecture demos (no root license; CaMeL and external models required)](https://github.com/anishathalye/ai-agent-security-lecture) - **[Meta SecAlign](https://github.com/facebookresearch/Meta_SecAlign)** ๐Ÿ”ฌโš ๏ธ โ€” Research code, training recipe, and evaluation harness for prompt-injection-resistant Meta SecAlign models across six security and eight utility benchmarks. *(Meta / UC Berkeley)* โ€” **note:** most repository code is non-commercial while published model weights use separate Llama community terms; reproducing training or the full evaluation requires substantial GPU capacity and optional provider credentials. *(โ˜… 70 ยท updated 2026-06-11)* - **Related:** [AgentDojo](https://github.com/ethz-spylab/agentdojo) ยท [Meta SecAlign models](https://huggingface.co/meta-llama) ยท [SecAlign (historical predecessor; CC BY-NC 4.0, noncommercial)](https://github.com/facebookresearch/SecAlign) ยท [StruQ predecessor research (CC BY-NC 4.0; trained models/GPU, OpenAI utility evaluation)](https://github.com/Sizhe-Chen/StruQ) ### Prompt-Injection Classifier Models - **[Wolf Defender Prompt Injection](https://huggingface.co/patronus-studio/wolf-defender-prompt-injection)** ๐ŸŸข โ€” Hugging Face text-classification model for prompt-injection detection in agents, chatbots, and CI workflows. *(Patronus Studio / Casdo Labs)* *license: Apache-2.0 ยท access: open ยท artifacts: Safetensors, ONNX.* - **[DeBERTa v3 Prompt Injection v2](https://huggingface.co/protectai/deberta-v3-base-prompt-injection-v2)** ๐ŸŸข โ€” Apache-licensed prompt-injection classifier usable via Transformers pipelines and ONNX. *(Protect AI)* *license: Apache-2.0 ยท access: open ยท artifacts: Safetensors, ONNX.* - **[PromptGuard](https://huggingface.co/codeintegrity-ai/promptguard)** ๐ŸŸขโš ๏ธ โ€” ModernBERT-based prompt-injection and jailbreak classifier. *(CodeIntegrity AI)* *license: Apache-2.0 ยท access: gated auto ยท artifacts: Safetensors.* - **[Prompt Guard 86M](https://huggingface.co/meta-llama/Prompt-Guard-86M)** ๐ŸŸ โš ๏ธ โ€” Meta prompt-injection and jailbreak classifier from the Llama Guard family. *(Meta)* *license: Llama 3.1 ยท access: gated manual ยท artifacts: Safetensors.* - **[prompt-injection-sentinel](https://huggingface.co/qualifire/prompt-injection-sentinel)** ๐Ÿ”ฌโš ๏ธ โ€” ModernBERT-large classifier for prompt-injection and jailbreak detection. *(Qualifire)* *license: other ยท access: gated auto ยท artifacts: Safetensors.* ### Specialty Security LLMs - **[Foundation-Sec-8B-Reasoning](https://huggingface.co/fdtn-ai/Foundation-Sec-8B-Reasoning)** ๐ŸŸ โš ๏ธ โ€” Downloadable Llama-3.1-derived 8B reasoning model for cybersecurity analysis, threat intelligence, and multi-step security workflows. *(Cisco Foundation AI)* *license: Llama 3.1 Community License (base materials); Apache-2.0 (Cisco changes) ยท access: open ยท artifacts: Safetensors.* **Caveat:** not an Apache-only model; retain the base-model license and NOTICE obligations. Local inference needs a separate runtime and substantial memory, but no mandatory Cisco service. Security analysis can hallucinate; this is not a verified scanner or guardrail, and any tool-executing wrapper needs independent authorization and isolation. - **[SecGPT](https://github.com/Clouditera/SecGPT)** ๐ŸŸข โ€” Open cybersecurity-tuned LLM family for vulnerability analysis, log/traffic investigation, anomaly detection, attack/defense reasoning, command analysis, and security Q&A. *(Clouditera)* *(โ˜… 3,099 ยท updated 2025-06-25)* - **Related:** [SecGPT model](https://huggingface.co/clouditera/secgpt) - **[Antares-1B](https://huggingface.co/fdtn-ai/antares-1b)** ๐ŸŸขโš ๏ธ โ€” Open-weight security SLM specialized for agentic vulnerability localization: it explores repository snapshots through a terminal-style loop and ranks likely vulnerable files for analyst review. *(Cisco Foundation AI)* *license: Apache-2.0 ยท access: gated manual ยท artifacts: Safetensors + CLI ZIP.* โ€” **note:** HF access is gated/manual; related smaller model: [Antares-350M](https://huggingface.co/fdtn-ai/antares-350m), benchmark: [VLoc Bench](https://cisco-foundation-ai.github.io/vulnerability-localization-benchmark/). - **[Trendyol Cybersecurity LLM v2 70B](https://huggingface.co/Trendyol/Trendyol-Cybersecurity-LLM-v2-70B-Q4_K_M)** ๐ŸŸข โ€” Defense-focused cybersecurity LLM based on Llama-3.3-70B, trained on an alignment-safe security instruction dataset for SOC, cloud, AppSec, detection, and vulnerability-management workflows. *(Trendyol Group Security Team)* *license: Apache-2.0 ยท access: open ยท artifacts: GGUF.* - **[WhiteRabbitNeo 2.5 Qwen Coder 7B](https://huggingface.co/WhiteRabbitNeo/WhiteRabbitNeo-2.5-Qwen-2.5-Coder-7B)** ๐ŸŸขโš ๏ธ โ€” Cybersecurity-oriented Qwen2.5-Coder fine-tune positioned for offensive and defensive security assistance. *(WhiteRabbitNeo)* *license: Apache-2.0 + WhiteRabbitNeo restrictions ยท access: open ยท artifacts: Safetensors.* - **[Lily-Cybersecurity-7B-v0.2](https://huggingface.co/segolilylabs/Lily-Cybersecurity-7B-v0.2)** ๐ŸŸข โ€” Mistral-7B-Instruct fine-tune for cybersecurity assistance, trained on hand-crafted security and hacking-related instruction pairs. *(Segolily Labs)* *license: Apache-2.0 ยท access: open ยท artifacts: Safetensors.* - **[RavenX CyberAgent 35B Q4_K_M](https://huggingface.co/deadbydawn101/RavenX-CyberAgent-Qwen3.6-35B-A3B-Opus-4.7-OpenMythos-Pentester-BugHunter-RATH-GGUF)** ๐ŸŸขโš ๏ธ โ€” GGUF security-specialized text-generation model positioned for pentest, bug-bounty, tool-calling, MCP, CVSS/CWE, and MITRE ATT&CK workflows. *(RavenX LLC / DeadByDawn101)* *license: Apache-2.0 ยท access: open ยท artifacts: GGUF.* โ€” **note:** built from an abliterated base model and marketed for autonomous security assessment; use only in authorized, sandboxed agent harnesses with tool-call validation. --- ## LLM Honeypots & Deception Honeypots and deception that use LLMs to simulate convincing systems. - **[Beelzebub](https://github.com/beelzebub-labs/beelzebub)** ๐ŸŸขโš ๏ธ โ€” Low-code honeypot using LLMs to simulate SSH/HTTP/MCP services (Go). โ€” **note:** GPL-3.0 licensed. *(โ˜… 2,148 ยท updated 2026-08-11)* - **[DECEIVE](https://github.com/splunk/DECEIVE)** ๐ŸŸข๐Ÿ”ฌ โ€” Proof-of-concept LLM-powered SSH honeypot that evaluates sessions as benign, suspicious, or malicious. *(Splunk)* *(โ˜… 288 ยท updated 2026-05-13)* - **[TRAP](https://github.com/parameterlab/trap)** ๐ŸŸข๐Ÿ”ฌ โ€” Research code for Targeted Random Adversarial Prompt honeypots that identify black-box LLM usage through model-specific prompt suffixes (ACL 2024 Findings). *(โ˜… 15 ยท updated 2024-11-20)* - **[shelLM](https://github.com/stratosphereips/shelLM)** ๐ŸŸข๐Ÿ”ฌ โ€” LLM-powered SSH honeypot (paper *"LLM in the Shell"*). *(โ˜… 64 ยท updated 2026-06-25)* - **Related:** [VelLMes](https://github.com/stratosphereips/VelLMes-AI-Honeypot) - **[VelLMes](https://github.com/stratosphereips/VelLMes-AI-Honeypot)** ๐ŸŸข๐Ÿ”ฌ โ€” Multi-protocol LLM honeypot framework (successor to shelLM). *(โ˜… 78 ยท updated 2025-02-18)* - **Related:** [shelLM](https://github.com/stratosphereips/shelLM) - **[llm-honeypot](https://github.com/PalisadeResearch/llm-honeypot)** ๐Ÿ”ฌโš ๏ธ โ€” Cowrie SSH honeypot extended with prompt-injection traps to detect LLM hacker agents. *(Palisade Research)* *(โ˜… 60 ยท updated 2026-01-23)* --- ## CTF / Exploit / Bug-Bounty Agents & Benchmarks Offensive agents and the benchmarks used to evaluate them. - **[OWASP PromptMe](https://github.com/OWASP/www-project-promptme)** ๐ŸŸข โ€” Intentionally vulnerable Flask/Ollama lab with hands-on challenges based on OWASP LLM risks. *(OWASP)* **Caveat:** use a disposable isolated environment without real data: examples bind to all interfaces, contain known training secrets, fetch URLs, and deliberately include fail-open behavior. Ollama and model weights are separate dependencies. *(โ˜… 43 ยท updated 2026-09-10)* - **[Sherpa](https://github.com/Azure-Samples/sherpa)** ๐ŸŸข โ€” Hands-on MCP security lab pairing vulnerable and hardened sample servers with authentication and authorization exercises. *(Microsoft Azure Samples)* **Caveat:** use an isolated teaching environment; demonstration tokens and fixed identities are not production authentication. Azure-backed exercises require appropriate account permissions and may incur costs; sample fixes are not a complete application-security boundary. *(โ˜… 29 ยท updated 2026-09-14)* - **[ExploitGym](https://github.com/sunblaze-ucb/exploitgym)** ๐ŸŸข๐Ÿ”ฌโš ๏ธ โ€” Research benchmark and evaluation harness for developing exploits for known vulnerabilities across userspace programs, V8, and Linux-kernel tasks, combining flag checks with model-based trajectory assessment. **Caveat:** executes real exploits against vulnerable containers or VMs and can require host mitigation changes; use dedicated isolated infrastructure. Providers and model judges receive task context, task artifacts retain separate upstream licenses, and judge verdicts need review. *(โ˜… 1,135 ยท updated 2026-09-29)* - **[Inspect Cyber](https://github.com/UKGovernmentBEIS/inspect_cyber)** ๐ŸŸข๐Ÿ”ฌ โ€” Installable Inspect extension for defining and running agentic cyber evaluations with standardized YAML configurations, adaptable sandboxes, scenario variants, scoring, and solvability verification. *(UK AI Security Institute)* โ€” **note:** evaluation framework rather than a benchmark result or defensive control; Docker, Kubernetes, or other configured sandboxes can contain intentionally vulnerable targets and should remain isolated from production networks. *(โ˜… 40 ยท updated 2026-06-18)* - **Related:** [inspect_evals](https://github.com/UKGovernmentBEIS/inspect_evals) - **[GenAI Red Team Lab](https://github.com/GenAI-Security-Project/GenAI-Red-Team-Lab)** ๐ŸŸข๐Ÿ”ฌ โ€” Collection of intentionally vulnerable GenAI sandboxes, exploitation examples, and tutorials for prompt injection, memory poisoning, orchestration attacks, guardrail bypass, and related red-team exercises. *(OWASP GenAI Security Project)* โ€” **note:** training and research lab, not a scanner; it contains exploitation code and deliberately vulnerable services that should run only in disposable isolated environments. *(โ˜… 53 ยท updated 2026-09-20)* - **Related:** [Vulnerable agent samples (CC BY-SA 4.0; isolated lab use, separate APIs/models)](https://github.com/GenAI-Security-Project/GenAI-Agent-Security-Initiative) ยท [Aembit agent-security control samples (teaching examples, not a production platform)](https://github.com/Aembit/agentic-ai-security-starter-kit) ยท [GenAI training labs (CC BY-NC notice plus government-use restriction; commercial/government reuse requires permission)](https://github.com/schwartz1375/genai-security-training) - **[SWE-agent (EnIGMA)](https://github.com/SWE-agent/SWE-agent)** ๐ŸŸข๐Ÿ”ฌ โ€” EnIGMA offensive-CTF mode; SOTA on NYU CTF, InterCode-CTF, and Cybench (v0.7 branch). *(โ˜… 20,069 ยท updated 2026-07-16)* - **Related:** [Cybench](https://github.com/andyzorigin/cybench) ยท [NYU CTF Bench](https://github.com/NYU-LLM-CTF/NYU_CTF_Bench) ยท [InterCode](https://github.com/princeton-nlp/intercode) - **[Cybench](https://github.com/andyzorigin/cybench)** ๐Ÿ”ฌ โ€” 40 professional CTF tasks across 4 competitions; widely used by AI safety institutes. *(โ˜… 311 ยท updated 2026-07-09)* - **[NYU CTF Bench](https://github.com/NYU-LLM-CTF/NYU_CTF_Bench)** ๐Ÿ”ฌ โ€” Dockerized CSAW CTF challenges for LLM-agent evaluation. *(โ˜… 170 ยท updated 2025-09-22)* - **[CTFTiny](https://github.com/NYU-LLM-CTF/CTFTiny)** ๐Ÿ”ฌโš ๏ธ โ€” Lightweight CTF benchmark from the NYU LLM CTF group; GPL-2.0 licensed. *(โ˜… 18 ยท updated 2026-03-10)* - **Related:** [NYU CTF Bench](https://github.com/NYU-LLM-CTF/NYU_CTF_Bench) - **[InterCode](https://github.com/princeton-nlp/intercode)** ๐Ÿ”ฌ โ€” Interactive-coding benchmark incl. InterCode-CTF. *(โ˜… 254 ยท updated 2024-05-05)* - **[inspect_evals](https://github.com/UKGovernmentBEIS/inspect_evals)** ๐ŸŸข๐Ÿ”ฌ โ€” Maintained Inspect AI evaluation suite containing multiple cyber benchmarks and tasks. *(UK AI Security Institute)* *(โ˜… 627 ยท updated 2026-08-17)* - **[BountyBench](https://github.com/bountybench/bountybench)** ๐Ÿ”ฌ โ€” 25 real systems / 40 bug bounties for Detect-Exploit-Patch evaluation. *(โ˜… 102 ยท updated 2025-06-22)* - **[Cyber-Zero](https://github.com/amazon-science/Cyber-Zero)** ๐Ÿ”ฌ โ€” Trains cybersecurity agents without runtime; ships an EnIGMA+ scaffold. *(Amazon Science)* โ€” **note:** archived research artifact retained as a historical reference for training cybersecurity agents without a live runtime. *(โ˜… 101 ยท updated 2025-09-02)* - **Sources:** [SWE-agent](https://github.com/SWE-agent/SWE-agent) - **Related:** [SWE-agent](https://github.com/SWE-agent/SWE-agent) - **[ExploitBench](https://github.com/exploitbench/exploitbench)** ๐Ÿ”ฌ โ€” Measures AI-agent progress on V8/Chromium exploit ladders. *(โ˜… 344 ยท updated 2026-07-04)* - **[AI Goat](https://github.com/dhammon/ai-goat)** ๐ŸŸข๐Ÿ”ฌโš ๏ธ โ€” Vulnerable-by-design local LLM CTF for learning prompt injection, insecure output handling, data leakage, excessive agency, and related LLM app risks. โ€” **note:** GPL-2.0 licensed. *(โ˜… 355 ยท updated 2024-08-22)* - **[AIGoat](https://github.com/AISecurityConsortium/AIGoat)** ๐ŸŸข๐Ÿ”ฌโš ๏ธ โ€” Local-first vulnerable LLM security playground with guided OWASP LLM Top 10 attack labs, CTF challenges, progressive defenses, and an Ollama-backed AI shopping-assistant target. โ€” **note:** platform code is Apache-2.0, but training/challenge content is CC BY-NC-SA-4.0 and requires permission for commercial workshops. *(โ˜… 76 ยท updated 2026-04-24)* - **[LLMVault](https://github.com/CyberSunil/LLMVault)** ๐ŸŸข๐Ÿ”ฌ โ€” Intentionally vulnerable LLM security-training platform with OWASP LLM Top 10 labs, CTF-style challenges, hints, scoring, and mitigation guidance. โ€” **note:** deliberately vulnerable training target; run only in an isolated, authorized environment. Live Mode optionally requires Ollama or provider credentials. *(โ˜… 296 ยท updated 2026-08-10)* - **[Damn Vulnerable LLM Agent](https://github.com/ReversecLabs/damn-vulnerable-llm-agent)** ๐ŸŸข๐Ÿ”ฌ โ€” Deliberately vulnerable LangChain ReAct agent for practicing prompt-injection and Thought/Action/Observation injection attacks. *(ReversecLabs / WithSecure)* *(โ˜… 505 ยท updated 2025-06-25)* - **Related:** [spikee](https://github.com/ReversecLabs/spikee) - **[claude-bug-bounty](https://github.com/shuvonsec/claude-bug-bounty)** ๐ŸŸข โ€” Claude Code plugin orchestrating recon โ†’ vuln classes โ†’ reporting. *(โ˜… 4,236 ยท updated 2026-08-10)* - **[Bug-Bounty-Agents](https://github.com/matty69v/Bug-Bounty-Agents)** ๐ŸŸข โ€” 43 AI agent personas for Claude Code / Copilot / Cursor across the bug-bounty lifecycle. *(โ˜… 369 ยท updated 2026-04-30)* - **[ai-exploits](https://github.com/protectai/ai-exploits)** ๐ŸŸข โ€” Real-world AI/ML exploits (Metasploit modules + Nuclei templates) for MLflow, Ray, H2O. *(Protect AI)* *(โ˜… 1,745 ยท updated 2024-10-23)* - **[CyberGym](https://github.com/sunblaze-ucb/cybergym)** ๐ŸŸข๐Ÿ”ฌ โ€” Large-scale evaluation framework for AI-agent vulnerability analysis on real-world tasks, with locally deployed challenge infrastructure, task generation, and proof-of-concept validation. *(UC Berkeley / Sunblaze)* โ€” **note:** deploy only in an isolated local environment; the full benchmark runtime is extremely large (documentation cites up to ~10 TB) and must not be exposed to the public internet. *(โ˜… 737 ยท updated 2026-08-04)* - **Related:** [Cybench](https://github.com/andyzorigin/cybench) ยท [CVE-Bench](https://github.com/uiuc-kang-lab/cve-bench) - **[CVE-Bench](https://github.com/uiuc-kang-lab/cve-bench)** ๐ŸŸข๐Ÿ”ฌ โ€” ICML 2025 benchmark that evaluates AI agents against reproducible Docker environments for real critical-severity web-application CVEs and exploit objectives. โ€” **note:** runs intentionally vulnerable services and should be used only in isolated environments; arm64 support is experimental. *(โ˜… 271 ยท updated 2026-01-14)* - **Related:** [CyberGym](https://github.com/sunblaze-ucb/cybergym) ยท [BountyBench](https://github.com/bountybench/bountybench) - **[SCAM](https://github.com/1Password/SCAM)** ๐ŸŸข๐Ÿ”ฌ โ€” Benchmark for measuring whether tool-using agents recognize and resist realistic scams, social engineering, credential requests, impersonation, and unsafe multi-turn instructions across 30 scenarios and nine threat categories. *(1Password)* โ€” **note:** evaluation requires at least one configured model provider; leaderboard results and reported safety improvements from the optional skill are project-reported, with raw result artifacts published for inspection. *(โ˜… 136 ยท updated 2026-02-12)* - **Related:** [AgentDojo](https://github.com/ethz-spylab/agentdojo) ยท [Agent Security Bench](https://github.com/agiresearch/ASB) - **[AIRTBench-Code](https://github.com/dreadnode/AIRTBench-Code)** ๐ŸŸข๐Ÿ”ฌ โ€” Research harness and dataset for evaluating autonomous AI red-team agents on AI/ML security CTF challenges, with published task definitions and result artifacts. *(Dreadnode)* โ€” **note:** the public code and dataset accompany arXiv 2506.14682, but reproducing the full evaluation depends on Dreadnode's hosted Strikes environment, platform access, model credentials, and containerized challenge infrastructure. *(โ˜… 107 ยท updated 2026-04-26)* - **Related:** [Cybench](https://github.com/andyzorigin/cybench) ยท [CyberGym](https://github.com/sunblaze-ucb/cybergym) - **[AI-Goat (Orca Security)](https://github.com/orcasecurity-research/AIGoat)** ๐ŸŸข๐Ÿ”ฌ โ€” Deliberately vulnerable AWS and Terraform lab for learning AI/ML infrastructure risks such as model supply-chain compromise, data poisoning, insecure output handling, and integrity failures. *(Orca Security Research)* โ€” **note:** distinct from the local LLM playgrounds with similar names; deploys intentionally vulnerable cloud resources and can incur AWS costs, so use only in an isolated authorized account and remove resources after each lab. *(โ˜… 284 ยท updated 2025-09-16)* - **Related:** [AIGoat](https://github.com/AISecurityConsortium/AIGoat) ยท [AI Goat](https://github.com/dhammon/ai-goat) - **[LivePI](https://github.com/leizhao7/livepi)** ๐Ÿ”ฌโš ๏ธ โ€” Reproducibility artifact for a production-like indirect prompt-injection benchmark spanning live but test-controlled email, chat, web, local-file, repository, and wallet surfaces. โ€” **note:** benchmark setup interacts with live test accounts and services and includes attack scenarios capable of exfiltration, unsafe execution, security-control changes, and value transfer; reproduce only in isolated infrastructure with synthetic credentials and explicit authorization. *(โ˜… 9 ยท updated 2026-06-08)* - **Related:** [LivePI paper](https://arxiv.org/abs/2605.17986) ยท [AgentDojo](https://github.com/ethz-spylab/agentdojo) --- ## Cloud / IaC / DFIR / OSINT / Phishing AI tooling for cloud/IaC security, digital forensics, OSINT, and phishing detection. - **[AWS AI/ML Security Assessment](https://github.com/aws-samples/sample-aiml-security-assessment)** ๐ŸŸข โ€” Security-posture assessment of AWS AI/ML services, including Bedrock, SageMaker, and AgentCore configurations. *(AWS Samples)* **Caveat:** requires AWS IAM permissions and deployed resources, with possible service costs; reports can expose sensitive infrastructure details. A configured prompt-attack filter is not evidence of attack resistance, and standards mappings are not compliance certification. *(โ˜… 47 ยท updated 2026-10-04)* - **[AIFT](https://github.com/FlipForensics/AIFT)** ๐ŸŸขโš ๏ธ โ€” Python forensic-triage application that parses disk images and collected Windows/Linux artifacts with Dissect, then produces model-assisted reports through GUI, CLI, REST, or MCP interfaces. **Caveat:** AGPL-3.0 licensed. Hosted providers receive parsed evidence; local models require separately supplied runtimes and weights. Use evidence copies and independent integrity controls, and keep API access trusted. Outputs are investigation leads requiring examiner verification, not final forensic conclusions. *(โ˜… 48 ยท updated 2026-06-13)* - **[Valhuntir](https://github.com/AppliedIR/Valhuntir)** ๐ŸŸข โ€” Analyst-directed incident-response platform combining a Python case-management CLI with SIFT MCP services for evidence processing, draft findings, human review, and report generation. **Caveat:** alpha; full deployment needs companion SIFT MCP packages and optional forensic VMs or OpenSearch. Loaded evidence may reach configured model providers. Lite has no sandbox or deny rules, tools can execute on the host, and evidence integrity and findings require independent examiner controls. *(โ˜… 106 ยท updated 2026-10-04)* - **Related:** [SIFT MCP backend](https://github.com/AppliedIR/sift-mcp) - **[EscalateGPT](https://github.com/tenable/EscalateGPT)** ๐ŸŸข โ€” GPT-based discovery of privilege-escalation paths in AWS IAM policies. *(Tenable)* *(โ˜… 122 ยท updated 2024-01-17)* - **[Cynative](https://github.com/cynative/cynative)** ๐ŸŸข โ€” Local AI security research agent for cloud, code, and runtime environments across GitHub, GitLab, AWS, GCP, Azure, and Kubernetes, with read-only action gates, sandboxed code execution, evidence-backed verification, and audit logs. *(โ˜… 190 ยท updated 2026-08-14)* - **Related:** [Fraim](https://github.com/fraim-dev/fraim) ยท [EscalateGPT](https://github.com/tenable/EscalateGPT) - **[Julius](https://github.com/praetorian-inc/julius)** ๐ŸŸข โ€” Local Go tool that fingerprints LLM service infrastructure on authorized endpoints, identifies 60+ serving, gateway, MCP, and RAG platforms, and can enumerate exposed models. *(Praetorian)* โ€” **note:** use only against endpoints you own or are authorized to assess. *(โ˜… 208 ยท updated 2026-08-06)* - **Related:** [AI-Infra-Guard](https://github.com/Tencent/AI-Infra-Guard) ยท [ai_osint](https://github.com/7WaySecurity/ai_osint) - **[MemoryInvestigator](https://github.com/jan-hendrik-lang/MemoryInvestigator)** ๐Ÿ”ฌ โ€” Volatility 3 + LLM + RAG for memory-forensic triage. *(โ˜… 13 ยท updated 2025-09-16)* - **Related:** [Volatility-MCP-Server](https://github.com/bornpresident/Volatility-MCP-Server) - **[Volatility-MCP-Server](https://github.com/bornpresident/Volatility-MCP-Server)** ๐ŸŸข โ€” MCP exposing Volatility 3 plugins for natural-language memory forensics. *(โ˜… 39 ยท updated 2025-07-07)* - **Related:** [MemoryInvestigator](https://github.com/jan-hendrik-lang/MemoryInvestigator) - **[llm_osint](https://github.com/sshh12/llm_osint)** ๐ŸŸข๐Ÿ”ฌ โ€” Proof-of-concept LLM OSINT framework using knowledge and web agents for internet research workflows. *(โ˜… 317 ยท updated 2024-11-02)* - **[ai_osint](https://github.com/7WaySecurity/ai_osint)** ๐ŸŸข โ€” Curated AI-OSINT dorks, queries, and techniques for discovering exposed LLM and AI infrastructure. *(โ˜… 154 ยท updated 2026-06-19)* - **[PhishLLM](https://github.com/code-philia/PhishLLM)** ๐Ÿ”ฌโš ๏ธ โ€” Reference-less phishing detection via LLM brand recognition (USENIX'24). *(โ˜… 39 ยท updated 2026-06-04)* - **Related:** [PhishVLM](https://github.com/code-philia/PhishVLM) - **[mcp-dnstwist](https://github.com/BurtTheCoder/mcp-dnstwist)** ๐ŸŸข โ€” MCP server for dnstwist DNS fuzzing to support typosquatting, phishing, and lookalike-domain analysis. *(โ˜… 51 ยท updated 2025-03-03)* - **[osintgpt](https://github.com/estebanpdl/osintgpt)** ๐ŸŸขโš ๏ธ โ€” OpenAI embeddings + Qdrant over OSINT corpora. *(โ˜… 525 ยท updated 2023-12-11)* --- ## Related Awesome Lists - **[Awesome GUI Agent Security](https://github.com/Yuxuan2003/Awesome-GUI-Agent-Security)** โ€” Focused bibliography of GUI, computer-use, and browser-agent security research, organized by attack surface and defense layer with Chinese-language annotations. *(โ˜… 73 ยท updated 2026-10-03)* - **[Jailbreak Observatory](https://github.com/GenggengSvan/Jailbreak-Observatory)** โ€” Classified bibliography and visualizations of LLM jailbreak research; no root license was found, and publication labels are not independently verified. *(โ˜… 23 ยท updated 2026-10-01)* - **[Awesome Agentic Security](https://github.com/kagnlp/Awesome-Agentic-Security)** โ€” Survey bibliography covering agent attacks, defenses, and security applications; an MIT badge is present but no root license file was found. *(โ˜… 57 ยท updated 2026-08-24)* - **[Awesome LLMSecOps](https://github.com/wearetyomsmnv/Awesome-LLMSecOps)** โ€” Resource directory covering LLM architecture, operations, threat modeling, and security; no root license was found. *(โ˜… 162 ยท updated 2026-09-27)* - **[Awesome GenAI Security](https://github.com/jassics/awesome-genai-security)** โ€” GPL-3.0 collection of GenAI and agentic-security references, tutorials, labs, and testing resources; linked artifacts retain their own terms. *(โ˜… 78 ยท updated 2026-09-29)* - **[Awesome Hacking with AI](https://github.com/frangelbarrera/Awesome-Hacking-with-AI)** โ€” Directory of AI-assisted security and AI-system security resources covering agent safety, MCP, red teaming, and detection engineering; linked projects require separate review. *(โ˜… 28 ยท updated 2026-10-02)* - **[AI/ML Security and Prompt-Injection Resources](https://github.com/anmolksachan/AI-ML-Free-Resources-for-Security-and-Prompt-Injection)** โ€” AI-security learning roadmap and resource directory that distinguishes historical and unverified material; no root license was found. *(โ˜… 801 ยท updated 2026-09-19)* - **[Awesome AI Security TR](https://github.com/fevziegeyurtsevenler/awesome-ai-security-tr)** โ€” Turkish-language annotated directory of prompt-injection, red-teaming, guardrail, agent, MCP, and RAG security resources under CC BY 4.0; link checks do not validate security claims. *(โ˜… 13 ยท updated 2026-08-28)* - **[Awesome Prompt Injection](https://github.com/Joe-B-Security/awesome-prompt-injection)** โ€” CC0 resource list focused on prompt injection and related machine-learning security references. *(โ˜… 666 ยท updated 2026-09-11)* - **[Awesome AI Red Teaming JP](https://github.com/HayatoFujihara/awesome-ai-red-teaming-jp)** โ€” CC0 collection of Japanese-language AI red-teaming and safety resources. *(โ˜… 13 ยท updated 2026-06-20)* - **[Arcanum AI Security Resources](https://github.com/Arcanum-Sec/ai-sec-resources)** โ€” Categorized directory of AI-security labs and competitions; no root license was found, and external availability or free access is not guaranteed. *(โ˜… 513 ยท updated 2026-07-03)* - **[Awesome Agentic MCP Security](https://github.com/mcp-security-project/awesome-agentic-mcp-security)** โ€” CC0 collection of agentic-AI and MCP security resources. *(โ˜… 21 ยท updated 2026-09-26)* - **[Awesome Agent Sandbox](https://github.com/fishman/awesome-agent-sandbox)** โ€” Directory of agent sandbox runtimes and comparisons; no root license was found, and performance or security comparisons are not independently verified. *(โ˜… 18 ยท updated 2026-07-07)* - **[Awesome Agent Sandboxes](https://github.com/arjan/awesome-agent-sandboxes)** โ€” Directory of self-hosted and cloud code-execution sandboxes for AI agents; no root license was found, and listed products offer differing isolation boundaries. *(โ˜… 49 ยท updated 2026-08-19)* - **[Awesome LM Safety, Security, and Privacy](https://github.com/CryptoAILab/Awesome-LM-SSP)** โ€” Apache-2.0 bibliography of language and multimodal model safety, security, and privacy research; linked publications, models, and datasets retain their own licenses. *(โ˜… 2,082 ยท updated 2026-09-02)* - **[Awesome AI Security (DeepSpaceHarbor)](https://github.com/DeepSpaceHarbor/Awesome-AI-Security)** โ€” Historical collection of adversarial examples, evasion, poisoning, and ML-security research code; no root license was found, and recent commits do not establish the currency of linked material. *(โ˜… 1,680 ยท updated 2026-03-08)* - **[Awesome LLM Agent Security](https://github.com/wearetyomsmnv/Awesome-LLM-agent-Security)** โ€” Unlicense bibliography of LLM-agent security research, attacks, and vulnerabilities; linked papers and code retain their own terms. *(โ˜… 58 ยท updated 2026-07-11)* - **[Awesome MLLM/LLM Guardrails](https://github.com/ant-research/awesome-mllm-guardrails)** โ€” Focused research index of text and multimodal guardrail models, moderation datasets, jailbreak studies, evaluation resources, and runtime frameworks; linked artifacts retain their own terms. *(โ˜… 31 ยท updated 2026-09-03)* - **[awesome-security-agent-harnesses](https://github.com/Ed-Marcavage/awesome-security-agent-harnesses)** โ€” Focused CC0 catalog of vulnerability-discovery and pentest agent harnesses, security skills, sandboxes, MCP integrations, benchmarks, and evaluation research. *(โ˜… 22 ยท updated 2026-09-04)* - **[awesome-llm-cybersecurity-tools](https://github.com/tenable/awesome-llm-cybersecurity-tools)** โ€” Tenable's list (archived but a strong reference). *(โ˜… 488 ยท updated 2024-04-08)* - **[Awesome-LLM4Cybersecurity](https://github.com/tmylla/Awesome-LLM4Cybersecurity)** โ€” 600+ papers on LLMs for cybersecurity. *(โ˜… 1,748 ยท updated 2026-07-08)* - **[awesome-ai-cybersecurity](https://github.com/ElNiak/awesome-ai-cybersecurity)** โ€” Broad AI-for-security collection. *(โ˜… 151 ยท updated 2026-08-13)* - **[awesome-genai-cyberhub](https://github.com/Ashfaaq98/awesome-genai-cyberhub)** โ€” GenAI-driven cybersecurity resources. *(โ˜… 54 ยท updated 2026-08-07)* - **[awesome-ai-security](https://github.com/gmh5225/awesome-ai-security)** โ€” For pentesters, bug hunters, and researchers. *(โ˜… 39 ยท updated 2026-08-17)* - **[awesome-ai-security](https://github.com/ottosulin/awesome-ai-security)** โ€” AI security resources. *(โ˜… 1,397 ยท updated 2026-08-16)* - **[Awesome-AI-Security](https://github.com/TalEliyahu/Awesome-AI-Security)** โ€” AI security resources. *(โ˜… 853 ยท updated 2026-07-18)* - **[Awesome-AI-For-Security](https://github.com/AmanPriyanshu/Awesome-AI-For-Security)** โ€” AI-for-security tools, papers, and datasets. *(โ˜… 146 ยท updated 2026-08-13)* - **[awesome-cybersecurity-agentic-ai](https://github.com/raphabot/awesome-cybersecurity-agentic-ai)** โ€” Agentic-AI cybersecurity tools and security MCP servers. *(โ˜… 536 ยท updated 2026-06-28)* - **[awesome-ai-agents-security](https://github.com/ProjectRecon/awesome-ai-agents-security)** โ€” Focused map of AI-agent security resources across runtime protection, red-teaming scanners, static analysis, sandboxing, guardrails, benchmarks, and identity. *(โ˜… 63 ยท updated 2026-06-12)* - **[Awesome-Offensive-AI-Agentic-Landscape](https://github.com/Yeti-791/Awesome-Offensive-AI-Agentic-Landscape)** โ€” Offensive AI-agent landscape covering open-source pentest/red-team agents, offensive/security-specialized models, papers, benchmarks, and commercial tools. *(โ˜… 214 ยท updated 2026-07-20)* - **[awesome-ai-agent-attacks](https://github.com/webpro255/awesome-ai-agent-attacks)** โ€” Sourced timeline of real AI-agent security incidents, breaches, vulnerabilities, and attack techniques. *(โ˜… 65 ยท updated 2026-08-11)* - **[AI Security Repository Radar](https://github.com/Zero0x00/Ai-Security-radar-)** โ€” Daily-updated AI/LLM/MCP/RAG security repository radar with category, license, stars, and quality/relevance metadata. *(โ˜… 5 ยท updated 2026-08-17)* - **[awesome-MLSecOps](https://github.com/RiccardoBiosas/awesome-MLSecOps)** โ€” Curated MLSecOps resources spanning adversarial ML, LLM security, AI red teaming, model scanning, supply-chain protection, and MLOps pipeline security. *(โ˜… 450 ยท updated 2026-08-15)* - **[awesome-ai-guardrails](https://github.com/enguard-ai/awesome-ai-guardrails)** โ€” Catalog of AI guardrail models, tools, organizations, datasets, and papers, with useful Hugging Face model coverage. *(โ˜… 64 ยท updated 2026-07-30)* - **[open-source-llm-scanners](https://github.com/psiinon/open-source-llm-scanners)** โ€” Open-source LLM scanners and testing tools. *(โ˜… 109 ยท updated 2026-02-05)* - **[awesome-mcp-security](https://github.com/Puliczek/awesome-mcp-security)** โ€” MCP security resources, tools, writeups, and server/client risk references. *(โ˜… 729 ยท updated 2026-03-03)* - **[awesome-ml-security](https://github.com/trailofbits/awesome-ml-security)** โ€” Trail of Bits' curated machine-learning security resources. *(โ˜… 170 ยท updated 2026-02-06)* - **[awesome-ml-privacy-attacks](https://github.com/stratosphereips/awesome-ml-privacy-attacks)** โ€” Machine-learning privacy-attack papers and resources. *(โ˜… 640 ยท updated 2024-03-18)* - **[awesome-ml-for-cybersecurity](https://github.com/jivoi/awesome-ml-for-cybersecurity)** โ€” Large classic list of machine-learning-for-cybersecurity resources (stale-ish but still useful). *(โ˜… 9,290 ยท updated 2024-04-11)* - **[Awesome-AI4DevSecOps](https://github.com/awsm-research/Awesome-AI4DevSecOps)** โ€” Taxonomy of AI-driven security solutions for DevSecOps. *(โ˜… 19 ยท updated 2025-07-02)* - **[awesome-llm-security](https://github.com/corca-ai/awesome-llm-security)** โ€” Securing LLMs. *(โ˜… 1,683 ยท updated 2025-08-20)* - **[awesome-gpt-security](https://github.com/cckuailong/awesome-gpt-security)** โ€” GPT/LLM security tools and cases. *(โ˜… 669 ยท updated 2026-07-24)* - **[awesome-threat-intelligence](https://github.com/hslatman/awesome-threat-intelligence)** โ€” Classic CTI list (pairs with the AI-CTI section). *(โ˜… 10,541 ยท updated 2026-05-31)* - **[awesome-threat-modelling](https://github.com/hysnsec/awesome-threat-modelling)** โ€” General-purpose threat modeling list โ€” methodologies and non-AI tools (Threat Dragon, pytm, Threagile); dormant since 2023 but a solid reference. *(โ˜… 1,793 ยท updated 2023-07-15)* - **[Awesome-LLMs-for-Vulnerability-Detection](https://github.com/huhusmang/Awesome-LLMs-for-Vulnerability-Detection)** โ€” Focused, continuously updated index of LLM-based software-vulnerability detection research across function, repository, agentic, and smart-contract analysis, including datasets, benchmarks, and surveys. *(โ˜… 1,239 ยท updated 2026-08-16)* --- ## Contributing Contributions are welcome! This README is generated from structured data. 1. Edit `data/sections.json` (validated by `data/schema.json`). 2. Use structured fields such as `status`, `flags`, and `license`; do not paste rendered emoji tags into the data. 3. Regenerate and check the README: ```bash python3 scripts/update_github_metrics.py python3 gen_readme.py python3 gen_readme.py --check ``` Example entry: ```json { "name": "ExampleTool", "repo": "OWNER/REPO", "status": ["open_source"], "license": "MIT", "flags": ["early_stage"], "desc": "One factual sentence about what the tool does.", "related": [ {"label": "Sibling tool", "url": "https://github.com/OWNER/SIBLING"} ] } ``` Status values: `open_source`, `research`, `commercial_open`. Common flags: `license_caveat`, `early_stage`, `archived`, `heavy_runtime`, `requires_api_key`, `authorized_testing_only`, `commercial_features`, `telemetry`, `model_dependent`, `research`, `no_license`, `noncommercial`, `copyleft`, `abliterated_or_uncensored`. Guidelines: link the canonical upstream repo (not a fork); verify the URL resolves; tag the correct type and add a caveat flag/note for non-permissive, non-commercial, unclear, missing, or restrictive licenses; prefer real, installable projects over blog-only references. For Hugging Face model entries, include the model id, license, access status (open/gated), and artifact formats (for example Safetensors or ONNX). ## Contact Maintained by **Sergey Gordeychik** โ€” [scadastrangelove@gmail.com](mailto:scadastrangelove@gmail.com) ยท [blog](https://scadastrangelove.blogspot.com/) ยท [@scadasl](https://x.com/scadasl). ## License To the extent possible under law, the contributors have waived all copyright and related rights to this list ([CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/)). Linked projects retain their own licenses โ€” check each before use.