nullpointer.studio

# agent-smith [![Quality Gate Status](https://sonarcloud.io/api/project_badges/measure?project=0x0pointer_agent-smith&metric=alert_status)](https://sonarcloud.io/summary/new_code?id=0x0pointer_agent-smith) [![Bugs](https://sonarcloud.io/api/project_badges/measure?project=0x0pointer_agent-smith&metric=bugs)](https://sonarcloud.io/summary/new_code?id=0x0pointer_agent-smith) [![Coverage](https://sonarcloud.io/api/project_badges/measure?project=0x0pointer_agent-smith&metric=coverage)](https://sonarcloud.io/summary/new_code?id=0x0pointer_agent-smith) **The pentest framework built for the tester who wants to think, not babysit.** You bring the expertise. agent-smith brings 50+ tools, the methodology, and the execution โ€” and the two of you close findings that neither could reach alone. > โš ๏ธ **Authorized testing only.** Use against systems you own or have explicit written permission to test. Unauthorized access is illegal.

agent-smith running a full /pentester scan from recon through reporting

--- ## Why agent-smith - ๐Ÿง  **The LLM is the brain, not a payload library.** Skills teach *methodology*; the LLM invents the actual attacks. No two scans look 100% alike. - ๐Ÿ”— **Skills chain themselves.** `/pentester` finds an injection point and pivots into `/web-exploit`; `/codebase` finds an LLM call site and pivots into `/ai-redteam`. The agent decides what runs next based on what it just found. - ๐Ÿ›  **Bring your own LLM.** Claude Code, OpenAI Codex, [OpenCode](https://opencode.ai) (any provider โ€” OpenAI, Gemini, Ollama, OpenRouter, local models), or any MCP-capable client. Smith auto-scales its context budget to small / local models, so it even runs **fully local on your own GPU** โ€” no API bills, nothing leaving your network. ([setup โ†’](docs/installation.md#self-hosted-local-model-dgx-spark--vllm)) - ๐Ÿ“ฆ **End-to-end deliverables.** Findings, PoCs (Burp-ready `.http` files), threat models, code patches, GitHub issues, and CVE submission packages โ€” all generated for you. - ๐Ÿณ **Sandboxed by default.** Every scanner runs inside an ephemeral Docker container. Hard cost / time / call-count limits enforced server-side. - ๐Ÿ” **Depth enforcement.** A background QA daemon watches Smith and pushes it to go deeper โ€” catching stalls, premature completion, and shortcut behaviour, escalating to you when it's genuinely stuck. ([details โ†’](docs/operating.md#qa-depth-enforcement)) - ๐Ÿงช **Evidence, not guesses.** Every finding is artifact-backed and passes a senior-review **adjudication gate**. Blind vulns are confirmed **out-of-band** via a callback server; multi-step attacks are recorded as **proven exploit chains**; white-box findings carry a **source trace** whose `file:line` is resolved against the repo โ€” a hallucinated location is rejected at the door. - ๐Ÿ“Š **A dashboard built for collaboration.** Watch findings, topology, coverage, and the threat model populate in real time at `localhost:7777` โ€” and steer Smith mid-scan, respond to intervention pauses, and fulfill its resource wishlist. ([API โ†’](docs/dashboard-api.md)) --- ## The new way: skills as pattern teachings Most pentest automation ships a giant payload library and runs it linearly. agent-smith does the opposite. **Skills are not scripts โ€” they are prompts that teach the LLM a way of *thinking*.** A skill describes the vulnerability class, the surface area, the verification logic, and the chaining rules, then leaves the actual attacks to the model. | Traditional security tools | agent-smith | |---|---| | Fixed payload list | LLM-generated payloads, contextual to each target | | One tool per phase | Skills compose โ€” `/codebase` enriches `/pentester`, which enriches `/post-exploit` | | Stops at first success | Keeps probing until the cost / time / coverage budget is hit | | Generates a PDF | Findings, PoCs, patches, threat models, coverage matrix, CVE packages, and more | | Same scan every time | Two runs against the same target produce different attack paths | The skills are inspiration. The LLM is the operator. --- ## See it in action

/pentester โ€” full autonomous engagement

pentester running from recon through reporting

Recon โ†’ fingerprint โ†’ exploit โ†’ loot โ†’ report. The agent decides every step.

/codebase โ€” white-box ASVS review

codebase skill performing an ASVS 5.0 review

Source โ†’ routes โ†’ sinks โ†’ ASVS chapters โ†’ enriched context for every downstream skill.

/ai-redteam โ€” OWASP LLM Top 10 + AITG

ai-redteam executing prompt injection and jailbreak chains

Prompt injection, jailbreaks, model extraction, MCP runtime attacks, and post-access infra checks.

/remediate โ€” auto-generated patches

remediate skill writing fixes for every confirmed finding

For every confirmed finding the agent writes a code or config patch and verifies it doesn't break the build.

--- ## What you can do Drop any of these the moment you start your client. `/pentester` orchestrates everything; the single-purpose skills give you laser focus. | Command | What it does | |---|---| | `/pentester scan https://staging.example.com depth=thorough` | Full hands-off engagement: OSINT โ†’ recon โ†’ web-exploit โ†’ post-exploit โ†’ report | | `/codebase path=./src` | White-box OWASP ASVS 5.0 review across 16 chapters / 427 requirements | | `/analyze-cve lodash 4.17.20 CVE-2021-23337` | Traces a CVE from user input to sink in your tree, decides if you're exploitable, writes a Burp PoC | | `/ai-redteam https://your-app.com/api/chat depth=thorough` | OWASP LLM Top 10 (2025) + AITG v1 + MCP Top 10 runtime attacks | | `/request-cves` | MITRE CVE form + GHSA draft + disclosure report + vendor email, per qualifying finding | | `/threat-modeling` | PASTA + STRIDE โ€” component map, data-flow diagram, attack tree, risk register | > ๐Ÿ’ก 35+ skills total โ€” full catalog and chaining map in **[docs/skills.md](docs/skills.md)**. --- ## Two ways to work with Smith The industry is racing toward full automation. We think that's the wrong finish line. The best pentests have always been about the *interplay* between a skilled tester and their tooling โ€” the tester brings context, intuition, and judgment; the tools provide speed, coverage, and consistency. agent-smith is built around that conviction. Autonomous mode exists and it's genuinely powerful, but **Augmented mode is where the framework does its best work**: Smith handles 50+ parallel tool runs, tracks coverage, and writes the deliverables, while you stay in the loop to redirect scope, respond to intervention pauses, and make the calls that no AI should make alone. The dashboard isn't a progress bar โ€” it's the collaboration interface.

โญ Augmented Mode โ€” recommended

A human expert drives the strategy. Smith handles the heavy lifting โ€” running 50+ tools in parallel, tracking coverage, writing the report. You steer via the dashboard: respond to intervention pauses, inject directives, shift scope mid-scan.

Best for: High-value engagements ยท Complex targets ยท When expert judgment shapes the approach

claude          # interactive โ€” you steer, Smith executes

Autonomous Mode

Give Smith a target โ€” it runs the full engagement end-to-end with no human input. Returns a report with every finding verified and a working proof-of-concept.

Best for: Continuous testing ยท CI/CD pipelines ยท Recurring red-team drills ยท Self-serve by dev teams

claude -p "/pentester scan https://staging.example.com depth=thorough"
--- ## Quick start | Requirement | Notes | |---|---| | [Docker Desktop](https://www.docker.com/products/docker-desktop/) | Must be running โ€” all scanners are sandboxed | | [Poetry](https://python-poetry.org) | `curl -sSL https://install.python-poetry.org \| python3 -` | | **One LLM client** | Claude Code ยท Codex ยท OpenCode (BYO LLM) ยท any MCP client | | [Node.js](https://nodejs.org) v18+ | Optional โ€” server-side Mermaid pre-rendering | ```bash git clone --recursive cd agent-smith ./installers/install.sh # Claude Code (or install_codex.sh / install_opencode.sh) ``` > โš ๏ธ **After install, fully restart your client** โ€” the MCP server connects at startup. **Full setup** โ€” other clients (Codex, OpenCode, custom MCP), self-hosted local models (vLLM / DGX Spark), Windows / PowerShell, and the optional Kali & Metasploit images โ†’ **[docs/installation.md](docs/installation.md)**. > ๐Ÿ›ก๏ธ **Running a real engagement?** Smith ingests attacker-controlled data and can run commands, so prompt injection is a design reality โ€” run it in an isolated, disposable VM. See **[docs/production-isolation.md](docs/production-isolation.md)**. --- ## How it works ``` You (/pentester scan target.com) โ””โ”€โ”€ Your LLM (Claude / GPT / Gemini / local โ€ฆ) โ””โ”€โ”€ MCP server (python -m mcp_server) โ”œโ”€โ”€ Lightweight scanners โ€” docker run --rm (nmap, nuclei, httpx, โ€ฆ) โ”œโ”€โ”€ Kali container โ€” persistent kali-mcp (nikto, sqlmap, ffuf, โ€ฆ) โ”œโ”€โ”€ Metasploit container โ€” exploit validation โ””โ”€โ”€ FastAPI dashboard โ€” live findings at localhost:7777 ``` The LLM decides what to run; each tool's output is aggregated and returned to the model, which chooses the next action. Hard cost / time / call-count limits are enforced server-side. Full component diagram, repository layout, and scan deliverables โ†’ **[docs/architecture.md](docs/architecture.md)**. --- ## Every scan builds training data Every pentest Smith runs is also a **structured, redacted dataset of the engagement** โ€” each *decision โ†’ action โ†’ result โ†’ finding* is captured as a schema-versioned event stream and retained per engagement (with the raw artifacts the model actually saw). It's a **byproduct: the capture is passive, read-only, leak-scanned, and never influences the scan** (opt out with `SMITH_EVENTS_DISABLED=1`). The value compounds: **the more pentests you run, the more data you accumulate โ€” and the better the model you can distill from it.** Pool the streams into a behaviour-cloning dataset and fine-tune a small open-weight base with QLoRA into a **LoRA adapter that runs *as Smith* locally**. Diversity of targets beats raw volume. โ†’ **[docs/training-data.md](docs/training-data.md)** โ€” what's captured, the safety model, and the exporter + DGX Spark QLoRA harness. --- ## Documentation **Getting started** - [installation.md](docs/installation.md) โ€” every client, self-hosted local models, Windows, optional images **Concepts** - [architecture.md](docs/architecture.md) โ€” component diagram, project layout, what a scan produces - [skills.md](docs/skills.md) โ€” full skill catalog, chaining map, per-skill reference - [training-data.md](docs/training-data.md) โ€” every scan builds a redacted dataset you can distill into a local LoRA adapter **Operating Smith** - [operating.md](docs/operating.md) โ€” scan modes, Human Intervention (HIR), lifecycle, QA depth enforcement - [production-isolation.md](docs/production-isolation.md) โ€” running Smith sandboxed in production - [dashboard-api.md](docs/dashboard-api.md) โ€” FastAPI endpoints and response shapes **Reference** - [tools.md](docs/tools.md) โ€” all MCP tools: parameters, purpose, examples - [kali-toolchain.md](docs/kali-toolchain.md) โ€” full `kali` command reference **Contributing** - [extending.md](docs/extending.md) โ€” adding new tools and skills - [testing.md](docs/testing.md) โ€” running the test suite, coverage, adding tests > **Adding a new skill?** Skills live in a separate repo ([github.com/0x0pointer/skills](https://github.com/0x0pointer/skills)) pulled in as a git submodule. After adding a skill there, update the submodule pointer (`git add skills && git commit`) and re-run the installer to deploy it. --- ## License GNU Affero General Public License v3.0 โ€” see [LICENSE](LICENSE). > Built for offensive-security professionals. Use it to make the internet safer.