--- name: oma-deepsec description: Set up and run Deepsec vulnerability scans, triage, and CI gates. Use for Deepsec work or an explicitly requested agent-powered vulnerability scan. --- # Deepsec: Agent-Powered Vulnerability Scanner Driver ## Scheduling ### Goal Operate Vercel's `deepsec` security scanner inside a target repository safely and cost-consciously: bootstrap the `.deepsec/` workspace, write a tight `INFO.md`, run the right scan/process/triage/revalidate/export sequence, gate PRs in CI via `process --diff`, and grow project-specific matchers, surfacing real, revalidated findings without runaway spend. ### Intent signature - User mentions `deepsec`, "deep security scan", `bunx deepsec`, `pnpm deepsec`, `npx deepsec`. - User asks an agent to scan a repository for vulnerabilities, security issues, or CVEs and the project has (or should have) a `.deepsec/` directory. - User asks how to add a deepsec PR / CI security gate, or about `process --diff`, `--diff-staged`, `--diff-working`, `--files-from`, `--comment-out`. - User mentions deepsec artefacts: `INFO.md`, `SETUP.md`, `data//files/`, `FileRecord`, `RunMeta`, `revalidation`, `triage`, custom matchers, `MatcherPlugin`, `noiseTier`, `priorityPaths`. - User asks about deepsec configuration: `deepsec.config.ts`, `defaultAgent`, `AI_GATEWAY_API_KEY`, `VERCEL_OIDC_TOKEN`, AI Gateway, Vercel Sandbox, `--agent codex`, `--agent claude`. - User asks how to lower deepsec cost, cut false-positive rate, or interpret severity / triage / revalidation verdicts. ### When to use - First-time deepsec install in a repo (`init`, `INFO.md` write, first calibration scan). - Running a full or scoped scan and processing findings. - Setting up a per-PR CI gate with `process --diff` and `--comment-out`. - Writing a project-specific matcher to cover entry points the default set misses. - Triaging a backlog of findings (severity bucketing, FP cuts via `revalidate`, exporting to issue tracker). - Diagnosing deepsec failures: missing credentials, AI Gateway quota stops, refusals, sandbox auth. ### When NOT to use - Generic OWASP / lint-style review without deepsec → use `oma-qa`. - Generic CVE / dependency advisories → use `oma-qa` or `oma-search`. - Architecting a brand-new SAST pipeline that is not deepsec → use `oma-architecture`. - Writing or auditing application code itself → route to `oma-backend` / `oma-frontend` / `oma-mobile`. - Cloud / IAM / Terraform hardening → use `oma-tf-infra` (deepsec only scans the IaC; remediation lives there). - Pure reasoning about a finding's fix in product code → use `oma-debug` once deepsec has produced the finding. ### Expected inputs - `target_repo_root`: absolute path of the codebase to scan (parent of `.deepsec/`). - `intent`: one of `setup` | `scan` | `pr-review` | `matchers` | `triage` | `config` | `troubleshoot`. - `credential_mode`: `ai-gateway-key` | `vercel-oidc` | `direct-anthropic` | `direct-openai` | `subscription`. - `agent_choice`: use the user-named backend, configured `defaultAgent`, or a choice within delegated scope; ask only when a material choice remains unresolved. - `severity_floor`: lowest severity worth surfacing (typically `HIGH`). - Optional: existing `.deepsec/data//`, `deepsec.config.ts`, custom matchers, CI provider. ### Expected outputs - A working `.deepsec/` workspace registered against the target repo. - A populated `data//INFO.md` (50-100 lines, project-specific, no line numbers). - One or more completed `scan` → `process` (→ `triage`/`revalidate`) runs with reproducible cost notes. - For PR mode: a CI workflow file using `process --diff ` with two-job split (no PR-write in PR-code job). - For matchers: new `.deepsec/matchers/.ts` files wired through the inline plugin in `deepsec.config.ts`. - A findings export (`md-dir` and/or `json`) plus a short summary of top severities and FP-rate notes. - Explicit, dollar-and-time-bounded plan before any pass that may cost more than ~$25. ### Dependencies - Node.js **22+**, plus a package manager: `bun` / `bunx` (preferred in this monorepo), `pnpm`, `npm`, or `yarn`. - A working AI credential: `AI_GATEWAY_API_KEY=vck_…`, or `VERCEL_OIDC_TOKEN`, or direct `ANTHROPIC_AUTH_TOKEN` + `ANTHROPIC_BASE_URL`, or a logged-in `claude` / `codex` CLI subscription. - Git (history is consulted by `revalidate` and `--diff` modes). - Optional: Vercel Sandbox auth for `deepsec sandbox …` distributed runs. - Reference resources under `resources/` (loaded only when the scenario requires them). ### Control-flow features - Branches by `intent` (setup vs scan vs pr-review vs matchers vs triage vs config vs troubleshoot). - Branches by repo size (calibrate with `--limit 50` before any large pass). - Branches by credential source (gateway key, OIDC, direct, subscription). - Stops on quota / credit exhaustion and resumes the same command after top-up. - Refuses to launch an unbounded `process` when no calibration has been done and the repo is large. - Reads codebase, writes `.deepsec/` files and CI configs, runs long-lived AI processes. ## Structural Flow ### Entry 1. Confirm whether `.deepsec/` already exists; if yes, treat the run as **incremental**, never re-init. 2. Resolve `intent` from the user prompt; if ambiguous (e.g. "scan this repo"), default to `setup` then `scan` (calibration mode). 3. Estimate scale: count source files (rough `rg --files | wc -l` excluding `node_modules`, `.git`, `dist`) to forecast cost before any AI pass. 4. Resolve the selected credential mode and backend. Check its required environment configuration or an existing `claude` / `codex` subscription login, without echoing secrets. A valid subscription session does not require an API-key environment variable. Resolve only missing configuration before AI calls. 5. Resolve backend, scope, and spend from existing instructions and configuration under `../_shared/core/execution-policy.md`. Before paid or custom-scope work, record the actual approved, limited, or declined action using `resources/decision-records.md`. A configured backend does not authorize additional spend; ask only for a material missing choice or new authorization. ### Transitions - If `.deepsec/` is missing and intent involves scanning → run `bunx deepsec init` (or `npx deepsec init`) and follow the printed prompt to populate `INFO.md` before any AI pass. - If `INFO.md` is empty or template-shaped → write it (50-100 lines, project-specific, 3-5 examples per section, no line numbers, no generic CWE enumeration). - If repo is > 500 files and no calibration has run → run a calibration pass first (deepsec docs recommend `--limit 50 --concurrency 5`) and report cost extrapolation before the full pass. - If a `process` / `revalidate` run halts on quota → leave file locks intact, surface the exact remediation URL, **re-run the same command after top-up**. - If the agent reports a refusal (`refused: true`) → never silently drop; document the affected files and either retry with the other backend or add the path to `config.json:ignorePaths` only if reproducible. - If the user wants a CI gate → emit the two-job pattern (PR-code job has no `pull-requests: write`, comment job has no PR code). - If the user wants more matcher coverage → run the matcher-authoring workflow against `data//files/` and the parent repo's entry points. ### Failure and recovery | Failure | Recovery | |---------|----------| | `Missing AI credentials for --agent claude` / `codex` | Check the selected mode per `resources/config.md`: environment configuration for key/OIDC/direct modes, or CLI login for subscription mode. Do not require `.env.local` credentials for a valid subscription session. | | `401 Unauthorized` from gateway | OIDC: re-run `vercel env pull` (12 h expiry). API key: regenerate. Confirm `.env.local` is in the cwd deepsec runs from. | | `Stopped: AI Gateway credits exhausted` | Top up via the printed URL; re-run the same command, files already done are skipped. | | `Stopped: Claude Pro/Max subscription exhausted` | Switch to AI Gateway; subscriptions don't carry full scans. | | Persistent refusal on a single file (>5% of batches) | Add the path to `data//config.json:ignorePaths`, or run that file alone with `--batch-size 1`. | | FP rate too high on `HIGH+` | Run `revalidate --min-severity HIGH`; tighten `INFO.md`'s threat model and FP notes; bias matchers to `precise`. | | `noisy` matcher wedges scanner on a 100k-file repo | Tighten `filePatterns` to language- or directory-anchored globs. | | Sandbox auth fails | OIDC: re-run `vercel env pull`. Access-token mode: verify `VERCEL_TOKEN` + `VERCEL_TEAM_ID` + `VERCEL_PROJECT_ID`. | | Full pass exceeds existing scope or spend authorization | Report the calibrated estimate and preserve completed free/limited work. Resolve only the missing authorization; record the actual approved, limited, or declined pass before executing it. | ### Exit - **Success**: planned passes ran, findings exist with verdicts (or no findings produced), files written are listed, residual cost / followups are explicit. - **Partial success**: some passes blocked on credentials/quota/refusal; the blocker, the safe-resume command, and the recommended next step are reported. - **Failure**: nothing destructive happened, the user has the exact next command to unblock the work. ## Logical Operations ### Tools and instruments - **Package manager**: `bun` / `bunx` (preferred), `pnpm`, `npm`, `yarn` are interchangeable. - **CLI commands**: `deepsec init`, `init-project`, `scan`, `process`, `process --diff`, `triage`, `revalidate`, `enrich`, `report`, `export`, `metrics`, `status`, `sandbox `. - **Diff sources for PR mode**: `--diff `, `--diff-staged`, `--diff-working`, `--files `, `--files-from ` (or `-` for stdin). - **Inspection**: `jq` over `data//files/**/*.json` for ad-hoc severity / TP queries. - **Credentials**: `AI_GATEWAY_API_KEY`, `VERCEL_OIDC_TOKEN`, `ANTHROPIC_AUTH_TOKEN` / `ANTHROPIC_BASE_URL`, `OPENAI_API_KEY` / `OPENAI_BASE_URL`, `claude login`, `codex login`. - **Resource files** under `resources/` for setup, scanning, PR review, matchers, triage, config, load on demand. ### Canonical workflow path 1. **Bootstrap** (one time per repo): ```bash cd bunx deepsec init cd .deepsec bun install # Edit .env.local: set AI_GATEWAY_API_KEY=vck_… (or VERCEL_OIDC_TOKEN via `vercel env pull`) ``` Then prompt the coding agent (this skill) to read `.deepsec/node_modules/deepsec/SKILL.md` and `.deepsec/data//SETUP.md`, skim `README` / `AGENTS.md` / `CLAUDE.md` and a handful of representative files, and replace each section of `data//INFO.md` (50-100 lines, 3-5 examples per section, no line numbers, no generic CWE rehash). 2. **Calibrate before any full pass.** First record the selected authorized calibration scope using `resources/decision-records.md`; `--limit` bounds files, not dollar spend. The deepsec docs (`getting-started.md`, `vercel-setup.md`, `faq.md`) recommend `--limit 50 --concurrency 5` as the calibration starting point. ```bash bunx deepsec scan bunx deepsec status bunx deepsec process --limit 50 --concurrency 5 ``` Read the run cost and extrapolate to the full repo using `resources/scanning.md`. Reuse existing authorization covering the backend, scope, and estimated spend. Record the actual scope decision before calibration and again before an expanded pass; resolve only missing authorization. If the user names different `--limit` / `--concurrency` values, use theirs. 3. **Full investigation, triage, revalidate, export**: ```bash bunx deepsec process --concurrency 5 bunx deepsec triage --severity HIGH bunx deepsec revalidate --min-severity HIGH ``` Record and verify every triaged finding's actual verdict using `resources/decision-records.md` before filtering or suppression, including findings that will not be surfaced. Then export: ```bash bunx deepsec export --format md-dir --out ./findings bunx deepsec metrics ``` 4. **PR mode** (CI gate, scoped to changed files, exit code = 0/1): ```bash bunx deepsec process \ --diff origin/${BASE_REF} \ --comment-out comment.md ``` Wire the two-job CI pattern from `resources/pr-review.md`. Never grant `pull-requests: write` to the job that runs PR-controlled code. 5. **Custom matchers** (close entry-point gaps surfaced in step 3): - Read the contract in `.deepsec/node_modules/deepsec/dist/config.d.ts` and the `samples/webapp/matchers/*` examples. - Write `.deepsec/matchers/.ts`, wire it through the inline plugin in `.deepsec/deepsec.config.ts`. - Verify hit rate: `bunx deepsec scan --matchers ` should land in 1-20 hits / 1k files (`precise`), 5-100 (`normal`), or roughly the framework entry-point count (`noisy`). 6. **Resume** after any quota stop, network blip, or Ctrl-C: re-run the same command. State is on disk under `.deepsec/data//`. ### Resource scope | Scope | Resource target | |-------|-----------------| | `CODEBASE` | Target repo source files, framework configs, route directories, `README` / `AGENTS.md` / `CLAUDE.md`. | | `LOCAL_FS` | `.deepsec/deepsec.config.ts`, `.deepsec/.env.local`, `.deepsec/matchers/`, `.deepsec/data//{project.json,INFO.md,config.json,files/,runs/,reports/}`, generated `findings/`, `comment.md`, CI workflow files. | | `PROCESS` | `bunx deepsec scan|process|triage|revalidate|export|metrics|status|sandbox`, `bun install`, optional `vercel link` / `vercel env pull`. | | `NETWORK` | Anthropic / OpenAI via Vercel AI Gateway (default) or direct provider endpoints; optional Vercel Sandbox microVM control plane. | | `CREDENTIALS` | `AI_GATEWAY_API_KEY`, `VERCEL_OIDC_TOKEN`, `ANTHROPIC_AUTH_TOKEN`, `OPENAI_API_KEY`, `VERCEL_TOKEN` / `VERCEL_TEAM_ID` / `VERCEL_PROJECT_ID`, `claude` / `codex` subscription tokens. Consume read-only; never echo secrets back to the user or commit them. | | `MEMORY` | User-stated budget cap, severity floor, and stop conditions for the current session. | ### Preconditions - Node.js 22+ is available. - Repo is a git checkout (deepsec uses git history for `revalidate` and `--diff`). - For any AI command: at least one credential mode is configured *before* the call, or the call is held until one is. - For `sandbox` mode: Vercel auth is wired; otherwise stay local. - For unbounded `process`: measured file scope and calibrated cost must fit existing authorization; otherwise use an authorized limited pass or resolve the missing spend/scope decision. ### Effects and side effects - Creates `.deepsec/` (config, lockfile, scaffolding) and `.deepsec/data//` (gitignored) inside the target repo. - Writes `.env.local` (never commit) and may run `vercel link` / `vercel env pull` (writes `.vercel/project.json` + token). - Spawns long-running AI processes that **cost real money**. Single full scans range from $25 to over $1,200 per the official cost guide and can climb to tens of thousands on very large repos. - Reads source code; sends snippets to the configured LLM (gateway = zero retention; direct provider = subject to that provider's policy). Never exfiltrates secrets; the gateway key stays outside the worker sandbox in `sandbox` mode. - May write `.github/workflows/deepsec.yml` (or analogue) when the user asks for a CI gate. - Edits `deepsec.config.ts` and adds `.deepsec/matchers/*.ts` when authoring matchers. - Does not commit, push, or open PRs unless the user explicitly authorizes a separate commit step (route via `oma-scm`). ### Guardrails 1. **Never launch an unbounded `process` on a repo whose size you have not measured.** Always run a calibration pass first when file count is unknown or > 500 (deepsec docs recommend `--limit 50 --concurrency 5`; defer to a user-named value if given). 2. **State cost and stopping condition before any AI pass.** Use the published bands (100 files ≈ $25-60, 500 ≈ $130-300, 2,000 ≈ $500-1,200; ×2-3 swing). 3. **Resume, do not reset.** After any network / quota / Ctrl-C interruption, re-run the same command. Never delete `data//` to "start clean" without explicit user instruction. 4. **`INFO.md` stays short and project-specific.** 50-100 lines, 3-5 examples per section. Name primitives but no line numbers. Skip generic CWE categories; built-in matchers cover those. 5. **For PR/CI gates, keep PR-controlled code in a no-write job.** Never grant `pull-requests: write` to a job that executes PR-controlled `pnpm install` / config-loading. Use the two-job pattern in `resources/pr-review.md`. 6. **Pin actions to full SHAs** in production CI; major-version tags are for examples only. 7. **Never silently drop refusals.** If the agent reports `refused: true`, log it, retry with the other backend, or add the file to `ignorePaths` only when reproducible. 8. **Bias matchers toward `precise` when the bug shape is exact.** Reserve `noisy` for entry-point coverage and tight globs. 9. **Never echo or commit credentials** (`vck_…`, `sk-ant-…`, `sk-…`, OIDC tokens). Treat `.env.local` as secret. Treat `data/` as gitignored by default. 10. **Treat deepsec like an agent with shell access.** Recommend `sandbox` for prompt-injection-prone repos (vendored code, untrusted deps). 11. **Findings need verdicts.** For any HIGH+ surfaced to the user, prefer `revalidate`-tagged verdicts (`true-positive` / `false-positive` / `fixed` / `uncertain`) over raw `process` output. 12. **Do not invent CLI flags, and trust the CLI over these notes.** Anything beyond `resources/scanning.md`'s flag list must be checked against `--help` first. Likewise, when the CLI's printed model names, defaults, or per-batch costs disagree with the values written in this skill, the CLI is right — upstream moves faster than these resources. 13. **Reuse existing backend, scope, and spend decisions.** Record consequential execution choices and per-finding verdicts through `resources/decision-records.md`; do not add another confirmation when existing authorization covers the action. ## References - L1 execution scope and per-finding verdict records: `resources/decision-records.md` (before paid/custom-scope work or filtering triaged findings). - Workspace install + `INFO.md` bootstrap: `resources/setup.md` - Full scan/process/triage/revalidate/export workflow + cost guide: `resources/scanning.md` - PR / CI gate via `process --diff` (two-job pattern, exit-code semantics): `resources/pr-review.md` - Authoring custom matchers (slugs, noise tiers, file globs, plugin wiring): `resources/matchers.md` - Reading findings, severities, triage / revalidation verdicts, FP cuts: `resources/triage.md` - `deepsec.config.ts` reference, env vars, plugin order, AI Gateway / Vercel Sandbox auth: `resources/config.md` - Upstream docs (load only when a resource file points at one): - Repo + README: https://github.com/vercel-labs/deepsec - Per-topic docs at https://github.com/vercel-labs/deepsec/tree/main/docs (`getting-started`, `reviewing-changes`, `writing-matchers`, `configuration`, `models`, `plugins`, `architecture`, `data-layout`, `vercel-setup`, `supported-tech`, `faq`) - Shared context loading: `../_shared/core/context-loading.md` - Shared quality principles: `../_shared/core/quality-principles.md`