--- name: triage description: Default pipeline scrutineer runs when a repository is added. Triggers a standard set of other skills in parallel, then writes a short summary of what was enqueued. Edit the list below to change the default scan coverage without touching scrutineer's Go code. license: MIT compatibility: Needs network access to the scrutineer API (http://host:port/api). Uses `brief` (github.com/git-pkgs/brief) for language and dependency gates; falls back to the general code and package scans if brief is unavailable. Repository modes use source evidence independently. metadata: scrutineer.version: 1 scrutineer.output_file: report.json scrutineer.output_kind: freeform --- # triage Scrutineer automatically queues `reflect` after a successful root default-branch triage invocation and waits for this invocation's child scans before running it. Do not enqueue `reflect` yourself or wait for it. It records operational lessons only and does not change scan selection or finding dispositions. Kick off the standard set of scans against a freshly-added repository. ## Workspace - `./context.json` — has `scrutineer.api_base`, `scrutineer.token`, and `scrutineer.repository_id`. Required. - `./report.json` — write a short summary of what you enqueued. ## Classify the repository Run `brief --json ./src` and read its JSON. These fields decide which scans are worth queueing: - `languages` — programming languages detected; `null` or empty means no source - `package_managers` — manifest/lockfile ecosystems detected; `null` or empty means nothing publishes or consumes packages here From those, set two flags: - `has_packages` = `package_managers` is present and non-empty - `has_code` = `languages` is present and non-empty A docs repo (markdown only) has neither. Anything with detected source gets the code scans, whether it is an application, a library, infra scripts, or a flat-layout package. The cost of running semgrep on a stray shell script is far lower than missing a real library because brief failed to populate `layout.source_dirs`. Do not gate on `layout.source_dirs`; it is a heuristic and routinely empty for legitimate codebases. Set `has_embedded_native` when Brief reports any of these signals: - `tools.native_extension` is non-empty - `tools.dependency_bot` contains `Git Submodules` - the detected languages include at least one of C, C++, Objective-C, Objective-C++, Rust, Go, Fortran, Zig, Assembly, or CUDA, and either another language or a CMake, Meson, or Autotools build tool is also present This covers native extensions and mixed-language repositories without assuming that native code means C or C++. If `brief` is not on PATH or exits non-zero, set `has_code`, `has_packages`, and `has_embedded_native` true and carry on so the follow-up scan can retry Brief after source preparation. `languages` also decides one language-specific scanner: - `has_python` = `languages` contains Python, matched case-insensitively This gates `bandit`, which reads Python and nothing else. It follows `has_code`: when `brief` is unavailable it is true too, since a missed Python codebase costs more than a scan that finds nothing. `brief` does not report CI configuration, so set a fourth flag from a direct filesystem check: - `has_workflows` = `./src/.github/workflows` exists and is a directory This gates `zizmor`, which only audits GitHub Actions workflows; with no workflows directory its scan immediately no-ops, so skipping it at triage avoids enqueueing a scan that can do nothing. Unlike the code/package flags this is a definitive check, not a heuristic, so do not default it true on error — if the directory is absent, `has_workflows` is false. ## The scan set Read [references/modes.md](references/modes.md) in this skill's directory and check each listed repository type against the source. Record matching modes and their file-level evidence in `modes`. Several modes can apply to one repository. Mode detection is independent of Brief's language and dependency detection, including when Brief fails. Add each matching mode's skills to the scan set below. Put skills belonging only to unmatched modes in `gated`. Apply the same skip set, request scope, and error handling to mode skills as to the standard scans, and enqueue each skill at most once. Mode scans are additive: keep the standard scans and the threat-model-driven deep dives. Before enqueueing anything, check what already ran so a re-trigger does not double-enqueue work that is already current. Get the commit you are running at: `git -C ./src rev-parse HEAD`. Then fetch `GET {api_base}/repositories/{repository_id}/scans`, which returns every scan on this repository with `skill_name`, `status`, `commit`, `ref` and `sub_path`. If that fetch fails, treat the skip set as empty and carry on. Consider only scans whose `ref` and `sub_path` match the current `scrutineer.scan_ref` and `scrutineer.scan_subpath`, treating absent fields as empty. Otherwise a scan of a sibling subproject could suppress this scope's audit. Build a set of skill names to skip: a skill goes in the skip set if it has a compatible scan with `status` in {`queued`, `running`}, or a compatible scan with `status="done"` whose `commit` equals the current HEAD. A `done` scan at any other commit does not count; the repository has moved since then and the skill should run again. `failed` scans are re-enqueued. Classify each skill in the list below into exactly one bucket, checking in this order and stopping at the first match: `gated` (its `has_code`/`has_packages`/`has_python`/`has_workflows`/`has_embedded_native` flag is false), `already_done` (it is in the skip set), `triggered` (enqueue it). Enqueue with `POST {api_base}/repositories/{id}/skills/{name}/run` and an `Authorization: Bearer {token}` header. Order does not matter; the scrutineer worker runs them as they come in. A 404 response moves the skill from `triggered` to `skipped`. If `scrutineer.scan_ref` is set in `context.json`, include it in the POST body as `{"ref": ""}` so child scans clone the same branch. If `scrutineer.scan_subpath` is set, also include `"sub_path": ""` in the same body so every child stays scoped to the same monorepo sub-package — a scan submitted as `repo#sub/dir` (or a `/tree//` URL) sets this, and without forwarding it the pipeline would silently widen back to the whole repository. Combine them when both are present, e.g. `{"ref": "main", "sub_path": "activesupport"}`. When both are empty, send an empty JSON body or omit it. Verify runs (below) always send `{}`; they are finding-scoped and take neither. Always: - `metadata` - `repo-overview` - `packages` - `advisories` Only when `has_workflows`: - `zizmor` Only when `has_python`: - `bandit` Only when `has_packages`: - `dependencies` - `sbom` `packages` and `advisories` query ecosyste.ms by repository URL rather than reading local manifests, so they run unconditionally even though they sound package-related. Only when `has_code`: - `subprojects` - `recon` - `history` - `threat-model` - `semgrep` - `betterleaks` Only when `has_embedded_native`: - `embedded-native` After `threat-model` finishes, Scrutineer reads its durable `scan_config` and enqueues one `security-deep-dive` per focus area. Do not enqueue `security-deep-dive` here: starting it before the threat model is complete would create an unscoped repository-wide audit and defeat the partition. For one third of eligible triage runs, Scrutineer also schedules at most one exploratory `security-deep-dive` after a successful threat model has queued a planned audit. Do not request this extra scan yourself. When the threat model names a usable out-of-scope source directory, half of selected root scans challenge that exclusion with an adversarial sweep; other selected scans use a context-free random dig. Both keep explicit path exclusions and subproject scope. Repeated completion notifications do not select another extra audit. If a skill name comes back `404 skill not found or inactive`, skip it and note which one in your report; the operator may have disabled it on purpose. ## Re-verify reported findings If this repository has been scanned before there may be findings already reported to the maintainer that have since been fixed upstream. For each of `status=reported` and `status=acknowledged`, fetch `GET {api_base}/repositories/{repository_id}/findings?status={status}` and collect the returned `id` values. For every finding id, enqueue a verify run: `POST {api_base}/findings/{id}/skills/verify/run` with the bearer header and an empty JSON body. Record the ids you enqueued in the `verify` field of your report; if there are none, write an empty list. If the verify endpoint returns `404 skill not found or inactive`, leave `verify` empty and carry on. Do not verify findings in `new`, `enriched`, `triaged`, `ready`, `published`, `rejected`, or `duplicate` states. The audit skills re-running above handle the first four; the last three are closed. ## Watch fixed findings for an upstream release When a finding reaches `fixed` the maintainer has landed a patch, but consumers cannot pin to a commit — they need a tagged release. For findings in `status=fixed`, enqueue release-watch the same way: `POST {api_base}/findings/{id}/skills/release-watch/run`. Record the ids in a `release_watch` field of your report; if there are none, write an empty list. If the endpoint returns `404 skill not found or inactive`, write an empty list and carry on. Release-watch is idempotent: a finding that already has a release recorded re-confirms the existing value rather than flapping. ## Output Write `./report.json` as: ```json { "has_code": true, "has_packages": true, "has_python": false, "has_workflows": false, "has_embedded_native": true, "modes": [], "brief": {"languages": ["Ruby", "Rust"], "package_managers": ["Bundler", "Cargo"], "native_signals": ["native_extension:rb-sys", "language:Rust"]}, "triggered": ["packages", "advisories", ...], "skipped": ["semgrep"], "gated": ["zizmor"], "already_done": ["metadata"], "verify": [12, 34], "release_watch": [55, 56], "errors": [] } ``` `gated` lists skills that were not enqueued because `has_code`, `has_packages`, `has_python`, `has_workflows`, or `has_embedded_native` was false. `already_done` holds skills that were skipped because a scan is currently running or already completed at this commit. `skipped` is for skills that came back `404 skill not found or inactive`. `brief` is the subset of brief's output the gates were derived from, including short `native_signals` entries for the embedded-native decision, so an operator can see why a repo got the short treatment and re-run triage manually if the classification was wrong. Do not wait for any of the scans to finish. The API returns a scan id immediately; your job is to fire them off and exit. Each `modes` entry has a `name` and a non-empty `evidence` list of source paths with short explanations. For example, a `package-manager` match should cite the install command and its implementation. An empty list means no mode matched; record unreadable source or inconclusive classification in `errors`. `gated` also includes skills whose repository type did not match. Do not fabricate scans or invent skill names. If the `api_base` or `token` is missing from context.json, write `{"error": "context.json missing scrutineer block"}` and exit 0 so the failure is visible on the scan page.