The open, recomputable evidence profile for privileged MCP tool actions.
Assay records what a privileged tool call decided, what was observed, and what stays unproven, so a reviewer can replay the claim offline instead of trusting the agent's account of itself. In enforcement mode, the proxy gate for routed MCP tools/call requests is deterministic and fail-closed, and the enforcing proxy is the reference producer rather than the contract itself. Optional eBPF/LSM instrumentation on supported Linux hosts adds kernel-level observations. CI-native, no backend, bounded by design.
Quickstart · How it works · See it work · MCP example · OWASP MCP Top 10 · Discussions
--- Agents got real tool access through MCP — and tool poisoning, rug pulls, and confused-deputy OAuth came with it. Assay sits at the tool-call boundary and does three things, in order. **One golden path:** the [release-pinned agent journey](docs/guides/agent-golden-path.md) records the nine driven CLI/MCP steps and their exit/stdout contracts. Its protected-action fixture lives in [examples/privileged-action-gate/](examples/privileged-action-gate/). ### Enforce, prove, stay honest - **Enforce.** In enforcement mode, the gate decides `tools/call` requests routed through it before forwarding, with the precise reason for each allow or deny. Separate controls on supported Linux hosts include cgroup eBPF IPv4/TCP connect filtering (whose hook error path allows connections) and Landlock TCP-connect port allowlisting. `assay sandbox --enforce` refuses an unavailable backend or conflicting filesystem policy unless `--allow-audit-fallback` is explicit. `--enforce-net` requires `--enforce`; active Landlock network enforcement rejects policies it cannot express as TCP-port allowlists, even with audit fallback enabled. - **Prove.** Configured producers can record decisions and bounded observations for export into offline-verifiable, tamper-evident evidence bundles. The privileged-action flow carries the verdict, pre-call establish journey, and declared-vs-observed conformance for CI review without a hosted backend. Basic `assay mcp wrap` does not automatically create a bundle; enable the required recording and export steps. - **Stay honest.** Trust Basis classifies supported claims as `verified`, `self_reported`, `inferred`, or `absent`; its gates check the declared claim boundaries. A tool returning "success" is the provider's assertion, never proof. Assay ships no single safety score; read each artifact’s source, coverage and non-claims before relying on it. ### Quickstart ```bash # Fast path: release installer for Linux and macOS. curl -fsSL https://getassay.dev/install.sh | sh # Confirm the command resolves; if setup fails, run `assay doctor`. assay --version # Source-build alternative (requires Rust): cargo install assay-cli --version 6.9.0 --locked python3 examples/mcp-quickstart/run.py ``` For v6.9.0, run the last command from a source checkout or an extracted published CLI archive. The installer is binary-only and does not carry the bounded quickstart assets. The live `getassay.dev` installer verifies the selected archive against its published SHA-256 sidecar before extraction. Set `ASSAY_REQUIRE_PROVENANCE=1` to additionally require GitHub artifact provenance; the default reports `provenance_not_requested` and strict success reports `provenance_verified`. A checksum proves byte equality with the published sidecar, not producer identity. Provenance identifies the source and build, not runtime safety or semantic correctness. Captured runner output (the bundled local mock performs no external action): ```text assay quickstart: PASS mcp_requests=initialize,tools/list,tools/call decision=allow tool=read_file decision_artifact=.assay/quickstart/decisions.ndjson non_claim=forwarded_to_local_mock_only ```  Released surfaces: - Static project manifests are shipped for Claude Code and Cursor; Codex uses the equivalent TOML entry documented in the [editor MCP recipe](docs/guides/editor-mcp-recipe.md). Manifest presence is not host-discovery proof. `assay mcp config-path` supports Claude Desktop and Cursor only. - Published v6.9.0 CLI archives cover Linux x86_64/arm64, macOS x86_64/arm64, and Windows x86_64. The Python wheels cover CPython 3.12, 3.13, and 3.14 on macOS x86_64/arm64 and Linux x86_64; other interpreters and platforms are not claimed. - Published `assay-mcp-server` archives cover Linux x86_64/arm64. MCPB and `server.json` package descriptors are also published; their presence is not host-discovery proof. - CI: [GitHub Action](https://github.com/marketplace/actions/assay-ai-agent-security). Core flows need no hosted backend or API key. New to the threat model? The [OWASP MCP Top 10 mapping](docs/security/OWASP-MCP-TOP10-MAPPING.md) states, per risk, what Assay covers and deliberately does not. ## What ships | Output | What it is | |--------|------------| | **Policy gate** | `assay mcp wrap` — deterministic allow/deny before tools run, with the reason. | | **Evidence bundle** | Offline-verifiable, tamper-evident archive for audit and replay. | | **Trust Basis / Trust Card** | Canonical `trust-basis.json` (bounded claim classification) plus review-friendly `trustcard.{json,md,html}`. | | **External receipts** | Eval outcomes, runtime decisions, and model inventory as bounded receipts with JSON Schema contracts. | | **Tool-decision logs** | For handled known-tool calls, the `assay-mcp-server` stdio server emits an info-level `tool_decision` event when enabled by its log filter; `decision` contains a JSON-encoded observed decision entry with projected target fields. | | **SARIF / CI** | GitHub Action, Security-tab integration, policy gates on PRs. | | **Attestation** | Sign an evidence bundle as a DSSE-wrapped in-toto v1 Statement with the evidence-bundle/v1 predicate. | ```text Agent ──► Assay ──► MCP Server ├─ ✅ ALLOW / ❌ DENY (policy, with reason) ├─► 📋 Evidence bundle (offline-verifiable) └─► 📊 Trust Basis → Trust Card → SARIF / CI ``` Current release: [`v6.9.0`](https://github.com/Rul1an/assay/releases/tag/v6.9.0). [CHANGELOG.md](CHANGELOG.md) and release notes remain the authority for released behavior; merged changes after the tag are `Unreleased`, and crates.io publication is separate from merge state. Launch definition and support commitment: [docs/LAUNCH.md](docs/LAUNCH.md). ## Is this for me? **Yes** if you already have eval output, runtime decisions, inventory artifacts, or MCP tool-call tests, and you want a small reviewable CI artifact instead of a dashboard — bounded auditability, not a scalar trust badge. **Not yet** if you need Assay to judge model correctness for you, want a hosted dashboard as the product, or want a compliance claim rather than a bounded evidence boundary. Assay is not a trust-score engine, a generic eval dashboard, or a hosted observability product — see [what it is and is not](docs/concepts/scope.md). ## See it work An agent tries a privileged action — `github.add_deploy_key` — through the enforcing proxy, decided per call **before it forwards**, offline against a local mock (no real credentials): ```bash cd examples/privileged-action-gate && ./run.sh ```  A deny is fail-closed caution, not a verdict on intent; an allow is the decision to forward, never proof the action happened. Declared-vs-observed conformance is recorded **beside** the verdict, never as a gate. Full walkthrough: [privileged-action-gate](examples/privileged-action-gate/). ## Pick your path | You have | What you get | Start here | |---|---|---| | Promptfoo JSONL from CI evals | Eval outcome receipts + verified bundle + Trust Basis diff | [Promptfoo JSONL](docs/use-cases/evidence-receipts-from-promptfoo-jsonl.md) | | OpenFeature `EvaluationDetails` | Decision receipt + verified bundle | [OpenFeature](docs/use-cases/openfeature-evaluationdetails-to-ci-review-artifact.md) | | CycloneDX ML-BOM model component | Inventory receipt + verified bundle | [CycloneDX ML-BOM](docs/use-cases/cyclonedx-mlbom-model-to-inventory-receipt.md) | | MCP tool calls | Allow/deny audit trail + observed-behavior evidence | [MCP Quick Start](examples/mcp-quickstart/) | | A GitHub PR gate | Trust Basis diff, gate status, SARIF/JUnit-ready output | [CI Guide](docs/guides/github-action.md) | | A Runner archive / coverage annotation | Coverage descriptors + claim-class cells + a claimed-vs-observed check | [Coverage-honesty walkthrough](examples/coverage-honesty-walkthrough/) | The workflow stays small: import or record a bounded outcome, bundle and verify it, compile `trust-basis.json`, gate the Trust Basis diff. Assay doesn't make the upstream tool the source of truth; it makes the evidence boundary inspectable. For privileged tool actions, the MCP proxy records each `tools/call` as a structured [tool-decision surface](docs/reference/tool-decision-surface.md) — keeping the asserted-versus-verified line honest. ## Policy is simple ```yaml version: "2.0" name: "my-policy" tools: allow: ["read_file", "list_dir"] deny: ["exec", "shell", "write_file"] schemas: read_file: type: object properties: path: { type: string, pattern: "^/app/.*" } required: ["path"] ``` `assay init --from-trace trace.jsonl` generates the runtime-observation policy used by the trace-generation flow (`files`, `network`, and `processes`); it is not an MCP authorization policy. Migrate a legacy MCP `constraints:` policy with `assay policy migrate`. See [Policy Files](docs/reference/config/policies.md). ## Why Assay | | | |---|---| | **Canonical evidence** | Assay's evidence model is the stable contract; OpenTelemetry and protocol adapters (ACP / A2A projection profile / UCP) map into it. | | **Deterministic** | The policy gate uses explicit rules; its decision depends on the request, policy and applicable session state. This does not make live evaluators or external effects deterministic. | | **Bounded claims** | Explicit about **verified** vs **visible** vs **absent** — no score-first UX. | | **Offline-first** | No backend required for core enforcement and bundle verification. | | **Checkable provenance** | Which piece of the source-class and coverage model shipped when, as commits you can `git log` rather than claims you have to take — [provenance](docs/PROVENANCE-SOURCE-CLASS.md), prior art credited first. | ## Learn more - [MCP Quickstart](examples/mcp-quickstart/) · [Editor MCP recipe](docs/guides/editor-mcp-recipe.md) — policy-enforcing MCP in Cursor / Claude Code / Codex - [MCP 2025/2026 protocol-era parity](https://docs.getassay.dev/mcp/protocol-era-parity/) — pinned `resultType` and interim-result compatibility corpus - [Coding-agent governance](docs/guides/coding-agent-governance.md) · [OpenTelemetry & Langfuse](docs/guides/otel-langfuse.md) — observed runs → evidence - [Evidence Receipts in Action](docs/notes/EVIDENCE-RECEIPTS-IN-ACTION.md) — Promptfoo / OpenFeature / CycloneDX receipt families - [CI Guide](docs/guides/github-action.md) · [Evidence Store](docs/guides/evidence-store-aws-s3.md) (S3 / B2 / MinIO) - [OWASP MCP Top 10 mapping](docs/security/OWASP-MCP-TOP10-MAPPING.md) · [Security experiments](docs/architecture/SYNTHESIS-TRUST-CHAIN-TRIFECTA-2026q2.md) - Positioning: [ADR-033](docs/architecture/ADR-033-OTel-Trust-Compiler-Positioning.md) · [RFC-005](docs/architecture/RFC-005-trust-compiler-mvp-2026q2.md)