# Native `reason` CLI The native Rust `reason` executable is the first supported product surface for Reasoning Harness. It invokes the same harness-owned validation, evidence, verification, decision, and finalization boundaries used by the core runtime. Research binaries and `eval*` commands remain useful, but they do not define the stable product contract. For the execution/trust model behind `--candidate` versus `--provider`, including how an AI-free run can still produce `accept | reject | unknown`, see [How Reasoning Harness works](how-it-works.md). ## Installation ### Current split release (`reason-v0.5.4`) The current product coordinate is **Reason CLI 0.5.4 on Harness Engine 0.6.1**. Published native installers are the normal end-user path; with Rust 1.88+ the CLI can also be installed directly: ```bash cargo install --git https://github.com/git-ksk/reasoning-harness --tag reason-v0.5.4 --locked reasoning-harness-cli --bin reason reason --version ``` This installs only the supported `reason` product binary, not the research binaries. Standalone 0.5.4 archives/installers are published for Linux x86_64, macOS arm64, macOS x86_64, and Windows x86_64 with `SHA256SUMS` and release provenance metadata. Use `main` only for intentionally unreleased development snapshots. The `reason-v0.5.x` line is the active split CLI line under the v0.x preview policy. `v0.4.2` remains the immutable final unified historical coordinate for Reason CLI 0.4.2 + Harness Engine 0.4.2. CLI and Engine versions now advance independently; see [versioning](versioning.md). ## Natural-language AI path The primary human-facing path accepts a task directly: ```bash reason "Analyze this incident" --fact http.status_code=503 ``` Provider/model can come from `reason-config-v1` or explicit flags. Human output is the default; use `--format json` for automation. `--plain` (also selected by `NO_COLOR`, `TERM=dumb`, or redirected human output) disables decorative progress without changing JSON or Harness semantics. Useful inputs are: | input | trust semantics | | --- | --- | | positional `TASK` | user request, not evidence | | `--file PATH` | untrusted model-readable context; no hard fact is inferred automatically | | piped stdin | same untrusted context semantics as `--file` | | `--fact KEY=VALUE` | explicit harness-owned structured fact eligible for deterministic verification | | `--hypothesis KEY=VALUE` | harness-owned proposition to evaluate/resolve | | `--resolver-fact KEY=VALUE` | explicit local fact available only through bounded resolution/admission/re-verification | | `--resolver-command PROGRAM` + `--resolver-arg ARG` | external stdio JSON acquisition adapter; output is untrusted acquired data by default | | `--resolver-timeout-ms N` / `--resolver-max-response-bytes N` | whole-invocation wall-clock deadline / accepted stdout bound for the external resolver; both must be > 0 | Example with context plus a typed target: ```bash cat error.log | reason "Determine whether the database is the root cause" \ --hypothesis incident.root_cause=database ``` If the context does not contain trusted structured support, the safe result may remain qualified/unknown. This is intentional. `--file` is not a shortcut that promotes document prose into verification authority. Introduced in v0.3.0 and retained in `v0.4.0`, the external resolver lane from #174/#175 uses `--resolver-command PROGRAM`, which launches one explicitly configured executable without shell interpolation and exchanges `reason-external-resolver-request-v1` / `reason-external-resolver-response-v1` JSON over stdio. The response can contain acquired evidence or a candidate revision, but cannot contain trusted metadata, verification receipts, verdicts, or final prose. External evidence remains fail-closed unless `resolution.external_command.admission` explicitly allowlists its source and supplies Harness-owned freshness/scope/authority policy; resolver-reported acquisition metadata cannot self-elevate. Admission rejection reasons are typed on resolution-attempt telemetry, and admitted evidence still re-enters ordinary qualification and verification. #178 added the original process timeout/response-size bounds, typed operational terminal states, actual call/latency/optional cost telemetry, and hashed adapter/admission config identities. #211 hardens that timeout into one absolute whole-invocation deadline spanning spawn, stdin write, response read, termination, and cleanup handoff across external-command, MCP, and trusted-command subprocess adapters. The command adapter does not generically retry authorization/policy/protocol failures. See [External resolver adapters](external-resolvers.md). A separately configured `resolution.mcp_readonly` lane uses the v0.4 `mcp_readonly_v3` operational successor (leaving frozen v1 untouched) for an allowlisted negotiated stdio session (`initialize` / `initialized` / `tools/list` / `tools/call`) under the same acquisition-only boundary; see [Read-only MCP resolver](mcp-resolver.md). An optional `resolution.trusted_command` lane provides explicitly trusted deterministic verification; see [Trusted verifier](trusted-verifier.md). `v0.4.0` adds an opt-in fourth acquisition lane for tasks that do not already provide an exact `--hypothesis`. The model first proposes closed-schema investigation targets and then selects only among Harness-configured read-only capability IDs. It cannot generate tool arguments, authority, receipts, or verdicts. Each acquisition is admitted under the normal source/freshness/scope/authority policy; admitted evidence causes candidate regeneration and then ordinary verification. The loop is bounded by target/round/action/no-progress limits and emits typed investigation telemetry. Static `--resolver-fact`, `external_command`, and `mcp_readonly` behavior remains unchanged when investigation is absent. Investigation external commands use the separate `investigation_external_command_v1` request protocol. See [Bounded investigation planning](investigation.md). `v0.4.0` Issue #213 adds `reason session start|inspect|resume|add|correct|fork|close` over the same native runtime and typed `ReasoningThread` checkpoint/replay model. Session files use `reason-session-v1`; inspect/resume/fork replay zero external acquisition calls, later file prose remains untrusted, premise correction records typed invalidation before re-evaluation, and finalized history can continue only by fork. Continuation turns deliberately do not replay the start turn's resolver/MCP/investigation configuration (`session-replay-only-acquisition-v1`). See [Resumable natural-language sessions](session.md). The final model-rendered answer is also untrusted until `finalize_answer` checks factual-claim coverage. Any newly introduced factual proposition is blocked; when an explicitly configured resolver can verify it, the proposition may re-enter bounded resolution and then be rendered again. If the artifact already authorizes an original requested hypothesis as exact `Known`/`Supported`, deterministic recovery may correct renderer-only omission, exact-key drift, or an exact-target downgrade from `grounded` to `uncertain`. The downgrade path is entered only when the renderer emitted that same exact requested proposition as `uncertain`; authority still comes exclusively from artifact state. Under artifact-global `Unknown`, recovery remains target-only `QualifiedPartialAnswer`. Under artifact-global `Reject`, the global verdict is never overridden, but the successor may expose a target-only `QualifiedPartialAnswer` if the target has direct evidence-bound trusted `Supported` verification and the typed artifact proves structural isolation from every problematic non-target claim (different key, typed blocker, no shared evidence, no inference/dependency path, and no target-local contradiction/qualification/hard adversarial signal). Any ambiguous dependency fails closed. No recovery parses prose, fuzzy-matches proposition keys, creates new authority, or skips the normal answer-safety gate. Provider transport reliability is kept outside Harness authority. Google/Gemini keeps bounded temporary-429 retry with `Retry-After`, but quota-classified 429 responses fail fast. HTTP 500/502/503/504 may retry at most twice, an otherwise valid success response with no model text may retry once, and the combined Google request is capped at four provider HTTP attempts. Credentials, quota, ordinary 4xx/provider errors, malformed successful responses, unsupported capability, transport interruption, and timeout remain fail-fast in this policy. `provider_attempts` reports the actual adapter HTTP-attempt count; if the Harness performs a separate structured-output fallback call, those adapter attempt counts are summed. Retry exhaustion remains a typed operational failure and never becomes semantic `unknown`, evidence, or abstention. The natural-language path also runs the current semantic + evidence-sufficiency safety checks before exposing grounded factual claims. These checks are **restrictive only**: they may require more verification, bounded resolution, or abstention, but they cannot turn model confidence into trusted evidence or an `accept` verdict. Supported partial facts can still be shown without requiring them to answer the whole task. The default is `--safety-profile current` (`verified-target-answer-gate-v1`). Use `rollback` to reproduce the previous claim-local gate (`d3-sufficiency-answer-gate-v2`); legacy `d3-sufficiency` / `d3-sufficiency-v2` selectors are aliases for that rollback. `legacy-v1` / `d3-sufficiency-v1` and `baseline` remain older testing/rollback surfaces; see [Semantic runtime product surface](#semantic-runtime-product-surface) and the [terminology guide](terminology.md) for exact machine identities. Natural JSON output declares `output_contract: reason-natural-output-v4` inside the normal `reason-cli-output-v1` envelope. It also reports `exposed_text.policy_id: harness-canonical-exposed-text-v1` and `renderer_text_exposed: false`; see [Exposed-text safety](exposed-text-safety.md). See [How Reasoning Harness works](how-it-works.md) and [product dogfood](product-dogfood.md). ## Which command should I use? | Goal | Command | | --- | --- | | Ask a person-facing natural-language question through the verified runtime | `reason "TASK"` | | Complete first-run setup on the 0.5.x product line | `reason setup` | | Manage provenance-verified update, explicit rollback, or uninstall | `reason update`, `reason update --rollback VERSION`, `reason uninstall` | | Manage provider credentials on the 0.5.x product line | `reason auth ...` | | Discover compatible models / set the user default on the 0.5.x product line | `reason models ...` / `reason model set ...` | | Persist, inspect, correct, resume, or fork a natural-language reasoning state | `reason session ...` | | Integrate an existing LLM/agent candidate with structured evidence | `reason run` | | Validate an already-materialized artifact | `reason verify` | | Run contradiction/counterexample/unsupported-premise/causal-gap semantic diagnostics | `reason semantic-check` | | Inspect the exact machine-readable JSON contracts | `reason schema` | For a human using the CLI directly, start with **`reason "TASK"`**. For application/CI integration and externally generated candidates, start with **`reason run`**. Neither path treats arbitrary prose as trusted evidence. ## Product commands - `reason auth login|status|list|logout` — manage provider credentials in the native OS credential store without secret-valued argv flags. - `reason models [provider]` / `reason model set ` — inspect the curated compatibility catalog and persist a general-use user default without silent fallback. See [Provider/model catalog](model-catalog.md). - `reason run` — execute the harness-owned correctness process from a recorded candidate or live provider candidate generation. - `reason verify` — deterministically validate a `ReasoningArtifact`. - `reason semantic-check` — execute the current semantic runtime as a soft diagnostic coordinate, with an explicit rollback profile. - `reason schema` — print the versioned JSON Schema for supported product wire contracts. `reason eval`, `reason eval-resolution`, and `reason eval-judges` are research/evaluation surfaces. Their JSON is intentionally not covered by the CLI product-envelope compatibility promise yet. ## Non-interactive input Following established CLI automation practice, `-` means standard input for supported JSON inputs. Only one source per command may consume stdin. ```bash # Harness input from stdin, candidate from a file. cat examples/input.json | reason run \ --input - \ --candidate examples/candidate.json \ --format json # Validate an artifact from stdin. cat examples/artifact.json | reason verify - --format json ``` For `reason run`, `--input`, `--candidate`, and `--receipts` may each use `-`, but no more than one of them may do so in the same invocation. Live `--provider` generation can be combined with `--input -` because the provider does not consume stdin. ## stdout and stderr For supported commands: - successful `--format json` output writes one JSON document to stdout; - human-readable progress/warnings and failure diagnostics go to stderr; - stdout is therefore safe to redirect or pipe when JSON mode succeeds; - human output is presentation only and is not a machine contract. Provider text, model confidence, and human rendering never acquire correctness authority merely by appearing in CLI output. ## JSON product envelope `reason run --format json`, `reason verify --format json`, and `reason schema` emit a versioned envelope: ```json { "schema_version": "reason-cli-output-v1", "command": "run", "cli_version": "0.2.0", "contracts": { "artifact": "reasoning-artifact-v1", "candidate": "reasoning-candidate-v1", "config": "reason-config-v1" }, "result": {} } ``` The envelope version is independent from the executable semver. A future incompatible machine output change requires a new envelope version rather than silently changing the meaning of `reason-cli-output-v1`. Inspect the current wire schemas directly: ```bash reason schema artifact reason schema candidate reason schema config ``` ## Exit semantics Exit status is process state, not epistemic state: | code | meaning | | ---: | --- | | `0` | command completed successfully; for `run`, this includes `accept`, `reject`, and `unknown` | | `1` | runtime, provider, I/O, JSON, harness-state, or artifact-validation failure | | `2` | CLI syntax/argument parsing failure emitted by `clap` | In particular, `unknown` or semantic abstention is not automatically a process failure. Scripts must inspect the JSON result when they care about epistemic outcome. When a supported product command is explicitly in JSON mode, process failures remain machine-readable. `run` and `verify` emit the same product envelope with a failed result such as: ```json { "schema_version": "reason-cli-output-v1", "command": "run", "result": { "status": "failed", "failure": { "failure_class": "input", "message": ": missing field `task` at line 1 column 2" } } } ``` Provider failures use normalized classes such as `credentials`, `rate_limit`, `quota`, `timeout`, `provider_unavailable`, and `protocol`. Configuration/input/harness failures use `configuration`, `input`, and `harness_state`. The process still exits 1. For automation, pass `--format json` explicitly (or use a valid config that resolves to JSON) so failures before normal command output can also be serialized rather than rendered as human diagnostics. ## Provider credentials and configuration Non-secret run configuration is schema-backed and layered. The supported precedence is: 1. explicit CLI flags; 2. an explicit `--config PATH` file; 3. project `.reason/config.json` in the current working directory; 4. user config; 5. compiled defaults. User config discovery uses `$REASON_HOME/config.json` when `REASON_HOME` is set, then `$XDG_CONFIG_HOME/reason/config.json`, `%APPDATA%/reason/config.json` on Windows, or `~/.config/reason/config.json`. Each config file must declare `"schema_version": "reason-config-v1"`; unknown fields fail closed instead of being silently ignored. Example: ```json { "schema_version": "reason-config-v1", "run": { "provider": "google", "model": "gemini-3.5-flash-lite", "max_tokens": 1024, "format": "json" } } ``` Use `reason schema config` for the current schema. `--no-config` ignores user/project config for a hermetic invocation, which is recommended in reproducible CI unless an explicit config is part of the job input. `--config` and `--no-config` are mutually exclusive. Project config has an additional activation boundary. A project config containing only safe run defaults can be read without trust; once it also contains `resolution.external_command`, `resolution.mcp_readonly`, or `resolution.investigation`, the project overlay fails closed until the exact canonical project/high-risk configuration is approved with `reason trust add`. `resolution.trusted_command` is never accepted from project config. Use `reason trust status|add|list|revoke` to inspect and manage this boundary; see [Project trust](project-trust.md). A configured live provider can supply the default provider/model pair. If a CLI `--provider` changes the configured provider, `--model` must also be supplied explicitly rather than accidentally reusing a model configured for another provider. A live provider with no explicit or configured model fails closed. Provider secrets are deliberately **not fields in `reason-config-v1`**. On the Reason CLI 0.5.x product line, runtime lookup uses an explicit environment variable first and the OS-native credential store only when that variable is absent: - Mistral: `MISTRAL_API_KEY` - Google / Gemma: `GEMINI_API_KEY` - NVIDIA Hosted NIM: `NVIDIA_API_KEY` - GroqCloud: `GROQ_API_KEY` The config parser rejects unknown secret-like fields such as `api_key`. Credentials are never serialized into effective run configuration, sessions, evidence, or authority state. No plaintext fallback is used when the OS credential service is unavailable. See [Secure provider credentials](secure-credentials.md). On the 0.5.x product line, use `reason auth login [provider]` for hidden TTY entry, `--stdin`/`--from-env` for automation, `--replace` for explicit rotation, and `reason auth status|list|logout` for secret-free inspection/removal. ## Semantic runtime product surface `reason semantic-check` is the supported product surface for the current semantic runtime. It is intentionally separate from `reason run`: a semantic finding is diagnostic evidence and never gains verification, trusted-evidence, or final-verdict authority merely because it was produced by the model-backed semantic runtime. The input contract is `semantic-check-input-v1` and contains exactly a `request` plus a harness-owned `artifact`. Inspect it with: ```bash reason schema semantic-check ``` Run the current profile: ```bash reason semantic-check \ --input examples/semantic-check.json \ --provider mistral \ --model ministral-8b-latest \ --format json ``` The current profile executes the stabilized materialization and deterministic typed-precondition path. The JSON result still exposes the exact machine runtime identity (`semantic-decidability-d3-v1`) for reproducibility. It includes the base decision, final semantic decision, decidability disposition, usage, model, and provider-attempt count. `force_abstain` can only make the semantic result more conservative. The characterized rollback remains explicit. Legacy `--profile v3` is still accepted as an alias: ```bash reason semantic-check \ --input examples/semantic-check.json \ --provider mistral \ --model ministral-8b-latest \ --profile rollback \ --format json ``` Operational failure is separate from semantic outcome. In JSON mode a provider/runtime failure emits a `semantic-check` product envelope containing `operational_failure.failure_class` and returns exit 1; it is never converted into `finding`, `no_finding`, or `abstain`. See [product roadmap](product-roadmap.md), [ADR-0001](adr/0001-interface-and-packaging-boundaries.md), and [semantic runtime stabilization](semantic-runtime-stabilization.md). ## Design references CLI ergonomics intentionally learn from mature terminal-first AI tools without copying their agent semantics: - OpenAI Codex treats non-interactive execution, machine-readable output, schema-constrained output, and layered/profile configuration as first-class automation concerns. Reasoning Harness adopts the separation of automation output from diagnostics, but not Codex's agent/sandbox authority model. See . - OpenCode exposes a dedicated non-interactive `run` path, stdin-friendly automation, JSON output, JSON-Schema-backed configuration, and explicit config discovery/precedence. Reasoning Harness uses these as UX references while keeping its own harness-owned evidence and verdict boundaries. See and . These projects are design references, not wire-compatibility targets. `reason` should stay narrower: its product value is a predictable evidence-grounded reasoning harness, not another general-purpose coding agent. ## `reason setup` (0.5.x product line) `reason setup` configures provider, credential, a curated general-use model default, and local readiness in one path. Credentials use hidden TTY input or `--credential-stdin` / `--from-env`; no secret-valued argv flag exists. `--non-interactive` requires `--provider` and selects the curated recommendation when `--model` is omitted. Local readiness is network-free. A live check may consume quota or incur cost, so it runs only with explicit `--live-check` or interactive consent. ## `reason doctor` (0.5.x product line) `reason doctor` reports Reason CLI and Harness Engine versions separately and inspects config sources, secret-free credential state, OS credential-store availability, provider/model readiness, managed-session paths, project trust, and configured user MCP state. Default diagnostics are local-only and non-mutating. `--live-check` opts into bounded provider/MCP/update checks and may consume provider quota or incur cost. JSON uses the stable `reason-doctor-v1` result surface. See [Reason doctor](reason-doctor.md). For recovery-oriented product failures, Reason preserves the existing typed failure class and adds explicit execution/trust/recovery guidance. Human errors show `What failed`, `Task execution`, `Result trust`, and `Next`; JSON product failures add a `remediation` object without changing `failure.failure_class`. See [Operational failure recovery](reason-recovery.md). For lifecycle management on the 0.5.x product line, see [Update, rollback, and uninstall](update-rollback-uninstall.md).