# Benchpack Format This is the initial manifest sketch. The schema can change until the first release. ## Example ```toml [pack] id = "smoke-chat" version = "0.1.0" description = "Tiny endpoint smoke test" [defaults] temperature = 0 max_tokens = 64 stream = true warmup = 0 repetitions = 1 [[cases]] id = "capital" kind = "chat" prompt = "What is the capital of France? Answer in one sentence." fixture_refs = ["synthetic-context"] [[fixtures]] id = "synthetic-context" kind = "context" path = "fixtures/context.md" description = "Portable synthetic context for future tasks" [scoring] mode = "contains" expected = "Paris" ``` ## Fields `pack.id` : Stable pack identifier used in result records. Must match the id grammar below. `pack.version` : Version of the workload. Change it when prompts, fixtures, fixture references, or scoring change. `defaults` : Request defaults shared by cases. `defaults.stream` : When true, adapters that support streaming may use a streaming request path. `openai-chat` honors this flag by requesting streamed chat completions and measuring TTFT from the first non-empty content delta. When false or absent, `openai-chat` keeps the non-streaming request shape. The example above is a schema example; individual packs such as `smoke-chat` may leave streaming off to preserve non-streaming smoke coverage. `defaults.warmup` : Number of unrecorded warmup executions per case. It defaults to `0` when absent and must be a non-negative integer. Warmups use the same adapter, endpoint, model, prompt, and defaults as measured executions, write raw request/response files under `raw/.warmup-NNN.*.json`, and are excluded from `run.jsonl`, scoring, and `summary.md`. `defaults.repetitions` : Number of measured executions per case. It defaults to `1` when absent and must be a positive integer. Each measured execution writes one `run.jsonl` record. Packs with `repetitions = 1` use legacy raw file names `raw/.request.json` and `raw/.response.json`. Packs with `repetitions > 1` use `raw/.rep-NNN.request.json` and `raw/.rep-NNN.response.json`, and each record includes a 1-based top-level `repetition` field owned by the reporter. `cases` : Ordered benchmark cases. Each case `id` must match the id grammar below and must be unique within the pack. `cases[].prompt` : Inline prompt text for a case. Inline prompts remain supported for compact smoke and runtime-measurement cases. `cases[].prompt_file` : Pack-relative path to a UTF-8 prompt file, for example `prompt_file = "prompts/wrap-plan-small.md"`. A case must define exactly one prompt source: either `prompt` or `prompt_file`, never both and never neither. Prompt files are resolved relative to the pack directory, must not be absolute paths, and must resolve inside the pack directory after following symlinks. The runner rejects paths that escape the pack directory through `..` traversal or symlinks. Loaded file contents become the case prompt at manifest-load time, so adapters and result records do not distinguish inline prompts from prompt files. `cases[].fixture_refs` : Optional list of fixture ids declared in the same pack's top-level `[[fixtures]]` inventory. It defaults to an empty list when absent. When present, it must be a TOML array of strings; every ref must match the id grammar, must point to an existing top-level fixture id in the same pack, and must not appear more than once in the same case. Referenced file fixtures are appended to the loaded case prompt in `fixture_refs` order with stable delimiters. Referenced directory fixtures remain metadata-only. Fixture refs do not execute fixtures, mutate repositories, template prompts, change adapter request or result schemas, extract patches, or run verifiers. The one current exception is `repo-task`: each measured execution copies exactly one referenced `kind = "repo"` directory fixture into a run-owned disposable workspace under the output directory and captures a deterministic patch from the source fixture to that workspace. `fixtures` : Optional top-level fixture inventory. Each `[[fixtures]]` entry declares a static pack-owned file or directory by `id`, `kind`, `path`, and optional `description`. Fixture ids use the same id grammar as packs and cases and must be unique within the pack. Fixture paths are source contracts: they must be strings, relative to the pack directory, resolve inside the pack after following symlinks, exist at manifest-load time, and point to either a file or a directory. The runner validates and exposes fixture metadata on loaded packs. Referenced file fixtures are read as UTF-8 and appended to `Case.prompt`; directory fixtures are not read into prompts. The runner does not mutate repositories, execute verifiers, or score from fixtures. For `repo-task` measured executions only, it copies one referenced `kind = "repo"` directory fixture into a disposable run-owned workspace and captures a deterministic source-vs-workspace patch artifact after the adapter call. `scoring` : Optional scoring configuration. May appear at pack level as a default and/or inline on individual cases as an override. Current executable deterministic modes are `none`, `contains`, `regex`, `json-schema`, and `verify-script` for measured `repo-task` executions. Other reserved modes still parse as manifest values but are not implemented by the scorer. See **Scoring** below. ### Prompt Files Use `prompt_file` when a prompt is long enough that keeping it in TOML would make the manifest hard to scan: ```toml [[cases]] id = "wrap-plan-small" kind = "chat" prompt_file = "prompts/wrap-plan-small.md" ``` Prompt files are static text in the current format. After the base prompt is loaded from `prompt` or `prompt_file`, referenced file fixtures may be appended as described below. There is no templating, variable substitution, globbing, include support, or multi-message loader in this slice. ### Fixtures Use top-level `[[fixtures]]` entries to declare pack-local static inputs that future workload slices can consume: ```toml [[fixtures]] id = "synthetic-django-app" kind = "context" path = "fixtures/synthetic-django-app.md" description = "Portable synthetic target app description" ``` Directory snapshots use the same manifest shape and can declare `kind = "repo"`: ```toml [[fixtures]] id = "synthetic-django-repo" kind = "repo" path = "fixtures/synthetic-django-repo" description = "Compact static synthetic Django source snapshot" ``` Required fields: - `id`: stable fixture identifier. It must match the id grammar below and be unique within the pack. - `kind`: explicit non-empty fixture type string, such as `context` or `repo`. - `path`: pack-relative file or directory path. Optional fields: - `description`: human-readable description. It defaults to an empty string when absent. The loader resolves `path` relative to the pack directory and rejects absolute paths, `..` traversal outside the pack, symlink targets outside the pack, paths that resolve to the pack directory itself, missing paths, and existing paths that are neither files nor directories. File and directory fixtures are both allowed so later repo-task work can introduce directory snapshots without changing this source contract. Fixture declarations remain available as metadata on loaded packs. Cases may reference fixtures by id with `fixture_refs`. When a referenced fixture path is a file, the loader reads it as UTF-8 and appends it to the loaded case prompt. When a referenced fixture path is a directory, including a repo snapshot, the loader validates the ref but does not read, execute, or inject the directory contents. The runner later copies the single `kind = "repo"` directory fixture only for measured `repo-task` executions. File fixture prompt assembly uses this stable plain-text shape: ```text --- BEGIN FIXTURE (, ) --- --- END FIXTURE --- ``` Multiple referenced file fixtures are appended in the exact `fixture_refs` order chosen by the case author. `Case.prompt` is the final assembled prompt that adapters receive. `Case.raw` preserves the original manifest fields, and `Case.fixture_refs` preserves the fixture id list. Fixture assembly does not add prompt templating, variable substitution, globbing, include support, fixture execution, repository mutation, patch extraction, verifier execution, adapter schema changes, or result schema changes. Repository copying is limited to the runner-owned measured repo-task workspace preparation described below. ### Case Fixture References Use `fixture_refs` on a case to declare which top-level fixture ids are relevant to that case: ```toml [[cases]] id = "wrap-plan-context" kind = "chat" prompt_file = "prompts/wrap-plan-context.md" fixture_refs = ["synthetic-django-app"] ``` `fixture_refs` defaults to `[]` when omitted. The loader rejects non-array values, non-string entries, ids that do not match the documented grammar, duplicate refs within one case, and refs that do not exist in the same pack's top-level fixture inventory. Top-level `[[fixtures]]` entries may appear before or after `[[cases]]` in TOML; refs are validated against the loaded inventory. Changing fixture refs alters the effective prompt for referenced file fixtures, so pack authors should bump `pack.version`. ## ID Grammar Pack, case, and fixture ids are used as stable record keys or source identifiers, and case ids are also used as filesystem path components (e.g. `raw/.request.json`). They must match `^[A-Za-z0-9][A-Za-z0-9_-]*$`: start with an alphanumeric, then any mix of alphanumerics, underscore, and hyphen. No dots, slashes, spaces, or empty strings. The runner rejects manifests that violate this at load time. ## Case Kinds `chat` : A direct prompt or message list sent to an adapter. `completion` : A raw prompt-completion case. `repo-task` : Partially implemented case kind for a task that prepares a disposable repository workspace and can verify it deterministically. Current runner support copies exactly one referenced `kind = "repo"` directory fixture into `workspace//rep-NNN/` under the run output directory before each measured adapter call, runs an internal task executor after the adapter call, and current CLI runs apply the first fenced `diff` or `patch` block from model output as a unified diff or explicitly marked full-file replacement inside that workspace through the default executor, captures a deterministic patch artifact at `patch//rep-NNN.diff` after the task phase, writes task logs at `task//rep-NNN.stdout.log` and `task//rep-NNN.stderr.log`, executes `verify-script` scoring when declared, and records workspace metadata, `patch.path`, `task`, `verify`, `repo_task`, and top-level `scoring` in the measured `run.jsonl` row. Public `external-agent` executions also receive a runner-owned context input at `task//rep-NNN.context.json`; that file is not a result row field. The context exposes an optional harness-owned model-call JSONL path at `task//rep-NNN.model-calls.jsonl`, also outside `run.jsonl`. A minimal internal agent-session harness path exists behind the same executor boundary for runner-side callers and tests. Current CLI runs use the fenced `diff`/`patch` executor by default, and repo-task cases may explicitly select that same public compatibility executor with `harness = { id = "fenced-patch" }` or route to a runner-owned external subprocess with `harness = { id = "external-agent" }` plus `BENCHPACK_EXTERNAL_AGENT_ARGV`. That harness table may also declare a narrow positive numeric `timeout_s` for the task executor phase. This format does not currently add full production external coding-agent integration, manifest task commands, retention options, repo-task warmups, task environment configuration, pack-level harness defaults, or production external agent-harness behavior. Bundled repo-task examples currently include `patch-from-failure`, `python-regression-fix`, `django-dashboard-regression-fix`, the opt-in project-completion prototype `mini-project-completion`, and the opt-in direct-edit external-agent `product-offer-matching` pack with Python and Rust matcher cases. `replay` : A recorded request sequence. ### `repo-task` Contract The current `repo-task` implementation prepares disposable measured workspaces, runs the task phase through an internal executor boundary, applies model output only through the executor's explicit fenced unified-diff or replacement-file contract, captures patch artifacts, writes deterministic task stdout/stderr logs for that task phase, and executes `verify-script` scoring when declared. Referenced file fixtures still append to `Case.prompt`; referenced non-repo directory fixtures are rejected for repo-task; the single referenced repo directory is copied into a run-owned workspace; the measured result record includes the prepared workspace metadata, `patch.path`, task log artifact paths, verifier artifact paths, final verifier status, and top-level `verify-script` scoring. The runner has internal harness paths and one public external subprocess harness id, but it still does not execute manifest-declared task commands. Repo-task cases should use this conservative shape: ```toml [[cases]] id = "wrap-repo" kind = "repo-task" prompt_file = "prompts/wrap-repo.md" fixture_refs = ["synthetic-django-repo", "synthetic-django-app"] scoring = { mode = "verify-script", script = "verify/wrap-repo.py", timeout_s = 30, environment = { PYTHONPATH = "src", BENCHPACK_CASE = "wrap-repo" } } ``` Fields: - `id`: case id using the normal id grammar. - `kind`: must be `repo-task`. - `prompt` or `prompt_file`: task instructions sent to the model or future agent harness. The one-prompt-source rule remains the starting point unless a later multi-message contract replaces it. - `fixture_refs`: must identify exactly one primary fixture with `kind = "repo"` whose path is a directory. That fixture is the source repository snapshot for the disposable workspace. Additional refs, if any, must be non-directory file fixtures and remain prompt/context inputs. - `scoring`: should use `mode = "verify-script"` for deterministic repo-task correctness. `contains` and `regex` remain prompt-output scoring modes and are not sufficient for repository correctness. `timeout_s` is optional and controls only the verifier subprocess timeout. `environment` is optional for `verify-script` and controls only the verifier subprocess environment. The manifest contract intentionally does not define broad generic blobs such as workspace configuration or `commands` yet. Verifier environment support is deliberately limited to the effective `verify-script` scoring table. Add explicit fields only when a future implementation needs them and the semantics are narrow enough to test. Agent-session harness inputs are runner-side concerns, not public manifest syntax in this slice. The current internal harness path may use the prepared workspace, case metadata, model output text, output directory, repetition, and deterministic task log paths internally, plus validated helpers to list workspace file paths and directory paths, check workspace file existence, read or write UTF-8 workspace text, and delete workspace files. File listings include regular files only, use sorted POSIX workspace-relative paths, and observe files created earlier in the same harness invocation. Symlinks to regular files are listed only when their target resolves inside the prepared workspace. Directory listings include nested directories, use sorted POSIX workspace-relative paths, exclude the workspace root, files, and symlinks including symlinks to directories, and observe directories created earlier in the same harness invocation. File existence checks return true only for existing regular files, including in-workspace symlinks to regular files, while missing paths and directories return false. File deletes use the same path boundary, return true after deleting an existing regular file or in-workspace symlink-to-file workspace entry, return false for missing paths and directories, and unlink symlink entries without deleting their targets. Future harnesses may add pack metadata and model/adapter/endpoint/default context. Pack authors should not rely on manifest-declared shell commands, task environment, pack-level harness defaults, or workspace retention because those fields do not exist yet. The only task timeout field is the narrow `harness.timeout_s` policy described below. #### Public Harness Selection Public harness selection is implemented narrowly as an explicit case-local table on `repo-task` cases: ```toml [[cases]] id = "fix-repo" kind = "repo-task" prompt_file = "prompts/fix-repo.md" fixture_refs = ["repo"] harness = { id = "fenced-patch", timeout_s = 5 } scoring = { mode = "verify-script", script = "verify/fix-repo.py" } ``` Rules: - `harness.id` names a runner-known public harness. `fenced-patch` routes to the existing fenced model-output `diff`/`patch` executor. `external-agent` routes to the runner-owned external subprocess harness and requires `BENCHPACK_EXTERNAL_AGENT_ARGV` when the CLI runs the pack. - Production external harnesses are public repo-task harnesses and must be selected by explicit case-local `harness.id` values. Selection must not be inferred from model names, adapters, endpoints, fixture shape, verifier choice, host environment, or pack id. - When `harness` is absent, the default/current behavior remains the fenced `diff`/`patch` executor for compatibility. - `harness` is accepted only on `repo-task` cases. The loader rejects non-table `harness` values, missing `id`, non-string `id`, unknown ids, and unexpected keys. The only supported keys are currently `id` and `timeout_s`. - `harness.timeout_s` is optional. When present, it must be a positive TOML integer or float. Booleans, strings, zero, negative values, arrays, and tables are rejected. It bounds the selected task harness/executor phase, not the adapter request and not the verifier subprocess. - Task timeout is enforced by subprocess-backed task executors. For `fenced-patch`, it is passed to the unified diff preflight and apply calls. Preflight tries `git apply --check --recount` first, then standard `git apply --check` if recount rejects the diff. The apply call uses the mode that passed preflight. The recount mode accepts otherwise valid unified diffs whose hunk line counts are inaccurate, which is common in model output. The timeout is applied independently to each subprocess call rather than as one shared phase budget. A timeout during `git apply --check` leaves the workspace unchanged, writes deterministic task stderr, and still allows patch capture and verifier execution. A timeout during the actual `git apply` after successful preflight is a runner failure because the workspace may be partially changed. For `external-agent`, a subprocess timeout is captured in the task stderr log after the runner stops the external process group with a bounded terminate-then-kill cleanup policy. - Runner-side internal in-process harness callables cannot be combined with task timeout. - Public harness selection does not change adapter request or result schemas. The generated external-agent context file is harness input and is not duplicated into `run.jsonl`. The context-provided `task//rep-NNN.model-calls.jsonl` path is optional and is not required, pre-created, or added to `run.jsonl` by the runner. If the file exists, the runner validates only allowlisted safe telemetry fields for aggregate `summary.md` and `benchpack report` summaries. Harness-owned model calls remain runner/harness concerns rather than normal adapter request fields. - Existing task logs remain `task//rep-NNN.stdout.log` and `task//rep-NNN.stderr.log`. - Patch capture still happens after the selected task phase and reflects the post-task workspace. Verifier execution still happens after patch capture. - Repo-task warmups remain rejected until a later warmup contract defines fresh workspace isolation and artifact behavior. - External harnesses may mutate only the prepared workspace and write only allowed run-output artifacts. Pack-owned fixtures, prompts, verifier scripts, source docs, and raw model artifacts remain immutable or runner-owned. - Task environment configuration, workspace retention, richer task status/reporting, pack-level harness defaults, full production external coding-agent integration, and repo-task warmups are separate future slices, not implicit support in `harness`. - This narrow selection adds no CLI flags, `run.jsonl` fields, adapter schema changes, raw artifact path changes, or task log path changes. The public external subprocess harness manifest shape stays as small as the current public table: ```toml [[cases]] id = "fix-repo" kind = "repo-task" prompt_file = "prompts/fix-repo.md" fixture_refs = ["repo"] harness = { id = "external-agent", timeout_s = 120 } scoring = { mode = "verify-script", script = "verify/fix-repo.py" } ``` The manifest does not name the command. The runner loads `BENCHPACK_EXTERNAL_AGENT_ARGV` when an `external-agent` case is selected. The value must be a JSON array of non-empty strings without NUL bytes; it is not shell-parsed and plain command strings are rejected. The runner appends `--workspace `, `--case `, `--output-dir `, `--repetition `, and `--context /task//rep-NNN.context.json` before executing the subprocess without a shell in the prepared workspace. On `harness.timeout_s`, timed-out external subprocesses are cleaned up as a POSIX process group with bounded termination and kill escalation before timeout task logs are written. Missing or malformed configuration fails before run output directory creation and before adapter calls. This shape keeps explicit case-local selection, optional task-phase timeout, no task command list, no task environment table, no shell expansion, no secrets handling, no workspace retention flag, and no pack-level harness default. External harness inputs are runner-side invocation context, not additional manifest syntax. The public external-agent subprocess receives appended workspace, case id, output directory, repetition, and context-file arguments. The context JSON has `version = 1` and includes case metadata, pack metadata, loaded prompt text, fixture refs and source repo metadata, task log paths, selected harness options, optional run metadata path, optional model-call JSONL path, model, adapter id, endpoint argument, request defaults, and compatibility options as explicit harness input. Those values do not alter normal adapter request/result schemas and do not create new `run.jsonl` fields by default. External harnesses may mutate only the prepared workspace. They may write the existing task stdout/stderr logs through runner-owned capture and may optionally write JSONL model-call telemetry to the context-provided `task//rep-NNN.model-calls.jsonl` path. The runner does not require, pre-create, or add that file to `run.jsonl`. When it exists, the runner parses only the recommended safe telemetry shape for summaries; invalid or unsafe lines are counted but not echoed. The recommended minimal line is: ```json {"schema_version":1,"sequence":1,"model":"test-model","ok":true} ``` Each line should describe one harness-owned model call. `schema_version` is currently integer `1` for this recommended shape, `sequence` is the positive call sequence within the external-agent task phase, `model` is the model id when known, and `ok` records whether the call completed successfully. Optional fields may include `started_at`, `ended_at`, `duration_s`, `adapter`, `endpoint`, `response_format`, `token_budget_field`, `finish_reason`, `prompt_tokens`, `output_tokens`, `cached_prompt_tokens`, and a short `error` string for failed calls. The summary allowlist is exactly those fields. Summary output reports aggregate counts, success/failure/error counts, unique model/adapter/endpoint/request-shape/finish labels, summed duration, and summed token fields; safe string fields must be short and must not contain control characters or Unicode separator characters other than plain spaces. `endpoint` is treated as a label rather than a URL. Values containing URL schemes, query strings, or userinfo markers are invalid for summaries. Summary output does not report full prompts, full responses, request bodies, headers, environment variables, API keys, bearer tokens, credentials, or short error text. Other additional explicit harness artifacts under the run output directory require a later schema slice. They must not write pack-owned fixtures, prompts, verifier scripts, source docs, normal adapter `raw/` artifacts, `hardware.json`, `run-metadata.json`, or paths outside the prepared workspace and allowed run-output artifacts. Patch capture still compares the source fixture to the post-task workspace, and verifier execution still follows patch capture. The source-controlled `examples/external-agent/reference-agent.py` script is a deterministic local reference for this handoff. It reads the appended context, checks the core public fields, writes a small workspace marker file, and writes one recommended model-call JSONL line. It is not a manifest command, does not make live model calls, and does not make the optional model-call log a runner schema. The sibling `examples/external-agent/model-call-agent.py` script demonstrates a deterministic harness-owned call boundary without changing the manifest format. It accepts an example-owned `--model-call-url` that must point to an HTTP loopback host without credentials or query string, sends one tiny local HTTP JSON request containing only safe identifiers such as case id, repetition, and model, writes the deterministic response content inside the prepared workspace, and writes one safe JSONL telemetry line to the context-provided model-call path. It remains example guidance; the runner summarizes only allowlisted safe fields from the optional JSONL file and keeps the full payload outside `run.jsonl`. Directory fixture semantics for repo-task cases: - The referenced `kind = "repo"` fixture is immutable source. The runner must never mutate files under `benchpacks//fixtures/`. - The runner prepares a fresh copy under the run output directory for every measured execution at `workspace//rep-NNN/`. The runner includes `rep-001` even when `defaults.repetitions = 1`. If the destination already exists, the run fails rather than merging into it. Repo-task warmups are rejected for now; if they are later supported, each warmup also gets a fresh copy and does not share mutations with measured repetitions. - Repo fixtures must not contain symlinks that would escape workspace isolation. Absolute symlinks and relative symlinks whose target resolves outside the source repo fixture are rejected before copying. Internal relative symlinks may be preserved. - Future mutation is allowed only inside the disposable workspace. Pack fixtures, prompts, verify scripts, and other source artifacts remain read-only by contract. - The internal agent-session harness path follows the same write boundary: it may inspect and mutate the prepared workspace through validated runner-owned helpers and write the existing task logs under the run output directory, but must not mutate pack-owned fixtures, prompts, verifier scripts, source docs, or public adapter/result schemas by default. - The current CLI mutation source is narrow and model-output-only: after the adapter call, the runner invokes the default internal task executor. It extracts the first fenced code block whose info string is exactly `diff` or `patch`, treats the block body as a unified diff or an explicit full-file replacement whose first content line is `*** Begin File: ` and whose final marker line is `*** End File`, and applies it from the prepared workspace root. Replacement blocks write UTF-8 text with LF-canonicalized line endings; trailing whitespace on the end marker is tolerated. Non-matching fenced blocks are ignored. If no matching block exists, if the block is empty, if paths are unsafe, if the replacement marker/content is invalid, or if the diff cannot be applied cleanly, the workspace remains unchanged, task stderr records a deterministic message, and the measured row is still written. - Repo-task execution must not write outside the run output directory and the prepared workspace, and pack contracts must not depend on implicit network access or private local host paths. - Multiple repo fixtures are not allowed until merge/copy rules are explicitly documented. Non-repo directory fixtures in repo-task `fixture_refs` are reserved and are rejected by the current repo-task runner. - Referenced file fixtures keep the existing prompt-assembly behavior. They are not copied into the workspace or exposed to verifiers unless a later explicit field says so. - Directory fixtures outside repo-task execution remain metadata-only. Workspace and artifact layout: - Workspaces live under the run output directory, for example `workspace//rep-NNN/`. - Measured repo-task `run.jsonl` records include a top-level `workspace` object with `path`, `source_fixture_id`, and `source_path`. The path is relative to the run output directory. The source path is the manifest-declared fixture path, not an absolute resolved path. - Patch capture writes a deterministic diff artifact at `patch//rep-NNN.diff`, including `rep-001` for single-repetition packs. Measured repo-task records include a top-level `patch` object with the run-relative `path`. Empty changes still create an empty patch file and record `patch.path`. - Task log capture writes deterministic stdout/stderr artifacts at `task//rep-NNN.stdout.log` and `task//rep-NNN.stderr.log`, including `rep-001` for single-repetition packs. Measured repo-task records include a top-level `task` object with run-relative `stdout_path` and `stderr_path`. For current CLI runs, these files record the default fenced unified-diff or explicit replacement-file extraction/application phase. On successful application stdout contains a short deterministic success message and stderr is empty; no-patch and rejected-patch outcomes are recorded in stderr. For internal harness runs, the same files record the harness task phase without adding result fields. - External-agent context capture writes deterministic runner-owned context JSON at `task//rep-NNN.context.json` for public `external-agent` executions only. The file is harness input and is not referenced from measured `run.jsonl` rows. - External-agent model-call logs, when a public external harness chooses to write them, use the context-provided `task//rep-NNN.model-calls.jsonl` path. The runner does not create or require the file, and measured `run.jsonl` rows do not reference it. Safe allowlisted telemetry is summarized in `summary.md` and `benchpack report` when the file exists. Harness authors should prefer the recommended per-call JSON object shape shown above, starting with `{"schema_version":1,"sequence":1,"model":"test-model","ok":true}`. - Verifier execution for `scoring.mode = "verify-script"` writes artifacts at `verify//rep-NNN.json`, `verify//rep-NNN.stdout.log`, and `verify//rep-NNN.stderr.log`, including `rep-001` for single-repetition packs. Measured repo-task records include a top-level `verify` object with run-relative `path`, `stdout_path`, and `stderr_path`. They also include `repo_task.status` (`"passed"` for exit code `0`, `"failed"` for nonzero or verifier timeout) and `repo_task.verify_exit_code` (the integer process exit code, or `null` on timeout). Timeout rows keep the same artifact and result object shape and set top-level scoring to `{"mode": "verify-script", "passed": false}`. - Model request/response payloads remain under `raw/`; repo-task workspace, patch, task logs, external-agent context/model-call artifacts, and verifier artifacts are conceptually separate. The current directory snapshot diff is deterministic and does not require the fixture or workspace to be a Git repository. It compares the immutable source fixture directory to the prepared workspace after the adapter call, orders paths lexicographically by POSIX-style relative path, emits unified diffs for UTF-8 text file additions, deletions, and changes, emits deterministic marker lines for binary additions, deletions, and changes, represents changed symlink targets as text diffs of the link target strings, normalizes UTF-8 text line endings to `\n` before comparison, and ignores empty directories. Cleanup should be deterministic. Keeping a workspace after a run should be an explicit runner option or future manifest-independent debug setting, not implicit behavior. Curated commits should normally include small summaries, `hardware.json`, compact `run.jsonl`, and small explanatory artifacts such as patch diffs or `verify.json` when useful; full workspaces and large logs should usually remain local or ignored. ## Scoring A pack can declare default scoring at the top level. Cases override it inline when they need a different mode. ### Pack-level default ```toml [scoring] mode = "contains" expected = "Paris" ``` ### Per-case override ```toml [[cases]] id = "json-output" kind = "chat" prompt = "Return a JSON object with key 'city'." scoring = { mode = "json-schema", schema = "fixtures/city.schema.json" } ``` ### Modes `none` : No scoring. The run is recorded but no pass/fail is computed. `contains` : Output must contain the `expected` string. `equals` : Reserved; not implemented by the scorer yet. Intended behavior is that output must equal `expected` exactly after trimming whitespace. `regex` : Output must match the regular expression in `pattern`. The scorer uses Python's standard `re.search(pattern, output)` with no implicit flags. Pack authors who need multiline, dotall, or other flag behavior should use inline regex flags or explicit character classes in the pattern. The runner raises a `ValueError` when `mode = "regex"` is evaluated without `pattern`. `json-schema` : Output must parse as JSON and validate against the pack-relative JSON schema file declared in `schema`. The runner resolves `schema` inside the pack directory at manifest-load time, reads it as a JSON object, and rejects absolute paths, `..` escapes, missing files, invalid JSON, and non-object schemas. The executable scorer implements the compact deterministic subset used by bundled packs: `type`, `required`, `properties`, `additionalProperties = false`, `items`, `enum`, `const`, string length bounds, and numeric bounds. `const` and `enum` use JSON-value equality: booleans are distinct from numbers, so `true` does not match `1` and `false` does not match `0`. Unknown schema keywords are ignored. Invalid JSON model output or schema-shape mismatches produce a normal scoring failure, not a runner crash. Failed `json-schema` scoring envelopes may include an `error` string describing the invalid JSON or first schema mismatch. This scoring mode checks raw adapter output text only; it does not request native tool calls from an endpoint. `verify-script` : Implemented for measured `repo-task` executions only. The runner resolves `script` as a pack-relative path, rejects absolute paths and paths that resolve outside the pack root, and requires an existing file. It runs the script with the current Python interpreter as `sys.executable