--- name: braintrust-agent-evals description: Inspect, query, compare, or explicitly export GitHits agent-eval history in Braintrust using the repository's verified workflow. metadata: internal: true --- # Braintrust agent evals Use this skill for read-only inspection of the GitHits agent-eval history in the Braintrust project `githits-cli-agent-evals`, or when the user explicitly asks to export a validated local suite. The normalized exporter stores one top-level eval span per scenario/workload cell plus structural tool children; see [`docs/implementation/agentic-eval-metrics.md`](../../../docs/implementation/agentic-eval-metrics.md) for the field contract. The persistence unit is one exporter invocation = one experiment, one scenario/workload cell = one eval row, and one normalized logical tool call = one structural tool child. Current exporter-owned experiment names are `main-r-a`, `pr--r-a`, and `local---`. Historical `github-*` experiments predate this identity contract and should be treated as historical evidence, not as current names or baseline candidates. ## Safety and interpretation - Read-only is the default. Never delete experiments or upload raw stdout, stderr, environment/configuration, provider events, or arbitrary artifacts. - Never read or print `BRAINTRUST_API_KEY`, `.bt/`, Keychain contents, or any credential/environment value. CI scopes the key only to its exporter step. - Do not treat an agent's self-reported confidence as result quality. This phase has no scorer or quality score. - A failed or partial cell can be valid history when its normalized evidence is complete; distinguish that from a rejected suite or failed preflight. ## Inspect experiments Use the exercised project-option placement for list/view: ```bash bt experiments --json --project githits-cli-agent-evals list bt experiments --json --project githits-cli-agent-evals view ``` Use the experiment ID returned by the view result for a bounded field query: ```bash bt sql --json --non-interactive "SELECT input, output, metrics, metadata, tags FROM experiment('') WHERE span_attributes.type = 'eval' LIMIT 100" bt sql --json --non-interactive "SELECT name, span_attributes.type, metrics, metadata FROM experiment('') WHERE span_attributes.type = 'tool' LIMIT 100" ``` The eval-root query is the verified path for prompts, neutral answers, hashes, statuses, native token/cost/duration metrics, and root metadata. Query tool children separately for native tool counts/errors and exact lifecycle timing. An unfiltered `count(*)` includes both eval roots and tool children, so it is not the workload-row count. A local proof experiment `poc-native-tool-spans-v2-20260831` (ID `e8480301-6622-4a06-a37b-0ebd0e42bb64`, ) read back two eval roots and 10 tool children. Native comparison reported `tool_calls` average `5.0` and `tool_errors` `0`; child durations totaled 30.970 seconds and ranged from 0.006 to 10.400 seconds. Native token and cost fields remain populated. Open the experiment permalink when row-level UI inspection is useful. The labeled CI path is proven by [run 33424857668](https://github.com/githits-com/githits-cli/actions/runs/33424857668) at code SHA `7195ccc56b9ac9288dfb3d8de854f2f0e7ae7cf0`. Its experiment is `github-33424857668-1` (ID `182ee9db-0df3-40f4-8987-6eeb6d91a89b`), source `github`, exporter/schema 2, metrics schema 3: 23 eval spans and 116 tool children, exactly matching 116 MCP calls, with zero CLI calls and zero failed tool spans. Totals were 513.911 seconds eval duration, 126.458999872 seconds tool duration, 2,686,094 prompt tokens, 20,172 completion tokens, 2,706,266 total tokens, and estimated cost `$0.22819038`. Compare averages were duration `22.343956532685652`, estimated cost `$0.009921320869565216`, tool calls `5.043478260869565`, tool errors `0`, and total tokens `117663.73913043478`. The first stable default-branch bootstrap is [run 33477846273](https://github.com/githits-com/githits-cli/actions/runs/33477846273) at SHA `40796bd0eabaf87afec5ea0e4460ff47e7448603`. Experiment `main-r33477846273-a1` (ID `6f3847fc-3816-4b32-b1f6-65019c2757b7`) read back 23 eval roots and 112 tool children, zero CLI calls, 3,024,404 tokens, 445.728 seconds cumulative agent duration, and estimated cost `$0.24188221`. Its null base is the expected one-time bootstrap result. Main pushes now temporarily run the same matrix, in addition to the daily/manual/label paths, to collect variance and workload-optimization evidence. The current `agent-eval-openrouter` label runs the shared main matrix on trusted same-repository PRs: `canary` discovery plus `stable-full` intent and full guidance. `eval/agentic/suites.json` owns the workload counts. The trial PR must commit credential-free `eval/agentic/openrouter.toml` selecting its exact candidate model; the repository's blank-model example deliberately selects none. Keep the active config out of main and use `OPENROUTER_API_KEY` as its provider `env_key`, the only wired provider execution credential. The shared `.github/workflows/agent-evals.yml` retains Codex 0.154.0/prompt-json and execution-only OpenRouter auth; other triggers retain Luna/low/schema. The old DeepSeek label no longer starts a run. The dedicated canary workflow is removed. Each trial exports all three scenarios into one PR experiment with actual model/report-format metadata and linked main Luna baseline. Configuration support does not prove compatibility or quality of an untried model. Account for every cell and verify actual linked base/stable inputs; a single preset comparison is not a quality or consistency score. The following DeepSeek runs remain historical measured evidence. The full comparison is live-proven by [`pr-401-r35099796991-a1`](https://www.braintrust.dev/app/GitHits/p/githits-cli-agent-evals/experiments/pr-401-r35099796991-a1) (ID `a6313674-e0cd-45b0-8d5b-037d885f1876`): exactly 50 eval roots and 495 tool children. Its persisted base is `13590571-39c1-4a33-831d-db144fb1fc7a` (`main-r35085880981-a1`), with all 50 stable inputs matching. DeepSeek validates 48 reports versus main Luna's 50, takes 2950.371 versus 787.752 cumulative seconds, and uses 495 versus 205 MCP calls. Two malformed JSON finals cause the summary to fail while complete failed-cell export succeeds; neither timed out. Do not repair their finals or confuse successful export with successful cells. See [the permanent comparison](../../../docs/implementation/agentic-eval-metrics.md#full-deepseek-matrix-comparison--2026-09-16) for scenario metrics, failed cells, source-path differences and interpretation. The OpenRouter DeepSeek two-workload canary is proven by [run 35093150512](https://github.com/githits-com/githits-cli/actions/runs/35093150512) on draft PR #401 at SHA `e3fe68c40b80ac74d0c9fa59b0009b28c0841660`. Experiment [`pr-401-r35093150512-a1`](https://www.braintrust.dev/app/GitHits/p/githits-cli-agent-evals/experiments/pr-401-r35093150512-a1) (ID `cf6ec867-e67a-4adb-86bf-ace617b30dc0`) read back two eval spans and 25 tool children, matching 25 completed MCP calls (package 5, router 20), zero failed calls and validated JSON finals. Metadata is DeepSeek/high/prompt-json, channel PR, exporter/schema 3. Actual base is `main-r35085880981-a1` (ID `13590571-39c1-4a33-831d-db144fb1fc7a`), sampled as Luna/low. This verifies PR linkage and integration. Cost remains unknown without a verified DeepSeek rate card; quality is ungraded and a single canary does not prove repeat consistency. For current comparisons, inspect experiment-level `metadata.channel` and `baseExperiment` in the safe exporter result or CI summary. A current main baseline has a `main-r...-a...` name and `channel: main`; PR and local exports resolve the newest such main experiment before initialization. The exporter reports the actual linked base `{id, name}` after `fetchBaseExperiment()`. Validate-only reports the base as unresolved/not queried and performs no discovery. The first main run is a one-time bootstrap; PR and default-local exports fail before initialization when no main baseline exists. Explicit local `--base-experiment` takes precedence and skips discovery. Live readback has proven the first main bootstrap and the PR linkage recorded below. For exports, use the returned experiment name from the SDK readback; it can differ from a reused explicit local name if Braintrust de-duplicates it. Validate-only reports the requested or generated name. The exercised comparison syntax is: ```bash bt experiments --json --project githits-cli-agent-evals compare ``` For custom cross-experiment SQL analysis, join eval rows by `metadata.cellId` and verify identical stable `input` values (including `promptSha256`), rather than joining only by `metadata.workloadId`: the same workload can appear in multiple scenarios. Braintrust's built-in experiment comparison already matches the stable row inputs and avoids this ambiguity. The prior custom-only experiments succeeded but reported only generic Braintrust trace metrics, which were zero and did not expose their custom eval telemetry. Treat that only as historical evidence about the older rows. The preceding native-root experiment is also historical: it set root `tool_calls=119` and `tool_errors=2`, so comparison reported zero before structural children were implemented. Use bounded SQL and the experiment UI for GitHits-specific/custom telemetry; the current exporter uses exact harness-observed lifecycle boundaries and never fabricates timing. ## Validate or explicitly export Credential-free validation maps complete suite artifacts without initializing Braintrust: ```bash bun run agent:e2e:braintrust \ --suite discovery=.agent-eval/suites//suite.json \ --suite intent=.agent-eval/suites//suite.json \ --project githits-cli-agent-evals \ --validate-only ``` An authenticated local subscription export uses the saved `bt` profile to run the same official entrypoint. The exporter, not `bt`, owns the experiment options and safe result file: ```bash bt eval --runner bun --no-auto-instrumentation scripts/agent-eval-braintrust.ts -- \ --suite discovery=.agent-eval/suites//suite.json \ --suite intent=.agent-eval/suites//suite.json \ --project githits-cli-agent-evals \ --source local \ --result-out .agent-eval/braintrust-result.json ``` This default local export lets the exporter derive its stable name and resolve the latest main baseline. Add `--branch ` only when the evaluated suite is detached or has no branch; add `--base-experiment ` to use an explicit local main override. `--experiment ` is also a local-only override. GitHub workflow exports supply their channel, branch, PR number, run identity, and URL through environment-bound arguments and never pass `--experiment`. The suite preflight rejects dry-run suites, suites with no workload cells, duplicate cells, mixed identity or schema contracts, and missing/unsafe child evidence before network setup. It does not reject a failed cell that retains complete report, metrics, workload, and contained prompt evidence. The result file is nonsecret and uses result-file `schemaVersion: 2`; it contains only mode, project, experiment, row count, suite summaries, an export URL when applicable, and `baseExperiment`. In validate-only mode `baseExperiment: null` means unresolved/not queried; in export mode `null` means the required Braintrust readback returned no actual linked base. Experiment metadata records exporter schema/version 3, including model, reasoning effort and Codex report format identity. Historical exporter/schema-2 experiments retain their recorded version. It never contains row bodies, prompts, answers, artifact paths, or credentials. Terminal tool-bearing rows lacking complete/valid observed lifecycle timing are rejected because they cannot produce accurate structural children; an observed started-only call remains an open child. Zero-tool legacy rows remain exportable. Do not create or upload a new experiment unless the user explicitly requests that export.