--- name: phoenix-cli description: Debug LLM applications using the Phoenix CLI. Fetch traces, spans, and sessions, annotate them, analyze errors, inspect datasets, review experiments, query annotation configs, and use the GraphQL API. Use whenever the user works with a Phoenix instance from the terminal. license: Apache-2.0 compatibility: Requires Node.js (for npx) or global install of @arizeai/phoenix-cli. Optionally requires jq for JSON processing. metadata: author: arize-ai version: "3.5.0" --- # Phoenix CLI ## Invocation ```bash px # if installed globally npx @arizeai/phoenix-cli # no install required ``` The CLI uses singular resource commands with subcommands like `list` and `get`: ```bash px trace list px trace get px trace annotate px trace add-note px trace delete px trace-annotations delete px span list px span annotate px span add-note px span delete px span-annotations delete px session list px session get px session annotate px session add-note px session delete px session-annotations delete px dataset list px dataset get px dataset delete px experiment list px experiment get px experiment delete px prompt list px prompt get px prompt delete px project list px project get px project delete px annotation-config list px annotation-config get px annotation-config create px annotation-config update px annotation-config delete px auth login px auth logout px auth status px profile list px profile show [name] px profile create px profile use px profile edit px profile delete px api graphql px docs fetch px setup px self update ``` Every `delete` above is gated: it requires `PHOENIX_CLI_DANGEROUSLY_ENABLE_DELETES=true` in the environment and prompts for confirmation unless `-y`/`--yes` is passed (`px profile delete` is local-only and takes `--yes` without the env gate). Without the env var the command exits without deleting anything. ## Setup ```bash export PHOENIX_ENDPOINT=http://localhost:6006 export PHOENIX_PROJECT=my-project export PHOENIX_API_KEY=your-api-key # if auth is enabled ``` `PHOENIX_ENDPOINT` is the base URL for API access. It usually holds the same URL as `PHOENIX_COLLECTOR_ENDPOINT`; when only the collector variable is set, the CLI uses it for API access too. For interactive local use, `px auth login` stores an OAuth session in the selected profile; the session acts with the permissions of the user who logged in. API keys take precedence over OAuth tokens when both are configured. OAuth access tokens are refreshed automatically for REST, GraphQL, and PXI requests, and rotated tokens are persisted to the selected profile. Always use `--format raw --no-progress` when piping to `jq`. ### `px setup` — onboarding `px setup` connects the app in the current directory to a Phoenix deployment and writes `.env.phoenix` (mode 0600, gitignored). The interactive flow is for humans — it prompts, launches coding agents, and polls for traces. **From an agent, always pass `--no-input`:** ```bash # Register only: connection + .env.phoenix, no source changes. px setup --no-input --endpoint http://localhost:6006 --project my-app --format raw ``` Headless requires a clean git repo and, by default, stops after writing the files — it will not touch source unless you ask. If auth is enabled, also set `PHOENIX_API_KEY`. The project doesn't need to exist — Phoenix creates it on first trace. Missing inputs exit `3` with exact remediation; cancel exits `2`. To also instrument the app, name the lane — headless has no prompt to pick one from, so `--instrument` requires `--agent`: ```bash px setup --no-input --instrument --agent claude --yolo --format raw ``` `--yolo` matters: a background agent has no terminal to approve its edits on, so without it the run stalls until trace verification times out. `--language python` skips the agent's language detection. `--docs-mcp` connects the Phoenix docs MCP server to the hand-off agent (`claude mcp add` for claude, config-file merge for cursor/opencode; codex unsupported) and skips the `.px/docs` download — the agent searches docs on demand instead; any failure falls back to the download. `--no-docs-mcp` suppresses the interactive offer. `--format raw` prints `{"endpoint","project","files","instrumentation","tracesVerified","tracesUrl"}` — check `tracesVerified`, which is set only when the API confirmed a trace arriving, not when the agent claims it finished. A run whose wait ran out with no trace exits `6`, not `0`: the configuration and edits are real, but tracing is not confirmed working. Treat that as a failure to report, not a success — and do not substitute the hand-off agent's own exit code or summary for the verdict. Registering without `--instrument`, and a human answering "verify later" at the timeout prompt, both exit `0`. `tracesVerified` is `false` for a registration-only run too, so it alone can't tell "nothing to verify" from "no trace arrived". Read `verification` (`verified` / `notVerified` / `deferred`, absent when there was nothing to verify) when you need the difference. Re-runnable slices, so an already-registered repo skips the questions: ```bash px setup instrument --agent claude # instrument + verify only px setup skills # install the Phoenix coding-agent skills ``` ### `px setup mcp` — register the remote MCP server Wire the Phoenix remote MCP server (`/mcp`) into a coding agent so it can query Phoenix data. The endpoint is inferred from `--endpoint`, the active profile, or `PHOENIX_ENDPOINT`. Bare command prompts for scope (global default) then agent; `--agent` skips both prompts. ```bash px setup mcp --agent codex --no-input --format raw px setup mcp --agent claude --local # write this repo's .mcp.json ``` Agents: `claude`, `codex`, `gemini`, `cursor`, `opencode`, `vscode`. Scope is `--global` (default) or `--local` (repo; Codex is global-only). Auth is OAuth by default (URL-only config, browser login on first use); pass `--header "Name: value"` (repeatable) for an API-key bearer fallback — for Codex a `Authorization: Bearer ${VAR}` header becomes `bearer_token_env_var`. `--format raw` prints `{"endpoint","url","serverName","agent","scope","auth","file?"}`. ## Auth ```bash px auth login # browser-based OAuth login px auth login --no-browser # print URL for SSH/headless use px auth logout # clear OAuth tokens; leaves API keys px auth status # check connection and authentication px auth status --endpoint http://other:6006 # check a specific endpoint px auth status --profile staging # check a named profile's connection px auth status --format raw # machine-readable credential source ``` `auth status` reports the credential source (`flag`, `env`, `profile-key`, `oauth`, or `none`). OAuth status includes the token expiry. When the stored credential source is `oauth` and the authenticated probe fails, `auth status` retries once without credentials and reports anonymous access only if the server explicitly says access is anonymous. This keeps a stale or expired profile token from being reported as an auth failure against a deployment that has since switched from OAuth to anonymous access. ## Profiles Named profiles let you switch between multiple Phoenix instances (local, staging, cloud) without juggling environment variables. Profiles are stored in `~/.px/settings.json` (or `$XDG_CONFIG_HOME/px/settings.json`). Configuration priority (highest to lowest): CLI flags > env vars > active profile > nearest `.env.phoenix` file > built-in defaults. The CLI also discovers the nearest `.env.phoenix` file at or above the current working directory (the same file `px setup` writes). Credentials are resolved as one group, so a process API key is never combined with file-provided headers. Set `PHOENIX_DISCOVER_CONFIG=false` to disable discovery. ```bash px profile list # list all profiles (shows active profile) px profile show # show the active profile's settings px profile show staging # show a named profile's settings px profile create prod --endpoint https://app.phoenix.arize.com --api-key --activate px profile create local --endpoint http://localhost:6006 --project my-app px profile use prod # switch the active profile px profile edit prod # open profile JSON in $EDITOR (validates on save) px profile delete prod --yes # delete a profile (--yes skips confirmation) ``` Use `--profile ` on any command to target a specific profile without changing the active one: ```bash px trace list --profile staging --limit 10 --format raw --no-progress | jq . px auth status --profile prod ``` `px profile create` options: `--endpoint `, `--project `, `--api-key `, `--header ` (repeatable), `--activate`. ## Projects ```bash px project list # list all projects (table view) px project list --format raw --no-progress | jq '.[].name' # project names as JSON px project list --name-contains prod # filter by name substring (case-insensitive) px project get my-project --format raw --no-progress # single record by exact name px project get my-project --format raw --no-progress | jq -r '.id' # extract project id ``` `project list` accepts `--limit ` (projects fetched per page) and `--name-contains `, which filters server-side on a case-insensitive name substring. Use it instead of piping `list` through `grep` when you only know part of a project's name. `project get` exits with `ExitCode.FAILURE` (1) on a name miss and writes a `StructuredError` `{error, code: "FAILURE", hint}` to stderr in `--format json|raw`. ## Traces ```bash px trace list --limit 20 --format raw --no-progress | jq . px trace list --last-n-minutes 60 --limit 20 --format raw --no-progress | jq '.[] | select(.status == "ERROR")' px trace list --since 2025-01-15T00:00:00Z --limit 50 --format raw --no-progress | jq . px trace list --since 2025-01-15T00:00:00Z --until 2025-01-16T00:00:00Z --limit 50 --format raw --no-progress | jq . # time range (until is exclusive) px trace list --format raw --no-progress | jq 'sort_by(-.duration) | .[0:5]' px trace list --include-notes --format raw --no-progress | jq '.[].notes' px trace get --format raw | jq . px trace get --format raw | jq '.spans[] | select(.status_code != "OK")' px trace get --include-notes --format raw | jq '.notes' px trace annotate --name reviewer --label pass px trace annotate --name reviewer --score 0.9 --format raw --no-progress px trace annotate --name reviewer --label pass --identifier "" # tag with a coding annotation identifier px trace add-note --text "needs follow-up" px trace add-note --text "needs follow-up" --identifier "" # tag + upsert on identifier px trace-annotations delete --identifier "" --all -y # nuke every annotation tied to this coding annotation identifier ``` `px -annotations delete` requires `--all` or both `--start-time` and `--end-time` and emits `{deleted: true, target, filter}` on success. ### Trace JSON shape ``` Trace traceId, status ("OK"|"ERROR"), duration (ms), startTime, endTime annotations[] (with --include-annotations, excludes note) name, result { score, label, explanation } notes[] (with --include-notes) name="note", result { explanation } rootSpan — top-level span (parent_id: null) spans[] name, span_kind ("LLM"|"CHAIN"|"TOOL"|"RETRIEVER"|"EMBEDDING"|"AGENT"|"RERANKER"|"GUARDRAIL"|"EVALUATOR"|"DECISION"|"UNKNOWN") status_code ("OK"|"ERROR"|"UNSET"), parent_id, context.span_id notes[] (with --include-notes) name="note", result { explanation } attributes input.value, output.value — raw input/output llm.model_name, llm.provider llm.token_count.prompt/completion/total llm.token_count.prompt_details.cache_read llm.token_count.completion_details.reasoning llm.input_messages.{N}.message.role/content llm.output_messages.{N}.message.role/content llm.invocation_parameters — JSON string (temperature, etc.) exception.message — set if span errored ``` ## Spans ```bash px span list --limit 20 # recent spans (table view) px span list --last-n-minutes 60 --limit 50 # spans from last hour px span list --since 2025-01-15T00:00:00Z --limit 50 # spans since a timestamp px span list --since 2025-01-15T00:00:00Z --until 2025-01-16T00:00:00Z --limit 50 # time range (until is exclusive) px span list --span-kind LLM --limit 10 # only LLM spans px span list --status-code ERROR --limit 20 # only errored spans px span list --name chat_completion --limit 10 # filter by span name px span list --trace-id --format raw --no-progress | jq . # all spans for a trace px span list --span-id --format raw --no-progress | jq . # fetch specific spans by ID (server >= 19.6.0) px span list --parent-id null --limit 10 # only root spans px span list --parent-id --limit 10 # only children of a span px span list --include-annotations --limit 10 # include annotation scores px span list --include-notes --limit 10 # include span notes px span list --attribute llm.model_name:gpt-4 --limit 10 # filter by string attribute px span list --attribute llm.token_count.total:500 --limit 10 # filter by numeric attribute px span list --attribute 'user.id:"12345"' --limit 10 # force string match for numeric-looking value px span list --attribute session.id:sess:abc:123 --limit 20 # colon in value OK (split on first colon only) px span list --attribute llm.model_name:gpt-4 --attribute session.id:abc --limit 10 # AND multiple filters px span list output.json --limit 100 # save to JSON file px span list --format raw --no-progress | jq '.[] | select(.status_code == "ERROR")' px span annotate --name reviewer --label pass px span annotate --name checker --score 1 --annotator-kind CODE px span annotate --name reviewer --label pass --identifier "" # tag with a coding annotation identifier px span add-note --text "verified by agent" px span add-note --text "verified by agent" --identifier "" # tag + upsert on identifier px span-annotations delete --identifier "" --all -y # nuke every annotation tied to this coding annotation identifier ``` `span list` orders by ingestion (newest first), not `start_time`; they diverge for late-arriving spans (backfills, replays). To sort by `start_time` (server >= 20.16.0), call REST and keep `sort`/`order` fixed across `next_cursor` pages: ```bash curl -s -H "Authorization: Bearer $PHOENIX_API_KEY" \ "$PHOENIX_ENDPOINT/v1/projects/my-project/spans?sort=start_time&order=desc&limit=20" ``` ### Span JSON shape ``` Span name, span_kind ("LLM"|"CHAIN"|"TOOL"|"RETRIEVER"|"EMBEDDING"|"AGENT"|"RERANKER"|"GUARDRAIL"|"EVALUATOR"|"DECISION"|"UNKNOWN") status_code ("OK"|"ERROR"|"UNSET"), status_message context.span_id, context.trace_id, parent_id start_time, end_time attributes input.value, output.value — raw input/output llm.model_name, llm.provider llm.token_count.prompt/completion/total llm.input_messages.{N}.message.role/content llm.output_messages.{N}.message.role/content llm.invocation_parameters — JSON string (temperature, etc.) exception.message — set if span errored annotations[] (with --include-annotations, excludes note) name, result { score, label, explanation } notes[] (with --include-notes) name="note", result { explanation } ``` ## Sessions ```bash px session list --limit 10 --format raw --no-progress | jq . px session list --order asc --format raw --no-progress | jq '.[].session_id' px session list --include-annotations --include-notes --format raw --no-progress | jq '.[].notes' px session get --format raw | jq . px session get --include-annotations --format raw | jq '.session.annotations' px session get --include-notes --format raw | jq '.session.notes' px session annotate --name reviewer --label pass px session annotate --name reviewer --score 0.9 --format raw --no-progress px session annotate --name reviewer --label pass --identifier "" # tag with a coding annotation identifier px session add-note --text "verified by agent" px session add-note --text "verified by agent" --identifier "" # tag + upsert on identifier px session-annotations delete --identifier "" --all -y # nuke every annotation tied to this coding annotation identifier px session delete -y # requires PHOENIX_CLI_DANGEROUSLY_ENABLE_DELETES=true ``` `session list` has no filter flag. To select sessions by shape — error counts, token totals, tool use, annotation labels — use the session filter expression language through GraphQL (see [Session filter expressions](#session-filter-expressions)). ### Session JSON shape ``` SessionData id, session_id, project_id start_time, end_time token_count_prompt, token_count_completion, token_count_total — cumulative across all LLM spans in the session (int, default 0) annotations[] (with --include-annotations, excludes note) name, result { score, label, explanation } notes[] (with --include-notes) name="note", result { explanation } traces[] id, trace_id, start_time, end_time ``` ## Datasets / Experiments / Prompts ```bash px dataset list --format raw --no-progress | jq '.[].name' px dataset get --format raw | jq '.examples[] | {input, output: .expected_output}' px dataset get --split train --format raw | jq . # filter by split px dataset get --version --format raw | jq . px experiment list --dataset --format raw --no-progress | jq '.[] | {id, name, failed_run_count}' px experiment get --format raw --no-progress | jq '.[] | select(.error != null) | {input, error}' px prompt list --format raw --no-progress | jq '.[].name' px prompt get --format text --no-progress # plain text, ideal for piping to AI ``` ## Annotation Configs Full CRUD: `list`, `get`, `create`, `update`, `delete`. Types are `CATEGORICAL` (labels + optional scores), `CONTINUOUS` (numeric range), `FREEFORM` (free text). ```bash px annotation-config list # all configs (table view) px annotation-config list --format raw --no-progress | jq -r '.[].name' # config names as JSON px annotation-config get response-quality --format raw --no-progress # one config by name or ID # create — categorical (scored labels), continuous (numeric range), or freeform (free text) px annotation-config create --type CATEGORICAL --name response-quality --value good=1 --value bad=0 px annotation-config create --type CONTINUOUS --name confidence --lower-bound 0 --upper-bound 1 px annotation-config create --type FREEFORM --name reviewer-notes --description 'Free-form reviewer feedback' # update by name or ID — only the fields you pass change; type is immutable px annotation-config update response-quality --name answer-quality --optimization-direction MAXIMIZE px annotation-config update response-quality --value good=1 --value acceptable=0.5 --value bad=0 px annotation-config update response-quality --description "Updated" --format raw --no-progress | jq -r '.id' # delete by ID — requires PHOENIX_CLI_DANGEROUSLY_ENABLE_DELETES=true; --yes skips the prompt px annotation-config delete QW5ub3RhdGlvbkNvbmZpZzoxMjM= --yes ``` Categorical values are specified the same way in `create` and `update`: repeatable `--value label[=score]` (score optional), or a single `--values ''` payload — mutually exclusive. `update` fetches the existing config, merges your flags, and writes the full body back via `PUT /v1/annotation_configs/{id}`; it requires at least one field flag. Other type-specific flags: `--lower-bound`/`--upper-bound` (CONTINUOUS/FREEFORM), `--threshold` (FREEFORM). Invalid input (bad flags, type mismatches, malformed values) exits `3` (`INVALID_ARGUMENT`) with a `{error, code, hint?}` JSON envelope on stderr in `raw`/`json` mode. `get`/`create`/`update` output the config object (single object in `raw`/`json`, not an array). ## GraphQL For ad-hoc queries not covered by the commands above. Output is `{"data": {...}}`. ```bash px api graphql '{ projectCount datasetCount promptCount evaluatorCount }' px api graphql '{ projects { edges { node { name traceCount tokenCountTotal } } } }' | jq '.data.projects.edges[].node' px api graphql '{ datasets { edges { node { name exampleCount experimentCount } } } }' | jq '.data.datasets.edges[].node' px api graphql '{ evaluators { edges { node { name kind } } } }' | jq '.data.evaluators.edges[].node' # evaluator kind values: "LLM" | "CODE" | "BUILTIN" # CODE = server-side code evaluator running in a sandbox; BUILTIN = pre-built server evaluator # Introspect any type px api graphql '{ __type(name: "Project") { fields { name type { name } } } }' | jq '.data.__type.fields[]' ``` Key root fields: `projects`, `getProjectByName(name:)`, `datasets`, `prompts`, `evaluators`, `projectCount`, `datasetCount`, `promptCount`, `evaluatorCount`, `viewer`. `getProjectByName(name:)` targets one project; `projects(first: 1)` picks an arbitrary one. There is no `traces` connection: to list traces, query `spans` with `filterCondition: "parent_span is None"`, which keeps root spans, as the UI's traces table does. See [Filter expressions](#filter-expressions) below. ### Filter expressions `spans`, `sessions`, and the project aggregates take filter conditions: Python boolean expressions compiled server-side. There are three languages, and the argument picks the language. Read [references/filter-expressions.md](references/filter-expressions.md) before writing a condition; it has the full vocabulary, operators, and compiled examples for each. | Argument | Matches | Names come from | | -------- | ------- | --------------- | | `filterCondition` | individual spans | the exhaustive table in the reference | | `traceFilterCondition` | whole traces | `traceFilterVocabulary` | | `sessionFilterCondition` | sessions | `sessionFilterVocabulary` | **Root spans.** There is no `traces` connection and no root-span argument. `filterCondition: "parent_span is None"` keeps root spans, including orphans whose parent was never received, and is what the UI's traces table runs; `parent_id is None` keeps only spans with no parent id. A root span is usually one per trace, and either clause composes with the rest of the filter: ```bash px api graphql '{ getProjectByName(name: "default") { spans( first: 20 filterCondition: "parent_id is None and status_code == \"ERROR\"" sort: { col: startTime, dir: desc } ) { edges { node { spanId name latencyMs } } } } }' | jq '.data.getProjectByName.spans.edges[].node' ``` **Annotations.** The accessor picks the level, and the wrong level matches nothing: | Accessor | Matches annotations on | Written by | | -------- | ---------------------- | ---------- | | `annotations["name"]` | the span itself | `px span annotate`, `px span add-note` | | `trace_annotations["name"]` | the span's parent trace | `px trace annotate`, `px trace add-note` | | `session_annotations["name"]` | the session (session filter only) | `px session annotate`, `px session add-note` | ```bash px api graphql '{ getProjectByName(name: "default") { spans( first: 20 filterCondition: "parent_id is None and trace_annotations[\"quality\"].label == \"poor\"" ) { edges { node { spanId name } } } } }' | jq '.data.getProjectByName.spans.edges[].node' ``` **Traces.** `traceFilterCondition` keeps the spans of matching traces and composes with `filterCondition`: ```bash px api graphql '{ getProjectByName(name: "default") { spans( first: 20 filterCondition: "parent_id is None" traceFilterCondition: "error_count > 0 and latency_ms > 1000" ) { edges { node { spanId name latencyMs } } } } }' | jq '.data.getProjectByName.spans.edges[].node' ``` **Sessions.** `px session list` has no filter flag, so selecting sessions by shape goes through GraphQL: ```bash px api graphql '{ projects(first: 1) { edges { node { sessions( first: 10 sessionFilterCondition: "num_traces > 5 and any(span.status_code == \"ERROR\" for span in spans)" ) { edges { node { sessionId numTraces numTracesWithError } } } } } } }' | jq '.data.projects.edges[0].node.sessions.edges[].node' ``` **Discover names and validate.** The vocabularies are generated from the compiler's own bindings, so they always match what compiles: ```bash px api graphql '{ projects(first: 1) { edges { node { traceFilterVocabulary { name type category description iterableName } } } } }' \ | jq '.data.projects.edges[0].node.traceFilterVocabulary[] | {name, type, category}' px api graphql '{ projects(first: 1) { edges { node { validateSpanFilterCondition(condition: "parent_id is None") { isValid errorMessage } validateTraceFilterCondition(condition: "error_count > 0") { isValid errorMessage } validateSessionFilterCondition(condition: "num_traces > 5") { isValid errorMessage } } } } }' ``` On fields that accept both levels (e.g. `Project.recordCount`), `sessionFilterCondition` and `filterCondition` are mutually exclusive. ## Docs Download Phoenix documentation markdown for local use by coding agents. ```bash px docs fetch # fetch default workflow docs to .px/docs px docs fetch --workflow tracing # fetch only tracing docs px docs fetch --workflow tracing --workflow evaluation px docs fetch --dry-run # preview what would be downloaded px docs fetch --refresh # clear .px/docs and re-download px docs fetch --output-dir ./my-docs # custom output directory ``` Key options: `--workflow` (repeatable, values: `tracing`, `evaluation`, `datasets`, `prompts`, `integrations`, `sdk`, `self-hosting`, `all`), `--dry-run`, `--refresh`, `--output-dir` (default `.px/docs`), `--workers` (default 10).