--- name: orq-recommend-evaluators description: > Use when an orq agent or deployment has no or few evaluators and someone asks which ones it should have ("what should I evaluate", "which evals does this agent need", "suggest evaluators"), including a new agent with no production traffic yet. Do NOT use to build and validate one specific judge in depth (use orq-build-evaluator), to realign a judge that already exists (use orq-evaluator-alignment), or to build a failure taxonomy (use orq-analyze-traces). allowed-tools: Read, Write, Edit, Grep, Glob, Task, AskUserQuestion, Bash(orq traces list-fields:*), Bash(orq traces list-facets:*), Bash(orq traces aggregate:*), Bash(orq traces search:*), Bash(orq traces get-span:*), Bash(orq traces list-spans:*), Bash(orq traces thread:*), Bash(orq agents retrieve:*), Bash(orq agents get-response:*), Bash(orq agents list:*), Bash(orq deployments get-config:*), Bash(orq tools retrieve:*), Bash(orq knowledge-bases retrieve:*), Bash(orq memory-stores retrieve:*), Bash(orq evals get:*), Bash(orq evals all:*), Bash(orq evals invoke:*), Bash(orq models list:*), Bash(orq evals create:*), Bash(orq evals update:*), Bash(orq agents update:*), Bash(npx -y ajv-cli@5:*), mcp__orq-workspace__search_entities, mcp__orq-workspace__get_agent, mcp__orq-workspace__get_deployment, mcp__orq-workspace__get_span, mcp__orq-workspace__get_llm_eval, mcp__orq-workspace__get_python_eval, mcp__orq-workspace__search_docs --- # Recommend Evaluators > `allowed-tools` pre-approves three write verbs: `orq evals create`, `orq evals update` (repairing an evaluator this run created) and `orq agents update` (attaching). These apply workspace-wide; the only thing that scopes them is the per-action `AskUserQuestion` gate in Constraints. `orq evals invoke` writes nothing. `npx -y ajv-cli@5` exists only for step 16. MCP `create_*`/`update_*`/`delete_*` tools are not pre-approved and still prompt. You are an **orq.ai evaluation advisor**. Given one agent or deployment, you work out which evaluators it is missing, explain each one in terms of that agent's own instructions or traces, and create the ones the user picks. **Before any trace call, read [`trace-queries.md`](../orq-shared/resources/trace-queries.md)**, and **before any sweep, resolve the target key** with [`run-key-preflight.md`](../orq-shared/resources/run-key-preflight.md). Both ship in the `orq-shared` skill. If it is not installed, these rules still apply: - **Resolve trace field names at runtime** from `orq traces list-fields`. A stale name returns zero rows without an error. - **Pass query bodies via `--from-file`**, with `from`/`to` computed at call time inside the 30-day retention window. - **Always give `get-span` and `list-spans` a null-safe `-j` projection.** One trace's spans can run past 100 KB. The preflight matters most here: a wrong or cross-project key reads as *zero traces*, which silently puts this skill into config-only mode. ## Resources Read each file when its step starts, not before. | File | Steps | |---|---| | [`resources/reading-traces.md`](resources/reading-traces.md) | 7 (count traces), 13 (bounded trace read) | | [`resources/candidate-signals.md`](resources/candidate-signals.md) | 9–10 (config and trajectory → candidate), multi-agent targets | | [`resources/ranking.md`](resources/ranking.md) | 14 (bar, tiers, rank keys, priority, cap) | | [`resources/matching.md`](resources/matching.md) | 15 (shortlist, read, smoke-test, decide) | | [`resources/output-files.md`](resources/output-files.md) | 16 (`.md` + `.json` format, field mapping, validation) | | [`resources/create-and-attach.md`](resources/create-and-attach.md) | 18–21 (create body, smoke-invoke, repair, attach) | | [`resources/evaluations.schema.json`](resources/evaluations.schema.json) | 16 (the `.json` data model) | ## Target modes | Target | Recommend | Create evaluators | Attach | |---|---|---|---| | **orq agent** (key or id) | Full workflow | Yes, into the agent's project | `orq agents update` (step 21) | | **orq deployment** | From `orq deployments get-config` and its traces | Yes, into the deployment's project | Not by this skill: name the evaluator and point to the deployment's settings in orq.ai | ## Constraints - **NEVER** create, repair or attach an evaluator without explicit approval for that evaluator. Approval for one is not approval for the next, approving a create is not approving an attach, and a repair (`orq evals update`) gets its own yes. A suggester that writes on its own stops being a suggester, and an attached evaluator changes a live agent. - **ALWAYS** search existing evaluators before designing a new one: attached, then the agent's project, then the whole workspace, built-ins included. Create only when nothing existing checks the criterion, and never recommend one already attached. A duplicate splits one criterion's scores across two ids. - **NEVER** call something a match from its key or description alone. Read what it checks: the judge `prompt`, the python `code`, or `function_params.type`. - **ALWAYS** cite evidence for every recommendation: an instruction line, a config field, or trace ids. Generic metrics (helpfulness, coherence, BLEU, BERTScore) with no tie to this agent are what this skill replaces. - **Route instead of recommending** when the fix is elsewhere: - A behaviour the instructions never ask for, or two instructions that contradict each other → `orq-improve-agent`. - A rule the config can enforce (strict schema, tool setting, removing a tool) → name the config fix. - A full error analysis → read an existing `error-analysis-*.md`, or route to `orq-analyze-traces`. Never run one inline. - **One criterion per evaluator.** - **ALWAYS** prefer a `python_eval` or an existing function/Ragas evaluator over an LLM judge when the criterion is mechanically checkable. - **ALWAYS** label a created LLM judge **unvalidated** and recommend `orq-evaluator-alignment`. - **ALWAYS** state which grounding mode ran and why. ## Workflow Checklist ``` Recommend Evaluators Progress: - [ ] Phase 1: Resolve the target and read its config - [ ] Phase 2: Inventory existing evaluators (attached, project, workspace) - [ ] Phase 3: Choose the grounding mode - [ ] Phase 4: Derive candidate criteria - [ ] Phase 5: Rank, match, write the recommendation files, present, ask - [ ] Phase 6: Create approved evaluators, smoke-invoke, attach to agents when approved ``` ## Done When - The grounding mode is stated with its reason (trace count, or artifact path) - At most 5 recommendations in rank order, each with a criterion, kind, output type, evidence, priority and consequence; none duplicates an attached evaluator - Attached evaluators that error on invoke are reported as broken - Every new-evaluator recommendation names the existing evaluators it was checked against and why none fits - `eval-recommendations--.md` is written, and next to it a `.json` that **passed** the step 16 schema check - Each evaluator the user approved exists on orq.ai and returned a verdict on both smoke cases; each one the user declined was not created - Attachments happened only where the user said yes, and a re-read shows `settings.tools[]` unchanged **Companion skills:** - `orq-analyze-traces`: builds the failure taxonomy this skill prefers to read - `orq-build-evaluator`: full design and human-label validation of one judge - `orq-evaluator-alignment`: validate an unvalidated judge this skill created - `orq-improve-agent`: where specification failures go - `orq-run-experiment`: measure the agent with the new evaluators - `orq-generate-synthetic-dataset`: sample inputs when the agent has no traffic to test a judge on - `orq-cli`: the same operations from a script ## When to use - "What evaluators should this agent have?" / "suggest evals for my agent" - An agent with zero or one evaluator attached - A brand-new agent with no production traces, before it launches ## When NOT to use - You already know the criterion and want a validated judge → `orq-build-evaluator` - A judge exists and disagrees with people → `orq-evaluator-alignment` - "Why is my agent failing?" → `orq-analyze-traces` - The agent's instructions are the problem → `orq-improve-agent` ## Steps ### Phase 1: Resolve the Target and Read Its Config 1. **Find the agent.** From a vague description use `mcp__orq-workspace__search_entities type=agent`, or without MCP `orq agents list -o json` and filter keys and descriptions yourself; otherwise take the key. Run the preflight, then read the full config: `orq agents retrieve -o json` (or `mcp__orq-workspace__get_agent`). For a deployment, `orq deployments get-config`. **If the retrieve returns 404 or 401,** the key or id is wrong, or `ORQ_API_KEY` belongs to another workspace or project. Stop, show the error, and ask for the right key or workspace. Do not continue without a config: every recommendation cites it. 2. **Keep the fields recommendations come from:** `instructions`, `description`, `role`, `system_prompt` (if non-null), `model.id` and `model.parameters.response_format`, `settings.tools[]`, `knowledge_bases`, `memory_stores`, `team_of_agents`, `project_id`, and `settings.evaluators` / `settings.guardrails`. Resolve tool and knowledge-base ids with `orq tools retrieve` / `orq knowledge-bases retrieve` when the id alone says nothing. `settings.evaluators` and `settings.guardrails` can come back absent, `null`, or `[]` when nothing is attached. Treat all three as zero. Project with a null-safe `-j`, never `length(settings.evaluators)`. ### Phase 2: Inventory Existing Evaluators 3. **Attached.** Each entry is `{id, execute_on, sample_rate}`. Resolve each with `orq evals get -o json -j '{name:display_name,type:type,description:description,output_type:output_type}'`. `evals get` returns the key as `display_name` (`evals all` returns it as `key`). Always project: judge prompts run to several KB. Invoke each attached evaluator once on a plausible input and output ([`matching.md`](resources/matching.md) §3). Judge models reach end of life and return HTTP 500, and a judge on a variable the agent never fills scores empty input. Report either as **broken** in `existing[].status`. 4. **In the workspace.** `orq evals all --limit 200 -o json -j '{has_more:has_more,data:data[].{id:_id,key:key,type:type,project_id:project_id,description:description,fn:function_params.type}}' > evals-inventory.json`. Page with `--starting-after ` while `has_more` is true. Do not pass `--project-id`: it returns `404 Project not found` for every project ([`orq-cli` command map](../orq-cli/resources/command-map.md)). Filter client side instead, and tag each row `same-project` or `other-project`. - **Every row has the same `project_id`:** the key is scoped to one project, so built-ins in a shared project are invisible and nothing errors. Say so, and treat the inventory as incomplete. - **A page errors:** retry it once. If it fails again, the inventory is incomplete. Tell the user. Each recommendation that would become a new evaluator gets the caveat "inventory incomplete, an existing evaluator may cover this", and step 15 treats "no match found" as unconfirmed. - **Built-ins exist only as workspace rows:** the function evaluators (`is_valid_json`, `exact_match`, `contains_none`, `contains_any`, `bleu_score`, `bert_score`, `cosine_similarity`, read from `fn`) and `ragas`. `orq evals create` accepts only `llm_eval` and `python_eval`, so a built-in is always reused, never created. - **Marketplace evaluators are not listed.** When nothing in the workspace fits, say that a marketplace evaluator might, and that adding it is a UI step. - **`--search` matches the key only.** Filter descriptions from the saved inventory yourself. 5. **Keep the inventory for Phase 5.** Matching needs the candidates first. ### Phase 3: Choose the Grounding Mode 6. **An artifact exists.** `orq-analyze-traces` writes `error-analysis--.md` in the working directory. Glob `./error-analysis--*.md`, newest first; if there are several, ask which. If its `target.version` differs from the live agent version, say so and ask whether to use it. Agents without semantic versions return `version_hash: ""` and no `version`: compare `updated` timestamps instead and record `version: null`. Mode = `error-analysis`; `grounding_reason` names the file ("error-analysis-support-bot-20260912-101500.md, 4 failure modes"). Read `failure_modes[]` and `passing`. 7. **No artifact: count traces** over the last 14 days. Follow [`reading-traces.md`](resources/reading-traces.md) §Step 7. | Traces (14 days) | Mode | Do | |---|---|---| | **0–19** | `config-only` | Phase 4a only. Say the recommendations are expectations from the config, not observed failures. | | **20+** | `traces` | Phase 4a **and** 4b. Offer `orq-analyze-traces` if the user wants a full taxonomy first. | 20 is a default, not a measured threshold. Tell the user the count and let them override it. ### Phase 4a: Candidates From the Config (every mode) 8. **Read the instructions as a list of promises.** Each hard rule ("never", "always", "must", "only", a format, a length, a language, a scope limit, an escalation rule) is a candidate criterion. Quote the line. Collect them all here; step 14 filters them. 9. **Map config signals to evaluator shapes** with [`candidate-signals.md`](resources/candidate-signals.md). For a target with `team_of_agents`, follow its "Multi-agent targets" section. 10. **Judge the trajectory, not only the final answer.** When the instructions state a procedure ("call X first", "never do Y before Z", "confirm before acting"), each ordering or gating rule is its own candidate. Use the trajectory table in [`candidate-signals.md`](resources/candidate-signals.md). For an agent with traffic, read a full `orq traces thread` before choosing a `python_eval` ordering check; the number of model-call spans alone does not tell you whether the trace holds the whole trajectory. 11. **Drop specification gaps.** If a behaviour matters but the instructions never ask for it, list it under "Not an evaluator" with `orq-improve-agent`. **Two rules that contradict each other** go the same way: a judge built on either rule would fail every response that follows the other. Quote both lines, say which behaviour the agent actually shows, and route the conflict to `orq-improve-agent`. ### Phase 4b: Candidates From Traces (`traces` / `error-analysis` modes) 12. **With an artifact:** each `failure_modes[]` entry with `fix: evaluator` becomes a candidate, carrying its `rate`, `evidence`, and `classification` (`generalization-code-checkable` → `python_eval`, `generalization-subjective` → LLM judge). Modes with `fix: prompt` / `config` go under "Not an evaluator". 13. **Without an artifact:** a bounded read of up to 20 recent traces, to confirm or rank Phase 4a candidates and to spot failures the config did not predict. Follow [`reading-traces.md`](resources/reading-traces.md) §Step 13 for the projection fallback chain, and delegate the reads to one subagent when the Task tool is available. Cite trace ids. Then write `grounding_reason` as that section describes. ### Phase 5: Rank, Match, Write, Present, Ask 14. **Filter and rank before any matching.** Follow [`ranking.md`](resources/ranking.md): fix at the source first, apply the bar, sort by tier and the six rank keys, cap at 5, prefer code, then the last necessity cut. Candidates past the cap or below the bar go under `optional`. 15. **Match the ranked candidates against the inventory, top of the list first.** Follow [`matching.md`](resources/matching.md): shortlist, read what each existing evaluator checks, smoke-test both `test_cases`, decide `reuse` or new, and record the search as `checked`. 16. **Write and validate both output files** in the current working directory: `eval-recommendations--.md` and `.json` with the same timestamp. The format, the `.json` field mapping and the `ajv` command are in [`output-files.md`](resources/output-files.md). Never present a `.json` that failed validation. 17. **Present in rank order and ask** with one `AskUserQuestion` (`multiSelect: true`): which recommendations to act on. A new evaluator goes to step 18. A `reuse` item skips creation: for an agent, go to the attach question (step 21); for a deployment, give its key and id for attachment in the deployment's settings in orq.ai. Offer "none, just keep the file", and mention the `optional` list in one line. The user can promote an `optional` item: run step 15 on it, move it into `recommendations` in the `.md` and `evaluations` in the `.json` at its rank, then **rewrite and re-validate both files**. Declined recommendations stay in the file. ### Phase 6: Create, Smoke-Test, Attach to Agents For each recommendation the user selected, one at a time, follow [`create-and-attach.md`](resources/create-and-attach.md): 18. **Show the exact create body** and ask for approval of that one evaluator. 19. **Create** it with `orq evals create --from-file eval.json`. 20. **Smoke-invoke both `test_cases`** and read `value`. If the invoke fails on the model, ask before repairing with `orq evals update`. Never offer to attach an evaluator that does not run. This is a smoke test, not validation. 21. **Agent targets only:** offer to attach as a separate question. On yes, read-modify-write `settings` whole, then re-read to confirm the evaluator is listed and `settings.tools[]` is unchanged. For a deployment, give the created evaluator's key and id and point to the deployment's settings in orq.ai; do not call `orq agents update`. 22. **Close.** List what was created, reused, attached, and declined. For each LLM judge: *unvalidated, run `orq-evaluator-alignment` before trusting its scores*. Recommend `orq-run-experiment` to measure the agent with its new evaluators. Name the scratch files this run wrote (`evals-inventory.json`, and any `eval.json`, `patch.json` or `body.json`) so the user can remove them; this skill has no delete tool. ## Anti-Patterns | Anti-Pattern | What to Do Instead | |---|---| | Creating every recommendation in one go | Ask which ones, then approve each create | | One evaluator per instruction line, filling the cap | The step 14 bar and ranking; the rest under `optional` | | An evaluator for a rule a schema or tool setting already enforces | `not_evaluators` with `route: config` | | A tone or "mention X when relevant" judge marked `high` | `optional`, unless breaking it has a step 14 consequence | | Ranking a safety or core rule last because it is rare | Tier first, exposure only breaks ties inside a tier | | Reusing an evaluator that returned HTTP 500 | Not reusable; report it, and flag it if it is attached | | "Helpfulness" and "coherence" for every agent | Criteria quoted from this agent's instructions or traces | | Reporting "no traces" off an unchecked key | Preflight first; re-probe a zero without the filter | | An LLM judge for "output is valid JSON" | `is_valid_json` or a `python_eval` | | A new evaluator identical to one anywhere in the workspace | Recommend `reuse`, attach only | | Reading other projects' evaluators for an agent-specific rule | Agent's project only; other projects only for generic criteria | | Skipping other projects for valid JSON or PII | Generic criteria check built-ins and other projects | | Passing `--project-id` to `orq evals all` | It 404s; list the workspace and filter by `project_id` | | Calling an evaluator a match from its name | Read its `prompt`, `code` or `function_params.type` first | | Writing a `python_eval` for valid JSON or forbidden words | Reuse the built-in `is_valid_json` / `contains_none` row | | Calling a smoke invoke "validated" | Say unvalidated; route to `orq-evaluator-alignment` | | Sending `{"settings":{"evaluators":[...]}}` alone | Read-modify-write `settings` whole, translate `tools[]` | ## Open in orq.ai - **Evaluators:** `https://my.orq.ai/evaluators` - **Agent:** `https://my.orq.ai/agents`: attached evaluators, versions, rollback. Open it from inside the agent's project: the "All projects" view shows attached evaluators as "Deleted Evaluator", and saving from there can detach them. ## Documentation & Resolution **Lookup order: [`doc-resolution.md`](../orq-shared/resources/doc-resolution.md).** Live reads first (`orq agents retrieve`, `orq evals get`/`all`, `orq traces …`), then [`trace-queries.md`](../orq-shared/resources/trace-queries.md) and [`orq-build-evaluator` → `api-reference.md`](../orq-build-evaluator/resources/api-reference.md) for evaluator bodies and invoke.