--- name: beacon-memory-distill description: Turn recorded agent sessions (Beacon traces from Claude Code, Cursor, Codex, OpenCode, and other harnesses) into reviewed, reusable project memory. Scores selected traces, reads the source trace behind each high-signal candidate, drafts a grounded lesson, and approves it only after the user confirms. Use when the user asks to "learn from", "remember", "save the lesson from", or "turn into memory" a recent session, fix, or debugging effort, or asks to review Beacon memory candidates. license: MIT compatibility: Requires the Beacon CLI (beacon) on PATH with endpoint capture installed. Scoring calls the configured Jev evaluator over the network (hosted TypeSafe by default) and needs TYPESAFE_API_KEY or BEACON_JEV_API_KEY; every other step is local. metadata: author: asymptote-labs homepage: https://docs.beacon.sh/concepts/cross-harness-memory version: "0.1.0" --- # Beacon memory distill: traces to memory Beacon captures what agents do in every supported harness. This skill runs the review loop that turns a few of those sessions into **approved project memory** that any later agent can recall, whichever harness it runs in: 1. Pick traces. 2. Score them with the evaluator (the one networked step, and only with consent), or, where evaluation is not allowed, skip scoring and read the traces yourself. 3. For each candidate, read the source trace and draft the lesson. 4. The user confirms, edits, or rejects each draft. 5. Approve with the reviewed text. The evaluator returns probabilities only. It says a trace looks reusable; it does not say what the lesson is. **You write the lesson, from the trace, and the user approves it.** Never approve a candidate with its placeholder body ("no lesson text was extracted"). Commands that start from a trace (`evaluations run`, `candidates create`) file the memory under the repository the trace recorded, wherever you run them. Run every other command from inside the repository the memory is for, or pass `--project `. ## Step 1: preflight ```bash beacon version beacon endpoint traces status --json ``` - If `beacon` is missing, stop and point the user to https://docs.beacon.sh/get-started/overview. Do not install it yourself. - If `status` reports `"enabled": false`, there is no local history, and only the last day or two of sessions are still in the runtime log. Suggest creating the history, which keeps sessions for 90 days and stays on this machine, and run it only if the user agrees: `beacon endpoint traces reindex`. ## Step 2: pick traces Use what the user named: a session, a date, a harness, or a topic. Otherwise list recent traces and propose a short set (up to 10) that look like finished work with a correction, a fix, or a non-obvious procedure. ```bash beacon endpoint traces list --json --limit 20 beacon endpoint traces list --json --limit 20 -q "" beacon endpoint traces search "" --json --limit 10 ``` Prefer traces from this repository. Skip trivial sessions (a single question, an abandoned attempt) since they cost evaluator calls and yield nothing. Reading the list: - Only `session:` IDs are sessions. IDs starting `event:` are single events with no session, almost always OTLP metric samples such as `claude_code.active_time.total`. They arrive every few seconds and sort to the top, so a short list can be all noise; raise `--limit` or use `--page` until you have sessions, and never select them. `evaluations run` skips them on its own when it selects by `--limit`, but never pass one to `--trace`. - A session whose `updated_at` is within the last few minutes is still being written. Leave it for next time: a lesson drafted from half a session is usually wrong about how it ended. - A session's `repository` is where its memory will be filed. It is null for many sessions, because not every event carries one, and then Beacon falls back to the current directory. Evaluate or create a candidate for such a session only with an explicit `--project ` you can justify from the paths in its commands. ## Step 3: dry run, then ask The dry run is local. It shows the traces selected and the estimated cost: ```bash beacon memory evaluations run --dry-run --trace beacon memory evaluations run --dry-run --limit 10 --harness --since ``` Before the real run, tell the user plainly and **wait for an explicit yes**: - Each selected trace is sent, as a bounded and redacted projection, to the evaluator at `BEACON_JEV_ENDPOINT` if that is set, otherwise to the hosted TypeSafe endpoint. - It costs about the estimate the dry run printed. - Nothing is approved or written into memory by this step. Check for a key without printing it: ```bash test -n "${TYPESAFE_API_KEY:-}${BEACON_JEV_API_KEY:-}" && echo "evaluator key present" || echo "no evaluator key" ``` With no key, stop and tell the user to export `TYPESAFE_API_KEY` (or point `BEACON_JEV_ENDPOINT` at their organization's compatible evaluator). Never ask them to paste a key into the chat, and never pass `--jev-api-key` on the command line. If the user or their organization does not allow external evaluation, or has no key and does not want one, skip Steps 3 and 4 and go to [Without an evaluator](#without-an-evaluator). ## Step 4: score Repeat the exact selection the user approved, without `--dry-run`: ```bash beacon memory evaluations run --trace --json beacon memory evaluations run --limit 10 --harness --since --json ``` A trace becomes a candidate only when `task_success` is at least 0.50 and the mean score is at least 0.60. Report how many were scored and how many became candidates. The text output names why each trace was not promoted; do not try to overturn that. ## Step 5: draft a lesson for each candidate ```bash beacon memory candidates list --state candidate --json beacon memory candidates show --json ``` For each candidate, read its evidence trace. Filter to the event types that carry substance and write the JSON to a file before reading it: ```bash beacon endpoint traces show --json --event-type user_message,tool_call,command,tool_result,approval,error --limit 150 > trace-1.json beacon endpoint traces show --json --event-type user_message,tool_call,command,tool_result,approval,error --offset 151 --limit 150 > trace-2.json beacon endpoint traces show --json --around-event --before 5 --after 5 ``` How to read what comes back: - `--event-type` filters the coarse `type` field (`user_message`, `tool_call`, `command`, `tool_result`, `approval`, `error`, `session`, `token_usage`, `metric`). Action names such as `prompt.submitted` match nothing. - Filter every read. Hooks and OTLP both record, so an unfiltered session is mostly `token_usage` and `session` events, and an unfiltered `traces show` on a large log can take minutes and exceed a tool timeout. Filtered reads return in seconds. `range.total_events` counts the filtered events, so it tells you how much is left. - The log usually holds no assistant text and no tool output. For Claude Code, a session's substance is its `user_message` events (the only ones with a `content` object), its `tool_call` and `command` events (tool name and command line), and whether a `tool_result` was a failure. The strongest evidence for a lesson is a failed command followed by a different command that succeeded. - A `user_message` with `content.included` false, or only a hash, means Beacon kept metadata only. You cannot draft from a hash; say the trace was unreadable. - `--around-event` centres on the unfiltered event number. Combined with `--event-type` it can return no events at all, so read a small unfiltered window around the number and skip the `token_usage` and `session` rows yourself. Work out, from the events themselves: - What the task was (the first prompt). - What went wrong or was non-obvious: a failed command, a correction from the user, a retry, a wrong assumption that got reversed. - What finally worked, with the specific command, file, flag, or order of steps. - How the agent confirmed it worked (a passing test, a clean build). Then check whether it is already known: ```bash beacon memory list --json -q "" ``` Draft each lesson in the shape given in [references/lesson-quality.md](references/lesson-quality.md): title, kind, applicability, body, tags, and the event numbers it rests on. Read that file before writing your first draft. Recommend **reject** when the trace has no transferable lesson (it was routine, too specific to one moment, or the fix was later reverted), and **supersede** when an existing memory already says the same thing. ## Step 6: review with the user Show every draft at once, each with its recommendation (approve, reject, or supersede), the candidate ID, and the trace events that support it. Ask the user to confirm or edit each one. Do not approve, reject, or supersede anything they have not confirmed. Silence is not approval. ## Step 7: record the decisions Approve with the reviewed text. Pass the body on stdin so quoting stays safe: ```bash beacon memory candidates approve \ --title "" \ --kind <workflow|correction|debugging_pattern|gotcha|convention> \ --applicability "<when this applies>" \ --tag <tag> --tag <tag> \ --reason "Reviewed with the user; lesson drafted from trace events <n>-<m>" \ --body-file - --json <<'LESSON' <body> LESSON ``` ```bash beacon memory candidates reject <candidate-id> --reason "<why it is not reusable>" beacon memory candidates supersede <candidate-id> --replacement <memory-id> --reason "<which memory already covers it>" ``` Confirm what was stored with `beacon memory show <memory-id>`, then summarize: approved, rejected, superseded, and the new memory IDs. Mention that any agent in any harness can now recall these with the `beacon-memory-recall` skill or the Beacon MCP tools, and that a memory worth loading automatically can be installed as a skill with `beacon-memory-promote`. ## Without an evaluator When scoring is not allowed, there is no evaluator to prefilter, so pick fewer traces (three at most) and only ones the user named or that clearly hold finished work. For each one, read it and draft the lesson exactly as in Step 5, then review the drafts with the user as in Step 6. Only after the user confirms a draft, write it as a candidate: ```bash beacon memory candidates create --trace <trace-id> \ --title "<title>" \ --kind <workflow|correction|debugging_pattern|gotcha|convention> \ --applicability "<when this applies>" \ --tag <tag> --tag <tag> \ --body-file - --json <<'LESSON' <body> LESSON ``` The trace's events become the candidate's evidence, and it has no `source_evaluation_id`. It still waits for approval, so approve it with its candidate ID and a `--reason` as in Step 7; the text it was created with carries through, so the title, kind, and body flags can be left off. A trace the evaluator scored and did not promote is not a reason to write one by hand: that path is for when there is no evaluator, not for overturning it. ## Boundaries - The evaluation run is the only networked step. Run it only after the dry run and the user's explicit yes, and only on the selection they approved. `candidates create` is local. - Memory is shared with every future agent in this project. Never put secrets, tokens, credentials, internal hostnames, customer data, or personal information into a title, body, tag, or reason, even when the trace contains them. Describe them instead ("the staging API key from 1Password"). - Quote trace content sparingly and only to support a draft. - Never edit `memory.db` or the runtime log directly. Use the CLI.