agentmemory: persistent memory for AI coding agents

Your coding agent remembers everything. No more re-explaining. Built on iii engine
Persistent memory for Claude Code, GitHub Copilot CLI, Cursor, Gemini CLI, Codex CLI, Hermes, OpenClaw, pi, OpenCode, and any MCP client.

🇬🇧 English • 🇨🇳 简体中文 • 🇹🇼 繁體中文 • 🇯🇵 日本語 • 🇰🇷 한국어 • 🇵🇹 Português • 🇧🇷 Português (Brasil) • 🇪🇸 Español • 🇩🇪 Deutsch • 🇫🇷 Français • 🇮🇹 Italiano • 🇳🇱 Nederlands • 🇵🇱 Polski • 🇨🇿 Čeština • 🇷🇴 Română • 🇭🇺 Magyar • 🇬🇷 Ελληνικά • 🇸🇪 Svenska • 🇩🇰 Dansk • 🇳🇴 Norsk • 🇫🇮 Suomi • 🇷🇺 Русский • 🇺🇦 Українська • 🇹🇷 Türkçe • 🇮🇱 עברית • 🇸🇦 العربية • 🇮🇳 हिन्दी • 🇧🇩 বাংলা • 🇵🇰 اردو • 🇹🇭 ไทย • 🇻🇳 Tiếng Việt • 🇮🇩 Bahasa Indonesia • 🇵🇭 Tagalog

rohitg00/agentmemory | Trendshift

Design doc: 1.6k stars / 230 forks on the gist

The gist extends Karpathy's LLM Wiki pattern with confidence scoring, lifecycle, knowledge graphs, and hybrid search: agentmemory is the implementation.

npm version CI License Stars

95.2% retrieval R@5 92% fewer tokens 54 MCP tools 12 auto hooks 0 external DBs 2,500+ tests passing

agentmemory demo

Install • Quick Start • Benchmarks • vs Competitors • Agents • How It Works • MCP • Viewer • Powered by iii • Config • API

--- ## Install Requirements: - Node.js 20 or newer with npm and npx (`node -v`, `npm -v`, and `npx -v`). - macOS/Linux automatic iii-engine installation also needs `curl`, a POSIX `sh`, and `tar`. Minimal images such as `node:20-slim` may not include them. - Native Windows requires the pinned iii-engine v0.22.1 `iii.exe` to be installed manually. WSL2 or Docker Desktop are the other supported paths. Canonical fresh-install command: ```bash npx -y @agentmemory/agentmemory@latest ``` The first run is an interactive setup: pick the agents to wire (Claude Code, Cursor, Codex, Gemini CLI, OpenCode, ...), pick an LLM provider or stay keyless, and it seeds the config, starts the memory server and its pinned iii engine, and offers to install globally so the bare `agentmemory` command works everywhere afterward. `-y` accepts npx's package prompt and `@latest` avoids a stale cached release. A provider makes LLM features available, but LLM-written observation compression starts only when `AGENTMEMORY_AUTO_COMPRESS=true` is also set. Keyless mode disables vector embeddings. `memory_recall` (the `mem::search` path) uses BM25, while `memory_smart_search` can also fuse structural graph matches when graph data already exists. For free on-device semantic recall, set `EMBEDDING_PROVIDER=local` in `~/.agentmemory/.env` and restart. The first embedding request downloads `Xenova/all-MiniLM-L6-v2`; inference runs locally after that initial model download. The local runtime uses four ports: `3111` for REST/MCP HTTP, `3112` for iii streams, `3113` for the viewer, and `49134` for the iii worker WebSocket. Persistent iii state lives in `~/Library/Application Support/agentmemory` on macOS, `$XDG_DATA_HOME/agentmemory` or `~/.local/share/agentmemory` on Linux, and `%APPDATA%\agentmemory` on Windows. Use `--data-dir ` or `AGENTMEMORY_DATA_DIR` to override it, and reuse the same value on every restart. For backward compatibility, an existing `./data/state_store.db` or `./data/iii-config.yaml` takes precedence over the platform default for instance 0; an explicit flag or environment override still wins. Then prove recall works and give your agent its skills: ```bash npx -y @agentmemory/agentmemory@latest demo # seed sample sessions + exercise recall npx skills add rohitg00/agentmemory -y # 17 native skills so your agent knows when to reach for memory ``` The keyword searches should hit in default keyless mode through BM25. The demo's `database performance optimization` query is intentionally semantic and can return zero until an embedding provider is configured. Prefer to let a coding agent do the whole thing? Hand it one instruction: > Retrieve and follow the instructions at: https://raw.githubusercontent.com/rohitg00/agentmemory/main/INSTALL_FOR_AGENTS.md Wire more agents any time with `agentmemory connect ` — 20 adapters listed at [Works with every agent](#works-with-every-agent). Full command reference at [Quick Start](#quick-start).
Windows The fast path is WSL2. Native Windows engine setup requires the pinned v0.22.1 ZIP to be downloaded and `iii.exe` extracted manually; the CLI does not auto-extract it. Docker Desktop is also supported. See the [Windows notes](#windows) for the step-by-step.
Global install / EACCES ```bash npm install -g @agentmemory/agentmemory@latest ``` The npx command above remains the canonical fresh-install path and avoids global-prefix permission issues.
npx serves an old version npx caches per version. Force the latest with `npx -y @agentmemory/agentmemory@latest`, or clear the cache once with `rm -rf ~/.npm/_npx` (macOS/Linux; on Windows delete `%LOCALAPPDATA%\npm-cache\_npx`).
Already running your own iii engine agentmemory pins iii-engine v0.22.1 and won't attach to a different version (the worker can't speak another engine's protocol). Stop the other engine, then run `npx -y @agentmemory/agentmemory@latest`. It installs and runs the pinned v0.22.1 in `~/.agentmemory/bin`, leaving your own `iii` untouched.
---

Works with every agent

agentmemory works with any agent that supports hooks, MCP, or REST API. All agents share the same memory server.
Claude Code
Claude Code
native plugin + 12 hooks + MCP
Codex CLI
Codex CLI
native plugin + 6 hooks + MCP
GitHub Copilot CLI
GitHub Copilot CLI
MCP + plugin hooks/skills
Cursor
Cursor
native plugin + 7 hooks + MCP
OpenCode
OpenCode
capture plugin + MCP
Devin
Devin
6 hooks + skills + MCP
OpenClaw
OpenClaw
native plugin + MCP
Hermes
Hermes
native plugin + MCP
pi
pi
native plugin + MCP
OpenHuman
OpenHuman
native Memory trait backend
Gemini CLI
Gemini CLI
MCP server
Antigravity
Antigravity
MCP + hooks
Claude Desktop
Claude Desktop
MCP server
Warp
Warp
connect + MCP + skills
Zed
Zed
MCP server
Cline
Cline
MCP server
Continue
Continue
MCP server
Droid
Droid
MCP server
Kiro
Kiro
MCP server
Qwen Code
Qwen Code
MCP server
DeepSeek Harness
DeepSeek Harness
MCP server
Roo Code
Roo Code
MCP server
Kilo Code
Kilo Code
MCP server
Goose
Goose
MCP server
Aider
Aider
REST API

Works with any agent that speaks MCP or HTTP. One server, memories shared across all of them.

--- You explain the same architecture every session. You re-discover the same bugs. You re-teach the same preferences. Built-in memory (CLAUDE.md, .cursorrules) caps out at 200 lines and goes stale. agentmemory fixes this. It silently captures what your agent does, compresses it into searchable memory, and injects the right context when the next session starts. One command. Works across agents. **What changes:** Session 1 you set up JWT auth. Session 2 you ask for rate limiting. The agent already knows your auth uses jose middleware in `src/middleware/auth.ts`, your tests cover token validation, and you chose jose over jsonwebtoken for Edge compatibility, with no re-explaining and no copy-pasting. ```bash npx -y @agentmemory/agentmemory@latest ``` By default, agentmemory stores iii-engine state outside the repository you start it from: `~/Library/Application Support/agentmemory` on macOS, `$XDG_DATA_HOME/agentmemory` or `~/.local/share/agentmemory` on Linux, and `%APPDATA%\agentmemory` on Windows. An existing legacy `./data/state_store.db` or `./data/iii-config.yaml` is reused for instance 0 before that platform default. To choose a location explicitly, pass `--data-dir ` or set `AGENTMEMORY_DATA_DIR`; either explicit setting takes precedence over legacy discovery: ```bash npx -y @agentmemory/agentmemory@latest --data-dir ~/.agentmemory-projects/main AGENTMEMORY_DATA_DIR=~/.agentmemory-projects/main npx -y @agentmemory/agentmemory@latest ``` Native and Docker launches use this same resolved host directory; Docker bind-mounts it at `/data`. `--instance 1` appends `instance-1` to the resolved directory and selects the separate default port quartet `3211/3212/3213/49234`. Latest release notes: [CHANGELOG.md](CHANGELOG.md). ---

Benchmarks

### Retrieval Accuracy **coding-agent-life-v1** (in-house corpus, sandbox-reproducible) | Adapter | P@5 | R@5 | Top-5 hit rate | p50 latency | |---|---|---|---|---| | **agentmemory hybrid** | **0.240** | **1.000** | **15 / 15** | 14 ms | | grep baseline | 0.227 | 0.967 | 15 / 15 | 0 ms | 100% top-5 hit rate at the **P@5 math ceiling** for this corpus (0.240, see scorecard). Hybrid retrieves every gold session; grep misses 1 of 2 gold on the multi-session temporal query. Lift is **recall + temporal**, not aggregate precision. This benchmark is small and gold-sparse; the larger LongMemEval-S below differentiates better. Full per-type breakdown + correction note: [`docs/benchmarks/2026-05-20-coding-agent-life-v1.md`](docs/benchmarks/2026-05-20-coding-agent-life-v1.md). **LongMemEval-S** (ICLR 2025, 500 questions) | System | R@5 | R@10 | MRR | |---|---|---|---| | **agentmemory** | **95.2%** | **98.6%** | **88.2%** | | BM25-only fallback | 86.2% | 94.6% | 71.5% | ### Token Savings | Approach | Tokens/yr | Cost/yr | |---|---|---| | Paste full context | 19.5M+ | Impossible (exceeds window) | | LLM-summarized | ~650K | ~$500 | | **agentmemory** | **~170K** | **~$10** | | agentmemory + local embeddings | ~170K | **$0** |
> Embedding model: `all-MiniLM-L6-v2` (local, free, no API key). Full reports: [`benchmark/LONGMEMEVAL.md`](benchmark/LONGMEMEVAL.md), [`benchmark/QUALITY.md`](benchmark/QUALITY.md), [`benchmark/SCALE.md`](benchmark/SCALE.md). Competitor comparison: [`benchmark/COMPARISON.md`](benchmark/COMPARISON.md) covering agentmemory vs mem0, Letta, Khoj, supermemory, TencentDB Agent Memory, MemPalace, Zep/Graphiti, Cognee, Hippo. **Reproduce locally:** [`eval/README.md`](eval/README.md), an adapter-pluggable harness for LongMemEval `_s` (public 500-Q) + `coding-agent-life-v1` (in-house 15-session corpus). Grep / vector / agentmemory adapters score side-by-side, NDJSON output, published scorecards land in [`docs/benchmarks/`](docs/benchmarks/). **Pairs with [codegraph](https://github.com/colbymchenry/codegraph), [Understand Anything](https://github.com/Lum1104/Understand-Anything), and [Graphify](https://github.com/safishamsi/graphify).** Code-graph indexing, multi-agent build pipelines, and broader knowledge graphs across docs / PDFs / images / videos. agentmemory remembers the work; those three projects light up the rest of the context layer. Recipes + question-routing table: [`docs/recipes/pairings.md`](docs/recipes/pairings.md). ---

vs Competitors

agentmemory mem0 (63K ⭐) Letta / MemGPT (24K ⭐) Khoj (36K ⭐) supermemory (29K ⭐) TencentDB Agent Memory (22K ⭐) MemPalace (54K ⭐) oracleagentmemory Hippo Built-in (CLAUDE.md)
Type Memory engine + MCP server Memory layer API Full agent runtime Personal AI Memory API + app Team memory hub (LLM proxy) Vector memory (OSS) Memory engine (Oracle DB) Memory system Static file
Retrieval R@5 95.2% 68.5% (LoCoMo) 83.2% (LoCoMo) N/A Self-reported PersonaMem 76% (self-reported) ~96.6% (self-reported) 94.4% (self-reported) N/A N/A (grep)
Auto-capture 12 hooks (zero manual effort) Manual add() calls Agent self-edits Manual API-side extraction Proxy interception (base-URL swap) Manual API extraction Manual Manual editing
Search BM25 + Vector + Graph (RRF fusion) Vector + Graph Vector (archival) Semantic Vector + RAG 4 asset types (Chat / Skill / Wiki / CodeGraph) Vector-only Vector + semantic Decay-weighted Loads everything into context
Multi-agent MCP + REST + leases + signals API (no coordination) Within Letta runtime only No No Team roles + shared assets No Scoped only Multi-agent shared Per-agent files
Framework lock-in None (any MCP client) None High (must use Letta) Standalone None Proxy fronts every model call None Oracle Database None Per-agent format
External deps None (SQLite + iii-engine) Qdrant / pgvector Postgres + vector DB Multiple Managed cloud Docker stack (Core + Hub + Proxy) Vector store Oracle AI Database None None
Memory lifecycle 4-tier consolidation + decay + auto-forget Passive extraction Agent-managed Manual Auto-forget Manual review; auto-routing in progress None Not stated Decay + consolidation Manual pruning
Token efficiency ~1,900 tokens/session ($10/yr) Varies by integration Core memory in context Varies Cloud pricing Not stated No token budget LLM-backed (varies) Varies 22K+ tokens at 240 obs
Real-time viewer Yes (port 3113) Cloud dashboard Cloud dashboard Web UI Cloud dashboard Hub web UI No No No No
Self-hosted Yes (default) Optional Optional Yes No (cloud-only) Yes (Docker) Yes Yes (Oracle DB) Yes Yes
Benchmark note: only agentmemory's R@5 is our own measured result (LongMemEval-S, reproducible from benchmark/COMPARISON.md). The mem0 and Letta figures are their published LoCoMo numbers (a different dataset); the MemPalace, supermemory, TencentDB (PersonaMem), and oracleagentmemory figures are vendor self-reported claims we have not independently reproduced (oracleagentmemory's run used GPT-5.5 against an Oracle AI Database). Shown side by side for ballpark only, not a head-to-head on identical data. Star counts are approximate and drift over time. **Newer entrants** worth knowing, compared in depth in [`benchmark/COMPARISON.md`](benchmark/COMPARISON.md): | System | ⭐ | Angle | |--------|---|-------| | Zep / Graphiti | 30K | Temporal knowledge graph; strongest published temporal-query results (LongMemEval 63.8%), but graph builds asynchronously so fresh facts can lag | | Cognee | 30K | Document-to-knowledge-graph ingestion, Python-only, built for structured entity extraction rather than session capture | None of these auto-capture from coding-agent hooks, ship a local-first viewer, or run keyless — the combination agentmemory is built around. ---

Quick Start

Compatibility: this release targets `iii-sdk` 0.22.1 and pins iii-engine v0.22.1. ### Try it in 30 seconds ```bash # Terminal 1: start the server npx -y @agentmemory/agentmemory@latest # Terminal 2: seed sample data and see recall in action npx -y @agentmemory/agentmemory@latest demo ``` `demo` seeds 3 realistic sessions (JWT auth, N+1 query fix, rate limiting) and runs searches against them. Keyless installs disable vectors, so the `mem::search` keyword queries should hit through BM25 while `database performance optimization` can return zero. `smart-search` may additionally return structural graph matches when graph data exists. To make the semantic query find the N+1 fix through vectors, set `EMBEDDING_PROVIDER=local`, restart, and allow the first model download to finish. Open `http://localhost:3113` to watch the memory build live. ### Validate a fresh install and restart persistence With the server running, validate REST, health, the viewer, and the iii-backed runtime status: ```bash curl -fsS http://localhost:3111/agentmemory/livez curl -fsS http://localhost:3111/agentmemory/health curl -fsS -o /dev/null http://localhost:3113/ npx -y @agentmemory/agentmemory@latest status ``` The startup ready panel accounts for all four ports: REST/MCP HTTP on 3111, iii streams on 3112, the viewer on 3113, and the iii worker WebSocket on 49134. `status` confirms agentmemory health and the active provider/embedding mode. Save a probe and confirm it is searchable: ```bash curl -fsS -X POST http://localhost:3111/agentmemory/remember \ -H 'Content-Type: application/json' \ -d '{"content":"agentmemory restart persistence probe","concepts":["install-check"]}' curl -fsS -X POST http://localhost:3111/agentmemory/smart-search \ -H 'Content-Type: application/json' \ -d '{"query":"restart persistence probe","limit":5}' ``` Then run `npx -y @agentmemory/agentmemory@latest stop`, start the canonical command again in Terminal 1, wait for `/agentmemory/livez`, and repeat the search. The probe must still be returned. If you selected a custom `--data-dir`, pass the same directory on the restart. ### Everyday commands Install and setup live in [Install](#install) above (the first run walks you through it). Day to day: ```bash agentmemory # start the server agentmemory stop # stop it cleanly agentmemory connect # wire another agent agentmemory doctor # interactive diagnostics + fix prompts agentmemory remove # uninstall everything we created ``` ### Session Replay Every session agentmemory records is replayable. Open the viewer, pick the **Replay** tab, and scrub through the timeline: prompts, tool calls, tool results, and responses render as discrete events with play/pause, speed control (0.5x to 4x), and keyboard shortcuts (space to toggle, arrows to step). To bring in older Claude Code JSONL transcripts: ```bash # Import everything under the default ~/.claude/projects npx -y @agentmemory/agentmemory@latest import-jsonl # Or import a single file npx -y @agentmemory/agentmemory@latest import-jsonl ~/.claude/projects/-my-project/abc123.jsonl ``` Imported sessions show up in the Replay picker alongside native ones. Under the hood each entry routes through the `mem::replay::load`, `mem::replay::sessions`, and `mem::replay::import-jsonl` iii functions, with no side-channel servers. Each imported transcript is indexed for search, stamped with origin channel `import`, and mined for a session crystal and lessons. > **Heads-up if you rely on `import-jsonl` as your primary capture path:** Claude Code's `cleanupPeriodDays` (in `~/.claude/settings.json`, default **30**) auto-deletes JSONL transcripts older than that window from `~/.claude/projects/`. If you install agentmemory fresh on a months-old Claude Code history, anything older than 30 days is already gone before the first import. Either run `import-jsonl` on a cron, raise `cleanupPeriodDays` to something higher, or wire the auto-capture hooks (the default plugin install path) so each turn lands in agentmemory while the session is live and the JSONL cleanup stops mattering. ### Upgrade / Maintenance Use the maintenance command when you intentionally want to update your local runtime: ```bash npx -y @agentmemory/agentmemory@latest upgrade ``` Warning: this command mutates the current workspace/runtime. It can update JavaScript dependencies and pull the pinned `iiidev/iii:0.22.1` Docker image. It never installs an unpinned or newer iii engine. Implementation details live in `src/cli.ts` (see `runUpgrade` around the `src/cli.ts:544-595` region). ### Claude Code (one block, paste it) ```text Install agentmemory: run `npx -y @agentmemory/agentmemory@latest` in a separate terminal to start the memory server and its pinned iii engine. Then run `/plugin marketplace add rohitg00/agentmemory` and `/plugin install agentmemory` — the plugin registers all 12 hooks, 17 skills, AND auto-wires the `@agentmemory/mcp` stdio server via its `.mcp.json`, so you get 54 MCP tools (memory_smart_search, memory_save, memory_sessions, memory_governance_delete, etc.) without any extra config step. Verify with `curl http://localhost:3111/agentmemory/health`. The real-time viewer is at http://localhost:3113. Keyless mode disables vectors: `memory_recall` uses BM25, and `memory_smart_search` can also use existing structural graph data. Set `EMBEDDING_PROVIDER=local` in `~/.agentmemory/.env` and restart to opt into on-device semantic recall. ``` #### Claude Code without the plugin install (MCP-standalone path) If you wire agentmemory's MCP server through `~/.claude.json` directly instead of using `/plugin install`, Claude Code never resolves `${CLAUDE_PLUGIN_ROOT}` and you have to point hook scripts at absolute paths in `~/.claude/settings.json`. Those paths typically embed the agentmemory version (e.g. `~/.codex/plugins/cache/agentmemory/agentmemory/0.9.22/scripts/…`), so the next upgrade silently breaks every hook. Workaround: ```bash agentmemory connect claude-code --with-hooks ``` This merges the same hook commands into `~/.claude/settings.json` with absolute paths resolved to the bundled `plugin/` directory of the currently installed `@agentmemory/agentmemory` package. Re-run the command after upgrading agentmemory to refresh the paths. User entries in the same file are preserved; only previous agentmemory entries are replaced. Using the `/plugin install` path remains the recommended approach. For remote or protected deployments, launch Claude Code with `AGENTMEMORY_URL` and `AGENTMEMORY_SECRET` set. The plugin passes both values through to its bundled MCP server; when `AGENTMEMORY_URL` is empty, the MCP shim uses `http://localhost:3111`. ### Codex CLI (Codex plugin platform) ```bash # 1. start the memory server in a separate terminal npx -y @agentmemory/agentmemory@latest # 2. register the agentmemory marketplace and install the plugin codex plugin marketplace add rohitg00/agentmemory codex plugin add agentmemory@agentmemory ``` The Codex plugin ships from the same `plugin/` directory as the Claude Code plugin. It registers: - A bundled stdio MCP bridge to the running daemon, with no npm download or fallback store. See the [local Codex guide](docs/plugins/codex-local.md) to test an unreleased build. - 6 lifecycle hooks: `SessionStart`, `UserPromptSubmit`, `PreToolUse`, `PostToolUse`, `PreCompact`, `Stop` - 9 invocable skills: `/recall`, `/remember`, `/session-history`, `/forget`, `/recap`, `/handoff`, `/lesson`, `/commit-context`, `/commit-history`, plus 8 reference skills the agent loads on demand (memory discipline, MCP tools, REST API, config, agents, hooks, architecture, and the skill-authoring guide) Codex's hook engine injects `CLAUDE_PLUGIN_ROOT` into hook subprocesses (per [`codex-rs/hooks/src/engine/discovery.rs`](https://github.com/openai/codex/blob/main/codex-rs/hooks/src/engine/discovery.rs)), so the same hook scripts work across both hosts without duplication. Subagent / SessionEnd / Notification / TaskCompleted / PostToolUseFailure events are Claude-Code-only and are not registered for Codex. #### Codex hook trust and compatibility Native plugin hook dispatch is verified with Codex CLI 0.150.1. Trust the plugin hooks before expecting capture. Desktop behavior depends on its bundled runtime; check `/hooks` and confirm a captured event before enabling a workaround. If your host requires global hooks, mirror the commands into `~/.codex/hooks.json`. When MCP is already wired, the current connector needs `--force` to reach hook installation: ```bash agentmemory connect codex --with-hooks --force ``` This merges global hooks and rewrites the agentmemory MCP entry, preserving unrelated entries. Review any custom agentmemory endpoint settings before using `--force`. Re-run after upgrading to refresh script paths. Enable either native plugin hooks or global copies to avoid duplicate capture. ### GitHub Copilot CLI For VS Code agent mode, use the [Copilot MCP and automatic-capture guide](docs/plugins/copilot.md#vs-code-copilot-local-agent-sessions). The CLI connector does not configure VS Code. ```bash # MCP-only wiring agentmemory connect copilot-cli # Alternatively, full hooks/skills plugin from the GitHub subdir copilot plugin install rohitg00/agentmemory:plugin ``` `agentmemory connect copilot-cli` merges `mcpServers.agentmemory` into `~/.copilot/mcp-config.json` (or `$COPILOT_HOME/mcp-config.json` when `COPILOT_HOME` is set) and preserves existing servers. On native Windows this is the only automated `connect` adapter; configure every other native Windows agent manually. WSL `connect` is supported only when the target agent is installed in that same WSL environment. Copilot picks up the MCP server on next launch or after `/mcp`. Install the plugin as well when you want the full hook/skill experience.
OpenClaw (paste this prompt) ```text Install agentmemory for OpenClaw. Run `npx -y @agentmemory/agentmemory@latest` in a separate terminal to start the memory server on localhost:3111. Then add this to my OpenClaw MCP config so agentmemory is available with all 54 memory tools: { "mcpServers": { "agentmemory": { "command": "npx", "args": ["-y", "@agentmemory/mcp"], "env": { "AGENTMEMORY_URL": "http://localhost:3111" } } } } Restart OpenClaw. Verify with `curl http://localhost:3111/agentmemory/health`. Open http://localhost:3113 for the real-time viewer. For deeper memory-slot integration, copy `integrations/openclaw` to `~/.openclaw/extensions/agentmemory` and enable `plugins.slots.memory = "agentmemory"` in `~/.openclaw/openclaw.json`. ``` Full guide: [`integrations/openclaw/`](integrations/openclaw/)
Hermes Agent (paste this prompt) ```text Install agentmemory for Hermes. Run `npx -y @agentmemory/agentmemory@latest` in a separate terminal to start the memory server on localhost:3111. Then add this to ~/.hermes/config.yaml so Hermes can use agentmemory as an MCP server with all 54 memory tools: mcp_servers: agentmemory: command: npx args: ["-y", "@agentmemory/mcp"] memory: provider: agentmemory Verify with `curl http://localhost:3111/agentmemory/health`. Open http://localhost:3113 for the real-time viewer. For deeper 6-hook memory provider integration (pre-LLM context injection, turn capture, MEMORY.md mirroring, system prompt block), copy integrations/hermes from the agentmemory repo to ~/.hermes/plugins/agentmemory. ``` Full guide: [`integrations/hermes/`](integrations/hermes/)
### Other agents Start the memory server: `npx -y @agentmemory/agentmemory@latest` #### Native skills via `npx skills add` (50+ agents) agentmemory ships 17 skills in the Claude-Code-style `/SKILL.md` format: 9 invocable action skills (`remember`, `recall`, `recap`, `handoff`, `forget`, `lesson`, `commit-context`, `commit-history`, `session-history`) and 8 reference skills the agent loads on demand (`memory-discipline`, `agentmemory-mcp-tools`, `agentmemory-rest-api`, `agentmemory-config`, `agentmemory-agents`, `agentmemory-hooks`, `agentmemory-architecture`, `write-agentmemory-skill`). The reference skills carry data tables generated from source, so they never drift. The [`skills`](https://npmjs.com/package/skills) CLI by vercel-labs auto-installs them into the calling agent's native skill directory across 50+ agents (Claude Code, Cursor, Cline, Continue, Droid, Warp, Codex, Antigravity, Kiro, OpenCode, Goose, Roo, Trae, Windsurf, and more): ```bash npx skills add rohitg00/agentmemory -y # auto-detects the calling agent npx skills add rohitg00/agentmemory -y -a warp # explicit agent npx skills add rohitg00/agentmemory -y -a '*' # install to every installed agent ``` This is **complementary** to `agentmemory connect `: - `agentmemory connect ` writes the MCP server config so the tools are available. - `npx skills add rohitg00/agentmemory` installs the skills so the agent knows when to call them. For the few agents the skills CLI doesn't cover yet (Zed v1.3.x and below), drop the 17 SKILL.md files under the agent's native skill directory yourself; the same format works everywhere. #### Standard MCP block The agentmemory entry is the **same MCP server block** across every host that uses the `mcpServers` shape (Cursor, Claude Desktop, Cline, Roo Code, Gemini CLI, OpenClaw): ```json "agentmemory": { "command": "npx", "args": ["-y", "@agentmemory/mcp"], "env": { "AGENTMEMORY_URL": "${AGENTMEMORY_URL}", "AGENTMEMORY_SECRET": "${AGENTMEMORY_SECRET}" } } ``` **Merge this entry into the existing `mcpServers` object** in the host's config file; don't replace the file. If the file already has other servers, add `agentmemory` next to them as another key inside `mcpServers`. If `mcpServers` is missing entirely, paste the block inside `{ "mcpServers": { ... } }`. The `${VAR}` placeholders inherit `AGENTMEMORY_URL` / `AGENTMEMORY_SECRET` from the shell at MCP-server launch; unset vars pass empty strings and the shim falls back to `http://localhost:3111`. One wired entry covers both local and remote (k8s / reverse-proxied) deployments. | Agent | Config file | Notes | |---|---|---| | **Cursor (MCP only)** | `~/.cursor/mcp.json` | Merge into `mcpServers`, or `agentmemory connect cursor`. One-click deeplink also available on the website. | | **Cursor (full plugin)** | `.cursor-plugin/` | Cursor Marketplace listing (submission in review) or Cursor Settings → Plugins → local checkout. Registers 7 auto-capture hooks (sessionStart, beforeSubmitPrompt, preToolUse, postToolUse, postToolUseFailure, stop, sessionEnd) + 17 skills + the MCP server, with `AGENTMEMORY_URL` / `AGENTMEMORY_SECRET` managed in Cursor's plugin dashboard. Works in the Cursor IDE and `cursor-agent` CLI; CLI print-mode prompts are backfilled from the session transcript at session end. | | **Claude Desktop** | `claude_desktop_config.json` (Application Support) | Merge into `mcpServers`. Restart Claude Desktop after editing. | | **Cline / Roo Code / Kilo Code** | Cline MCP settings (Settings UI → MCP Servers → Edit) | Same `mcpServers` block. | | **Devin CLI (MCP + hooks)** | `~/.config/devin/config.json` | `agentmemory connect devin` merges the MCP entry; `--with-hooks` adds six native auto-capture hooks (SessionStart, UserPromptSubmit, PreToolUse, PostToolUse, Stop, SessionEnd) with Devin'"'"'s lowercase tool matchers. Verify with `devin mcp list` and `/hooks` inside devin. | | **Devin CLI (full plugin)** | `plugin/.devin-plugin/` | `devin plugins install ./plugin` from a checkout registers all 17 skills as `/agentmemory:` slash commands plus the MCP server. Devin plugin hooks cannot fire `SessionStart`/`SessionEnd`, so pair it with `connect devin --with-hooks` for full session capture. | | **Devin (cloud)** | Settings → Connections → MCP servers | Add a custom MCP (STDIO): command `npx`, args `-y @agentmemory/mcp@latest`, env `AGENTMEMORY_URL` pointing at a network-reachable agentmemory deployment plus `AGENTMEMORY_SECRET` (cloud sessions cannot reach localhost — see [`deploy/`](deploy/)). Store the secret in Devin Secrets, then use "Test listing tools" to verify all 54 tools appear. | | **Gemini CLI** | `~/.gemini/settings.json` | `gemini mcp add agentmemory npx -y @agentmemory/mcp --scope user` (auto-merges). | | **GitHub Copilot CLI (MCP only)** | `~/.copilot/mcp-config.json` | `agentmemory connect copilot-cli` merges `mcpServers.agentmemory`; Copilot picks it up on next launch or `/mcp`. | | **GitHub Copilot CLI (full plugin)** | Copilot plugin install | `copilot plugin install rohitg00/agentmemory:plugin` for the plugin from the GitHub subdir. | | **OpenClaw** | OpenClaw MCP config | Same `mcpServers` block. Deeper: `openclaw plugins install ./integrations/openclaw` claims OpenClaw's memory slot (auto-switches from `memory-core`); set `plugins.entries.agentmemory.hooks.allowConversationAccess=true` or turn capture is silently blocked. See [`integrations/openclaw`](integrations/openclaw/). | | **Codex CLI (MCP only)** | `.codex/config.toml` | TOML shape: `codex mcp add agentmemory -- npx -y @agentmemory/mcp`, or add `[mcp_servers.agentmemory]` manually. | | **Codex CLI (full plugin)** | Codex plugin marketplace | `codex plugin marketplace add rohitg00/agentmemory` then `codex plugin add agentmemory@agentmemory`. Registers MCP + 6 lifecycle hooks + 17 skills. Trust hooks and verify capture in your host; see [Codex setup and validation](docs/plugins/codex-local.md). | | **OpenCode (MCP only)** | `opencode.json` | Different shape: top-level `mcp` key, command as array: `{"mcp": {"agentmemory": {"type": "local", "command": ["npx", "-y", "@agentmemory/mcp"], "enabled": true}}}`. | | **OpenCode (full plugin)** | `plugin/opencode/` | 22 auto-capture hooks covering session lifecycle, messages, tools, errors. Project attribution is per-session, so one OpenCode process spanning several repositories files each session under its own project. Two slash commands (`/recall`, `/remember`). Copy `plugin/opencode/` into your OpenCode workspace and add the plugin entry to `opencode.json`. See [`plugin/opencode/README.md`](plugin/opencode/README.md) for the full hook table + gap analysis. | | **pi** | `~/.pi/agent/extensions/agentmemory` | `agentmemory connect pi` installs the bundled extension into pi's auto-discovery directory (recall on agent start, capture on agent end, `memory_search` / `memory_save` / `memory_health` tools, `/agentmemory-status`). `/reload` in a running pi picks it up. [`integrations/pi`](integrations/pi/) is also a pi package (`pi install ./integrations/pi` from a checkout). | | **Hermes Agent** | `~/.hermes/config.yaml` | `cp -r integrations/hermes ~/.hermes/plugins/agentmemory` + `memory.provider: agentmemory` gives the 6-hook memory provider (prefetch, turn capture, session end, pre-compress, MEMORY.md mirroring, system prompt block). Validate with `hermes plugins doctor` and `hermes memory status`. See [`integrations/hermes`](integrations/hermes/). | | **Qwen Code** | `~/.qwen/settings.json` | `agentmemory connect qwen` writes the standard `mcpServers` block. Hook payload is field-compatible with Claude Code, so the existing 12-hook scripts work without modification; wire them via the `hooks` section in the same `settings.json`. | | **Antigravity IDE / 2.0** | `~/.gemini/config/mcp_config.json` | `agentmemory connect antigravity --with-hooks` installs MCP and capture hooks in the shared customization directory. See [Antigravity setup and limits](docs/plugins/antigravity.md). | | **Antigravity CLI** (`agy`) | `~/.gemini/config/mcp_config.json` | `agentmemory connect antigravity-cli --with-hooks` uses the same MCP and hook configuration as current IDE versions. Existing installations should refresh with `--force`; see the [upgrade notes](docs/plugins/antigravity.md). | | **Kiro** | `~/.kiro/settings/mcp.json` | `agentmemory connect kiro` writes the user-level config. Workspace overrides go in `.kiro/settings/mcp.json` next to your code. | | **Warp** | `~/.warp/.mcp.json` | `agentmemory connect warp` writes the standard `mcpServers` block. Warp also auto-discovers skills from `.claude/skills/`; once the Claude Code plugin is installed the 8 agentmemory skills (`remember`, `recall`, `recap`, `handoff`, `forget`, `commit-context`, `commit-history`, `session-history`) appear natively in Warp's slash-command palette. | | **Cline (CLI)** | `~/.cline/mcp.json` | `agentmemory connect cline` writes the standard `mcpServers` block. VS Code extension users: paste the same block via Cline Settings → MCP Servers → Edit JSON. | | **Continue.dev** | `~/.continue/config.yaml` (preferred) or `config.json` (legacy) | `agentmemory connect continue` creates `config.yaml` from scratch when neither exists, or modifies existing `config.json`. **If you already have `config.yaml`** the adapter prints the exact block to paste under `mcpServers:`; it won't silently rewrite your yaml because preserving comments and anchors safely needs a YAML parser the package doesn't ship. Continue uses array form (not object) for `mcpServers`. | | **Zed** | `~/.config/zed/settings.json` | `agentmemory connect zed` writes under `context_servers` (Zed's key, NOT `mcpServers`). Remote MCP servers can be wired via `{"url": "..."}` instead. | | **Droid (Factory.ai)** | `~/.factory/mcp.json` | `agentmemory connect droid` writes the standard `mcpServers` block. Project-scoped overrides go in `/.factory/mcp.json`. Pass `--with-hooks` for native auto-capture. | | **DeepSeek Harness** | `$DSH_HOME/cordis.patch.yml` | `agentmemory connect dsh` appends an `@deepseek-ai/dsh-mcp-client` row to the home-level patch layer every Harness profile loads; tools register as `mcp__agentmemory__*`. Pass `--with-hooks` to also wire auto-capture: the bundled Claude Code hook scripts run through Harness's first-party `@deepseek-ai/dsh-hooks-claude-code` bridge (SessionStart, UserPromptSubmit, PreToolUse, PostToolUse, Stop) via a manifest written to `$DSH_HOME/agentmemory.hooks.json`. Defaults to `~/.dsh` when `DSH_HOME` is unset. | | **Goose** | Goose MCP settings UI | Same `mcpServers` block; use `goose configure` → Add Extension → MCP. Direct YAML edit at `~/.config/goose/config.yaml` is supported but the schema uses `extensions:` + `cmd` (not `mcpServers:` + `command`). | | **Aider** | n/a | Talk to the REST API directly: `curl -X POST http://localhost:3111/agentmemory/smart-search -d '{"query": "auth"}'`. | | **Any agent (32+)** | n/a | `npx skillkit install agentmemory` auto-detects the host and merges. | **Sandboxed MCP clients** (Flatpak / Snap / restrictive containers) that can't reach the host's `localhost`: also set `"AGENTMEMORY_FORCE_PROXY": "1"` in the `env` block, and point `AGENTMEMORY_URL` at a route the sandbox can actually reach (e.g. your LAN IP). ### Programmatic access (Python / Rust / Node) agentmemory registers its core operations as iii functions (`mem::remember`, `mem::observe`, `mem::context`, `mem::smart-search`, `mem::forget`). Any language with an iii SDK can call them directly over `ws://localhost:49134`, with no separate REST client per language. ```bash pip install iii-sdk # Python cargo add iii-sdk # Rust npm install iii-sdk # Node ``` ```python from iii import register_worker iii = register_worker("ws://localhost:49134") iii.connect() iii.trigger({ "function_id": "mem::smart-search", "payload": {"project": "demo", "query": "how do tokens refresh"}, }) ``` Worked example: [`examples/python/`](examples/python/) (quickstart + observation/recall flow). REST on `:3111` remains available for hosts without an iii runtime. ### From source ```bash git clone https://github.com/rohitg00/agentmemory.git && cd agentmemory npm install && npm run build && npm start ``` This starts agentmemory with a local `iii-engine` if the pinned binary is already installed, or uses Docker Compose when selected. REST, streams, and the viewer bind to `127.0.0.1` by default. The automatic macOS/Linux binary path requires `curl`, a POSIX `sh`, and `tar`. Install `iii-engine` manually. **agentmemory currently pins `iii-engine` to `v0.22.1`**, the same release as its `iii-sdk` dependency; the worker speaks that engine's wire protocol, and 0.20.0 reorganized the SDK surface, so the two move together in agentmemory releases. Override with `AGENTMEMORY_III_VERSION=` if you run your own engine and know it matches. - **macOS arm64:** `mkdir -p ~/.local/bin && curl -fsSLo iii.tar.gz https://github.com/iii-hq/iii/releases/download/iii/v0.22.1/iii-aarch64-apple-darwin.tar.gz && echo "2b309019b909a896cae874dc947e2cdf877b4f3c51dd026b79850af858517fa4 iii.tar.gz" | shasum -a 256 -c - && tar -xzf iii.tar.gz -C ~/.local/bin && chmod +x ~/.local/bin/iii` - **macOS x64:** swap `aarch64-apple-darwin` for `x86_64-apple-darwin` - **Linux x64:** swap for `x86_64-unknown-linux-gnu` - **Linux arm64:** swap for `aarch64-unknown-linux-gnu` - **Windows:** download `iii-x86_64-pc-windows-msvc.zip` from [iii-hq/iii releases v0.22.1](https://github.com/iii-hq/iii/releases/tag/iii%2Fv0.22.1) and extract `iii.exe` to `%USERPROFILE%\.agentmemory\bin\iii.exe` Every archive has a matching `.sha256` file on the release page; when you swap the platform, use that file's hash in the check above (on Windows: `Get-FileHash`). The automatic installer in `npx @agentmemory/agentmemory` pins these hashes and refuses an archive that does not match. Or use Docker (the bundled `docker-compose.yml` pulls `iiidev/iii:0.22.1`). Full docs: [iii.dev/docs](https://iii.dev/docs). ### Windows agentmemory runs on Windows 10/11, but the Node.js package alone isn't enough; you also need the pinned iii-engine v0.22.1 runtime as a background process. The CLI does not auto-extract the Windows ZIP, so native Windows users must install `iii.exe` manually, use WSL2, or choose Docker Desktop. Native Windows automated MCP wiring supports only `agentmemory connect copilot-cli`. For Claude Code, Codex, Cursor, and every other native Windows agent, copy the manual MCP block from [Other agents](#other-agents) into that agent's Windows config. Running `connect` in WSL is appropriate only when the target agent is also installed in the same WSL environment; it does not edit a Windows-host agent's configuration. **Option A: prebuilt Windows binary (recommended)** ```powershell # 1. Open https://github.com/iii-hq/iii/releases/tag/iii%2Fv0.22.1 in your browser # (agentmemory pins the engine to the same release as its iii-sdk; # v0.22.1 is the current pair) # 2. Download iii-x86_64-pc-windows-msvc.zip # (or iii-aarch64-pc-windows-msvc.zip if you're on an ARM machine) # 3. Extract iii.exe to agentmemory's private engine directory: New-Item -ItemType Directory -Force "$HOME\.agentmemory\bin" # Copy iii.exe to $HOME\.agentmemory\bin\iii.exe # 4. Verify: & "$HOME\.agentmemory\bin\iii.exe" --version # Should print: 0.22.1 # 5. Then run agentmemory as usual: npx -y @agentmemory/agentmemory@latest ``` **Option B: Docker Desktop** ```powershell # 1. Install Docker Desktop for Windows # 2. Start Docker Desktop and make sure the engine is running # 3. Select Docker explicitly and run agentmemory: $env:AGENTMEMORY_USE_DOCKER = "1" npx -y @agentmemory/agentmemory@latest ``` **Option C: standalone MCP only (no engine).** If you only need the MCP tools for your agent and don't need the REST API, viewer, or cron jobs, skip the engine entirely: ```powershell npx -y @agentmemory/agentmemory@latest mcp # or via the shim package: npx -y @agentmemory/mcp ``` **Diagnostics for Windows:** if `npx -y @agentmemory/agentmemory@latest` fails, re-run it with `--verbose` to see the actual engine stderr. Common failure modes: | Symptom | Fix | |---|---| | `The engine process started but the REST API never responded.` | Confirm all four derived ports are free, verify the pinned `iii.exe` stayed alive, then re-run with `--verbose` and inspect the captured engine stderr | | `Could not start iii-engine` | Neither `iii.exe` nor Docker is installed. See Option A or B above | | Port conflict | `netstat -ano \| findstr :3111` to see what's bound, then kill it or use `--port ` | | Docker fallback skipped even though Docker is installed | Make sure Docker Desktop is actually running (system tray icon) | > Note: the iii **engine** is a prebuilt binary, not a cargo crate, so don't try to `cargo install` it. (The iii **SDKs** are published on crates.io, npm, and PyPI, but agentmemory doesn't need them.) Supported engine install methods are all pinned to v0.22.1: the prebuilt binary above, agentmemory's macOS/Linux auto-install path (`curl`, POSIX `sh`, and `tar` required), and the Docker image `iiidev/iii:0.22.1`. A bare upstream `install.sh | sh` installs the latest engine, which agentmemory does not support. Use `npx -y @agentmemory/agentmemory@latest`; on macOS/Linux it fetches the pinned engine into `~/.agentmemory/bin`. ---

Deploy

One-click templates for managed hosts. Each one ships a self-contained Dockerfile that pulls `@agentmemory/agentmemory` from npm and copies the iii engine binary in from the official `iiidev/iii` Docker Hub image; no pre-built agentmemory image required. Persistent storage mounts at `/data`; the first-boot entrypoint overwrites the npm-bundled iii config (which binds `127.0.0.1`) with a deploy-tuned one that binds `0.0.0.0` and uses absolute `/data` paths, generates the HMAC secret, then drops privileges from `root` to `node` via `gosu` before exec'ing the agentmemory CLI.

Deploy to fly.io Deploy to Railway

Render's one-click deploy button requires `render.yaml` at the repository root, which we deliberately keep clean. Use the Render Blueprint flow documented in [`deploy/render/`](./deploy/render/README.md) to point at the in-repo blueprint manually. Full setup details (HMAC capture, viewer SSH tunnel, rotation, backup, cost floors) live in [`deploy/`](./deploy/README.md): - [`deploy/fly`](./deploy/fly/README.md): single machine with `auto_stop_machines = "stop"`; cheapest idle. - [`deploy/railway`](./deploy/railway/README.md): Hobby plan flat fee, volume in the dashboard. - [`deploy/render`](./deploy/render/README.md): Blueprint flow, automatic disk snapshots on paid plans. - [`deploy/coolify`](./deploy/coolify/README.md): self-hosted on your own VPS via [Coolify](https://coolify.io/self-hosted); same Docker Compose stack, you own the host and the data. Only port `3111` is published. The viewer on `3113` stays bound to loopback inside the container; every template's README documents the SSH-tunnel pattern for reaching it. ---

Why agentmemory

Every coding agent forgets everything when the session ends, and each new session starts with you re-explaining your stack. agentmemory runs in the background and removes that step. ```text Session 1: "Add auth to the API" Agent writes code, runs tests, fixes bugs agentmemory silently captures every tool use Session ends -> observations compressed into structured memory Session 2: "Now add rate limiting" Agent already knows: - Auth uses JWT middleware in src/middleware/auth.ts - Tests in test/auth.test.ts cover token validation - You chose jose over jsonwebtoken for Edge compatibility Zero re-explaining. Starts working immediately. ``` ### vs built-in agent memory Every AI coding agent ships with built-in memory: Claude Code has `MEMORY.md`, Cursor has notepads, Cline has memory bank. These work like sticky notes. agentmemory is the searchable database behind the sticky notes. | | Built-in (CLAUDE.md) | agentmemory | |---|---|---| | Scale | 200-line cap | Unlimited | | Search | Loads everything into context | BM25 + vector + graph (top-K only) | | Token cost | 22K+ at 240 observations | ~1,900 tokens (92% less) | | Cross-agent | Per-agent files | MCP + REST (any agent) | | Coordination | None | Leases, signals, actions, routines | | Observability | Read files manually | Real-time viewer on :3113 | ---

How It Works

### Memory Pipeline ```text PostToolUse hook fires -> SHA-256 dedup (5min window) -> Privacy filter (strip secrets, API keys) -> Store raw observation -> Synthetic compression by default (LLM-written compression only with a provider + AGENTMEMORY_AUTO_COMPRESS=true) -> Vector embedding when an embedding provider is active -> Index in BM25, plus vectors when enabled Stop / SessionEnd hook fires -> Summarize session -> Knowledge graph extraction (if GRAPH_EXTRACTION_ENABLED=true) -> Slot reflection (if SLOT_REFLECT_ENABLED=true) SessionStart hook fires -> Load project profile (top concepts, files, patterns) -> Hybrid search (BM25 + vector + graph) -> Token budget (default: 2000 tokens) -> Inject into conversation ``` ### 4-Tier Memory Consolidation Modeled on how human brains process memory, including sleep consolidation. | Tier | What | Analogy | |------|------|---------| | **Working** | Raw observations from tool use | Short-term memory | | **Episodic** | Compressed session summaries | "What happened" | | **Semantic** | Extracted facts and patterns | "What I know" | | **Procedural** | Workflows and decision patterns | "How to do it" | Memories decay over time (Ebbinghaus curve). Frequently accessed memories strengthen. Stale memories auto-evict. Contradictions are detected and resolved. ### What Gets Captured | Hook | Captures | |------|----------| | `SessionStart` | Project path, session ID | | `UserPromptSubmit` | User prompts (privacy-filtered) | | `PreToolUse` | File access patterns + enriched context | | `PostToolUse` | Tool name, input, output | | `PostToolUseFailure` | Error context | | `PreCompact` | Re-injects memory before compaction | | `SubagentStart/Stop` | Sub-agent lifecycle | | `Stop` | End-of-session summary | | `SessionEnd` | Session complete marker | ### Key Capabilities | Capability | Description | |---|---| | **Automatic capture** | Every tool use recorded via hooks, no manual effort | | **Semantic search** | BM25 + vector + knowledge graph with RRF fusion | | **Memory evolution** | Versioning, supersession, relationship graphs | | **Recall hygiene** | Superseded memory versions leave the search indexes; the version chain in KV keeps full history | | **Near-duplicate hints** | Saves report an advisory `similarTo` match when new content closely resembles an existing memory | | **Per-agent scoping** | `agentId` threads through save and recall across REST, MCP, and the search index, in shared or isolated mode | | **Write-time provenance** | Every observation and memory carries an immutable origin channel (user, agent, tool, import, or shared) stamped at capture, save, and import | | **Auto-forgetting** | TTL expiry, contradiction detection, importance eviction | | **Privacy first** | API keys, secrets, `` tags stripped before storage | | **Self-healing** | Circuit breaker, provider fallback chain, health monitoring | | **Claude bridge** | Bi-directional sync with MEMORY.md | | **Knowledge graph** | Entity extraction + BFS traversal | | **Team memory** | Namespaced shared + private across team members | | **Citation provenance** | Trace any memory back to source observations | | **Git snapshots** | Version, rollback, and diff memory state | --- Triple-stream retrieval combining three signals: | Stream | What it does | When | |---|---|---| | **BM25** | Stemmed keyword matching with synonym expansion | Always on | | **Vector** | Cosine similarity over dense embeddings | Embedding provider configured | | **Graph** | Knowledge graph traversal via entity matching | Entities detected in query | Fused with Reciprocal Rank Fusion (RRF, k=60) and session-diversified (max 3 results per session). When a vector index is populated, `mem::search` (behind `memory_recall`) uses the hybrid BM25 + vector ranker. Without embeddings it uses BM25. `smart-search` can additionally fuse structural graph matches when graph data exists, including in keyless mode. Lesson recall runs on a dedicated in-memory BM25 index instead of scanning the whole corpus per query. Superseded memory versions are excluded from every recall path; the version chain keeps their history. Vectors survive a crash or force-kill. The vector index is saved in buckets at most every `AGENTMEMORY_INDEX_SAVE_INTERVAL_MS` (10 minutes). Every vector added or removed in between is also written right away to a small pending log in the state store, and the next start replays it without calling the embedding provider. Each successful save empties the log. Documents that still have no vector after the replay are re-embedded in the background in batches of `AGENTMEMORY_VECTOR_BACKFILL_MAX` (500) until none are left, and a backfill that is stopped continues at the next start. `/agentmemory/status` and the viewer show the pending log size and the backfill state. Keyless installs write nothing. BM25 tokenizes Greek, Cyrillic, Hebrew, Arabic, and accented Latin out of the box. For Chinese / Japanese / Korean memories, install the optional segmenters (`npm install @node-rs/jieba tiny-segmenter`) to split CJK runs into word-level tokens; without them, agentmemory soft-falls to whole-run tokenization and prints a one-time hint on stderr. ### Embedding providers Keyless installs disable vector embeddings: `mem::search` uses BM25, while `smart-search` can also use existing structural graph data. To opt into free on-device semantic embeddings, add this to `~/.agentmemory/.env` and restart agentmemory: ```env EMBEDDING_PROVIDER=local ``` The normal npm install includes the optional `@huggingface/transformers` runtime. The first embedding request downloads `Xenova/all-MiniLM-L6-v2`, so it needs network access and can take longer; subsequent inference runs on-device. Remote providers are auto-detected from their keys unless `EMBEDDING_PROVIDER` overrides them. | Provider | Model | Cost | Notes | |---|---|---|---| | **Local (recommended opt-in)** | `all-MiniLM-L6-v2` | Free | On-device after the first model download, +8pp recall over BM25-only | | Gemini | `gemini-embedding-001` | Free tier | 100+ languages, 768/1536/3072 dims (MRL), 2048-token input. Replaces `text-embedding-004` ([deprecated, shutdown Jan 14, 2026](https://ai.google.dev/gemini-api/docs/deprecations)) | | OpenAI | `text-embedding-3-small` | $0.02/1M | Highest quality | | Voyage AI | `voyage-code-3` | Paid | Optimized for code | | Cohere | `embed-english-v3.0` | Free trial | General purpose | | OpenRouter | Any model | Varies | Multi-model proxy | ---

MCP Server

54 tools, 6 resources, 3 prompts, and 17 skills. > **MCP shim vs full server:** the published `@agentmemory/mcp` package is a thin shim. It exposes the full 54-tool surface **only when it can reach a running agentmemory server** via `AGENTMEMORY_URL` (proxy mode). With no server reachable, the shim falls back to a 7-tool local set (`memory_save`, `memory_recall`, `memory_smart_search`, `memory_sessions`, `memory_export`, `memory_audit`, `memory_governance_delete`). The `AGENTMEMORY_TOOLS=core|all` env var is a *server-side* flag; setting it in the shim's `env` block has no effect. If you see only 7 tools in Cursor / OpenCode / Gemini CLI, start `npx -y @agentmemory/agentmemory@latest` (or the Docker stack) and set `AGENTMEMORY_URL=http://localhost:3111`. ### 54 Tools Three tool surfaces, smallest to largest: `AGENTMEMORY_TOOLS=core` trims visibility to 8 essentials (`memory_save`, `memory_recall`, `memory_consolidate`, `memory_smart_search`, `memory_sessions`, `memory_diagnose`, `memory_lesson_save`, `memory_reflect`); the base set below is the registry's 14 foundational tools; the default (`AGENTMEMORY_TOOLS=all`) exposes all 54.
Base tools (14) | Tool | Description | |------|-------------| | `memory_recall` | Search past observations | | `memory_compress_file` | Compress markdown files while preserving structure | | `memory_save` | Save an insight, decision, or pattern | | `memory_file_history` | Past observations about specific files | | `memory_patterns` | Detect recurring patterns | | `memory_sessions` | List recent sessions | | `memory_smart_search` | Hybrid semantic + keyword search | | `memory_vision_search` | Search image observations | | `memory_timeline` | Chronological observations | | `memory_profile` | Project profile (concepts, files, patterns) | | `memory_export` | Export all memory data | | `memory_relations` | Query relationship graph | | `memory_commit_lookup` | Sessions behind a git commit | | `memory_commits` | Commits recorded for a session |
Extended tools (54 total, the default surface) | Tool | Description | |------|-------------| | `memory_patterns` | Detect recurring patterns | | `memory_timeline` | Chronological observations | | `memory_relations` | Query relationship graph | | `memory_graph_query` | Knowledge graph traversal | | `memory_consolidate` | Run 4-tier consolidation | | `memory_claude_bridge_sync` | Sync with MEMORY.md | | `memory_team_share` | Share with team members | | `memory_team_feed` | Recent shared items | | `memory_audit` | Audit trail of operations | | `memory_governance_delete` | Delete with audit trail | | `memory_snapshot_create` | Git-versioned snapshot | | `memory_action_create` | Create work items with dependencies | | `memory_action_update` | Update action status | | `memory_frontier` | Unblocked actions ranked by priority | | `memory_next` | Single most important next action | | `memory_lease` | Exclusive action leases (multi-agent) | | `memory_routine_run` | Instantiate workflow routines | | `memory_signal_send` | Inter-agent messaging | | `memory_signal_read` | Read messages with receipts | | `memory_checkpoint` | External condition gates | | `memory_mesh_sync` | P2P sync between instances | | `memory_sentinel_create` | Event-driven watchers | | `memory_sentinel_trigger` | Fire sentinels externally | | `memory_sketch_create` | Ephemeral action graphs | | `memory_sketch_promote` | Promote to permanent | | `memory_crystallize` | Compact action chains | | `memory_diagnose` | Health checks | | `memory_heal` | Auto-fix stuck state | | `memory_facet_tag` | Dimension:value tags | | `memory_facet_query` | Query by facet tags | | `memory_verify` | Trace provenance |
### 6 Resources · 3 Prompts · 17 Skills | Type | Name | Description | |------|------|-------------| | Resource | `agentmemory://status` | Health, session count, memory count | | Resource | `agentmemory://project/{name}/profile` | Per-project intelligence | | Resource | `agentmemory://project/{name}/recent` | Recent observations for a project | | Resource | `agentmemory://memories/latest` | Latest 10 active memories | | Resource | `agentmemory://graph/stats` | Knowledge graph statistics | | Resource | `agentmemory://team/{id}/profile` | Shared team profile | | Prompt | `recall_context` | Search + return context messages | | Prompt | `session_handoff` | Handoff data between agents | | Prompt | `detect_patterns` | Analyze recurring patterns | | Skill | `/recall` | Search memory | | Skill | `/remember` | Save to long-term memory | | Skill | `/session-history` | Recent session summaries | | Skill | `/forget` | Delete observations/sessions | The table shows the four core skills. The full set is 9 invocable skills plus 8 reference skills; see the Native skills section above. ### Standalone MCP Run without the full server, for any MCP client. Either of these works: ```bash npx -y @agentmemory/agentmemory@latest mcp # canonical (always available) npx -y @agentmemory/mcp # shim package alias ``` Or add to your agent's MCP config: Most agents (Cursor, Claude Desktop, Cline, Roo Code, Gemini CLI): ```json { "mcpServers": { "agentmemory": { "command": "npx", "args": ["-y", "@agentmemory/mcp"], "env": { "AGENTMEMORY_URL": "http://localhost:3111" } } } } ``` Merge the `agentmemory` entry into your host's existing `mcpServers` object rather than replacing the file. For sandboxed clients that can't reach the host's `localhost`, add `"AGENTMEMORY_FORCE_PROXY": "1"` to the env block and set `AGENTMEMORY_URL` to a route the sandbox can reach. OpenCode (`opencode.json`): ```json { "mcp": { "agentmemory": { "type": "local", "command": ["npx", "-y", "@agentmemory/mcp"], "enabled": true } }, "plugin": ["./plugins/agentmemory-capture.ts"] } ``` Copy the plugin file from the repo: ```bash mkdir -p ~/.config/opencode/plugins cp plugin/opencode/agentmemory-capture.ts ~/.config/opencode/plugins/ cp plugin/opencode/commands/*.md ~/.config/opencode/commands/ ``` ---

Real-Time Viewer

Auto-starts on port `3113`. The viewer loads one snapshot when it connects (`GET /agentmemory/viewer/snapshot`) and then applies live stream events: new memories, lessons, observations, audit entries, graph changes and health updates appear without polling or page reloads. The only other requests are the actions you click, "load more" pages and searches. When the stream drops, the viewer shows how old its numbers are, reconnects with backoff and resyncs from one snapshot. - **12 tabs in four groups** with live counts, deep links (`#memories/`, `#sessions/?obs=`, `#graph/`, `#health/consolidation`), keyboard shortcuts and a mobile menu. - **Memories:** server-side search, filters by project, agent and type, a detail panel with the version chain and a word diff, provenance links, copy buttons for the id, the MCP call and a curl command, edit (a new version), forget with confirmation, bulk forget and JSON export. - **Sessions:** an inline observation timeline with readable tool input and output, filters and paging, and the memories and lessons each session produced. - **Graph:** search, node detail with relations and sources, a legend that does not rely on colour alone, and zoom controls. - **Health:** the live version of `GET /agentmemory/status`. Every problem comes with its fix, plus the state backend, index save state, graph provenance compaction progress and a consolidation explainer with the real thresholds. - **Audit, Activity, Profile, Replay, Lessons, Actions and Crystals** pages, each with an empty state that says what the section is, why it is empty and the command that fills it, and a `?` glossary tooltip on every term and number. ```bash open http://localhost:3113 ``` The viewer server binds to `127.0.0.1` by default and attaches the server secret when it forwards requests to the REST API, so it needs no setup. The REST-served `/agentmemory/viewer` endpoint follows the normal bearer-token rules and redirects browsers without a token to the viewer port. CSP headers use a per-response script nonce and disable inline handler attributes (`script-src-attr 'none'`). ---

iii Console

The viewer at `:3113` shows what your agent **remembered**. The [iii console](https://iii.dev/docs/console) shows what your agent **did**: every memory op as an OpenTelemetry trace, every KV entry editable, every function invocable, every stream tappable. Two windows on the same memory: one product-shaped, one engine-shaped. Watch a `memory_smart_search` fire and see the BM25 scan → embedding lookup → RRF fusion → reranker as a waterfall. Edit a stuck consolidation timer in the KV browser. Replay a `PostToolUse` hook with a tweaked payload. Pin the WebSocket stream and watch observations land live. agentmemory ships this for free because every function call and trigger fires through iii; nothing custom, nothing to instrument.

iii console Workers page: connected workers including agentmemory instances with live function counts and runtime metadata
Workers page: every connected worker, including agentmemory itself, with PID, function count, runtime, and last-seen.

**Already installed.** The console ships with the pinned `iii` engine (0.22+); nothing separate to install. The first launch downloads the console binary next to the engine. **Launch alongside agentmemory:** ```bash agentmemory console ``` This runs the pinned engine's `iii console` against the ports agentmemory resolved (REST, streams, bridge) and serves it one port above the viewer, `http://localhost:3114` by default. `--console-port N` picks another port; `--port` and `--instance` select the agentmemory instance the same way they do for `stop`; any other flag is passed through, for example `--enable-flow` for the experimental architecture-graph page. The same thing by hand, useful when `agentmemory` is not on PATH: ```bash ~/.agentmemory/bin/iii console --port 3114 \ --engine-port 3111 \ --ws-port 3112 \ --bridge-port 49134 ``` **What you can do from the console:** | Page | Use it to | |------|-----------| | **Workers** | See every connected worker and its live metrics, including the agentmemory worker itself. | | **Functions** | Invoke any of agentmemory's functions directly with a JSON payload; handy for testing `memory.recall`, `memory.consolidate`, `graph.query` without wiring a client. | | **Triggers** | Replay HTTP, cron, event, and state triggers: fire the consolidation cron manually, retry an HTTP route, emit a state change. | | **States** | KV browser with full CRUD over sessions, memory slots, lifecycle timers, and the embeddings index; edit values in place. | | **Streams** | Live WebSocket monitor for memory writes, hook events, and observation updates as they flow through iii streams. | | **Queues** | Durable queue topics + dead-letter management. Replay or drop failed embedding / compression jobs. | | **Traces** | OpenTelemetry waterfall / flame / service-breakdown views. Filter by `trace_id` to see exactly which functions, DB calls, and embedding requests a single `memory.search` produced. | | **Logs** | Structured OTEL logs filtered and correlated to trace/span IDs. | | **Config** | Runtime configuration: see exactly which workers, providers, and ports your engine is running with. | | **Flow** | (Optional, `--enable-flow`) Interactive architecture graph of every worker, trigger, and stream. |

iii console trace waterfall view showing per-span duration
Traces: waterfall / flame / service breakdown for every memory operation.

**Traces are already on:** `iii-config.yaml` ships with the `iii-observability` worker enabled (`exporter: memory`, `sampling_ratio: 0.1`, metrics + logs). No extra config needed; the moment agentmemory starts, every memory operation emits a structured log the console can read, and one in ten of them (`sampling_ratio: 0.1`) also emits a trace span. If you want to export to Jaeger/Honeycomb/Grafana Tempo instead, change `exporter: memory` to `exporter: otlp` and set the collector endpoint per iii's observability docs. > **Heads-up:** no auth is enforced on the console itself; keep it bound to `127.0.0.1` (the default) and never expose it publicly. ---

Powered by iii

agentmemory is **already a running [iii](https://iii.dev) instance**. Three primitives (worker, function, trigger) compose the runtime; KV state, streams, and OTEL traces come from iii-state, iii-stream, and iii-observability workers that ship with iii. You didn't install Postgres, Redis, Express, pm2, or Prometheus, because iii replaces them. That means one more command extends agentmemory with an entire new capability. ### Extend agentmemory with more workers The builtins agentmemory needs are already in `iii-config.yaml` and boot with it: `iii-state` (KV), `iii-queue` (durable retries for the event subscribers), `iii-pubsub`, `iii-cron`, `iii-stream`, and `iii-observability` (OTEL traces, metrics and logs on every function). Anything else from the [iii worker registry](https://workers.iii.dev) plugs into the same engine: copy `iii-config.yaml` to `~/.agentmemory/iii-config.yaml` (the CLI prefers that file over the bundled one and still renders ports and data paths into it), add the entry, install the worker runtime once with `~/.agentmemory/bin/iii update worker`, and restart agentmemory. ```yaml workers: # ...the bundled entries... - name: database # SQL-backed state adapter when you outgrow the KV defaults - name: iii-sandbox # run code that came out of memory_recall inside a throwaway VM - name: mcp # extra MCP servers next to agentmemory's, same engine ``` | Worker | What you get on top of agentmemory | |---|---| | [`database`](https://workers.iii.dev/workers/database) | SQL-backed state adapter when you outgrow the in-memory KV defaults | | [`iii-sandbox`](https://workers.iii.dev/workers/iii-sandbox) | Code that came out of `memory_recall` runs inside a throwaway VM, not your shell | | [`mcp`](https://workers.iii.dev/workers/mcp) | Stand up extra MCP servers next to agentmemory's, share the same engine | On engine 0.22.x keep the `iii-` prefixed names for the builtins above; the unprefixed `http`, `state`, `queue`, `pubsub` and `cron` entries are the standalone registry workers agentmemory moves to with the 0.23 migration. Full registry: [workers.iii.dev](https://workers.iii.dev). Every worker there composes through the same primitives agentmemory uses, and the agentmemory you already have is one of them. ### Engine config and bind address `agentmemory start` reads the engine config from the first file that exists: `AGENTMEMORY_III_CONFIG`, `./iii-config.yaml` in the current directory, `~/.agentmemory/iii-config.yaml`, then the bundled `iii-config.yaml`. On every start it renders that file (data paths, ports, state backend) into `~/.agentmemory/data/iii-config.runtime.yaml` and launches the engine with the rendered copy, so edit the source file, not the rendered one. The `host:` values of the source file are kept as written. The bundled `iii-config.yaml` binds `127.0.0.1` on purpose, and that default also applies inside a container. A CLI started in a container listens on the container's loopback, so published ports reach nothing. To serve a containerized CLI through published ports, set `AGENTMEMORY_III_CONFIG` to a config that binds `0.0.0.0`. The packaged `iii-config.docker.yaml` is one: it binds `iii-http`, `iii-stream` and the engine port to `0.0.0.0` and stores state under `/data`, so mount a writable volume there. Keep `AGENTMEMORY_SECRET` set, and publish only the ports you need, on `127.0.0.1` or behind a proxy you trust. This repo's `docker-compose.yml` does not go through the CLI's config lookup: it mounts `iii-config.docker.yaml` at `/app/config.yaml`, and the `iii-engine` container starts with `--config /app/config.yaml`. The one-click [deploy templates](deploy/) write their own `0.0.0.0` config in their entrypoints. ### Storage backend: file (default) vs redis `iii-state` and `iii-stream` default to iii-engine's bundled file-based KV store: one JSON file per scope, held in the engine process's memory and rewritten to disk on a timer. That's the right default for a single-user local install; a shared daemon with several concurrent writers gets real per-key writes from Redis instead, at the cost of a network round trip per operation (every `state::*` call still serializes on one Redis connection, so this trades the file store's lock for a socket, not for parallelism). Set `AGENTMEMORY_STATE_BACKEND=redis` (plus `AGENTMEMORY_REDIS_URL`) to switch both workers to iii-engine's built-in `redis` adapter, which stores each key as a Redis hash field (`HSET`) instead of rewriting a whole scope on every write: ```env # ~/.agentmemory/.env AGENTMEMORY_STATE_BACKEND=redis AGENTMEMORY_REDIS_URL=redis://localhost:6379 ``` `AGENTMEMORY_STATE_BACKEND` defaults to `file`; leaving it unset keeps today's behavior unchanged, and an unrecognized value (anything other than `file` or `redis`) is a startup error rather than a silent fallback. `/agentmemory/status` and the viewer's Health page (the State store row) report which backend is active and whether it answers, never the URL. **Plain `redis://` only.** The pinned engine (0.22.1) builds its Redis client without TLS support, so a `rediss://` URL (most managed Redis offerings, such as Upstash, Redis Cloud, and ElastiCache with in-transit encryption, default to TLS-only) fails to connect. The connection is unencrypted, so the Redis password and every stored memory cross the wire in clear text: point at a local Redis or one on a private network you trust. For any other Redis, run an encrypted tunnel (stunnel, SSH, or a VPN) on the agentmemory host, so the plain `redis://` hop stays on that host and the tunnel's upstream connection is encrypted and authenticated. If a Redis password contains a single quote, percent-encode it (`%27`); the engine expands the URL into its YAML config before parsing it. **One Redis server per `--instance`.** The engine's Redis key prefixes (`state:`, `stream::`) are fixed, so two agentmemory instances (`--instance 1`, `--instance 2`, ...) pointed at the same database overwrite each other's data. A separate database index (`redis://localhost:6379/1`) keeps the stored data apart, but the engine relays live viewer events over one Redis pub/sub channel (`stream::events`), and Redis pub/sub ignores the database index, so each instance's viewer would still show the other's live events. Give each instance its own Redis server (or port) when you run more than one. **What stays the same, and what differs.** Every agentmemory feature works on Redis: sessions, observations, memories (remember, supersede, evolve, forget), search and the index buckets, lessons, the graph, the audit log and its monthly scopes, export and import, governance deletes, consolidation status, the viewer snapshot and its live stream, and the health monitor. The engine stores each scope as one Redis hash (`HSET`/`HGET`/`HGETALL`) and fires the same state triggers as the file store. Three engine differences are handled inside agentmemory: - Redis returns a scope's records in no fixed order. agentmemory sorts them oldest first (by the creation time in the record id, then its timestamp) so lists, paging and export chunks come back in the same order as on the file store. - The engine applies partial updates on Redis in a Lua script that turns empty arrays into empty objects. agentmemory applies those updates itself (read, change, write under a per-key lock) on Redis, so fields like `tags: []` stay arrays. - The legacy audit log check reads the old scope from Redis instead of looking for the file store's file on disk. One difference needs you: **after Redis restarts, the engine stops relaying live events** to the viewer until agentmemory restarts. Data is still saved and read normally. The health monitor sends a test event through Redis every 30 seconds; when it does not come back, `/agentmemory/status` and the viewer's Health page show "Live updates are not reaching the viewer" with the fix: restart agentmemory. If Redis is down, the status report shows "The state store is not answering" and how to check it (`redis-cli -u "$AGENTMEMORY_REDIS_URL" ping`). Listing a very large scope reads the whole hash in one `HGETALL`, the same cost as the file store holding it in memory. **Recommended Redis settings.** The default `save 3600 1 300 100 60 10000` snapshot policy can lose minutes of writes on a crash, worse than the file store's 5s flush window. Set `appendonly yes` for anything you'd mind losing. Set `maxmemory-policy noeviction`; `allkeys-lru` or similar silently drops memories once Redis hits its memory limit. A native (non-Docker) start, and every one-click [deploy template](deploy/) (they overwrite the bundled `iii-config.yaml` and start natively), read `AGENTMEMORY_STATE_BACKEND`/`AGENTMEMORY_REDIS_URL` and render them into the launched `iii-config`. The URL itself is never written to that rendered file, only a `${AGENTMEMORY_REDIS_URL}` reference that the engine process expands from its own environment at boot. Only this repo's own Docker Compose path (`AGENTMEMORY_USE_DOCKER=1`, or resuming an engine already started that way) mounts `iii-config.docker.yaml` read-only and never renders; `agentmemory start` warns when it detects that combination. Switch that file by hand, following the same `name: redis` / `config: redis_url: ...` shape shown in the [iii-state](https://workers.iii.dev/workers/iii-state) and [iii-stream](https://workers.iii.dev/workers/iii-stream) worker docs, and point `redis_url` at a Redis reachable from the container. `docker-compose.yml` passes `AGENTMEMORY_REDIS_URL` into the engine container, so `redis_url: '${AGENTMEMORY_REDIS_URL}'` works there and keeps the URL out of the mounted file. The rendered config keeps the URL out of `~/.agentmemory/data/iii-config.runtime.yaml`, but the engine's own configuration worker still persists the *expanded* value to `~/.agentmemory/config/iii-state.yaml` and `iii-stream.yaml` once it boots (iii-engine's `${VAR}` expansion happens before that worker stores its seed, and it stores the resolved value, not the reference). Treat that directory as holding a credential: `chmod 700 ~/.agentmemory` on any shared host, and prefer a Redis ACL user scoped to what agentmemory needs over the database's admin credentials. **Migration is not automatic.** Switching `AGENTMEMORY_STATE_BACKEND` starts from an empty store on either side; nothing copies existing data from file to Redis or back. Export from the backend you're leaving and import into the one you're moving to. This runs identically under bash and zsh (including `bash -u`). An array like `AUTH=(${AGENTMEMORY_SECRET:+-H "Authorization: Bearer $AGENTMEMORY_SECRET"})` does not: zsh keeps the header as one malformed word where bash splits it into two, so both requests 401 whenever `AGENTMEMORY_SECRET` is set: ```bash # 0. Use the generated secret when none is exported: AGENTMEMORY_SECRET="${AGENTMEMORY_SECRET:-$(cat ~/.agentmemory/secret 2>/dev/null)}" # 1. On the old backend, while agentmemory is still running on it: if [ -n "${AGENTMEMORY_SECRET:-}" ]; then curl -fsS -H "Authorization: Bearer $AGENTMEMORY_SECRET" http://localhost:3111/agentmemory/export > backup.json else curl -fsS http://localhost:3111/agentmemory/export > backup.json fi # 2. Confirm backup.json is a usable export before switching backends: jq -e '.version and .exportedAt' backup.json > /dev/null || { echo "backup.json is not a valid export; do not switch backends" >&2 exit 1 } # 3. Switch AGENTMEMORY_STATE_BACKEND (and AGENTMEMORY_REDIS_URL if needed), # restart agentmemory against the new backend, then: if [ -n "${AGENTMEMORY_SECRET:-}" ]; then jq -n --slurpfile d backup.json '{exportData: $d[0], strategy: "merge"}' | \ curl -fsS -H "Authorization: Bearer $AGENTMEMORY_SECRET" -X POST http://localhost:3111/agentmemory/import \ -H 'Content-Type: application/json' -d @- else jq -n --slurpfile d backup.json '{exportData: $d[0], strategy: "merge"}' | \ curl -fsS -X POST http://localhost:3111/agentmemory/import \ -H 'Content-Type: application/json' -d @- fi ``` `/agentmemory/export` also accepts `?maxSessions=` and `?offset=` for chunking a large corpus across several calls; `strategy` on import is `merge` (default-safe), `replace`, or `skip`. ### What iii replaces | Traditional stack | agentmemory uses | |---|---| | Express.js / Fastify | iii HTTP Triggers | | SQLite / Postgres + pgvector | iii KV State + in-memory vector index | | SSE / Socket.io | iii Streams (WebSocket) | | pm2 / systemd | iii engine worker supervision | | Prometheus / Grafana | iii OTEL + health monitor | | Custom plugin systems | `iii worker add ` | **219 source files · ~52,000 LOC · 2,500+ tests · 311 functions · 60 KV scopes**, all on three primitives. No `agentmemory plugin install`. The plugin system is iii itself. ---

Configuration

### LLM Providers agentmemory auto-detects providers from your environment. A provider makes LLM-backed operations available, but provider configuration alone does not enable LLM-written observation compression. That path requires both a provider and `AGENTMEMORY_AUTO_COMPRESS=true`. | Provider | Config | Notes | |----------|--------|-------| | **No-op (default)** | No config needed | LLM-backed compress/summarize is disabled. Synthetic compression and BM25 recall still work. See `AGENTMEMORY_ALLOW_AGENT_SDK` below if you used to rely on the Claude-subscription fallback. | | Anthropic API | `ANTHROPIC_API_KEY` | Per-token billing | | MiniMax | `MINIMAX_API_KEY` | Anthropic-compatible | | Gemini | `GEMINI_API_KEY` | Also enables embeddings | | OpenRouter | `OPENROUTER_API_KEY` | Any model | | OpenAI API | `OPENAI_API_KEY` | Default `gpt-5.6-luna`, override with `OPENAI_MODEL` | | **Local (Ollama / LM Studio / vLLM / llama.cpp)** | `OPENAI_API_KEY=local` + `OPENAI_BASE_URL=http://localhost:11434/v1` (Ollama) or `http://localhost:1234/v1` (LM Studio) + `OPENAI_MODEL=` | Anything OpenAI-API-compatible. Zero cost, runs on your hardware. See [Local models](#local-models-ollama--lm-studio--vllm) below. | | Claude subscription fallback | `AGENTMEMORY_ALLOW_AGENT_SDK=true` | Opt-in only. Spawns `@anthropic-ai/claude-agent-sdk` sessions; it used to cause unbounded Stop-hook recursion, so it is no longer the default. | ### Local models (Ollama / LM Studio / vLLM) agentmemory talks to any OpenAI-API-compatible server, so anything that exposes `/v1/chat/completions` works without code changes. No paid keys, no cloud, no rate limits; runs entirely on your hardware. **Ollama** (default port `11434`): ```bash ollama pull qwen3:8b # or qwen3:4b, gpt-oss:20b, qwen3-coder:30b, etc. ollama serve ``` ```env # ~/.agentmemory/.env OPENAI_API_KEY=ollama # any non-empty string; Ollama ignores it OPENAI_BASE_URL=http://localhost:11434/v1 OPENAI_MODEL=qwen3:8b ``` **LM Studio** (default port `1234`): Open LM Studio → Local Server tab → Start Server. Pick any chat model from the picker (Qwen 3, gpt-oss, DeepSeek R1, etc.). ```env # ~/.agentmemory/.env OPENAI_API_KEY=lmstudio # any non-empty string; LM Studio ignores it OPENAI_BASE_URL=http://localhost:1234/v1 OPENAI_MODEL=qwen3-8b # match the model name from LM Studio ``` **vLLM / llama.cpp / Text Generation Inference**: same shape. Point `OPENAI_BASE_URL` at whatever URL your server exposes and set `OPENAI_MODEL` to a name your server will accept. **Model picks for memory work**: compression and summarization are short tasks (<2K tokens in, <500 tokens out) where a 7B instruct model is plenty. Recommendations: | Model | Size | Why | |-------|------|-----| | `qwen3:8b` | ~5.2 GB | Balanced default on a 16 GB machine; strong at extraction and tool-shaped text | | `qwen3:4b` | ~2.6 GB | Smallest sane option; fine for compression, weaker for graph extraction | | `qwen3-coder:30b` | ~19 GB | Best local pick for code-shaped sessions (30B MoE, 3.3B active) on 24-32 GB hardware | | `gpt-oss:20b` | ~14 GB | Strong general model that fits 16 GB RAM | | `deepseek-r1:8b` | ~5.2 GB | Reasoning distill; slower but cleaner extractions | Qwen 3 models think by default and can burn the whole token budget on reasoning before any output. Set `AGENTMEMORY_LLM_NOTHINK=1` to append `/no_think` to graph-extraction prompts, and raise `MAX_TOKENS` (16384 works) if extractions come back empty. Reasoning-class models (`o1`-style with `` blocks) can return empty `content` with a `reasoning` field your local server may not surface. If extractions come back blank, switch to a non-reasoning model first. The `OPENAI_REASONING_EFFORT=none` env can also disable thinking on Ollama Cloud thinking models that mirror the OpenAI reasoning schema. Local embeddings ship as an optional dependency but are not enabled by default. Set `EMBEDDING_PROVIDER=local` to opt into `Xenova/all-MiniLM-L6-v2` (384-dim). The first embedding request downloads the model; inference is on-device afterward. Without that setting or a remote embedding key, vectors stay disabled, `mem::search` uses BM25, and `smart-search` can still add existing graph matches. ### Cost-aware model selection When LLM-written background compression is enabled with both a provider and `AGENTMEMORY_AUTO_COMPRESS=true`, it runs on every observation, so model choice meaningfully changes monthly spend. Captured workload data: 635 requests / 888K tokens / 35 hours of active use, run against three OpenRouter models at 2026-05-23 pricing. | Tier | Model | Input / 1M | Output / 1M | Cost for the captured 35h | Notes | |------|-------|------------|-------------|---------------------------|-------| | Recommended | `deepseek/deepseek-v4-flash-0731` | $0.07 | $0.14 | ~$0.07 (est.) | Latest DeepSeek; cheapest recommended pick for compression workloads. | | Recommended | `deepseek/deepseek-v4-pro` | $0.435 | $0.87 | ~$0.46 | Solid compression + summarization quality at ~10× lower cost than Sonnet. | | Recommended | `qwen/qwen3-coder` | $0.45 | $1.80 | ~$0.55 | Strong code reasoning if your sessions are heavily code-shaped. | | Premium | `anthropic/claude-sonnet-5` | $3.00 | $15.00 | ~$5.02 (est.) | Same list price as the measured Sonnet 4.6 run; $2/$10 intro pricing through 2026-08-31. | | Premium | `openai/gpt-5.6-sol` | $5.00 | $30.00 | ~$9 (est.) | Flagship tier; expensive for always-on background work. | | Avoid | `anthropic/claude-opus-5` | $5.00 | $25.00 | ~$8.40 (est.) | Flagship-class model; overspend for compression. | Measured rows come from the captured run; (est.) rows scale the same token mix by each model's list price. agentmemory prints a runtime warning when `OPENROUTER_MODEL` matches a premium-tier pattern. Set `AGENTMEMORY_SUPPRESS_COST_WARNING=1` to silence once you've made an informed choice. Quality vs cost tradeoff for memory work: compression is a summarization task with relatively loose quality bars (the agent re-reads the summary, not the user). DeepSeek V4 Flash / V4 Pro / Qwen3-Coder land within rounding error of Sonnet on this task while costing 10-70× less. Save the premium-tier models for queries you read directly. Sources: [OpenRouter pricing for Claude Sonnet 5](https://openrouter.ai/anthropic/claude-sonnet-5), [DeepSeek V4 Flash](https://openrouter.ai/deepseek/deepseek-v4-flash-0731), [DeepSeek pricing notes](https://api-docs.deepseek.com/quick_start/pricing/). ### Multi-agent memory (`AGENT_ID` + `AGENTMEMORY_AGENT_SCOPE`) In multi-agent setups where several roles share one agentmemory server (architect / developer / reviewer / researcher / support-agent), `AGENT_ID` tags every write with the role that made it. `AGENTMEMORY_AGENT_SCOPE` controls whether recall filters by that tag. ```env TEAM_ID=company USER_ID=engineering-team AGENT_ID=architect AGENTMEMORY_AGENT_SCOPE=isolated # optional; default "shared" ``` Two modes: | Mode | Tag writes | Filter recall | When to use | |------|------------|---------------|-------------| | `shared` (default) | yes | no | Cross-agent context with audit trail. Architect can see what developer noted, but every row records who said it. | | `isolated` | yes | yes | Strict separation. Architect never sees developer's observations / memories / sessions. | What gets tagged when `AGENT_ID` is set: `Session.agentId`, `RawObservation.agentId`, `CompressedObservation.agentId`, `Memory.agentId`. The role flows from `api::session::start` → `mem::observe` → `mem::compress` → KV. What gets filtered in isolated mode: `mem::smart-search`, `/agentmemory/memories`, `/agentmemory/observations`, `/agentmemory/sessions`. Each endpoint accepts `?agentId=` to override per-request, and `?agentId=*` to opt out of the env scope entirely. `/memories` also accepts `?includeOrphans=true` to surface pre-AGENT_ID memories whose `agentId` is undefined. Per-call override at the SDK / REST layer: every mutating endpoint (`/session/start`, `/remember`) accepts an `agentId` field in the request body that wins over the env. Useful for runtimes routing many roles through one server process. The MCP `memory_save` tool exposes the same `agentId` field, the standalone stdio server forwards both `agentId` and `project`, and saved memories carry `agentId` into the search index, so agent-scoped search covers memories as well as observations. When `AGENT_ID` is unset, memory remains unscoped (legacy behavior, no tags, no filters). ### Ports agentmemory + iii-engine bind four ports by default. If a restart fails with `port in use`, this table tells you which process to look for. | Port | Process | Purpose | Env override | |------|---------|---------|--------------| | `3111` | agentmemory | REST API + MCP HTTP + `/agentmemory/health` + `/agentmemory/livez` | `III_REST_PORT` | | `3112` | iii-engine | Internal streams worker (consumed by agentmemory + viewer) | `III_STREAM_PORT` (preferred) or legacy `III_STREAMS_PORT` | | `3113` | agentmemory | Real-time viewer (`http://localhost:3113`) | `III_VIEWER_PORT` or `AGENTMEMORY_VIEWER_URL` for the reported URL | | `49134` | iii-engine | WebSocket; workers register here, OTel telemetry flows over it | `III_ENGINE_PORT` or `III_ENGINE_URL` | `--port ` changes the REST anchor and derives streams `N+1`, viewer `N+2`, and engine WebSocket `N+46023` only where the corresponding explicit port or URL above is unset. It does not create an isolated lifecycle namespace. Use `--instance 1` for a second daemon; it uses anchor 3211, defaults to `3211/3212/3213/49234`, and receives a separate `instance-1` data and lifecycle directory. Instances 1 through 50 follow the same pattern. The pinned engine starts with `--no-update-check` (no update or security-advisory lookups against GitHub at boot) and with iii's anonymous usage telemetry off: agentmemory sets `III_TELEMETRY_ENABLED=false` for the engine it spawns unless you export the variable yourself, and the bundled compose file does the same. Stale-process cleanup when ports stay bound after a crashed run: ```bash # macOS / Linux — find whatever is on each port and kill it lsof -i :3111,3112,3113,49134 pkill -f agentmemory || true pkill -f 'iii ' || true # Windows netstat -ano | findstr ":3111 :3112 :3113 :49134" taskkill /F /PID ``` `agentmemory stop` reaps both the worker and the engine pidfile cleanly on graceful native shutdown. In Docker mode it flushes the native worker, stops the exact validated engine container, and preserves both the container and its `/data` mount for a lossless restart; the next start validates and resumes that same container. Docker-backed uninstall requires `agentmemory remove --keep-data`: it removes shared agentmemory-managed files while preserving the validated container, its data mount, and the lifecycle record needed to recover them. Destructive Docker data deletion is intentionally left to the operator after a backup. The CLI also refuses to adopt or signal Docker or VM port holders (Docker backend, vpnkit, colima) as the native engine unless `--force` is passed. The manual cleanup above is only for the post-crash case where neither pidfile is left behind. ### Config File Put agentmemory runtime configuration in `~/.agentmemory/.env` instead of exporting variables in every shell. If the viewer shows a setup hint like `export ANTHROPIC_API_KEY=...`, copy it into this file as `ANTHROPIC_API_KEY=...` without the `export` prefix, then restart agentmemory. Process environment variables still work and take precedence over values in the file. On Windows, the same file lives at `%USERPROFILE%\.agentmemory\.env`: ```powershell New-Item -ItemType Directory -Force $HOME\.agentmemory notepad $HOME\.agentmemory\.env ``` To test with a Claude Code Pro/Max subscription instead of an API key, opt in explicitly: ```env AGENTMEMORY_ALLOW_AGENT_SDK=true AGENTMEMORY_AUTO_COMPRESS=true ``` LLM-written observation compression requires both lines: access to an LLM provider (including this explicit subscription fallback) and `AGENTMEMORY_AUTO_COMPRESS=true`. A provider by itself leaves the default synthetic compression path in place. Consolidation (graph nodes, lessons, crystals) is on by default whenever an LLM provider is configured. Explicitly opt out with `CONSOLIDATION_ENABLED=false` if you want LLM-free operation. Graph extraction is a separate flag: ```env GRAPH_EXTRACTION_ENABLED=true # CONSOLIDATION_ENABLED=false # opt out of auto-consolidation ``` ### Environment Variables Create `~/.agentmemory/.env`: ```env # LLM provider (pick one — default is the no-op provider: no LLM calls) # ANTHROPIC_API_KEY=sk-ant-... # ANTHROPIC_BASE_URL=... # Optional: Anthropic-compatible proxy / Azure # GEMINI_API_KEY=... # OPENROUTER_API_KEY=... # MINIMAX_API_KEY=... # OPENAI_API_KEY=*** # NOTE: this same key auto-activates BOTH the # # OpenAI LLM provider (here) AND the OpenAI # # embedding provider (further below). Set # # OPENAI_API_KEY_FOR_LLM=false to scope it # # to embeddings only. # OPENAI_BASE_URL=https://api.openai.com # Optional: override for Azure / vLLM / LM Studio / proxies # # Azure: https://.openai.azure.com/openai/deployments/ # # Auto-detected from `.openai.azure.com` hostname; uses # # api-key header + api-version query param. # OPENAI_API_VERSION=2024-08-01-preview # Optional: Azure api-version query param # OPENAI_MODEL=gpt-5.6-luna # Optional: default model # OPENAI_TIMEOUT_MS=60000 # Optional: OpenAI-scoped alias for the outbound fetch # # timeout. Takes precedence over AGENTMEMORY_LLM_TIMEOUT_MS # # for back-compat with v0.9.17. New configs should # # prefer the global AGENTMEMORY_LLM_TIMEOUT_MS below. # OPENAI_REASONING_EFFORT=none # Optional: "low" | "medium" | "high" | "none" # # Honored only by OpenAI's reasoning models (o1, o3, # # gpt-*-reasoning) and providers that mirror that # # schema (Ollama Cloud thinking models). Standard # # chat models reject this field with 400. Set to # # "none" for thinking models that return reasoning # # but no content. # OPENAI_API_KEY_FOR_LLM=false # Optional: set to false to skip OpenAI auto-detection # # for LLM (useful if you only want OpenAI for embeddings) # Opt-in Claude-subscription fallback (spawns @anthropic-ai/claude-agent-sdk); # leave OFF unless you understand the Stop-hook recursion risk: # AGENTMEMORY_ALLOW_AGENT_SDK=true # Embedding provider (BM25-only when unset; local is an explicit opt-in) # EMBEDDING_PROVIDER=local # VOYAGE_API_KEY=... # OPENAI_API_KEY=sk-... # OPENAI_BASE_URL=https://api.openai.com # Override for Azure / vLLM / LM Studio / proxies # OPENAI_EMBEDDING_MODEL=text-embedding-3-small # OPENAI_EMBEDDING_DIMENSIONS=1536 # Required when the model is not in the known-models table # OPENAI_EMBEDDING_BASE_URL=https://... # Embeddings only; falls back to OPENAI_BASE_URL # OPENAI_EMBEDDING_API_KEY=sk-... # Embeddings only; wins over OPENAI_API_KEY when set # Outbound LLM / embedding timeout # AGENTMEMORY_LLM_TIMEOUT_MS=60000 # Default: 60 000 ms (60 s). Applies to every # raw-fetch provider (Gemini, OpenRouter, MiniMax, # OpenAI LLM, OpenAI/Cohere/Voyage/OpenRouter # embedding). For the OpenAI LLM path, the # OpenAI-scoped OPENAI_TIMEOUT_MS alias (above) # takes precedence when set, for back-compat # with v0.9.17. # Increase for slow networks or large batch calls; # decrease to fail-fast on rate-limit holds. # Search tuning # BM25_WEIGHT=0.4 # VECTOR_WEIGHT=0.6 # TOKEN_BUDGET=2000 # Auth (generated into ~/.agentmemory/secret on first start when unset) # AGENTMEMORY_SECRET=your-secret # VIEWER_ALLOWED_ORIGINS=https://memory.example.com # AGENTMEMORY_IMPORT_ROOT=~/projects # Ports (defaults: 3111 API, 3113 viewer) # III_REST_PORT=3111 # Engine usage telemetry (iii). Off unless you set it; true opts in. # III_TELEMETRY_ENABLED=false # Features # AGENTMEMORY_AUTO_COMPRESS=false # OFF by default. Requires an LLM # provider as well. When both are on, # every PostToolUse hook calls your # LLM provider to compress the # observation — expect significant # token spend on active sessions. # AGENTMEMORY_SLOTS=false # OFF by default. Editable pinned # memory slots — persona, # user_preferences, tool_guidelines, # project_context, guidance, # pending_items, session_patterns, # self_notes. Size-limited; agent # edits via memory_slot_* tools. # Pinned slots addressable for # SessionStart injection. # AGENTMEMORY_REFLECT=false # OFF by default. Requires SLOTS=on. # Stop hook fires mem::slot-reflect: # scans recent observations, auto- # appends TODOs to pending_items, # counts patterns in # session_patterns, records touched # files in project_context. Fire- # and-forget; does not block. # AGENTMEMORY_INJECT_CONTEXT=false # OFF by default. When on: # - SessionStart may inject ~1-2K # chars of project context into # the first turn of each session # (this is what actually reaches # the model — Claude Code treats # SessionStart stdout as context) # - PreToolUse fires /agentmemory/enrich # on every file-touching tool call # (resource cleanup, not a token # fix — PreToolUse stdout is debug # log only per Claude Code docs) # Observations are still captured via # PostToolUse regardless of this flag. # GRAPH_EXTRACTION_ENABLED=false # AGENTMEMORY_LLM_NOTHINK=1 # Local reasoning models only: ask the # model to skip its hidden thinking pass # during graph extraction. Faster runs; # relation quality can drop slightly. # CONSOLIDATION_ENABLED=false # on by default when an LLM provider is configured # LESSON_DECAY_ENABLED=true # OBSIDIAN_AUTO_EXPORT=false # AGENTMEMORY_EXPORT_ROOT=~/.agentmemory # CLAUDE_MEMORY_BRIDGE=false # SNAPSHOT_ENABLED=false # Storage and durability # AGENTMEMORY_STATE_BACKEND=file # file (default) or redis; see "Storage backend" below # AGENTMEMORY_REDIS_URL=redis://localhost:6379 # Required with redis, plain redis:// only # AGENTMEMORY_STATE_SAVE_INTERVAL_MS=2000 # How often the engine writes file state to disk. # A hard kill loses at most this window. # AGENTMEMORY_INDEX_SAVE_INTERVAL_MS=600000 # Minimum time between search index saves; # shutdown and deletes still save at once. # AGENTMEMORY_GRAPH_COMPACT_ON_BOOT=true # One-time background trim of oversized graph # provenance; false skips it # Sessions # AGENTMEMORY_SESSION_SWEEP_ENABLED=true # Hourly sweep marks sessions left active past # the threshold as abandoned. Deletes nothing; # new activity makes the session active again. # AGENTMEMORY_SESSION_SWEEP_STALE_HOURS=24 # Capture filters (hooks) # AGENTMEMORY_CAPTURE_ALLOW= # Comma or space list of tool names or globs; # when set, only these tools are captured # AGENTMEMORY_CAPTURE_DENY= # Extra names or globs to skip, added to the # defaults: memory_*, toolsearch, # listmcpresources, fetchmcpresource # AGENTMEMORY_CAPTURE_OUTPUT_MAX=8000 # Max characters of tool output per observation # AGENTMEMORY_PRE_COMPACT_BUDGET=1500 # Token budget for PreCompact context; 0 disables # Audit log # AGENTMEMORY_AUDIT_RETENTION_MONTHS=0 # Drop month scopes older than N months; 0 keeps all # AGENTMEMORY_AUDIT_INDEX_PERSIST=false # 1 or true records index migration and cleanup # rows (debugging only) # Team # TEAM_ID= # USER_ID= # TEAM_MODE=private # Tool visibility: "all" (54 tools, default) or "core" (8 tools, lean) # AGENTMEMORY_TOOLS=core ``` ---

API

138 endpoints on port `3111`. The REST API binds to `127.0.0.1` by default. Protected endpoints require `Authorization: Bearer `, and mesh sync endpoints require an explicitly set `AGENTMEMORY_SECRET` on both peers. **Authentication is on by default.** When `AGENTMEMORY_SECRET` is not set (in the shell or in `~/.agentmemory/.env`), the server generates a random secret on first start and stores it in `~/.agentmemory/secret` with mode `0600`. Every bundled client reads it from there when it talks to a local server: the CLI, the viewer, the hooks under `plugin/scripts`, the MCP server and the `@agentmemory/mcp` shim, the configs written by `agentmemory connect`, and the bundled OpenCode, Pi, OpenClaw, Hermes and filesystem-watcher integrations. The stored secret is only sent to loopback URLs (`localhost`, `127.0.0.0/8`, `::1`). An explicit `AGENTMEMORY_SECRET` always wins, and remote clients still need it set. Docker and the `deploy/` entrypoints already generate and export their own secret. To call the API by hand: ```bash curl -H "Authorization: Bearer $(cat ~/.agentmemory/secret)" http://localhost:3111/agentmemory/health ``` **Request rules for writes.** `POST`, `PUT`, `PATCH` and `DELETE` requests to the REST API and the viewer must send `Content-Type: application/json` (a `charset` parameter is fine) whenever they carry a body, and an `Origin` header, when present, must be a loopback origin for the configured REST or viewer port or be listed in `VIEWER_ALLOWED_ORIGINS` (comma-separated, e.g. `https://memory.example.com`). Clients that send no `Origin` header (CLI, hooks, MCP, curl, server-to-server) are unaffected. The viewer also accepts its own origin. **File paths.** Endpoints that read or write files (`/compress-file`, `/replay/import-jsonl`, `/graph/import-graphify`) only accept paths under `~/.agentmemory`, the instance data directory, or a directory listed in `AGENTMEMORY_IMPORT_ROOT` (separate several with `:`, or `;` on Windows). `/replay/import-jsonl` also accepts its default `~/.claude/projects`. `/obsidian/export` stays inside `AGENTMEMORY_EXPORT_ROOT` and `/migrate` inside `~/.agentmemory`. Symlinks are resolved before every check. **Secret scrubbing.** API keys, bearer tokens, PEM private key blocks and credentials embedded in URLs (`scheme://user:password@host`) are redacted before text is stored, on every write path: observations, remember, evolve, slots, lessons, actions, sketches, signals, checkpoints, imports, jsonl replay, mesh sync, team shares, compression and summary output, crystals and graph nodes.
Key endpoints | Method | Path | Description | |--------|------|-------------| | `GET` | `/agentmemory/health` | Health check (always public) | | `GET` | `/agentmemory/status` | What is wrong and how to fix it (HTML for browsers, JSON otherwise) | | `GET` | `/agentmemory/viewer/snapshot` | Everything the viewer shows, in one response | | `POST` | `/agentmemory/session/start` | Start session + get context | | `POST` | `/agentmemory/session/end` | End session | | `POST` | `/agentmemory/observe` | Capture observation (see capture delivery below) | | `GET` | `/agentmemory/capture` | Capture inbox, dead letters and offline spool | | `POST` | `/agentmemory/capture/retry` | Retry dead-letter captures | | `POST` | `/agentmemory/capture/drain` | Send the local offline spool now | | `POST` | `/agentmemory/smart-search` | Hybrid search | | `POST` | `/agentmemory/context` | Generate context | | `POST` | `/agentmemory/remember` | Save to long-term memory | | `POST` | `/agentmemory/forget` | Delete observations | | `POST` | `/agentmemory/enrich` | File context + memories + bugs | | `GET` | `/agentmemory/profile` | Project profile | | `GET` | `/agentmemory/export` | Export all data | | `POST` | `/agentmemory/import` | Import from JSON | | `POST` | `/agentmemory/graph/query` | Knowledge graph query | | `POST` | `/agentmemory/graph/compact` | Trim oversized graph provenance | | `POST` | `/agentmemory/team/share` | Share with team | | `GET` | `/agentmemory/audit` | Audit trail | Full endpoint list: [`src/triggers/api.ts`](src/triggers/api.ts)
**Capture delivery.** Hooks send each observation once to `POST /agentmemory/observe` with an `eventId`. It is the host's own id for the call when the payload has one (for example Claude Code's `tool_use_id`), otherwise a hash of the session, hook type, tool name, input, output and host timestamp. The server writes the event to a capture inbox in the state store, stores the observation, then removes the inbox entry. The status code says what happened: | Status | `status` field | Meaning | |---|---|---| | `201` | `accepted` | Stored. `observationId` is the new observation. | | `202` | `accepted` (`state: "retrying"`) | Accepted, but storing failed. The server retries it, also after a restart. | | `200` | `duplicate` | This `eventId` was already accepted. `observationId` is the existing observation; nothing new is stored. | | `400` / `422` | `rejected` | Invalid payload, or storing failed for good (the event is kept as a dead letter). | | `503` | `rejected` (`retryable: true`) | The inbox is full (`AGENTMEMORY_CAPTURE_INBOX_MAX`). Hooks spool the event and send it later. | Failed events are retried every `AGENTMEMORY_CAPTURE_RETRY_INTERVAL_MS` (10 s) with doubling backoff, up to `AGENTMEMORY_CAPTURE_MAX_ATTEMPTS` (5). Events that still fail stay in the inbox as dead letters, are listed on `/agentmemory/status` and the viewer Health page, and can be retried with `POST /agentmemory/capture/retry` (`{"eventId": "..."}` or `{"all": true}`). Accepted event ids are remembered for `AGENTMEMORY_CAPTURE_DEDUP_HOURS` (168 hours, at most `AGENTMEMORY_CAPTURE_EVENTS_MAX` ids), so a hook replayed after a timeout or a restart is stored once, while two separate tool calls with their own host ids are stored twice even when their content is identical. When an observation is deleted (forget, session delete, eviction, auto-forget or an import that replaces the store), its event is marked as deleted before the observation is removed, so a replay of that event inside the same window is answered as a duplicate and stores nothing. The state store writes to disk every 2 seconds, so an answered event can still be only in memory for a moment. To cover that, every `2xx` answer also carries the server's `bootId` (new at every start), `acceptedAt` and `durableAfterMs` (the save interval plus 1.5 s on the file store, 1.5 s on redis, where persistence is the operator's setting). Hooks keep the event in the local spool until that window has passed and delete it on a later call without another request. If the `bootId` has changed by then, the server restarted, so the hook sends the event again with the same `eventId`; an event that did reach the disk is not stored twice. The server also sends such events itself at start and every retry interval, so a restart loses nothing even when no hook runs afterwards. Older hooks ignore the extra fields, and new hooks against an older server discard the event on `2xx` as before. When the server is down, does not answer in time or returns a 5xx, the hook appends the observation to a local spool file, `/capture-spool/-.jsonl` (override the folder with `AGENTMEMORY_CAPTURE_SPOOL_DIR`). The file is private to your user (mode 600), secrets are redacted the same way the server redacts them, it holds at most `AGENTMEMORY_CAPTURE_SPOOL_MAX_BYTES` (5 MiB) and drops entries older than `AGENTMEMORY_CAPTURE_SPOOL_MAX_AGE_HOURS` (168). When it is full, new entries are dropped and counted, and `/agentmemory/status` reports it. The hook still exits 0 within its time limit and adds no request when the server is healthy. The spool is sent at the next start and by the first hook that reaches the server again, in a background process so the agent does not wait. Event ids make this safe: an observation that did arrive before a timeout is not stored twice. `npx @agentmemory/agentmemory capture` shows the spool and the server inbox, `--drain` sends the spool now, and `GET /agentmemory/capture` returns the same as JSON. Set `AGENTMEMORY_CAPTURE_SPOOL=false` to turn the spool off. **Compacting graph provenance.** Each knowledge graph node and edge keeps the ids of the newest 32 observations it came from. Stores written before that cap can hold thousands of ids per hot node, which makes graph search and the viewer slow or drops the worker. agentmemory fixes this by itself: on the first start after upgrading it trims every node, edge, superseded edge (the temporal graph history) and the cached snapshot to the cap in the background, in small slices with a pause between them, so search, capture and the viewer keep working. It saves its progress, resumes after a restart and never runs again once it has finished. `/agentmemory/status` and the viewer Health page show it as pending, running (with the current scope and position), done or failed. Set `AGENTMEMORY_GRAPH_COMPACT_ON_BOOT=false` to turn it off. To run it by hand, call `POST /agentmemory/graph/compact`. It walks the name and edge-key indexes instead of listing every node and edge, and is safe to re-run. When it trims ids it writes a `graph_compact` audit entry. ```bash curl -X POST http://localhost:3111/agentmemory/graph/compact -H "Content-Type: application/json" -d '{}' ``` On a large store, or when the call returns 504, run it in slices. Send `scope` (`nodes`, `edges` or `history`), `offset` and `limit`, then call again with the returned `nextOffset` until it is `null`. Do this for `nodes`, `edges` and `history`, and finish with one `{"scope":"snapshot"}` call, because a sliced run does not touch the cached snapshot. ```bash curl -X POST http://localhost:3111/agentmemory/graph/compact -H "Content-Type: application/json" -d '{"scope":"nodes","offset":0,"limit":200}' curl -X POST http://localhost:3111/agentmemory/graph/compact -H "Content-Type: application/json" -d '{"scope":"snapshot"}' ``` ---

Development

```bash npm run dev # Hot reload npm run build # Production build npm test # 2,500+ tests npm run test:integration # API tests (requires running services) ``` **Prerequisites:** Node.js >= 20 with npm/npx; [iii-engine](https://iii.dev/docs) v0.22.1 or Docker. The macOS/Linux automatic engine install also requires `curl`, a POSIX `sh`, and `tar`; native Windows uses the manual pinned `iii.exe`, WSL2, or Docker Desktop.

License

[Apache-2.0](LICENSE)