A temporal knowledge engine for markdown vaults.
It indexes, links, remembers, forgets, and dreams — locally, over files you own.
Quickstart · What it does · CLI · Agent memory · Benchmarks · How it works · Research
---
[-k] [--since] [--until] [--tag] [--folder] [--json]` | hybrid retrieval with provenance | | `lore ask` | extractive answer: current facts + top passages (no LLM needed) | | `lore facts [--subject] [--predicate] [--as-of] [--as-known-at] [--history]` | query the fact store | | `lore timeline[--since] [--until]` | chronological history: fact changes merged with dated mentions | | `lore resume [--since]` | what changed since the last resume: notes, facts, supersessions | | `lore review [--threshold] [--limit]` | important-but-fading knowledge to revisit or archive | | `lore assert
[--valid-from]` | record a fact (journalled, supersedes) | | `lore invalidate ` | close the current fact in a slot | | `lore count [--predicate] [--group-by] [--since]` | aggregate over fact history | | `lore capture
` | append a timestamped line to `lore/inbox.md` | | `lore dream [--apply]` | consolidation pass + optional digest/review queue | | `lore watch` | reindex automatically as the vault changes | | `lore mark-used [anchor]` | reinforce a passage that proved useful | | `lore graph export --format json\|graphml\|dot` | export the graph | | `lore doctor` | health check: broken links, integrity, coverage | | `lore stats` | vault statistics and top entities | | `lore serve --mcp` | start the MCP server on stdio | ## Use it as agent memory (MCP) ```jsonc // Claude Code: .mcp.json (or claude_desktop_config.json) { "mcpServers": { "loreweave": { "command": "npx", "args": ["-y", "loreweave", "--vault", "/path/to/vault", "serve", "--mcp"] } } } ``` | Tool | What the agent gets | |---|---| | `lore_search` | hybrid retrieval with provenance and temporal filters | | `lore_context_pack` | one-call session context: relevant passages + current facts + recent changes | | `lore_read_note` | full text of a note by path | | `lore_assert_fact` / `lore_invalidate_fact` | write/close facts — journalled, superseding, never destructive | | `lore_query_facts` | point-in-time fact queries (`asOf`, `asKnownAt`, history) | | `lore_timeline` | an entity's merged fact + prose chronology | | `lore_resume` | exactly what changed since the agent last connected | | `lore_review` | important-but-fading passages worth resurfacing | | `lore_aggregate_facts` | count/group-by over fact history | | `lore_capture` | append a timestamped line to the inbox | | `lore_mark_used` | reinforcement signal: this passage actually helped | | `lore_propose_facts` | extraction candidates for the agent to review and assert | | `lore_dream_report` | the consolidation report (duplicates, contradictions, stale, orphans) | | `lore_index` | trigger a reindex | Facts asserted through MCP are written back to `lore/journal/YYYY-MM-DD.md` as readable markdown, so an agent's memory is something you can open, edit, and `git diff`: ```markdown - [fact] Ledger Format :: status :: final {valid_from=2026-08-01, confidence=0.9, source=stated} ``` Delete `.lore/` and reindex — every fact and edge is reconstructed from those files. **Compared to hosted memory services** (Mem0, Zep): those run LLMs at write time to extract and summarize into their own store; loreweave runs no model in the core, keeps memory in your files under your version control, and makes every retrieval reproducible. The trade: they do abstractive summarization, this engine deliberately does not. Retrieval quality against their published benchmarks is below — with the caveats stated, because most published agent-memory numbers measure end-to-end QA with an LLM, which is a different quantity than retrieval. ## Benchmarks These are third-party benchmarks with relevance labels nobody here chose. All loreweave numbers are **retrieval** metrics — it finds the evidence, it does not write the answer — so they are not comparable to the end-to-end QA accuracy quoted by systems that put a language model after retrieval. Reproduce any number: [`docs/benchmarks.md`](docs/benchmarks.md). Field comparison with the traps called out: [`docs/scoreboard.md`](docs/scoreboard.md). **LongMemEval_S** (ICLR 2025) — 500 questions, each with ~50 sessions of chat history; find the sessions holding the evidence. Session-level recall, all 500 questions: | configuration | R@1 | R@5 | R@10 | |---|---|---|---| | model-free (no embeddings, no network) | 0.552 | 0.899 | 0.943 | | + embeddings & 8-turn chunking | **0.590** | **0.959** | **0.983** | [A third-party benchmark][lme] of the same task reports BM25 alone at **86.2%** R@5, BM25+vector hybrid at **95.2%**, and a vector-only system at **96.6%**. One caveat before the comparison: their metric scores a question 1 if *any* gold session is retrieved, ours scores the *fraction* of gold sessions found — ours is the stricter definition, so treat cross-system gaps of under a point as noise. With that stated: **at 95.9% loreweave is past the hybrid and just short of the vector-only system**, with a local model, no network, and a lexical index it can fall back to. [lme]: https://github.com/rohitg00/agentmemory/blob/main/benchmark/LONGMEMEVAL.md **BEIR / SciFact** — 5 183 scientific abstracts, 300 claims, scored by nDCG@10 (ranking quality: rewards putting relevant documents nearer the top): | configuration | nDCG@10 | Recall@10 | |---|---|---| | BM25 baseline (BEIR paper, Anserini) | 0.665 | — | | loreweave, model-free | 0.681 | 0.817 | | loreweave + nomic-embed-text | 0.727 | 0.865 | | loreweave + mxbai-embed-large | **0.742** | **0.884** | **LoCoMo** — 10 long conversations, 1 982 evidence-labelled questions, turn-level recall. No comparable third-party *retrieval* number exists (published LoCoMo results are LLM QA accuracy), so these are offered as a target rather than a comparison: | | R@1 | R@5 | R@10 | R@20 | |---|---|---|---|---| | model-free | 0.337 | 0.538 | 0.610 | 0.656 | | + embeddings | 0.318 | 0.532 | **0.627** | **0.705** | **What the numbers say, plainly.** Recall is strong and rank-1 is the weakness — LongMemEval R@10 0.983 vs R@1 0.590 — which is why figures here are quoted at R@5 and why an agent consuming these results should read a top-5 list, not trust rank 1. The single biggest quality lever measured is the embedding model itself: mxbai over nomic is worth more than every downstream tuning combined on document (BEIR) and session (LongMemEval) retrieval — and measures flat on LoCoMo's single-turn passages, so it is a scale effect, not magic. A cross-encoder reranker is available (`rerank` in config) but earns its keep only for rank-1 consumers without embeddings — stacked on embeddings it *loses* recall on both public benchmarks, so leave it off unless that trade is yours. Full analysis, including the failed experiments and the defect that benchmarking caught: [`docs/evaluation.md`](docs/evaluation.md). Loreweave also ships an internal regression benchmark (`npm run eval`) over three purpose-built vaults — including a temporal-perturbation test where the shipped config scores **100% consistency vs BM25's 0%**, gated in CI on every push. Internal corpora are good for regression and worthless as proof, so the details live in [`docs/evaluation.md`](docs/evaluation.md) rather than here. ## Scale Measured, like the quality numbers — `npm run scale` reproduces this on your own machine (synthetic vaults, 3 blocks per note, dense interlinking): | notes | blocks | entities | edges | full index | incremental | search p50 | p95 | dream | heap | |---|---|---|---|---|---|---|---|---|---| | 1 000 | 3 000 | 2 766 | 17 k | 1.2 s | 32 ms | 3 ms | 4 ms | 0.2 s | 63 MB | | 5 000 | 15 000 | 13 766 | 85 k | 6.1 s | 165 ms | 9 ms | 12 ms | 1.0 s | 145 MB | | 20 000 | 60 000 | 55 016 | 339 k | 22.5 s | 656 ms | 38 ms | 53 ms | 4.8 s | 333 MB | Full index scales at **0.93× per note** from 5 k to 20 k — linear or better, no superlinear step hiding in the middle. "Incremental" is one changed note, which is what `lore watch` actually does all day. Everything here is one process, one SQLite file, no daemon — and search at 3-38 ms p50 is fast enough to sit inside an agent loop. ## How it works ``` vault/*.md ──parse──▶ notes · blocks · wiki-links · tags · entities │ (incremental: mtime + content hash) ▼ SQLite .lore/index.db ── disposable cache, rebuildable │ ┌───────────────────┼────────────────────┐ ▼ ▼ ▼ graph (CSR) retrieval facts blocks ∪ entities BM25 + dense + PPR bitemporal, supersession, 2-iteration PPR → weighted RRF deterministic freshness, α = 0.5 → FSRS boosts aggregates └─────────┬─────────┴──────────┬─────────┘ ▼ ▼ dream (idle-time) CLI · MCP ``` Design rules the code enforces: - **Files win.** User markdown is never mutated. The engine only appends, and only under `lore/`. - **Invariants in code, not prompts.** Schema, migrations, graph construction, and supersession are typed, versioned, and tested — no LLM re-specifies them at runtime. - **No LLM required anywhere in the core.** Indexing and retrieval use zero tokens. Language models are consumers of this engine, not dependencies of it. - **Everything is re-derivable.** A full rebuild reproduces byte-identical derived state (there's a test for that). ## Research lineage Every significant choice traces to 2024-2026 literature; the survey lives in [`docs/research/`](docs/research/). | Choice | Source | |---|---| | Dense-sparse fusion + PPR with dense reset probabilities | HippoRAG 2 (ICML 2025), [2502.14802](https://arxiv.org/abs/2502.14802) | | Shallow 2-iteration PPR, heterogeneous nodes | NodeRAG (2025), [2504.11544](https://arxiv.org/abs/2504.11544) | | Relation-free graph — no LLM triple extraction | LinearRAG (ICLR 2026), [2510.10114](https://arxiv.org/abs/2510.10114) | | No index-time community summarization | LazyGraphRAG (Microsoft, 2024) — same quality at 0.1% index cost | | Route/fuse instead of graph-everything | GraphRAG-Bench (ICLR 2026), [2506.05690](https://arxiv.org/abs/2506.05690) | | Bitemporal facts, invalidate-never-delete | Zep/Graphiti (2025), [2501.13956](https://arxiv.org/abs/2501.13956) | | Power-law forgetting, use-gated reinforcement | FSRS; RMM (ACL 2025), [2503.08026](https://arxiv.org/abs/2503.08026) | | Consolidation as idle-time work | Sleep-time compute (Letta, 2025), [2504.13171](https://arxiv.org/abs/2504.13171) | | Never let an LLM rewrite whole memory files | ACE (2025), [2510.04618](https://arxiv.org/abs/2510.04618) | | Computable facts for aggregation | User as Code (2026), [2606.16707](https://arxiv.org/abs/2606.16707) | | Fine-grained indexing + fact-augmented keys | LongMemEval (ICLR 2025), [2410.10813](https://arxiv.org/abs/2410.10813) | ## Library use ```ts import { openContext, indexVault, search, assertFact, queryFacts, dream } from 'loreweave'; const ctx = openContext('/path/to/vault'); await indexVault(ctx.store, ctx.root); const hits = await search(ctx, 'streaming compaction', { k: 5 }); assertFact(ctx, { subject: 'Atlas', predicate: 'status', object: 'shipped', validFrom: '2026-08-01' }); const asOfMarch = queryFacts(ctx.store, { subject: 'Atlas', asOf: '2026-03-01' }); const report = dream(ctx); ctx.close(); ``` ## Development ```bash npm install npm test # 440 tests npm run eval # retrieval benchmark vs BM25 baseline npm run typecheck npm run build ``` Requires Node ≥ 20. Single native dependency (`better-sqlite3`). Tested in CI on Linux, macOS and Windows across Node 20 and 22. ## License MIT © Ambuj Upadhyay