--- name: wiki-export description: > Export the Obsidian wiki graph to JSON, GraphML, Neo4j Cypher, Postgres/SQL, HTML, or OKF bundles. Use when transferring or visualizing wiki data in external tools; wiki-import handles the reverse direction. --- # Wiki Export — Knowledge Graph Export You are exporting the wiki's wikilink graph to structured formats so it can be used in external tools (Gephi, Neo4j, custom scripts, browser visualization). ## Before You Start 1. **Resolve config** — follow the Config Resolution Protocol in `llm-wiki/SKILL.md` (inline `@name` override → walk up CWD for `.env` → global config → prompt setup). This gives `OBSIDIAN_VAULT_PATH` 2. Confirm the vault has pages to export — if fewer than 5 pages exist, warn the user and stop ## Project Filter (optional) If the user's invocation includes a project name — e.g. `/wiki-export prismor`, `"export the prismor project"`, `"export project:security"` — activate **project filter mode**: 1. **Extract the project name** from the argument or phrase. Normalise: lowercase, strip the word "project". 2. Keep only pages where **either** condition holds: - The page `id` starts with `projects//` (path-based match) - The page's `tags` array contains `` (tag-based match) 3. Drop any edge where either endpoint was excluded. 4. Note the filter in the summary: `(filtered: project: — X of Y pages)` 5. Set `graph.graph.filter = "project:"` in the JSON output. If both a project filter and a visibility filter are active, apply both (project filter first, then visibility filter on the remaining set). ## Visibility Filter (optional) By default, **all pages are exported** regardless of visibility tags. This preserves existing behavior. If the user requests a filtered export — phrases like **"public export"**, **"user-facing export"**, **"exclude internal"**, **"no internal pages"** — activate **visibility filtered mode**: - Build a **blocked tag set**: `{visibility/internal, visibility/pii}` - Skip any page whose frontmatter tags contain a blocked tag when building the node list - Skip any edge where either endpoint was excluded - Note the filter in the summary: `(filtered: visibility/internal, visibility/pii excluded)` Pages with no `visibility/` tag, or tagged `visibility/public`, are always included. ## Step 1: Build the Node and Edge Lists Glob all `.md` files in the vault (excluding `_archives/`, `_raw/`, `_readouts/`, `.obsidian/`, `index.md`, `log.md`, `_insights.md`). Apply any active filters (project and/or visibility) after collecting the full file list. For each page, extract from frontmatter: - `id` — relative path from vault root, without `.md` extension (e.g. `concepts/transformers`) - `label` — `title` field from frontmatter, or filename if missing - `category` — directory prefix (`concepts`, `entities`, `skills`, `references`, `synthesis`, `projects`, or `journal`) - `tags` — array from frontmatter tags field - `summary` — frontmatter `summary` field if present This is your **node list**. For each page, Grep the body for `\[\[.*?\]\]` to extract all wikilinks: - Parse each `[[target]]` or `[[target|display]]` — use the target part only - Resolve the target to a node id (normalize: lowercase, spaces→hyphens, strip `.md`) - Skip links that point outside the node list (broken links) - Each resolved link becomes an edge: `{source: page_id, target: linked_id, relation: "wikilink", confidence: "EXTRACTED"}` - If the linking sentence ends with `^[inferred]` or `^[ambiguous]`, override `confidence` accordingly **Typed edge enrichment:** After building the wikilink edge list, read each page's `relationships:` frontmatter block. For each `{target, type}` entry: - The `target` YAML value is a quoted wikilink string such as `"[[concepts/lstm]]"`. Strip the surrounding `[[` and `]]` characters, then apply the same normalization (lowercase, spaces→hyphens, strip `.md`) to get the node id. - Skip entries whose resolved target is not in the node list (broken link) - If an edge for this `(source, target)` pair already exists, override its `relation` field with the typed value (e.g., `"contradicts"`) and set `typed: true` - If no edge exists yet for this pair, add one: `{source: page_id, target: target_id, relation: , confidence: "EXTRACTED", typed: true}` This means `relation: "wikilink"` is the default for plain untyped links; a `relationships:` entry promotes it to a named semantic type. Edges that originated from both a body wikilink and a `relationships:` entry keep a single record — the typed version wins. This is your **edge list**. ## Step 2: Assign Community IDs Group pages into communities by tag clustering: - Pages sharing the same dominant tag belong to the same community - Dominant tag = the first tag in the page's frontmatter tags array - Pages with no tags get community id `null` - Number communities starting from 0, ordered by size descending (largest community = 0) This enables community-based coloring in the HTML visualization and tools like Gephi. ## Step 3: Write the Output Files Create `wiki-export/` at the vault root if it doesn't exist. Write all five files: --- ### 3a. `graph.json` NetworkX node_link format — standard for graph tools and scripts: ```json { "directed": false, "multigraph": false, "graph": { "exported_at": "", "vault": "", "total_nodes": N, "total_edges": M }, "nodes": [ { "id": "concepts/transformers", "label": "Transformer Architecture", "category": "concepts", "tags": ["ml", "architecture"], "summary": "The attention-based architecture introduced in Attention Is All You Need.", "community": 0 } ], "links": [ { "source": "concepts/transformers", "target": "entities/vaswani", "relation": "wikilink", "confidence": "EXTRACTED" }, { "source": "concepts/transformers", "target": "concepts/lstm", "relation": "contradicts", "confidence": "EXTRACTED", "typed": true } ] } ``` --- ### 3b. `graph.graphml` GraphML XML format — loadable in Gephi, yEd, and Cytoscape: ```xml Transformer Architecture concepts ml, architecture 0 wikilink EXTRACTED contradicts contradicts EXTRACTED ``` Write one `` per page and one `` per link. For typed edges (those where `typed: true` in the edge list), emit both `` with the semantic type value **and** `` with the same value — this keeps `relation` readable for tools that already consume it while letting type-aware tools filter on the dedicated `type` key. Untyped wikilinks omit the `` element entirely. --- ### 3c. `cypher.txt` Neo4j Cypher `MERGE` statements — paste into Neo4j Browser or run with `cypher-shell`: ```cypher // Wiki knowledge graph export — // Load with: cypher-shell -u neo4j -p password < cypher.txt // Nodes MERGE (n:Page {id: "concepts/transformers"}) SET n.label = "Transformer Architecture", n.category = "concepts", n.tags = ["ml","architecture"], n.community = 0; MERGE (n:Page {id: "entities/vaswani"}) SET n.label = "Ashish Vaswani", n.category = "entities", n.tags = ["person","ml"], n.community = 0; MERGE (n:Page {id: "concepts/lstm"}) SET n.label = "LSTM", n.category = "concepts", n.tags = ["ml","rnn"], n.community = 0; // Relationships // Untyped wikilinks use [:WIKILINK] MATCH (a:Page {id: "concepts/transformers"}), (b:Page {id: "entities/vaswani"}) MERGE (a)-[:WIKILINK {relation: "wikilink", confidence: "EXTRACTED"}]->(b); // Typed edges use the relationship type as the label (UPPERCASE) MATCH (a:Page {id: "concepts/transformers"}), (b:Page {id: "concepts/lstm"}) MERGE (a)-[:CONTRADICTS {relation: "contradicts", confidence: "EXTRACTED"}]->(b); ``` Write one `MERGE` node statement per page, then one `MATCH`/`MERGE` relationship statement per edge. For typed edges, use the `type` value uppercased as the Cypher relationship label (e.g., `contradicts` → `[:CONTRADICTS]`, `derived_from` → `[:DERIVED_FROM]`). Untyped wikilinks always use `[:WIKILINK]`. --- ### 3d. `postgres.sql` Plain SQL — loadable into any Postgres database (local, Supabase, RDS, Neon, …) with `psql -f postgres.sql` or a migration runner. Two tables: `wiki_pages` (nodes) and `wiki_edges` (links), with `ON CONFLICT` upserts so re-running the export is safe and idempotent, mirroring the `MERGE` semantics of `cypher.txt`. ```sql -- Wiki knowledge graph export — -- Load with: psql -d yourdb -f postgres.sql CREATE TABLE IF NOT EXISTS wiki_pages ( id TEXT PRIMARY KEY, label TEXT NOT NULL, category TEXT, tags JSONB NOT NULL DEFAULT '[]'::jsonb, summary TEXT, community INT ); CREATE TABLE IF NOT EXISTS wiki_edges ( source TEXT NOT NULL REFERENCES wiki_pages(id) ON DELETE CASCADE, target TEXT NOT NULL REFERENCES wiki_pages(id) ON DELETE CASCADE, relation TEXT NOT NULL DEFAULT 'wikilink', confidence TEXT, typed BOOLEAN NOT NULL DEFAULT false, PRIMARY KEY (source, target, relation) ); CREATE INDEX IF NOT EXISTS wiki_edges_source_idx ON wiki_edges(source); CREATE INDEX IF NOT EXISTS wiki_edges_target_idx ON wiki_edges(target); -- Nodes INSERT INTO wiki_pages (id, label, category, tags, summary, community) VALUES ('concepts/transformers', 'Transformer Architecture', 'concepts', '["ml","architecture"]'::jsonb, 'The attention-based architecture introduced in Attention Is All You Need.', 0) ON CONFLICT (id) DO UPDATE SET label = EXCLUDED.label, category = EXCLUDED.category, tags = EXCLUDED.tags, summary = EXCLUDED.summary, community = EXCLUDED.community; -- Edges -- Untyped wikilink INSERT INTO wiki_edges (source, target, relation, confidence, typed) VALUES ('concepts/transformers', 'entities/vaswani', 'wikilink', 'EXTRACTED', false) ON CONFLICT (source, target, relation) DO UPDATE SET confidence = EXCLUDED.confidence, typed = EXCLUDED.typed; -- Typed edge from relationships: block INSERT INTO wiki_edges (source, target, relation, confidence, typed) VALUES ('concepts/transformers', 'concepts/lstm', 'contradicts', 'EXTRACTED', true) ON CONFLICT (source, target, relation) DO UPDATE SET confidence = EXCLUDED.confidence, typed = EXCLUDED.typed; ``` Write one `INSERT ... ON CONFLICT (id) DO UPDATE` statement per page (values escaped: single quotes doubled, `tags` serialized as a JSON array literal cast to `jsonb`), then one `INSERT ... ON CONFLICT (source, target, relation) DO UPDATE` per edge. `relation` stays lowercase here (unlike the uppercased Cypher relationship label) since it's a plain column value, not a schema identifier — this keeps it directly filterable with `WHERE relation = 'contradicts'` or joinable without case-folding. `typed` is `true` only for edges promoted by a `relationships:` frontmatter entry; plain wikilinks stay `false`. Skip pages whose `id` collides only after the `ON CONFLICT` clause fires from a prior run — do not attempt to deduplicate synthetic multi-edges (e.g. same source/target with both a `wikilink` and a typed relation) since the composite primary key `(source, target, relation)` already keeps them as distinct rows, matching the "typed version wins" merge behavior of `graph.json`/`graph.graphml` at the query layer (`SELECT * FROM wiki_edges WHERE source=$1 AND target=$2 ORDER BY typed DESC LIMIT 1`). --- ### 3e. `graph.html` A self-contained interactive visualization using the vis.js CDN (no local dependencies). The user opens this file in any browser — no server needed. Build the HTML file by: 1. Generating a JSON array of node objects for vis.js: ```js {id: "concepts/transformers", label: "Transformer Architecture", color: {background: "#4E79A7"}, size: , title: "concepts | #ml #architecture", community: 0} ``` - Color by community (cycle through: `#4E79A7`, `#F28E2B`, `#E15759`, `#76B7B2`, `#59A14F`, `#EDC948`, `#B07AA1`, `#FF9DA7`, `#9C755F`, `#BAB0AC`) - Size by degree (incoming + outgoing link count): `size = degree * 3 + 8`, capped at 60 - `title` = tooltip text shown on hover: category, tags, summary (if available) 2. Generating a JSON array of edge objects for vis.js: ```js // Untyped wikilink {from: "concepts/transformers", to: "entities/vaswani", dashes: false, width: 1, color: {color: "#666", opacity: 0.6}, title: "wikilink"} // Typed edge {from: "concepts/transformers", to: "concepts/lstm", dashes: false, width: 2, color: {color: "#E15759", opacity: 0.8}, label: "contradicts", font: {size: 9, color: "#ccc"}, title: "contradicts"} ``` - `dashes: true` for INFERRED edges - `dashes: [4,8]` for AMBIGUOUS edges - **Typed edges** (`typed: true`): set `width: 2`, add a `label` field showing the type, and apply a type-specific color: | Type | Edge color | |---|---| | `extends` | `#59A14F` (green) | | `implements` | `#4E79A7` (blue) | | `contradicts` | `#E15759` (red) | | `derived_from` | `#F28E2B` (orange) | | `uses` | `#76B7B2` (teal) | | `replaces` | `#B07AA1` (purple) | | `related_to` | `#BAB0AC` (grey — same as untyped) | Untyped `wikilink` edges keep the existing `#666` grey color and no label. 3. Writing the full HTML file: ```html Wiki Knowledge Graph
``` Replace `/* NODES_JSON */` and `/* EDGES_JSON */` with the actual JSON arrays you generated in step 1. --- ## Step 3.5: OKF Bundle Export (optional) Run this step **only** when the user asks for OKF / a markdown bundle (phrases like "export to OKF", "OKF bundle", "open knowledge format", "export as markdown bundle"). It is additive — the five graph files above are always produced; this writes an extra full-fidelity markdown bundle. The five graph files are a *lossy* projection (graph skeleton only). An **OKF bundle is the actual page bodies**, so an export→`wiki-import` round-trip through OKF preserves full content, and the bundle drops straight into MkDocs, Notion, Hugo, GitHub's renderer, or any [Open Knowledge Format](https://github.com/GoogleCloudPlatform/knowledge-catalog/blob/main/okf/SPEC.md) consumer. ### Canonical frontmatter mapping (obsidian-wiki ⇄ OKF) This table is the single source of truth for the mapping; `wiki-import` references it for the reverse direction. | OKF key | ← export from | → import to | Notes | |---------------|-------------------------|--------------------------|-------| | `type` (required) | `category`, title-cased | `category` (lower-cased) | `concepts`→`Concept`, `entities`→`Entity`, `skills`→`Skill`, `references`→`Reference`, `synthesis`→`Synthesis`, `projects`→`Project`, `journal`→`Journal`. OKF requires `type`; consumers tolerate any string. | | `title` | `title` | `title` | Verbatim. | | `description` | `summary` | `summary` | Our one-line `summary:` is exactly OKF's `description` (used in `index.md` entries). | | `tags` | `tags` | `tags` | Verbatim list. `visibility/*` system tags pass through unchanged. | | `generated` | `updated` → `generated.at`; `by` is the producer | `updated` ← `generated.at` | Write `generated: { by: obsidian-wiki/, at: }`, taking `` from `obsidian-wiki --version` (plain `obsidian-wiki` if the CLI isn't installed). OKF §5 requires every timestamp to be a datetime with an explicit offset, so a date-only `updated: 2026-04-12` becomes `2026-04-12T00:00:00Z`. Replaces v0.1's `timestamp`, which is no longer written. | | `sources` | `sources` (list of strings) | `sources` (list of strings) | OKF `sources` is a list of objects with a required `resource`, so each native string becomes `- resource: `. Opaque strings like `conversation:2026-04-12` are valid: §5.1 allows scope descriptors a consumer can't follow. Emit nothing else per entry, so import recovers the exact strings. | | `status` | `lifecycle` | — (native `lifecycle` is preserved) | `draft` → `draft`; `reviewed`, `verified`, `disputed` → `stable`; `archived` → `deprecated`. Omit `status` when `lifecycle` is absent (absent means `stable` in OKF). | | `verified` | `_meta/trust-ledger.json` entry for the page | — (preserved verbatim) | Only when the page has a ledger entry: `verified: { by: human:vault-owner, at: }`. `trust-record` only writes entries a human approved with `--approved`, so the `human:` actor is accurate and gives the page OKF's *human-reviewed* tier (§5.3). Never derive `verified` from `lifecycle` alone. | | `resource` | first `sources:` entry **iff** it is an `http(s)://` URL | — | Optional; omit when no source URL. Most pages describe abstract knowledge and have none. | | *(extensions)* | `category`, `created`, `updated`, `relationships`, `lifecycle`, `lifecycle_changed`, `tier`, `base_confidence`, … | preserved verbatim | OKF §4.1 requires consumers to preserve unknown keys. **Writing our native keys as OKF extension frontmatter is what makes the round-trip lossless** — on import, preserved `category`/`created`/`updated`/`lifecycle` are preferred over re-deriving from `type`/`generated`/`status`. (`updated` rides along because `generated.at` must be a full datetime, so a date-only `updated` would otherwise come back as midnight UTC.) `sources` is *not* an extension any more: it's a spec field (row above). | ### Steps Reuse the node list from Step 1 (with any active project/visibility filters already applied). Write a directory tree under `wiki-export/okf/`: 1. **One file per in-scope page.** For each page, parse its frontmatter, apply the mapping table above to build the OKF frontmatter (required `type` first, then `title`, `description`, `tags`, `generated`, `status`, `verified`, `sources`, optional `resource`, then the preserved extension keys), transform the body links (below), and write to `wiki-export/okf//.md` — same relative path the page has in the vault. 2. **Body link transform** (`[[wikilinks]]` → standard markdown links): - `[[concepts/transformers]]` → `[](.md)`, e.g. from `entities/foo.md` a link to `concepts/transformers` becomes `[Transformer Architecture](../concepts/transformers.md)`. Link text = the target page's `title` (fall back to the target id if unknown). - `[[target|display]]` → `[display](.md)`. - Use **file-relative** paths (`../concepts/x.md`), **never** `/`-absolute — `/`-rooted links break GitHub rendering. (This matches knowledge-catalog's own production agent.) - Compute that relative path from the **target file path**, not the bare page id: normalize the wikilink target to its page id, append `.md`, then compute `relpath(, )`. Never call `relpath()` on the id before adding `.md`. - This is required for the common "folder note" layout where a page id exists both as a file and as a directory prefix, e.g. `projects/social-twitter.md` plus `projects/social-twitter/...`. From `projects/social-twitter/concepts/mem0-memory-analysis.md`, a link to `[[projects/social-twitter]]` must export to `../../social-twitter.md`, not `...md`. - Resolve link targets with the same normalization used in Step 1 (lowercase, spaces→hyphens, strip `.md`). Handle unresolved targets by form, so forward-references survive the round-trip: - **Resolves to an in-scope page** → relative markdown link to it. - **Path-form target** (contains a `/`, e.g. `[[concepts/attention-mechanism]]`) with **no page yet**, and not excluded by an active filter → still emit the relative markdown link. OKF §11 forbids rejecting a bundle over a broken cross-link, and keeping the link makes the user's forward-references lossless on re-import. (Verified on st3ve: dropping these silently deleted real `[[wikilinks]]`.) - **Excluded by an active project/visibility filter** → plain text. Do not emit a path pointing into filtered-out content. - **Bare-title target** with no match (e.g. `[[tractorex]]` when no such page exists in scope) → plain text; there is no reliable path to write. - Leave existing external `http(s)://` links and `# Citations` sections untouched. 3. **Generate `index.md` files** (OKF §8 progressive disclosure; these contain no per-entry frontmatter): - Bundle root `wiki-export/okf/index.md` — a `# Subdirectories` section listing each category folder: `* [](/index.md) - `. This is the **only** index permitted frontmatter: add a single key `okf_version: "0.2"` (OKF §8, §12). - One `index.md` per category folder listing its pages: `* [](<slug>.md) - <description from the page's summary>`. 4. **Write `log.md`** from the vault root's `log.md`, reshaped to OKF §9, which requires `## YYYY-MM-DD` date headings, newest first. Our log is a flat, oldest-first list of `- [<timestamp>] VERB key=value …` lines, so: start with `# Wiki Update Log`; group lines by the date part of `<timestamp>`; emit groups newest first under `## <date>`; render each line as `* **<Verb>**: <rest of line>`, with the verb title-cased (`INGEST` → `Ingest`). Keep lines that don't match the pattern under the date of the line above them, verbatim after `* `. `wiki-import` ignores `log.md`, so this costs nothing on the round-trip. 5. **Filters.** Honor the same project/visibility filters as the graph export — filtered pages are omitted from the bundle and their inbound links degrade to plain text per step 2. Excluded from the bundle (same as the graph export): `_archives/`, `_raw/`, `_readouts/`, `.obsidian/`, `_insights.md`, `_meta/*.base`, and the vault root `index.md` (regenerated above). --- ## Step 4: Print Summary ``` Wiki export complete → wiki-export/ graph.json — N nodes, M edges (NetworkX node_link format) graph.graphml — N nodes, M edges (Gephi / yEd / Cytoscape) cypher.txt — N MERGE nodes + M MERGE relationships (Neo4j) postgres.sql — N upsert rows (wiki_pages) + M upsert rows (wiki_edges) (any Postgres) graph.html — interactive browser visualization (open in any browser) ``` Append this line only when the OKF bundle was produced (Step 3.5): ``` okf/ — OKF v0.2 markdown bundle (N pages, lossless; import via wiki-import) ``` Append filter notes when active: ``` (filtered: project:prismor — 19 of 67 pages) (filtered: X of Y pages excluded — visibility/internal, visibility/pii) ``` Only include lines for filters that were actually applied. ## Notes - **Re-running is safe** — all output files (and the `okf/` bundle) are overwritten on each run - **Broken wikilinks are skipped** — only edges to pages that exist in the vault are exported; in the OKF bundle, a wikilink to a missing/filtered page degrades to plain text - **OKF is the lossless format** — `graph.json` reconstructs only stubs on import, while the `okf/` bundle preserves full page bodies. Use OKF for vault-to-vault transfer and external markdown tools (MkDocs, Notion, GitHub); use the graph files for analysis tools (Gephi, Neo4j) - **The `wiki-export/` directory should be gitignored** if the vault is version-controlled — these are derived artifacts - **`graph.json` is the primary format** — the others are derived from it. If a future tool supports graph queries natively, point it at `graph.json`