--- name: autorag-lite-setup description: Install and register the model-free AutoRAG Lite MCP server, configure approved roots and datasources, build indexes, and verify MCP discovery and search. Use when autorag-mcp is missing, MCP connection fails, indexes are stale, or the user wants document search without configuring a search model. license: MIT --- # AutoRAG Lite setup Use this skill only to bootstrap, register, or repair AutoRAG Lite MCP: initialize trusted config, connect the host, build indexes, and verify search. Normal Lite use is MCP tool calling, not a search skill or shell command. Use `autorag-setup` instead when the model-backed librarian must be configured. ## Safety - Inspect only non-secret config metadata: `searchPaths`, `workspacePath`, `memoryPath`, `minSync`, `jikji`, `datasources`. - Never print, copy, migrate, compare, or persist credential values. Store only environment-variable names such as `tokenEnv` or `apiKeyEnv`. - Never move, rename, edit, or delete source documents. Lite commands write indexes only under the configured workspace `.autorag/` directory and Jikji's per-source `.jikji/` caches. - Do not index system trees, app bundles, caches, credential stores, `node_modules`, `.git`, `dist`, `build`, `target`, `.cache`, `.autorag`, or `.jikji`. ## Install the package if needed `@autorag/librarian` ships both `autorag` (bootstrap/maintenance) and `autorag-mcp` (stdio server). The executable requires Node.js >= 24 on PATH. ```bash command -v autorag-mcp >/dev/null || bun install -g @autorag/librarian command -v autorag command -v autorag-mcp autorag lite --help ``` If Bun is unavailable, `npm install -g @autorag/librarian` is acceptable. If only `autorag` exists, upgrade the package rather than substituting CLI retrieval for MCP. For a source checkout, run `bun run build` and register the absolute `dist/mcp/index.js` path instead of the installed executable. ## Initialize a model-free config ```bash autorag lite init \ --search-paths "/path/to/documents,/path/to/notes" \ --workspace "/path/to/workspace" \ --memory-path "/path/to/memory.json" ``` `lite init` writes the configured model-free lifecycle config. Use `--force` only when intentionally replacing an existing config, and target it with an explicit path (`--config` or `AUTORAG_CONFIG`); `--force` refuses to replace the implicit `~/.autorag/config.json`. Explicit user paths win; otherwise propose one to three document-dense roots and get approval before indexing. Supported parsed formats are `md`, `markdown`, `txt`, `text`, `pdf`, `docx`, `pptx`, `xlsx`, `xls`, `hwp`, `hwpx`, and `eml`. Legacy `.doc` is not a supported parsed format. Config resolution follows the usual order: `--config`, `AUTORAG_CONFIG`, `$AUTORAG_HOME/config.json`, or `~/.autorag/config.json`. Environment overrides include `AUTORAG_HOME`, `AUTORAG_CONFIG`, `AUTORAG_SEARCH_PATHS`, `AUTORAG_WORKSPACE`, and `AUTORAG_MEMORY_PATH`. ## Register and verify the MCP connection After config and datasource setup, register the server with the host. Use absolute config, search root, workspace, and executable paths: hosts may start the server from a different working directory or with a restricted PATH. Resolve `command -v autorag-mcp` and substitute its absolute path below. Inspect an existing `autorag` registration first; keep a working registration and update only an outdated command or config path. ```bash # Claude Code: project-local registration claude mcp add --transport stdio --scope local autorag \ --env AUTORAG_CONFIG=/absolute/path/to/.autorag/config.json \ -- /absolute/path/to/autorag-mcp # Codex: user registration codex mcp add autorag \ --env AUTORAG_CONFIG=/absolute/path/to/.autorag/config.json \ -- /absolute/path/to/autorag-mcp ``` For other hosts, use their stdio MCP configuration with the same command, empty arguments, and `AUTORAG_CONFIG` environment variable. The host starts the subprocess; do not run it as a background HTTP service. Pass any required credential environment-variable names through the host's secret mechanism; never put credential values in registration examples or logs. Reload/reconnect the host, then use MCP `tools/list` to discover the actual tools and schemas. Verify `autorag.status`, `autorag.refresh`, `autorag.search`, `autorag.search.files`, and `autorag.datasources.list` are available. Configured integrated datasources add their own search tools; do not assume a fixed count. Run `autorag.status` and `autorag.datasources.list` through MCP, then refresh and search a known phrase from an approved document. Registration alone is not proof of a connected or searchable server. After config changes, restart the server so datasource tools are rebuilt. `AUTORAG_MCP_READ_ONLY=1` omits refresh; build indexes with the CLI before connecting that mode. `AUTORAG_MCP_TOOLS` is an optional comma-separated exact tool allowlist; omitted tools cannot be called. Use unrestricted tools for initial setup unless the user intentionally requests a restricted connection. ## Probe and configure datasources (setup wizard) A datasource setup UI is not shipped in this build — do not recommend it for datasource setup. Configure datasources directly in trusted config, wizard-style: 1. Probe every datasource for setup feasibility before asking the user anything: the backing CLI exists (`lazykatok`, `discrawl`, `slacrawl`, `wacrawl`, `telecrawl`, `notcrawl`, `qmd`, `mailcrawl`, `rclone`) and its local store or archive is present. CLI-backed datasources own their own archive, index, and authentication, so environment credentials (such as bot tokens) are never required or checked for them. Non-CLI connectors (such as `github`) require their credential environment variable (`GITHUB_TOKEN`). 2. Auto-configure every datasource that probes feasible — write its trusted `datasources` entries without asking. For example, when Slack (`slacrawl`) and Discord (`discrawl`) are installed with local stores present, set both up automatically. Discord uses discrawl's local desktop wiretap archive; no Discord bot token is configured or needed. 3. Skip every datasource that probes infeasible (for example Notion or Telegram when their CLIs are not installed) and always report the skipped list to the user, with what is missing for each. 4. Set up a skipped datasource only when the user explicitly asks for it: install or authenticate the backing CLI first, then configure it. 5. E-mail datasources (`mail-export`, `mailcrawl`) matter to most users — always probe them and report their status, even when they end up skipped. Datasource skills belong in trusted config. Config keys may be builtin template names (`kakao`, `whatsapp`, `telegram`, `slack`, `discord`, `clawgallery`, `notion`, `github`, `cloud-drive`, `mail-export`, `mailcrawl`, `obsidian`, `rss`, `spotlight`, `lark`) or connection aliases with `"type": "