English  ·  简体中文

# MOLT 🦀 > **The coding agent that grows its own tools.** Most agents solve your task and forget everything they learned. **MOLT doesn't forget.** After every job it reflects on what it did, and when it spots a reusable pattern, it **writes itself a new tool**, **tests it**, and **keeps it** in a skill library. Next time the pattern shows up, the tool is already there. Watch the toolbox compound: ``` $ molt run "tidy the config files" --learn --mock ── molt run — task ───────────────────────────────────────────── task: tidy the config files model: mock ── final answer ──────────────────────────────────────────────── All done — I inspected the workspace and the job is complete. ── learn ─────────────────────────────────────────────────────── ✔ grew a new tool: `parse_csv_line` → .molt/skills/parse_csv_line $ molt skills list parse_csv_line [project] Parse a CSV line of key=value pairs into a dict. $ molt run "parse some key=value config" --mock # the tool is already there ``` That's the whole idea: **your AI doesn't just code — it molts.** It sheds the old, grows a new tool, and keeps it. --- ## Why MOLT - **It compounds.** Every task can make the next one cheaper and more reliable. Your agent gets better at *your* codebase, not just smarter in general. - **It's honest self-modification.** A skill only enters the library after **passing its own test**. The agent can grow itself new tools, but they have to prove they work first. - **It's shareable.** `molt export` bundles your toolbox; `molt install` pulls someone else's. Your agent's hard-won skills become a library you can fork, star, and share. - **It's yours.** MIT, zero tracking, runs on any OpenAI-compatible model — DeepSeek by default. ## Quickstart ```bash # From PyPI (the package is published as `molt-agent`): pip install molt-agent # …or straight from source: git clone https://github.com/houyongsheng/deepseek-harness-molt cd deepseek-harness-molt pip install -e . export DEEPSEEK_API_KEY=sk-... ``` Run a task and let it learn: ```bash molt run "add retry logic to the http client" --learn ``` No API key? Run the exact same loop with a built-in fake model: ```bash molt run "tidy the config files" --learn --mock molt evolve "polish the project" -n 5 --mock # watch the toolbox grow ``` ## Commands | Command | What it does | |---|---| | `molt run "task" [--learn]` | Run one task. `--learn` keeps any reusable tool it finds. | | `molt run --stream --sandbox` | Stream the answer as it's generated; run skill code in a subprocess. | | `molt evolve "goal" -n N` | Loop subtask + learn N times; watch the toolbox accumulate. | | `molt learn` | Re-run reflection over the last transcript. | | `molt skills list / show / remove / dedup` | Inspect and de-duplicate the library. | | `molt eval [--cases FILE]` | Benchmark a skill's accuracy against labeled cases. | | `molt export [--out DIR]` | Bundle your skills for sharing. | | `molt install ` | Pull someone else's skills into your toolbox. | | `molt publish [--remote URL]` | Publish the toolbox to a git registry (community hub). | ## How it works ``` task ─▶ agent loop ─▶ answer │ └──▶ reflect: "did I hit a reusable pattern?" │ yes ▼ author a skill (python + schema + test) │ test it ── fail ─▶ discard │ pass ▼ commit to .molt/skills/ │ ▼ next run loads it as a tool ``` - **Run** — an agent loop with `shell`, `read_file`, `write_file`, plus every skill you've grown. - **Reflect** — a second LLM pass asks: was anything here reusable? If yes, it returns a skill as JSON. - **Test** — the skill's own test runs in a fresh subprocess; no test, no commit. - **Reuse** — skills become callable tools, project skills shadowing your home library. Skills live at `.molt/skills//` (project) and `~/.molt/skills/` (home). Each is plain files: `skill.json` (name/description/inputs), `skill.py` (`def run(**kwargs)`), `test.py`. ## Measure & publish A skill only earns its place if it works. `molt eval` scores a skill against labeled cases (a `cases.json` next to the skill, or any file you pass with `--cases`): ```bash molt eval parse_kv_records # accuracy: 5/5 (100%) ``` When you're happy, share it — or publish the whole toolbox to a git registry: ```bash molt export --out ./my-toolbox # plain copy + manifest molt publish --remote git@github.com:you/toolbox.git # pushes, prints the install line # others: molt install git@github.com:you/toolbox.git ``` That's the seed of an ecosystem: **your agent's skills are a library, and libraries get forked, starred, and shared.** ## Trust & safety - **Tested before committed.** Untested or failing skills never enter the library. - **Skills run in-process by default.** Use `--sandbox` to run skill code in a subprocess (process isolation). Either way, treat others' skills like any code you `pip install` — evaluate before trusting. - **Plain files, no lock-in.** Your toolbox is just directories; delete it, share it, move it. ## Roadmap - [x] Skill **dedup/merge** — `molt skills dedup` removes shadowed copies, flags identical code - [x] Skill **evals** — `molt eval` benchmarks accuracy against `cases.json` - [x] **Registry publish** — `molt publish` pushes the toolbox to a git registry - [x] **Streaming + richer tools + sandbox** — `--stream`, `list_dir`/`search`, `--sandbox` - [ ] Auto-run a skill's evals before every commit - [ ] A discoverable **registry hub** (`molt search`) — find skills others published ## Codex & Claude ecosystem MOLT's grown tools are **plain Python** — so they travel. Two directions: **MOLT → Claude Code / Codex.** Export your toolbox in a format they already understand: ```bash molt export --format claude --out .claude # → .claude/skills//SKILL.md molt export --format codex --out . # → AGENTS.md reusable-procedure block ``` Drop `.claude/` into a project and Claude Code picks up each skill; append the `AGENTS.md` block and Codex can implement the same procedure. Ready-made templates live in [`examples/integrations/`](examples/integrations/). **Claude Code / Codex → MOLT.** Delegate the "grow a tool" job back to MOLT with a one-file skill: `examples/integrations/claude/molt-skill.md` tells Claude Code to run `molt run "" --learn` whenever it spots a repeatable pattern — then the tool it grew is available to everyone. ## Inspiration MOLT's core bet — *an agent that writes, tests, and keeps its own tools* — is the idea at the heart of [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness), whose self-referential toolset lets the model inspect and mount plugins inside its own running runtime. MOLT is the lightweight, standalone take on that idea: one pip-installable package, any OpenAI-compatible model, no framework to learn. Want the full plugin-everything harness? Go there. Want the idea in a few hundred lines you can read in an afternoon? Stay here. ## License MIT — go build something that grows itself.