---
**The agent loop as a library, not as an app.**
astorlm gives you the machinery behind a coding agent โ the loop, tools, providers, sessions, hooks and executors โ as plain TypeScript objects you compose inside your own process. No CLI, no TUI, no daemon to shell out to and no output to scrape.
- **Embed it.** `await createAgent({ ... })` returns an object you own. Your process, your logging, your permission prompts, your UI.
- **Swap every part.** `Provider`, `Executor`, `SessionManager`, `SkillSource` and `CodeRunner` are interfaces. The built-ins are conveniences, not requirements.
- **Bring your own model.** Anthropic, or anything OpenAI-compatible โ OpenAI, Groq, OpenRouter, Together, vLLM, Ollama, a local proxy. Small local models are a supported target, not an afterthought.
- **Runtime-agnostic core.** The loop, providers, tools, sessions, skills and embeddings import zero Node built-ins. Everything OS-bound lives behind `astorlm/core`.
## โก Quickstart
```bash
pnpm add astorlm # Node >= 20
```
```typescript
import { createAgent, OpenAIProvider, tool } from 'astorlm'
import { z } from 'zod'
const getWeather = tool({
name: 'get_weather',
description: 'Returns the current temperature for a city',
schema: z.object({ city: z.string() }),
execute: async ({ city }) => `Weather in ${city}: 22ยฐC, sunny.`,
})
// createAgent is async: it loads any persisted state before returning.
const agent = await createAgent({
provider: new OpenAIProvider({
model: 'qwen2.5-coder',
baseURL: 'http://localhost:11434/v1', // any OpenAI-compatible endpoint
apiKey: 'ollama',
}),
tools: [getWeather],
})
agent.on('text', (chunk) => process.stdout.write(chunk))
await agent.run('How is the weather in Buenos Aires?')
console.log(agent.getUsage()) // { inputTokens, outputTokens, ... }
```
That is the whole setup. Want it to touch the filesystem? Swap `createAgent` for [`createLocalAgent`](#-2-full-coding-agent-nodejs) and pass `createCodingTools()`.
## ๐ค How it compares
| | astorlm |
| --- | --- |
| **vs. a coding CLI** (Claude Code, Codex CLI, Aider) | Those are applications you drive. astorlm is the machinery they are built out of, running in *your* process โ so the UI, the audit log and the approval flow are yours to write. |
| **vs. an agent framework** (LangGraph, Mastra) | No graph DSL and no workflow engine to learn. One loop, five hooks, and interfaces you implement. The whole public surface is `src/index.ts` and `src/core.ts`. |
| **vs. a vendor SDK** | Provider-neutral by construction. Local and weak models get first-class support through [`experimental/edge-boost`](#12-astorlmexperimentaledge-boost-experimental--hardening-for-weak-models), not a "best effort" disclaimer. |
## ๐ What's in the box
| | |
| --- | --- |
| **Loop** | Multi-turn, parallel tool execution, token streaming, thinking blocks, cooperative cancellation, `REACT` or `PLAN_EXECUTE` [patterns](#-loop-patterns) |
| **Tools** | `read`, `write`, `edit`, `bash` (plus [background spawn/poll/kill](#-background-processes-bash_spawn--bash_get_output--bash_kill)), `ls`, `grep`, `glob` โ or [define your own](#-1-minimal-usage-custom-tool) from a Zod schema |
| **Isolation** | Swappable [executors](#-executors-sandboxing--swappable-backends): local, Docker, or your own. Plus a [WASM code sandbox](#-wasm-code-sandbox-astorlmexperimentalwasm-runner) that needs no daemon |
| **Control** | Five [hooks](#-control-hooks-sessionhooks) covering permissions, mocking, prompt rewriting and output sanitising; [steering](#-steering-redirect-without-aborting) at tool boundaries |
| **Memory** | Session persistence and native branching, [context auto-compaction](#-context-optimizer-auto-compaction), [token accounting](#-token-usage-tracking) |
| **Interop** | [MCP](#-mcp-connectivity-model-context-protocol) over stdio and HTTP (including MCP Apps UI), and [Agent Skills](#-skills-loadable-knowledge-packs) in the same filesystem format Claude Code and Codex use |
| **Composition** | [Subagents as tools](#-subagents-agent-as-tool), [goal loops](#-goal-loops-rungoalloop), [structured output](#-structured-output-generateobject), [heartbeats](#-heartbeat-proactive-loop) |
| **Observability** | Tracing with an OTLP exporter, metrics with cost accounting, deterministic record & replay, and an offline eval harness โ all under `astorlm/experimental/*` |
## ๐ Table of contents
- [โก Quickstart](#-quickstart)
- [๐ค How it compares](#-how-it-compares)
- [๐ What's in the box](#-whats-in-the-box)
- [๐ฆ Module Layout (Entrypoints)](#-module-layout-entrypoints)
- [๐ Quick Use Examples](#-quick-use-examples)
- [๐งฌ Subagents (agent-as-tool)](#-subagents-agent-as-tool)
- [๐ช Control Hooks (`SessionHooks`)](#-control-hooks-sessionhooks)
- [๐ MCP Connectivity (Model Context Protocol)](#-mcp-connectivity-model-context-protocol)
- [๐ Skills (loadable knowledge packs)](#-skills-loadable-knowledge-packs)
- [๐ณ Executors (sandboxing & swappable backends)](#-executors-sandboxing--swappable-backends)
- [๐งช WASM code sandbox (`astorlm/experimental/wasm-runner`)](#-wasm-code-sandbox-astorlmexperimentalwasm-runner)
- [โฑ๏ธ Background processes (`bash_spawn` / `bash_get_output` / `bash_kill`)](#-background-processes-bash_spawn--bash_get_output--bash_kill)
- [๐งญ Loop patterns](#-loop-patterns)
- [๐ Goal loops (`runGoalLoop`)](#-goal-loops-rungoalloop)
- [๐งฑ Structured output (`generateObject`)](#-structured-output-generateobject)
- [๐ Context optimizer (auto-compaction)](#-context-optimizer-auto-compaction)
- [๐ Retry policy for transient provider errors](#-retry-policy-for-transient-provider-errors)
- [๐ Token usage tracking](#-token-usage-tracking)
- [๐ซ Heartbeat (proactive loop)](#-heartbeat-proactive-loop)
- [๐ ๏ธ Development Commands](#-development-commands)
---
## ๐ฆ Module Layout (Entrypoints)
AstorLM ships clearly separated entrypoints. The runtime-agnostic pieces โ loop, providers, tools, sessions, skills, embeddings โ live in the main `astorlm` barrel and import zero Node built-ins. Everything OS-bound lives behind `astorlm/core`.
```text
astorlm agnostic createAgent ยท Anthropic/OpenAIProvider ยท tool
ToolRegistry ยท SessionHooks ยท InMemorySessionManager
createSubagentTool ยท createSteeringController
โโ re-exports ./core for convenience โ ๏ธ pulls Node deps into the main barrel
astorlm/core Node createLocalAgent ยท FileSessionManager ยท AstorAgent
mountMcpServer ยท Local/DockerExecutor
create{FileSystem,Layered}SkillSource
astorlm/tools Node createCodingTools() / createReadOnlyTools()
read write edit bash bash_spawn ls grep glob
astorlm/embeddings agnostic createOpenAIEmbedder ยท createSemanticIndex
astorlm/experimental/* varies error-registry ยท tracing ยท metrics ยท replay
evals ยท wasm-runner ยท edge-boost
```
### 1. `astorlm` (main barrel)
* **Description**: One-stop import. Exposes the agnostic core (`createAgent`, the providers, `tool`, `ToolRegistry`, sessions, skills, hooks, optimizer, retry, the subagent/steering helpers) **and** re-exports everything from `astorlm/core` for unified imports.
* **Key exports**:
- `createAgent` (async โ `Promise`)
- `InMemorySessionManager`, `SessionManager`
- `AnthropicProvider`, `OpenAIProvider`
- `createOpenAIEmbedder`, `cosineSimilarity`, `createSemanticIndex`, `withEmbeddingCache` (embeddings โ also at `astorlm/embeddings`)
- `tool`, `ToolRegistry`
- `EventBus`, `buildSystemPrompt`, `estimateTokens`, `optimizeContext`, `isTransientError`, `computeBackoffDelay`
- `createSubagentTool` (subagents / agent-as-tool), `createSteeringController` (out-of-hook steering)
- `generateObject` (schema-constrained output), `runGoalLoop` (iterate until a condition holds)
- `createNoopExecutor` (the safe default executor of the agnostic core)
- `SkillRegistry`, `createInMemorySkillSource`, `parseSkillFrontmatter`, `renderSkillsBlock`, `createLoadSkillTool`
- `validateSkillName`, `validateSkillDescription`, `validateSkillSpec`, `SkillValidationError`, `SKILL_VALIDATION_LIMITS`
- `parseAllowedTools`, `restrictToolsHook` (per-skill tool gating)
- Base types: `Agent`, `CreateAgentOptions`, `SessionHooks`, `Message`, `ContentBlock`, `AgentEvent`, `Skill`, `SkillMetadata`, `SkillSource`, `SkillMode`, etc.
> โ ๏ธ Because the main barrel re-exports `./core`, importing from `astorlm` pulls in Node-only dependencies. If you target Edge / Workers / browser, import from the agnostic modules directly (the loop, providers and abstractions don't import Node) and avoid the Node runner.
### 2. `astorlm/core` (Node.js runner)
* **Description**: Extensions that require Node.js OS-native APIs (`node:fs`, `node:path`, `node:child_process`).
* **Key exports**:
- `createLocalAgent` (session factory wired with `LocalExecutor` + a local file reader by default).
- `FileSessionManager` (history persistence as JSONL plus JSON metadata).
- `AstorAgent` (simplified execution and branching facade).
- `mountMcpServer` (adapter and Stdio/HTTP transports for Model Context Protocol clients).
- `LocalExecutor`, `DockerExecutor` (shell-command backends; see "Executors" below).
- `createFileSystemSkillSource` (reads skills from a `dir//SKILL.md` layout).
- `createLayeredSkillSource` (hierarchical skill discovery โ user โ project โ repo).
### 3. `astorlm/tools` (Built-in tools)
* **Description**: A bundle of filesystem-manipulation and analysis tools tuned for coding agents, with path-traversal protection.
* **Key exports**:
- `createCodingTools()` (`read`, `write`, `edit`, `bash`, `bash_spawn`, `bash_get_output`, `bash_kill`, `ls`, `grep`, `glob`).
- `createReadOnlyTools()` (safe variant โ no writes or execution: `read`, `ls`, `grep`, `glob`).
- Individual tools: `readTool`, `writeTool`, `editTool`, `bashTool`, `bashSpawnTool`, `bashGetOutputTool`, `bashKillTool`, `lsTool`, `grepTool`, `globTool`.
### 4. `astorlm/embeddings` (Embeddings & semantic search)
* **Description**: First-class, runtime-agnostic embeddings primitives โ sit next to the providers in the main barrel and are also reachable via this dedicated subpath. `fetch`-based, zero extra dependencies, work anywhere `fetch` exists (Node, Deno, browsers, edge).
* **Key exports**:
- `createOpenAIEmbedder({ baseURL, model, apiKey?, dimensions? })` โ OpenAI-compatible `Embedder` with `embed` (single) and `embedMany` (batch, one round-trip) plus token `usage`.
- `createSemanticIndex({ embedder })` โ in-memory vector store: `add` / `addMany` / `query(text, { topK, threshold })` / `queryByVector` / `remove` / `clear`. The reusable primitive behind semantic search, RAG retrieval and dedupe.
- `withEmbeddingCache(embedder)` โ memoizes identical inputs so repeated lookups don't re-embed (or re-bill); `embedMany` only requests the cache misses.
- `cosineSimilarity` (standard `[-1..1]`), `dotProduct`, `euclideanDistance`.
- Types: `Embedder`, `EmbedResult`, `EmbedManyResult`, `SemanticIndex`, `SemanticHit`, etc.
```typescript
import { createOpenAIEmbedder, createSemanticIndex } from 'astorlm'
const embedder = createOpenAIEmbedder({
baseURL: 'http://localhost:11434/v1',
model: 'nomic-embed-text',
apiKey: 'ollama',
})
const index = createSemanticIndex({ embedder })
await index.addMany([
{ id: 'oom', text: 'Node process crashes with out of memory during the build.' },
{ id: 'tls', text: 'TLS handshake fails: expired certificate in production.' },
])
const hits = await index.query('ran out of RAM while compiling', { topK: 1 })
// โ [{ id: 'oom', score: 0.8โฆ, text: 'โฆ' }]
```
### 5. `astorlm/prompt` (System-prompt building & composition)
* **Description**: The system-prompt layer, runtime-agnostic. `buildSystemPrompt` is what the loop uses internally; the prompt *compiler* is the opt-in half โ it merges prompt fragments coming from different owners (a base identity, a product policy, a per-skill instruction) into a single string, deduping repeats and **reporting contradictions** instead of silently concatenating them.
* **Key exports**:
- `buildSystemPrompt(opts)`, `DEFAULT_SYSTEM_PROMPT` โ also reachable from the main barrel.
- `compilePrompts({ modules, layout?, deduplicate?, detectConflicts? })` โ `{ systemPrompt, report, print() }`. Modules are emitted in `layout` order (default `identity โ context โ constraint โ policy โ format`); modules sharing an `id` resolve by `priority` (highest wins).
- `formatPromptReport(report)` โ readable audit of what was kept, overridden, deduped or flagged as conflicting.
```typescript
import { compilePrompts, formatPromptReport } from 'astorlm/prompt'
const { systemPrompt, report } = compilePrompts({
modules: [
{ id: 'identity', kind: 'identity', content: 'You are a release engineer.' },
{ id: 'lang', kind: 'policy', content: 'Always answer in English.' },
{ id: 'lang', kind: 'policy', content: 'Always answer in Spanish.', priority: 10 }, // wins
],
})
console.log(formatPromptReport(report)) // shows the override and any conflicts
```
### 6. `astorlm/experimental/error-registry` (Experimental โ Federated Error Registry)
> โ ๏ธ **Experimental**. Lives under a dedicated subpath, not the main barrel. The import path itself is the signal that the API is volatile and may change between minor releases.
* **Description**: A registry of agent-encountered errors and human-approved resolutions. When an agent hits an error that another agent (or a previous run) has already resolved, the registry injects the fix as a hint into the next `tool_result` โ the agent applies the known solution instead of fighting through it again. Honest single-org PoC; federation across organizations and full secret sanitization are out of scope.
* **Key exports**:
- `createErrorRegistry(opts)` โ JSONL append-only store (or in-memory) with optional OpenAI-compatible embeddings and a Jaccard fallback.
- `errorRegistryHooks({ registry, context, successWindow? })` โ returns a `SessionHooks` object that wires the session to the registry.
```typescript
import { createLocalAgent, OpenAIProvider } from 'astorlm'
import { createCodingTools } from 'astorlm/tools'
import { createErrorRegistry, errorRegistryHooks } from 'astorlm/experimental/error-registry'
const registry = createErrorRegistry({
storePath: '.astorlm/error-registry.jsonl',
// Optional โ if omitted, falls back to Jaccard over tokens:
// embeddings: { baseURL: 'http://localhost:11434/v1', model: 'nomic-embed-text' },
})
await registry.init()
const agent = await createLocalAgent({
provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }),
tools: createCodingTools(),
hooks: errorRegistryHooks({
registry,
context: { cwd: process.cwd(), osPlatform: process.platform, nodeVersion: process.version, tags: [] },
}),
})
// Note: errorRegistryHooks returns a SessionHooks object โ if you already
// have your own hooks, merge them into a single object before passing.
```
A human approves pending resolutions asynchronously (e.g. `registry.approveResolution(id, approver)`). Until approved, a candidate resolution is not suggested to other sessions.
### 7. `astorlm/experimental/tracing` (Experimental โ Observability)
> โ ๏ธ **Experimental**. Volatile API behind a dedicated subpath. Runtime-agnostic (reads only the event bus; uses Web Crypto for ids).
* **Description**: Derives a hierarchical span tree (`session โ turn โ provider_call | tool_execution`) from the agent's event bus **without touching the loop** โ attaching a tracer is pure subscription. Spans carry OpenTelemetry GenAI semantic-convention attributes (`gen_ai.*`) plus astorlm-specific ones (TTFT, tool duration/errors, retries).
* **Key exports**:
- `attachTracer(agent, { exporter })` โ subscribes to the bus and builds spans; returns `{ detach(), currentTraceId() }`.
- `createInMemoryExporter()` โ collects spans in an array for tests / local inspection.
- `createTracer(opts?)` โ low-level span factory (usually managed by `attachTracer`).
- `createOtlpSpanExporter(opts)` (from `astorlm/experimental/tracing/otel`) โ OTLP/HTTP (JSON) exporter built on `fetch` alone, **no OpenTelemetry SDK dependency**. Ships spans to any OTLP collector (OpenTelemetry Collector, Tempo, Jaeger, Honeycomb, Arize Phoenix, Langfuse).
```typescript
import { createLocalAgent, OpenAIProvider } from 'astorlm'
import { createCodingTools } from 'astorlm/tools'
import { attachTracer, createInMemoryExporter } from 'astorlm/experimental/tracing'
import { createOtlpSpanExporter } from 'astorlm/experimental/tracing/otel'
const agent = await createLocalAgent({
provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }),
tools: createCodingTools(),
})
// In-memory (inspect locally) ...
const memory = createInMemoryExporter()
// ... and/or export to an OTLP collector:
const otlp = createOtlpSpanExporter({ endpoint: 'http://localhost:4318/v1/traces', serviceName: 'my-agent' })
const tracer = attachTracer(agent, {
exporter: { export: (spans) => { memory.export(spans); otlp.export(spans) } },
})
await agent.run('List the .ts files and count them.')
tracer.detach()
await otlp.shutdown() // final flush
for (const s of memory.spans) console.log(s.kind, s.name, s.endTime! - s.startTime, 'ms')
```
The OTLP exporter buffers spans and flushes by batch size (`maxBatch`, default 256) or on a timer (`flushIntervalMs`, default 5s, `unref()`-ed). Hex `trace_id`/`span_id` are forwarded verbatim per the OTLP/JSON convention.
### 8. `astorlm/experimental/metrics` (Experimental โ Cost & metrics)
> โ ๏ธ **Experimental**. Runtime-agnostic (reads only the event bus).
* **Description**: Aggregates operational metrics from the agent's event bus and, given a pricing table, the USD cost of a run. No prices are hardcoded โ you supply the table (USD per 1M tokens).
* **Key exports**:
- `attachMetrics(agent, { pricing? })` โ subscribes to the bus; returns `{ snapshot(), reset(), detach() }`. The snapshot has counts (`runs`, `turns`, `providerCalls`, `toolCalls`, `toolErrors`, `providerRetries`), `tokens`, optional `costUsd`, and latency stats (`ttftMs`, `turnMs`, `toolMs`).
- `computeCost(usage, pricing)`, `resolvePricing(table, model)` โ usable standalone (exact then longest-prefix model match).
```typescript
import { attachMetrics } from 'astorlm/experimental/metrics'
const metrics = attachMetrics(agent, {
pricing: { 'gpt-4o': { inputPer1M: 2.5, outputPer1M: 10, cacheReadPer1M: 1.25 } },
})
await agent.run('...')
const m = metrics.snapshot() // { costUsd, latency: { ttftMs: { avg, ... } }, tokens, ... }
```
### 9. `astorlm/experimental/replay` (Experimental โ Record & replay)
> โ ๏ธ **Experimental**. Runtime-agnostic; the `Recording` is a plain serializable object.
* **Description**: Captures exactly the provider events a run produced and replays them later with no network and no token spend โ the deterministic debugging primitive. Capture is at the provider boundary, so it's independent of tools, hooks and timing.
* **Key exports**:
- `createRecordingProvider(inner)` โ wraps a real provider, passes events through and records them; `getRecording()` returns a serializable object.
- `createReplayProvider(recording, { onExhausted? })` โ a provider that replays the recording turn by turn.
```typescript
import { createRecordingProvider, createReplayProvider } from 'astorlm/experimental/replay'
const rec = createRecordingProvider(realProvider)
const agent = await createLocalAgent({ provider: rec, tools })
await agent.run('...')
fs.writeFileSync('run.json', JSON.stringify(rec.getRecording()))
// later โ same events, no model call
const recording = JSON.parse(fs.readFileSync('run.json', 'utf8'))
const replay = await createLocalAgent({ provider: createReplayProvider(recording), tools })
await replay.run('...')
```
### 10. `astorlm/experimental/evals` (Experimental โ Offline evaluation)
> โ ๏ธ **Experimental**. Runtime-agnostic core; `llmJudge` needs a `Provider` (point it at a local OpenAI-compatible endpoint).
* **Description**: Runs a dataset of cases through fresh agents, applies scorers, and aggregates a report (overall pass rate + per-scorer stats). Built for CI gating; pair the agent factory with the replay provider for fast, network-free regression runs.
* **Key exports**:
- `runEval({ dataset, createAgent, scorers, concurrency?, onResult? })` โ `EvalReport`.
- Scorers: `exactMatch`, `contains`, `regexMatch`, `toolTrajectory` (tool-call sequence: `'exact' | 'ordered-subset' | 'set'`), `llmJudge` (LLM-as-judge with a rubric โ normalized 0..1).
```typescript
import { runEval, contains, toolTrajectory, llmJudge } from 'astorlm/experimental/evals'
const report = await runEval({
dataset: [{ id: 'q1', input: 'How many .ts files?', expected: 'a number' }],
createAgent: () => createLocalAgent({ provider: provider(), tools: createReadOnlyTools() }),
scorers: [contains('.ts'), toolTrajectory(['ls'], { mode: 'set' }), llmJudge({ provider: provider(), rubric: '...' })],
})
if (report.summary.passRate < 0.8) process.exit(1) // CI gate
```
### 11. `astorlm/experimental/wasm-runner` (Experimental โ WASM code sandbox)
> โ ๏ธ **Experimental**. Dedicated subpath, volatile API.
* **Description**: `CodeRunner` โ runs untrusted *source code* (not shell commands) inside a memory-safe WebAssembly sandbox with no host access. The daemon-free, runtime-agnostic isolation tier, complementary to `DockerExecutor`. See ["WASM code sandbox"](#-wasm-code-sandbox-astorlmexperimentalwasm-runner) below.
* **Key exports**: `QuickJsCodeRunner` (JS via QuickJS-wasm), `createCodeRunnerTool({ runner })` (opt-in `run_code` tool), plus the `CodeRunner` / `RunCodeOptions` / `RunCodeResult` types.
### 12. `astorlm/experimental/edge-boost` (Experimental โ hardening for weak models)
> โ ๏ธ **Experimental**. Dedicated subpath, volatile API. Not re-exported from the main barrel.
* **Description**: `edgeBoost(options, tuning?)` โ a pure options-transformer that hardens an agent for weak / local / free-tier models (free OpenRouter tiers, small Ollama/vLLM models, heavily quantized checkpoints). It defends against the failure where a model, on a long multi-turn tool context (typically the synthesis turn after several tool round-trips), degenerates into an infinite repetition loop or stalls without producing content until a timeout fires. Four opt-in defenses, none of which change default behavior unless applied:
- **Stream guard** (`createGuardedProvider`) โ detects degeneration (identical-delta repetition, no-content stall, reasoning-budget blowout) and retries cheaply against the same endpoint. Owns *degeneration* retries; the loop-level `retry` policy still owns *network/HTTP* retries. Model failover between endpoints is out of scope by design (that belongs to a routing proxy).
- **Context diet** โ per-call truncation of old `tool_result` blocks (idempotent, coexists with the core optimizer's markers) + an optional compact system prompt; never mutates session history. Optional LLM `digest` summarizes an evicted block with an injected cheap provider instead of truncating (cached per block, falls back to truncation on failure/timeout).
- **Forced synthesis** โ after N tool rounds in the current request, appends a synthesis instruction and forces `toolChoice: 'none'` (or strips tools) so the model stops calling tools and answers.
- **Sampling defaults** โ injects `max_tokens` and a `frequencyPenalty` (repetition penalty) when the call omits them. Plus optional **semantic tool pruning** (top-K relevant tools per turn via an embedder; disabled by default, degrades gracefully if the embedder is down).
* **Key exports**: `edgeBoost`, `edgeBoostHooks`, `createGuardedProvider`, `DegenerationError`, `mergeSessionHooks` (generally useful hook composition), `buildContextDietHook` / `buildSynthesisHook` / `buildToolPruningHook`, `resolveEdgeBoostTuning`, plus the `EdgeBoostTuning` / `GuardOptions` / `ContextDietOptions` / `SynthesisOptions` / `SamplingDefaults` / `ToolPruningOptions` types.
```ts
import { createLocalAgent } from 'astorlm/core'
import { OpenAIProvider } from 'astorlm'
import { createCodingTools } from 'astorlm/tools'
import { edgeBoost } from 'astorlm/experimental/edge-boost'
// Wrap the same options you already pass to createLocalAgent โ one call, no layers.
const agent = await createLocalAgent(edgeBoost({
cwd: process.cwd(),
provider: new OpenAIProvider({ model: 'myproxyllm', baseURL: 'http://127.0.0.1:11434/v1', apiKey: 'not-needed' }),
tools: createCodingTools(),
}))
// Tune any section; pass `false` to disable one. Example: force synthesis earlier.
const tuned = edgeBoost(options, { synthesis: { forceAfterToolRounds: 2 } })
```
A runnable `34-edge-boost` example (`pnpm start 34`, with optional `--prune` / `--digest` flags) lives in the companion examples repository, which is published separately from this one.
### 13. `astorlm/experimental/contract` (Experimental โ Agent Contract)
> โ ๏ธ **Experimental**. Dedicated subpath, volatile API. Node-bound (uses `node:path` for path matching).
* **Description**: A declarative budget-and-permission envelope for a run, enforced through hooks instead of trust. You declare what the agent may spend and touch; `createContractHooks` turns that into a `SessionHooks` object that stops the run with a `ContractViolationError` the moment a rule is crossed. Complementary to `restrictToolsHook` (tools only) and to `DockerExecutor` (isolates, but does not budget).
* **Key exports**:
- `createContractHooks(contract)` โ `SessionHooks`, ready to pass to `createAgent` / `createLocalAgent`.
- `ContractValidator` โ the enforcement engine on its own, if you'd rather drive it yourself.
- `ContractViolationError` (carries the `rule` that failed), `globToRegex`, plus the `AgentContract` / `AgentContractBudget` / `AgentContractTools` / `AgentContractSandbox` types.
```typescript
import { createContractHooks } from 'astorlm/experimental/contract'
const hooks = createContractHooks({
budget: { maxTurns: 12, maxTotalTokens: 200_000, maxDurationMs: 5 * 60_000 },
tools: { allow: ['read', 'ls', 'grep', 'glob', 'bash'] },
sandbox: {
allowedPaths: ['src/**', 'tests/**'],
deniedPaths: ['**/.env', '**/secrets/**'],
bash: { deniedCommands: ['rm', 'curl', 'git push'] },
},
})
// pass into createLocalAgent({ hooks })
```
---
## ๐ Quick Use Examples
> **All snippets use `OpenAIProvider`** pointed at a local OpenAI-compatible endpoint. The examples assume [Ollama](https://ollama.com) (`http://localhost:11434/v1`, model `qwen2.5-coder`, `apiKey: 'ollama'` โ a placeholder local endpoints ignore), but any OpenAI-compatible server works (LM Studio `http://localhost:1234/v1`, vLLM `http://localhost:8000/v1`, โฆ).
>
> **Hosted providers:** point `baseURL` at the vendor and pass a real key โ e.g. OpenAI (`https://api.openai.com/v1`, `gpt-4o-mini`), Groq, OpenRouter or Together. If you omit `apiKey`, the provider reads it from `OPENAI_API_KEY` (override the env var name with `envVar`). `AnthropicProvider` exists in the public API with the same shape โ swap it in if you prefer Anthropic. It additionally takes `thinking` (`{ type: 'adaptive', display? }` on current models, or `{ budget_tokens }` on older ones), `effort` (`'low'` โฆ `'max'`), and `contextLimit`, which defaults to a conservative 200k โ raise it to match the model you target.
### ๐ 1. Minimal Usage (custom tool)
```typescript
import { createAgent, OpenAIProvider, tool } from 'astorlm'
import { z } from 'zod'
// 1. Define a custom tool (Zod schema โ JSON Schema under the hood)
const getWeather = tool({
name: 'get_weather',
description: 'Returns the current temperature for a city',
schema: z.object({ city: z.string() }),
execute: async ({ city }) => `Weather in ${city}: 22ยฐC, sunny.`,
})
// 2. Create the session.
// Note: createAgent is async โ it loads any persisted state from the
// session manager before returning. Always await it.
const agent = await createAgent({
provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }),
tools: [getWeather],
})
// 3. Subscribe to the token stream
agent.on('text', (text) => process.stdout.write(text))
// 4. Run the prompt
await agent.run('How is the weather in Buenos Aires?')
```
---
### ๐ป 2. Full Coding Agent (Node.js)
The standard setup for building an autonomous backend coding agent with local filesystem access.
```typescript
import { createLocalAgent, OpenAIProvider } from 'astorlm'
import { createCodingTools } from 'astorlm/tools'
const agent = await createLocalAgent({
cwd: process.cwd(), // safe working directory
provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }),
tools: createCodingTools(), // read, write, edit, bash, bash_spawn, bash_get_output, bash_kill, ls, grep, glob
})
agent.on('text', (text) => process.stdout.write(text))
agent.on('tool-start', (t) => console.log(`\n๐ ๏ธ [tool: ${t.name}]`, t.input))
await agent.run('Refactor src/utils.ts to use arrow functions.')
console.log(agent.getUsage()) // { inputTokens, outputTokens, cacheReadTokens?, cacheCreationTokens? }
```
---
### ๐๏ธ 3. Session & History Persistence (FileSessionManager)
Persist conversation history on disk to resume the agent's work or branch it at any point.
```typescript
import { createLocalAgent, FileSessionManager, OpenAIProvider } from 'astorlm'
// 1. On-disk persister (creates a .jsonl history file + .meta.json per session)
const sessionManager = new FileSessionManager({ dir: './.astor-sessions' })
// 2. Load or create the persistent session. createLocalAgent is async โ it
// loads prior history before resolving, so by the time you have `agent`
// it's ready to run(). There is no separate initPromise.
const agent = await createLocalAgent({
sessionId: 'my-refactor-session',
sessionManager,
provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }),
})
await agent.run('Write an optimized fibonacci function.')
```
#### ๐ฟ Session Branching
Create a child session by copying the messages of an existing session (or truncating up to a given message ID):
```typescript
// Either via the session manager directly...
const childState = await sessionManager.create({
parentId: 'my-refactor-session',
branchFromMessageId: 'optional-message-id-cutoff', // omit to clone full history
})
// ...or fork from the live agent:
const child = await agent.fork({ branchFromMessageId: 'optional-message-id-cutoff' })
await child.run('Now rewrite it in TypeScript with strict types?')
```
---
### ๐ญ 4. Simplified Facade with `AstorAgent`
To streamline recurring flows, `AstorAgent` wraps lifecycle management, output subscription, and branching.
```typescript
import { AstorAgent, FileSessionManager, OpenAIProvider } from 'astorlm'
const agent = new AstorAgent({
provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }),
sessionManager: new FileSessionManager({ dir: './.astor-sessions' }),
defaultOutputMode: 'verbose', // 'silent' | 'console' | 'verbose' | (event) => void
})
const { sessionId, text } = await agent.runTask('Create a test.js script that adds 2 + 2')
await agent.runBranchTask({
parentId: sessionId,
promptText: 'Change that script so it subtracts instead of adding',
outputMode: 'console',
})
```
---
## ๐งฌ Subagents (agent-as-tool)
Expose a whole child agent to a parent as a single tool. When the parent calls it, `createSubagentTool` spins up an independent session with its own (typically narrower) system prompt and tool set, runs ONE prompt to completion, and returns the child's final text as the `tool_result`. The parent never sees the child's intermediate turns โ only the distilled answer.
The child inherits the parent's `cwd` / `executor` from the `ToolContext`, and the parent's abort signal propagates (cancelling the parent cancels the child mid-flight). Pure composition over the public API โ no loop changes.
```typescript
import { createLocalAgent, createSubagentTool, OpenAIProvider } from 'astorlm'
import { createReadOnlyTools, createCodingTools } from 'astorlm/tools'
const provider = () =>
new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' })
// A focused subagent with a read-only tool surface.
const explorer = createSubagentTool({
name: 'repo_explorer',
description: 'Delegate repository exploration: list files, read them, summarise. Pass the task in `task`.',
provider: provider(),
systemPrompt: 'You explore repositories with read-only tools and return a concise summary.',
tools: createReadOnlyTools(),
maxTurns: 8,
})
const orchestrator = await createLocalAgent({
provider: provider(),
tools: [explorer, ...createCodingTools()],
})
await orchestrator.run('Understand this project: list the root .ts files and summarise each in one line.')
```
> Returns the child's final text only โ it does not stream the child's intermediate tokens up to the parent.
---
## ๐ช Control Hooks (`SessionHooks`)
Hooks let you intercept the agent loop. Five optional interception points, all can be async:
| Hook | When | Can |
|---|---|---|
| `beforeTurn` | start of each turn | observe `{ turn, messages, ... }` |
| `beforeProviderCall` | before `provider.stream` | **mutate** `{ messages, systemPrompt, tools }` sent to the model |
| `beforeToolExecution` | before each tool | return `{ authorize, mockResult?, steer?, feedback? }` โ permission + mocking + steering in one |
| `afterToolExecution` | after each tool | return the final `output` string the model sees (sanitisation / wrapping) |
| `afterTurn` | end of each turn | observe `{ turn, lastMessage, ... }` |
```typescript
import { createLocalAgent, OpenAIProvider } from 'astorlm'
import { createCodingTools } from 'astorlm/tools'
const agent = await createLocalAgent({
provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }),
tools: createCodingTools(),
hooks: {
beforeProviderCall: async ({ messages, systemPrompt }) => {
return { messages, systemPrompt: `${systemPrompt}\nAlways answer in English.` }
},
beforeToolExecution: async ({ toolName, input }) => {
if (toolName === 'bash') {
const approved = await askUserForPermission((input as any).command)
return { authorize: approved, mockResult: approved ? undefined : 'Command canceled by the operator.' }
}
return { authorize: true }
},
afterToolExecution: async ({ toolName, output, durationMs }) => {
console.log(`[metric] ${toolName} took ${durationMs}ms`)
return output
},
},
})
```
### ๐ฏ Steering (redirect without aborting)
A `beforeToolExecution` hook returning `{ steer: true, feedback }` cancels the turn's tool calls and feeds the model a `[User Steering Feedback]` note on the next turn โ redirecting it without aborting the run. To drive that from *outside* a hook (e.g. a UI button), use `createSteeringController`:
```typescript
import { createLocalAgent, createSteeringController, OpenAIProvider } from 'astorlm'
import { createCodingTools } from 'astorlm/tools'
const controller = createSteeringController() // optionally wraps an existing SessionHooks
const agent = await createLocalAgent({
provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }),
tools: createCodingTools(),
hooks: controller.hooks,
})
// From anywhere (button handler, watcher, another process):
controller.steer('Stop โ do not create files, just list the existing ones.')
await agent.run('Create a file BORRAR.txt with "temp".')
// The queued feedback is consumed at the next tool boundary.
```
> Steering takes effect at the next tool-call boundary, not mid-token.
---
## ๐ MCP Connectivity (Model Context Protocol)
Mount external MCP servers (local stdio or remote HTTP). Their tools are adapted to the agent standard automatically and prefixed `__`.
```typescript
import { createLocalAgent, mountMcpServer, OpenAIProvider } from 'astorlm'
import { createCodingTools } from 'astorlm/tools'
const mcpServer = await mountMcpServer({
name: 'local-fs',
transport: {
type: 'stdio',
command: 'npx',
args: ['-y', '@modelcontextprotocol/server-filesystem', '/allowed/path'],
},
})
const agent = await createLocalAgent({
provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }),
tools: [...createCodingTools(), ...mcpServer.tools],
})
```
### MCP Apps โ interactive UI (SEP-1865)
A mounted tool can return, alongside the text for the model, a sandboxed UI
component (a `ui://` resource, mimeType `text/html;profile=mcp-app`) linked via
`_meta.ui.resourceUri`. astorlm surfaces it on a side channel: the text still
goes to the model, the UI travels separately (it never pollutes the context).
```typescript
const mcp = await mountMcpServer({
name: 'docs',
transport: { type: 'stdio', command: 'npx', args: ['tsx', 'server.ts'] },
// Fires when a tool result carries _meta.ui.resourceUri.
onToolUi: (ui) => {
// ui = { toolName, resourceUri, structuredContent, content }
console.log(ui.toolName, ui.resourceUri, ui.structuredContent)
},
})
// Read the ui:// template to render it host-side.
const view = await mcp.readUiResource('ui://semantic/results')
// view.text = component HTML ยท view.mimeType = 'text/html;profile=mcp-app'
```
`onToolUi` receives `{ toolName, resourceUri, structuredContent, content }`.
`readUiResource(uri)` / `readResource(uri)` read resources from the mounted
server. A full server (semantic search over the embeddings module) plus a host that
renders the component is available as `35-mcp-apps-semantic` in the companion
examples repository, published separately from this one.
---
## ๐ Skills (loadable knowledge packs)
A **skill** is a self-contained piece of instructions the agent can consult โ a markdown body plus metadata. The SDK has no opinion about where skills come from: the `SkillSource` interface is the seam (filesystem, HTTP registry, in-memory, database).
Three activation modes (`skillMode`):
* `'filesystem'` โ **the canonical Agent Skills pattern** (Claude Code, OpenAI Codex, Gemini CLI). The system prompt lists each skill's name, description and the absolute path of its `SKILL.md`; the agent reads it with the standard `read` tool when triggered. No meta-tool. Most robust across models; bundled `scripts/`, `references/`, `assets/` are reachable for free via `read`/`bash`. Requires every skill to expose a path (use `createFileSystemSkillSource`) and the session to include a `read` tool.
* `'on-demand'` (default) โ only `{name, description}` go into the system prompt; the session registers a `load_skill` meta-tool the model calls to materialise a body. Scales to many skills; depends on the model invoking a meta-tool.
* `'all'` โ every body is concatenated into the system prompt up front. Cheapest at runtime; eats context. Use for a small, always-relevant set.
### From the filesystem
Convention: one directory per skill, each with a `SKILL.md` whose YAML frontmatter declares `name` and `description`. The frontmatter `name` is the source of truth and must match the folder name (drift throws).
```
./.astor-skills/
pptx/SKILL.md
refactor/SKILL.md
```
```markdown
---
name: pptx
description: Build PowerPoint decks when the user asks for slides.
# Optional extra fields โ captured into Skill.metadata, ignored by the SDK,
# available for your own policy code (filter by tag, gate by version, etc.).
version: 1.2.0
tags: documents, presentation
allowed-tools: read, write, bash
---
# How to build a deck
Use pptx-genjs. Prefer one slide per concept; keep titles under 60 chars.
```
#### Frontmatter rules (Agent Skills spec)
| Field | Required | Rule |
| ------------- | -------- | ---- |
| `name` | โ | 1โ64 chars, regex `^[a-z0-9][a-z0-9-]*$`. Reserved: `anthropic`, `claude`. |
| `description` | โ | 1โ1024 chars. No XML tags (``) โ they confuse models that emit native tool-call syntax. |
| any other key | โ | Captured into `Skill.metadata: Record` verbatim. |
```typescript
import { createLocalAgent, createFileSystemSkillSource, OpenAIProvider } from 'astorlm'
import { createCodingTools } from 'astorlm/tools'
const agent = await createLocalAgent({
provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }),
tools: createCodingTools(),
skillSources: [createFileSystemSkillSource({ dir: './.astor-skills' })],
skillMode: 'filesystem',
})
await agent.run('Build a deck about climate change.')
```
### Hierarchical discovery (user โ project โ repo)
Passing several `createFileSystemSkillSource` directly throws on duplicate names (no silent override). `createLayeredSkillSource` resolves precedence internally (last-wins, loudly via `onOverride`) and presents itself as a single source, so the registry's conflict-throw stays intact for genuine clashes.
```typescript
import { createLayeredSkillSource } from 'astorlm/core'
const layered = createLayeredSkillSource({
layers: [userSkillsDir, projectSkillsDir], // LOW โ HIGH precedence; project wins
onOverride: ({ name, winner, loser }) =>
console.log(`"${name}": layer "${winner}" shadows "${loser}"`),
})
// pass in skillSources: [layered]
```
### Per-skill `allowed-tools`
The frontmatter `allowed-tools` field is stored verbatim. Interpret and enforce it with two helpers:
```typescript
import { parseAllowedTools, restrictToolsHook } from 'astorlm'
const allowed = parseAllowedTools(skill) ?? [] // CSV โ string[] | null
const hooks = restrictToolsHook(allowed, {
alwaysAllow: ['read', 'load_skill'], // keep these usable regardless
denyMessage: (t) => `Tool "${t}" is not in the active skill's allowed-tools.`,
})
// pass `hooks` into createLocalAgent({ hooks })
```
Deciding *which* skill is active (and therefore which allowlist applies) is left to the consumer โ combine `parseAllowedTools` with your own logic and merge it into the session hooks. The SDK deliberately does not track an "active skill" in the core.
### From an in-memory bundle
```typescript
import { createInMemorySkillSource } from 'astorlm'
const source = createInMemorySkillSource({
name: 'shipped-skills',
skills: [
{
name: 'sql-review',
description: 'Review SQL migrations for safety on a live DB.',
body: '# SQL review\nCheck for table locks, NOT NULL without default, ...',
metadata: { version: '1.0.0', 'allowed-tools': 'read, grep' },
},
],
})
```
Conflicting names across sources throw at session creation โ there is no silent override.
---
## ๐ณ Executors (sandboxing & swappable backends)
Bash-family tools (`bash`, `bash_spawn`, `bash_get_output`, `bash_kill`) never talk to `child_process` directly โ they delegate to an `Executor`, making the execution backend pluggable.
* **`LocalExecutor`** (`astorlm/core`) โ runs commands in the host process. Default for `createLocalAgent`.
* **`DockerExecutor`** (`astorlm/core`) โ runs every command inside a container. Real sandboxing, not a command allowlist.
* **Custom** โ implement the `Executor` interface (`exec`, `spawn`, `getOutput`, `kill`, `dispose`) and pass it via `executor`.
```typescript
import { createLocalAgent, DockerExecutor, OpenAIProvider } from 'astorlm'
import { createCodingTools } from 'astorlm/tools'
const agent = await createLocalAgent({
provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }),
tools: createCodingTools(),
executor: new DockerExecutor({ image: 'node:20-alpine' }),
})
```
The agnostic core ships `createNoopExecutor()` as default โ it throws a clear error if a bash tool tries to use it without explicit configuration, so the SDK never silently runs commands on the host.
---
## ๐งช WASM code sandbox (`astorlm/experimental/wasm-runner`)
`CodeRunner` is a sibling primitive to `Executor`, not a replacement for it. Where `Executor` runs **shell commands** with the host toolchain (isolated by Docker, or not at all), `CodeRunner` runs a **self-contained code snippet** inside a memory-safe WebAssembly runtime โ no filesystem, no network, no host syscalls unless explicitly granted (capability-based, default-deny).
The key difference: it needs no daemon and no `child_process`, so it runs anywhere WASM does (Node, Deno, the browser, edge) โ exactly where `DockerExecutor` cannot reach.
* **`QuickJsCodeRunner`** โ JavaScript via QuickJS compiled to WASM (`quickjs-emscripten`, an optional dependency loaded lazily). Per-run fresh context, enforced wall-clock deadline and memory limit, captured `console`, read-only JSON `globals`.
* **`createCodeRunnerTool({ runner })`** โ wraps a runner as a `run_code` tool. Opt-in: it is *not* part of `createCodingTools()`; register it explicitly.
```typescript
import { createLocalAgent, OpenAIProvider } from 'astorlm'
import { createReadOnlyTools } from 'astorlm/tools'
import { QuickJsCodeRunner, createCodeRunnerTool } from 'astorlm/experimental/wasm-runner'
const runner = new QuickJsCodeRunner({ timeoutMs: 3_000 })
const agent = await createLocalAgent({
provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }),
tools: [...createReadOnlyTools(), createCodeRunnerTool({ runner })],
})
```
| | `Executor` (Docker) | `CodeRunner` (WASM) |
|---|---|---|
| Runs | shell commands + toolchain | self-contained code snippets |
| Where | Node + Docker daemon only | any runtime (Node/Deno/browser/edge) |
| Isolation | OS-level (read-write cwd mount) | capability-empty, default-deny |
| Cold start | hundreds of msโs per command | ~ms |
> โ ๏ธ Experimental API. Python (`PyodideCodeRunner`) is planned behind the same `CodeRunner` interface.
---
## โฑ๏ธ Background processes (`bash_spawn` / `bash_get_output` / `bash_kill`)
Beyond the synchronous `bash` tool, the agent can manage long-running processes:
* `bash_spawn { command }` โ returns an opaque `pid`.
* `bash_get_output { pid }` โ drains stdout/stderr buffered since the last call, plus status / exit code.
* `bash_kill { pid, signal? }` โ terminates the process.
This is what lets the agent launch a dev server, inspect logs, and tear it down without blocking the loop.
---
## ๐งญ Loop patterns
`createAgent` / `createLocalAgent` accept `pattern: 'REACT' | 'PLAN_EXECUTE'` (default `'REACT'`).
`'PLAN_EXECUTE'` auto-registers `add_plan_item` and `update_plan_item` tools that mutate a `PlanItem[]`. Each turn the loop injects the plan state into the system prompt (same idea as a visible, mutable to-do list). The plan persists in session metadata and survives resume/fork; read it with `agent.getPlan()`.
## ๐ Goal loops (`runGoalLoop`)
A single `agent.run()` ends when the model stops calling tools โ which is not the same as the job being *done*. `runGoalLoop` repeats the attempt until a condition you control returns true, giving each iteration a **fresh agent** (and therefore a clean context window) instead of letting one conversation grow unbounded.
The stop condition is yours and should be cheap and deterministic โ run the test suite and check the exit code, assert a file exists โ which is what separates this from "ask the model if it's finished". `maxIterations` (default 10) is a mandatory fuse, so a condition that never holds cannot run forever.
```typescript
import { createLocalAgent, runGoalLoop, OpenAIProvider } from 'astorlm'
import { createCodingTools } from 'astorlm/tools'
const result = await runGoalLoop({
goal: 'Make the test suite pass. Run `pnpm test` to check your work.',
createIterationAgent: () =>
createLocalAgent({
cwd: process.cwd(),
provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }),
tools: createCodingTools(),
}),
isDone: async () => (await runTests()).exitCode === 0, // your own check
maxIterations: 5,
onIteration: ({ iteration, lastText }) => console.log(`#${iteration}: ${lastText.slice(0, 80)}`),
})
console.log(result) // { iterations, done, lastText, stopReason: 'done' | 'max_iterations' | 'aborted' }
```
> Iterations share on-disk state (same `cwd`), not conversation history. That is the point: the work accumulates in the repo, the context does not.
## ๐งฑ Structured output (`generateObject`)
When you want data back rather than prose, `generateObject` constrains the answer to a Zod schema and returns a parsed, typed object โ with repair attempts if the model emits something invalid.
```typescript
import { generateObject, OpenAIProvider } from 'astorlm'
import { z } from 'zod'
const { object } = await generateObject({
provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }),
schema: z.object({
severity: z.enum(['low', 'medium', 'high']),
files: z.array(z.string()),
summary: z.string(),
}),
prompt: 'Triage the failure described in the log below: ...',
})
object.severity // typed as 'low' | 'medium' | 'high'
```
`mode` picks how the constraint is applied:
* `'tool'` โ a synthetic terminal tool (`provide_final_answer`) whose schema *is* the shape. Runs a real agent, so it works on any endpoint with tool calls and reuses the loop's own repair behaviour (`maxTurns`, default 8). This is the mode that lets you pass `tools`, so the model can gather what it needs before delivering the object.
* `'native'` โ the provider's `response_format: json_schema`. Output failing validation is retried up to `maxRepairAttempts` (default 2), then throws `GenerateObjectError`.
* `'auto'` (default) โ `'native'` when no `tools` are passed and the provider is OpenAI-compatible; `'tool'` otherwise. The mode actually used comes back on the result.
## ๐ Context optimizer (auto-compaction)
When the provider exposes a `contextLimit`, the loop runs a structural optimizer between turns that prunes / dedupes once the conversation crosses a threshold. It's structural (prune / dedupe), not LLM-based summarization.
```typescript
const agent = await createLocalAgent({
provider: /* ... */,
contextOptimizer: {
maxTokens: 200_000,
compressThreshold: 0.8, // optimize when usage > 80% of maxTokens
keepRecentTurns: 3, // always keep the last N turns verbatim
},
})
// Disable entirely: contextOptimizer: false
// Implicit default: enabled if provider.contextLimit is set, off otherwise.
```
## ๐ Retry policy for transient provider errors
Opt-in retries for transient failures (HTTP 429, 5xx, network timeouts, streams cut before any chunk). Already-streamed events are never duplicated โ once any event is emitted on an attempt, the loop will not retry that turn.
```typescript
const agent = await createLocalAgent({
provider: /* ... */,
retry: { maxAttempts: 3, baseDelayMs: 500, maxDelayMs: 10_000, jitter: true },
})
// Default (omitted): no retries โ errors propagate and the session closes with session_end: error.
```
## ๐ Token usage tracking
```typescript
await agent.run('...')
console.log(agent.getUsage())
// { inputTokens, outputTokens, cacheReadTokens?, cacheCreationTokens? }
```
No pricing layer โ raw counts only. `cacheReadTokens` / `cacheCreationTokens` are populated when the provider reports them (Anthropic always; OpenAI's `cached_tokens` when applicable; many OpenAI-compat endpoints leave them undefined).
## ๐ซ Heartbeat (proactive loop)
`agent.startHeartbeat(opts?)` / `agent.stopHeartbeat()` (or pass `heartbeat` to the factory) run a check prompt on an interval. A `localCondition(cwd)` enables the **latent heartbeat**: the check runs locally in TypeScript and only wakes the LLM when it returns `true` โ zero token cost until the trigger fires.
---
## ๐ ๏ธ Development Commands
```bash
pnpm install # install dependencies
pnpm build # build the library (dist/ in ESM, CJS, and d.ts) โ respects every entrypoint
pnpm dev # interactive watch-mode build
pnpm test # run the unit test suite (Vitest)
pnpm typecheck # run TypeScript type checking without emitting
```