astorlm

CI npm version MIT license Node >= 20 TypeScript

๐Ÿ“– Documentation site ยท npm ยท Changelog

--- **The agent loop as a library, not as an app.** astorlm gives you the machinery behind a coding agent โ€” the loop, tools, providers, sessions, hooks and executors โ€” as plain TypeScript objects you compose inside your own process. No CLI, no TUI, no daemon to shell out to and no output to scrape. - **Embed it.** `await createAgent({ ... })` returns an object you own. Your process, your logging, your permission prompts, your UI. - **Swap every part.** `Provider`, `Executor`, `SessionManager`, `SkillSource` and `CodeRunner` are interfaces. The built-ins are conveniences, not requirements. - **Bring your own model.** Anthropic, or anything OpenAI-compatible โ€” OpenAI, Groq, OpenRouter, Together, vLLM, Ollama, a local proxy. Small local models are a supported target, not an afterthought. - **Runtime-agnostic core.** The loop, providers, tools, sessions, skills and embeddings import zero Node built-ins. Everything OS-bound lives behind `astorlm/core`. ## โšก Quickstart ```bash pnpm add astorlm # Node >= 20 ``` ```typescript import { createAgent, OpenAIProvider, tool } from 'astorlm' import { z } from 'zod' const getWeather = tool({ name: 'get_weather', description: 'Returns the current temperature for a city', schema: z.object({ city: z.string() }), execute: async ({ city }) => `Weather in ${city}: 22ยฐC, sunny.`, }) // createAgent is async: it loads any persisted state before returning. const agent = await createAgent({ provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', // any OpenAI-compatible endpoint apiKey: 'ollama', }), tools: [getWeather], }) agent.on('text', (chunk) => process.stdout.write(chunk)) await agent.run('How is the weather in Buenos Aires?') console.log(agent.getUsage()) // { inputTokens, outputTokens, ... } ``` That is the whole setup. Want it to touch the filesystem? Swap `createAgent` for [`createLocalAgent`](#-2-full-coding-agent-nodejs) and pass `createCodingTools()`. ## ๐Ÿค” How it compares | | astorlm | | --- | --- | | **vs. a coding CLI** (Claude Code, Codex CLI, Aider) | Those are applications you drive. astorlm is the machinery they are built out of, running in *your* process โ€” so the UI, the audit log and the approval flow are yours to write. | | **vs. an agent framework** (LangGraph, Mastra) | No graph DSL and no workflow engine to learn. One loop, five hooks, and interfaces you implement. The whole public surface is `src/index.ts` and `src/core.ts`. | | **vs. a vendor SDK** | Provider-neutral by construction. Local and weak models get first-class support through [`experimental/edge-boost`](#12-astorlmexperimentaledge-boost-experimental--hardening-for-weak-models), not a "best effort" disclaimer. | ## ๐Ÿ“‹ What's in the box | | | | --- | --- | | **Loop** | Multi-turn, parallel tool execution, token streaming, thinking blocks, cooperative cancellation, `REACT` or `PLAN_EXECUTE` [patterns](#-loop-patterns) | | **Tools** | `read`, `write`, `edit`, `bash` (plus [background spawn/poll/kill](#-background-processes-bash_spawn--bash_get_output--bash_kill)), `ls`, `grep`, `glob` โ€” or [define your own](#-1-minimal-usage-custom-tool) from a Zod schema | | **Isolation** | Swappable [executors](#-executors-sandboxing--swappable-backends): local, Docker, or your own. Plus a [WASM code sandbox](#-wasm-code-sandbox-astorlmexperimentalwasm-runner) that needs no daemon | | **Control** | Five [hooks](#-control-hooks-sessionhooks) covering permissions, mocking, prompt rewriting and output sanitising; [steering](#-steering-redirect-without-aborting) at tool boundaries | | **Memory** | Session persistence and native branching, [context auto-compaction](#-context-optimizer-auto-compaction), [token accounting](#-token-usage-tracking) | | **Interop** | [MCP](#-mcp-connectivity-model-context-protocol) over stdio and HTTP (including MCP Apps UI), and [Agent Skills](#-skills-loadable-knowledge-packs) in the same filesystem format Claude Code and Codex use | | **Composition** | [Subagents as tools](#-subagents-agent-as-tool), [goal loops](#-goal-loops-rungoalloop), [structured output](#-structured-output-generateobject), [heartbeats](#-heartbeat-proactive-loop) | | **Observability** | Tracing with an OTLP exporter, metrics with cost accounting, deterministic record & replay, and an offline eval harness โ€” all under `astorlm/experimental/*` | ## ๐Ÿ“– Table of contents - [โšก Quickstart](#-quickstart) - [๐Ÿค” How it compares](#-how-it-compares) - [๐Ÿ“‹ What's in the box](#-whats-in-the-box) - [๐Ÿ“ฆ Module Layout (Entrypoints)](#-module-layout-entrypoints) - [๐Ÿš€ Quick Use Examples](#-quick-use-examples) - [๐Ÿงฌ Subagents (agent-as-tool)](#-subagents-agent-as-tool) - [๐Ÿช Control Hooks (`SessionHooks`)](#-control-hooks-sessionhooks) - [๐Ÿ”Œ MCP Connectivity (Model Context Protocol)](#-mcp-connectivity-model-context-protocol) - [๐Ÿ“š Skills (loadable knowledge packs)](#-skills-loadable-knowledge-packs) - [๐Ÿณ Executors (sandboxing & swappable backends)](#-executors-sandboxing--swappable-backends) - [๐Ÿงช WASM code sandbox (`astorlm/experimental/wasm-runner`)](#-wasm-code-sandbox-astorlmexperimentalwasm-runner) - [โฑ๏ธ Background processes (`bash_spawn` / `bash_get_output` / `bash_kill`)](#-background-processes-bash_spawn--bash_get_output--bash_kill) - [๐Ÿงญ Loop patterns](#-loop-patterns) - [๐Ÿ”„ Goal loops (`runGoalLoop`)](#-goal-loops-rungoalloop) - [๐Ÿงฑ Structured output (`generateObject`)](#-structured-output-generateobject) - [๐Ÿ“‰ Context optimizer (auto-compaction)](#-context-optimizer-auto-compaction) - [๐Ÿ” Retry policy for transient provider errors](#-retry-policy-for-transient-provider-errors) - [๐Ÿ“Š Token usage tracking](#-token-usage-tracking) - [๐Ÿซ€ Heartbeat (proactive loop)](#-heartbeat-proactive-loop) - [๐Ÿ› ๏ธ Development Commands](#-development-commands) --- ## ๐Ÿ“ฆ Module Layout (Entrypoints) AstorLM ships clearly separated entrypoints. The runtime-agnostic pieces โ€” loop, providers, tools, sessions, skills, embeddings โ€” live in the main `astorlm` barrel and import zero Node built-ins. Everything OS-bound lives behind `astorlm/core`. ```text astorlm agnostic createAgent ยท Anthropic/OpenAIProvider ยท tool ToolRegistry ยท SessionHooks ยท InMemorySessionManager createSubagentTool ยท createSteeringController โ””โ”€ re-exports ./core for convenience โš ๏ธ pulls Node deps into the main barrel astorlm/core Node createLocalAgent ยท FileSessionManager ยท AstorAgent mountMcpServer ยท Local/DockerExecutor create{FileSystem,Layered}SkillSource astorlm/tools Node createCodingTools() / createReadOnlyTools() read write edit bash bash_spawn ls grep glob astorlm/embeddings agnostic createOpenAIEmbedder ยท createSemanticIndex astorlm/experimental/* varies error-registry ยท tracing ยท metrics ยท replay evals ยท wasm-runner ยท edge-boost ``` ### 1. `astorlm` (main barrel) * **Description**: One-stop import. Exposes the agnostic core (`createAgent`, the providers, `tool`, `ToolRegistry`, sessions, skills, hooks, optimizer, retry, the subagent/steering helpers) **and** re-exports everything from `astorlm/core` for unified imports. * **Key exports**: - `createAgent` (async โ€” `Promise`) - `InMemorySessionManager`, `SessionManager` - `AnthropicProvider`, `OpenAIProvider` - `createOpenAIEmbedder`, `cosineSimilarity`, `createSemanticIndex`, `withEmbeddingCache` (embeddings โ€” also at `astorlm/embeddings`) - `tool`, `ToolRegistry` - `EventBus`, `buildSystemPrompt`, `estimateTokens`, `optimizeContext`, `isTransientError`, `computeBackoffDelay` - `createSubagentTool` (subagents / agent-as-tool), `createSteeringController` (out-of-hook steering) - `generateObject` (schema-constrained output), `runGoalLoop` (iterate until a condition holds) - `createNoopExecutor` (the safe default executor of the agnostic core) - `SkillRegistry`, `createInMemorySkillSource`, `parseSkillFrontmatter`, `renderSkillsBlock`, `createLoadSkillTool` - `validateSkillName`, `validateSkillDescription`, `validateSkillSpec`, `SkillValidationError`, `SKILL_VALIDATION_LIMITS` - `parseAllowedTools`, `restrictToolsHook` (per-skill tool gating) - Base types: `Agent`, `CreateAgentOptions`, `SessionHooks`, `Message`, `ContentBlock`, `AgentEvent`, `Skill`, `SkillMetadata`, `SkillSource`, `SkillMode`, etc. > โš ๏ธ Because the main barrel re-exports `./core`, importing from `astorlm` pulls in Node-only dependencies. If you target Edge / Workers / browser, import from the agnostic modules directly (the loop, providers and abstractions don't import Node) and avoid the Node runner. ### 2. `astorlm/core` (Node.js runner) * **Description**: Extensions that require Node.js OS-native APIs (`node:fs`, `node:path`, `node:child_process`). * **Key exports**: - `createLocalAgent` (session factory wired with `LocalExecutor` + a local file reader by default). - `FileSessionManager` (history persistence as JSONL plus JSON metadata). - `AstorAgent` (simplified execution and branching facade). - `mountMcpServer` (adapter and Stdio/HTTP transports for Model Context Protocol clients). - `LocalExecutor`, `DockerExecutor` (shell-command backends; see "Executors" below). - `createFileSystemSkillSource` (reads skills from a `dir//SKILL.md` layout). - `createLayeredSkillSource` (hierarchical skill discovery โ€” user โ†’ project โ†’ repo). ### 3. `astorlm/tools` (Built-in tools) * **Description**: A bundle of filesystem-manipulation and analysis tools tuned for coding agents, with path-traversal protection. * **Key exports**: - `createCodingTools()` (`read`, `write`, `edit`, `bash`, `bash_spawn`, `bash_get_output`, `bash_kill`, `ls`, `grep`, `glob`). - `createReadOnlyTools()` (safe variant โ€” no writes or execution: `read`, `ls`, `grep`, `glob`). - Individual tools: `readTool`, `writeTool`, `editTool`, `bashTool`, `bashSpawnTool`, `bashGetOutputTool`, `bashKillTool`, `lsTool`, `grepTool`, `globTool`. ### 4. `astorlm/embeddings` (Embeddings & semantic search) * **Description**: First-class, runtime-agnostic embeddings primitives โ€” sit next to the providers in the main barrel and are also reachable via this dedicated subpath. `fetch`-based, zero extra dependencies, work anywhere `fetch` exists (Node, Deno, browsers, edge). * **Key exports**: - `createOpenAIEmbedder({ baseURL, model, apiKey?, dimensions? })` โ€” OpenAI-compatible `Embedder` with `embed` (single) and `embedMany` (batch, one round-trip) plus token `usage`. - `createSemanticIndex({ embedder })` โ€” in-memory vector store: `add` / `addMany` / `query(text, { topK, threshold })` / `queryByVector` / `remove` / `clear`. The reusable primitive behind semantic search, RAG retrieval and dedupe. - `withEmbeddingCache(embedder)` โ€” memoizes identical inputs so repeated lookups don't re-embed (or re-bill); `embedMany` only requests the cache misses. - `cosineSimilarity` (standard `[-1..1]`), `dotProduct`, `euclideanDistance`. - Types: `Embedder`, `EmbedResult`, `EmbedManyResult`, `SemanticIndex`, `SemanticHit`, etc. ```typescript import { createOpenAIEmbedder, createSemanticIndex } from 'astorlm' const embedder = createOpenAIEmbedder({ baseURL: 'http://localhost:11434/v1', model: 'nomic-embed-text', apiKey: 'ollama', }) const index = createSemanticIndex({ embedder }) await index.addMany([ { id: 'oom', text: 'Node process crashes with out of memory during the build.' }, { id: 'tls', text: 'TLS handshake fails: expired certificate in production.' }, ]) const hits = await index.query('ran out of RAM while compiling', { topK: 1 }) // โ†’ [{ id: 'oom', score: 0.8โ€ฆ, text: 'โ€ฆ' }] ``` ### 5. `astorlm/prompt` (System-prompt building & composition) * **Description**: The system-prompt layer, runtime-agnostic. `buildSystemPrompt` is what the loop uses internally; the prompt *compiler* is the opt-in half โ€” it merges prompt fragments coming from different owners (a base identity, a product policy, a per-skill instruction) into a single string, deduping repeats and **reporting contradictions** instead of silently concatenating them. * **Key exports**: - `buildSystemPrompt(opts)`, `DEFAULT_SYSTEM_PROMPT` โ€” also reachable from the main barrel. - `compilePrompts({ modules, layout?, deduplicate?, detectConflicts? })` โ†’ `{ systemPrompt, report, print() }`. Modules are emitted in `layout` order (default `identity โ†’ context โ†’ constraint โ†’ policy โ†’ format`); modules sharing an `id` resolve by `priority` (highest wins). - `formatPromptReport(report)` โ€” readable audit of what was kept, overridden, deduped or flagged as conflicting. ```typescript import { compilePrompts, formatPromptReport } from 'astorlm/prompt' const { systemPrompt, report } = compilePrompts({ modules: [ { id: 'identity', kind: 'identity', content: 'You are a release engineer.' }, { id: 'lang', kind: 'policy', content: 'Always answer in English.' }, { id: 'lang', kind: 'policy', content: 'Always answer in Spanish.', priority: 10 }, // wins ], }) console.log(formatPromptReport(report)) // shows the override and any conflicts ``` ### 6. `astorlm/experimental/error-registry` (Experimental โ€” Federated Error Registry) > โš ๏ธ **Experimental**. Lives under a dedicated subpath, not the main barrel. The import path itself is the signal that the API is volatile and may change between minor releases. * **Description**: A registry of agent-encountered errors and human-approved resolutions. When an agent hits an error that another agent (or a previous run) has already resolved, the registry injects the fix as a hint into the next `tool_result` โ€” the agent applies the known solution instead of fighting through it again. Honest single-org PoC; federation across organizations and full secret sanitization are out of scope. * **Key exports**: - `createErrorRegistry(opts)` โ€” JSONL append-only store (or in-memory) with optional OpenAI-compatible embeddings and a Jaccard fallback. - `errorRegistryHooks({ registry, context, successWindow? })` โ€” returns a `SessionHooks` object that wires the session to the registry. ```typescript import { createLocalAgent, OpenAIProvider } from 'astorlm' import { createCodingTools } from 'astorlm/tools' import { createErrorRegistry, errorRegistryHooks } from 'astorlm/experimental/error-registry' const registry = createErrorRegistry({ storePath: '.astorlm/error-registry.jsonl', // Optional โ€” if omitted, falls back to Jaccard over tokens: // embeddings: { baseURL: 'http://localhost:11434/v1', model: 'nomic-embed-text' }, }) await registry.init() const agent = await createLocalAgent({ provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }), tools: createCodingTools(), hooks: errorRegistryHooks({ registry, context: { cwd: process.cwd(), osPlatform: process.platform, nodeVersion: process.version, tags: [] }, }), }) // Note: errorRegistryHooks returns a SessionHooks object โ€” if you already // have your own hooks, merge them into a single object before passing. ``` A human approves pending resolutions asynchronously (e.g. `registry.approveResolution(id, approver)`). Until approved, a candidate resolution is not suggested to other sessions. ### 7. `astorlm/experimental/tracing` (Experimental โ€” Observability) > โš ๏ธ **Experimental**. Volatile API behind a dedicated subpath. Runtime-agnostic (reads only the event bus; uses Web Crypto for ids). * **Description**: Derives a hierarchical span tree (`session โ†’ turn โ†’ provider_call | tool_execution`) from the agent's event bus **without touching the loop** โ€” attaching a tracer is pure subscription. Spans carry OpenTelemetry GenAI semantic-convention attributes (`gen_ai.*`) plus astorlm-specific ones (TTFT, tool duration/errors, retries). * **Key exports**: - `attachTracer(agent, { exporter })` โ€” subscribes to the bus and builds spans; returns `{ detach(), currentTraceId() }`. - `createInMemoryExporter()` โ€” collects spans in an array for tests / local inspection. - `createTracer(opts?)` โ€” low-level span factory (usually managed by `attachTracer`). - `createOtlpSpanExporter(opts)` (from `astorlm/experimental/tracing/otel`) โ€” OTLP/HTTP (JSON) exporter built on `fetch` alone, **no OpenTelemetry SDK dependency**. Ships spans to any OTLP collector (OpenTelemetry Collector, Tempo, Jaeger, Honeycomb, Arize Phoenix, Langfuse). ```typescript import { createLocalAgent, OpenAIProvider } from 'astorlm' import { createCodingTools } from 'astorlm/tools' import { attachTracer, createInMemoryExporter } from 'astorlm/experimental/tracing' import { createOtlpSpanExporter } from 'astorlm/experimental/tracing/otel' const agent = await createLocalAgent({ provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }), tools: createCodingTools(), }) // In-memory (inspect locally) ... const memory = createInMemoryExporter() // ... and/or export to an OTLP collector: const otlp = createOtlpSpanExporter({ endpoint: 'http://localhost:4318/v1/traces', serviceName: 'my-agent' }) const tracer = attachTracer(agent, { exporter: { export: (spans) => { memory.export(spans); otlp.export(spans) } }, }) await agent.run('List the .ts files and count them.') tracer.detach() await otlp.shutdown() // final flush for (const s of memory.spans) console.log(s.kind, s.name, s.endTime! - s.startTime, 'ms') ``` The OTLP exporter buffers spans and flushes by batch size (`maxBatch`, default 256) or on a timer (`flushIntervalMs`, default 5s, `unref()`-ed). Hex `trace_id`/`span_id` are forwarded verbatim per the OTLP/JSON convention. ### 8. `astorlm/experimental/metrics` (Experimental โ€” Cost & metrics) > โš ๏ธ **Experimental**. Runtime-agnostic (reads only the event bus). * **Description**: Aggregates operational metrics from the agent's event bus and, given a pricing table, the USD cost of a run. No prices are hardcoded โ€” you supply the table (USD per 1M tokens). * **Key exports**: - `attachMetrics(agent, { pricing? })` โ€” subscribes to the bus; returns `{ snapshot(), reset(), detach() }`. The snapshot has counts (`runs`, `turns`, `providerCalls`, `toolCalls`, `toolErrors`, `providerRetries`), `tokens`, optional `costUsd`, and latency stats (`ttftMs`, `turnMs`, `toolMs`). - `computeCost(usage, pricing)`, `resolvePricing(table, model)` โ€” usable standalone (exact then longest-prefix model match). ```typescript import { attachMetrics } from 'astorlm/experimental/metrics' const metrics = attachMetrics(agent, { pricing: { 'gpt-4o': { inputPer1M: 2.5, outputPer1M: 10, cacheReadPer1M: 1.25 } }, }) await agent.run('...') const m = metrics.snapshot() // { costUsd, latency: { ttftMs: { avg, ... } }, tokens, ... } ``` ### 9. `astorlm/experimental/replay` (Experimental โ€” Record & replay) > โš ๏ธ **Experimental**. Runtime-agnostic; the `Recording` is a plain serializable object. * **Description**: Captures exactly the provider events a run produced and replays them later with no network and no token spend โ€” the deterministic debugging primitive. Capture is at the provider boundary, so it's independent of tools, hooks and timing. * **Key exports**: - `createRecordingProvider(inner)` โ€” wraps a real provider, passes events through and records them; `getRecording()` returns a serializable object. - `createReplayProvider(recording, { onExhausted? })` โ€” a provider that replays the recording turn by turn. ```typescript import { createRecordingProvider, createReplayProvider } from 'astorlm/experimental/replay' const rec = createRecordingProvider(realProvider) const agent = await createLocalAgent({ provider: rec, tools }) await agent.run('...') fs.writeFileSync('run.json', JSON.stringify(rec.getRecording())) // later โ€” same events, no model call const recording = JSON.parse(fs.readFileSync('run.json', 'utf8')) const replay = await createLocalAgent({ provider: createReplayProvider(recording), tools }) await replay.run('...') ``` ### 10. `astorlm/experimental/evals` (Experimental โ€” Offline evaluation) > โš ๏ธ **Experimental**. Runtime-agnostic core; `llmJudge` needs a `Provider` (point it at a local OpenAI-compatible endpoint). * **Description**: Runs a dataset of cases through fresh agents, applies scorers, and aggregates a report (overall pass rate + per-scorer stats). Built for CI gating; pair the agent factory with the replay provider for fast, network-free regression runs. * **Key exports**: - `runEval({ dataset, createAgent, scorers, concurrency?, onResult? })` โ†’ `EvalReport`. - Scorers: `exactMatch`, `contains`, `regexMatch`, `toolTrajectory` (tool-call sequence: `'exact' | 'ordered-subset' | 'set'`), `llmJudge` (LLM-as-judge with a rubric โ†’ normalized 0..1). ```typescript import { runEval, contains, toolTrajectory, llmJudge } from 'astorlm/experimental/evals' const report = await runEval({ dataset: [{ id: 'q1', input: 'How many .ts files?', expected: 'a number' }], createAgent: () => createLocalAgent({ provider: provider(), tools: createReadOnlyTools() }), scorers: [contains('.ts'), toolTrajectory(['ls'], { mode: 'set' }), llmJudge({ provider: provider(), rubric: '...' })], }) if (report.summary.passRate < 0.8) process.exit(1) // CI gate ``` ### 11. `astorlm/experimental/wasm-runner` (Experimental โ€” WASM code sandbox) > โš ๏ธ **Experimental**. Dedicated subpath, volatile API. * **Description**: `CodeRunner` โ€” runs untrusted *source code* (not shell commands) inside a memory-safe WebAssembly sandbox with no host access. The daemon-free, runtime-agnostic isolation tier, complementary to `DockerExecutor`. See ["WASM code sandbox"](#-wasm-code-sandbox-astorlmexperimentalwasm-runner) below. * **Key exports**: `QuickJsCodeRunner` (JS via QuickJS-wasm), `createCodeRunnerTool({ runner })` (opt-in `run_code` tool), plus the `CodeRunner` / `RunCodeOptions` / `RunCodeResult` types. ### 12. `astorlm/experimental/edge-boost` (Experimental โ€” hardening for weak models) > โš ๏ธ **Experimental**. Dedicated subpath, volatile API. Not re-exported from the main barrel. * **Description**: `edgeBoost(options, tuning?)` โ€” a pure options-transformer that hardens an agent for weak / local / free-tier models (free OpenRouter tiers, small Ollama/vLLM models, heavily quantized checkpoints). It defends against the failure where a model, on a long multi-turn tool context (typically the synthesis turn after several tool round-trips), degenerates into an infinite repetition loop or stalls without producing content until a timeout fires. Four opt-in defenses, none of which change default behavior unless applied: - **Stream guard** (`createGuardedProvider`) โ€” detects degeneration (identical-delta repetition, no-content stall, reasoning-budget blowout) and retries cheaply against the same endpoint. Owns *degeneration* retries; the loop-level `retry` policy still owns *network/HTTP* retries. Model failover between endpoints is out of scope by design (that belongs to a routing proxy). - **Context diet** โ€” per-call truncation of old `tool_result` blocks (idempotent, coexists with the core optimizer's markers) + an optional compact system prompt; never mutates session history. Optional LLM `digest` summarizes an evicted block with an injected cheap provider instead of truncating (cached per block, falls back to truncation on failure/timeout). - **Forced synthesis** โ€” after N tool rounds in the current request, appends a synthesis instruction and forces `toolChoice: 'none'` (or strips tools) so the model stops calling tools and answers. - **Sampling defaults** โ€” injects `max_tokens` and a `frequencyPenalty` (repetition penalty) when the call omits them. Plus optional **semantic tool pruning** (top-K relevant tools per turn via an embedder; disabled by default, degrades gracefully if the embedder is down). * **Key exports**: `edgeBoost`, `edgeBoostHooks`, `createGuardedProvider`, `DegenerationError`, `mergeSessionHooks` (generally useful hook composition), `buildContextDietHook` / `buildSynthesisHook` / `buildToolPruningHook`, `resolveEdgeBoostTuning`, plus the `EdgeBoostTuning` / `GuardOptions` / `ContextDietOptions` / `SynthesisOptions` / `SamplingDefaults` / `ToolPruningOptions` types. ```ts import { createLocalAgent } from 'astorlm/core' import { OpenAIProvider } from 'astorlm' import { createCodingTools } from 'astorlm/tools' import { edgeBoost } from 'astorlm/experimental/edge-boost' // Wrap the same options you already pass to createLocalAgent โ€” one call, no layers. const agent = await createLocalAgent(edgeBoost({ cwd: process.cwd(), provider: new OpenAIProvider({ model: 'myproxyllm', baseURL: 'http://127.0.0.1:11434/v1', apiKey: 'not-needed' }), tools: createCodingTools(), })) // Tune any section; pass `false` to disable one. Example: force synthesis earlier. const tuned = edgeBoost(options, { synthesis: { forceAfterToolRounds: 2 } }) ``` A runnable `34-edge-boost` example (`pnpm start 34`, with optional `--prune` / `--digest` flags) lives in the companion examples repository, which is published separately from this one. ### 13. `astorlm/experimental/contract` (Experimental โ€” Agent Contract) > โš ๏ธ **Experimental**. Dedicated subpath, volatile API. Node-bound (uses `node:path` for path matching). * **Description**: A declarative budget-and-permission envelope for a run, enforced through hooks instead of trust. You declare what the agent may spend and touch; `createContractHooks` turns that into a `SessionHooks` object that stops the run with a `ContractViolationError` the moment a rule is crossed. Complementary to `restrictToolsHook` (tools only) and to `DockerExecutor` (isolates, but does not budget). * **Key exports**: - `createContractHooks(contract)` โ†’ `SessionHooks`, ready to pass to `createAgent` / `createLocalAgent`. - `ContractValidator` โ€” the enforcement engine on its own, if you'd rather drive it yourself. - `ContractViolationError` (carries the `rule` that failed), `globToRegex`, plus the `AgentContract` / `AgentContractBudget` / `AgentContractTools` / `AgentContractSandbox` types. ```typescript import { createContractHooks } from 'astorlm/experimental/contract' const hooks = createContractHooks({ budget: { maxTurns: 12, maxTotalTokens: 200_000, maxDurationMs: 5 * 60_000 }, tools: { allow: ['read', 'ls', 'grep', 'glob', 'bash'] }, sandbox: { allowedPaths: ['src/**', 'tests/**'], deniedPaths: ['**/.env', '**/secrets/**'], bash: { deniedCommands: ['rm', 'curl', 'git push'] }, }, }) // pass into createLocalAgent({ hooks }) ``` --- ## ๐Ÿš€ Quick Use Examples > **All snippets use `OpenAIProvider`** pointed at a local OpenAI-compatible endpoint. The examples assume [Ollama](https://ollama.com) (`http://localhost:11434/v1`, model `qwen2.5-coder`, `apiKey: 'ollama'` โ€” a placeholder local endpoints ignore), but any OpenAI-compatible server works (LM Studio `http://localhost:1234/v1`, vLLM `http://localhost:8000/v1`, โ€ฆ). > > **Hosted providers:** point `baseURL` at the vendor and pass a real key โ€” e.g. OpenAI (`https://api.openai.com/v1`, `gpt-4o-mini`), Groq, OpenRouter or Together. If you omit `apiKey`, the provider reads it from `OPENAI_API_KEY` (override the env var name with `envVar`). `AnthropicProvider` exists in the public API with the same shape โ€” swap it in if you prefer Anthropic. It additionally takes `thinking` (`{ type: 'adaptive', display? }` on current models, or `{ budget_tokens }` on older ones), `effort` (`'low'` โ€ฆ `'max'`), and `contextLimit`, which defaults to a conservative 200k โ€” raise it to match the model you target. ### ๐Ÿ”Œ 1. Minimal Usage (custom tool) ```typescript import { createAgent, OpenAIProvider, tool } from 'astorlm' import { z } from 'zod' // 1. Define a custom tool (Zod schema โ†’ JSON Schema under the hood) const getWeather = tool({ name: 'get_weather', description: 'Returns the current temperature for a city', schema: z.object({ city: z.string() }), execute: async ({ city }) => `Weather in ${city}: 22ยฐC, sunny.`, }) // 2. Create the session. // Note: createAgent is async โ€” it loads any persisted state from the // session manager before returning. Always await it. const agent = await createAgent({ provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }), tools: [getWeather], }) // 3. Subscribe to the token stream agent.on('text', (text) => process.stdout.write(text)) // 4. Run the prompt await agent.run('How is the weather in Buenos Aires?') ``` --- ### ๐Ÿ’ป 2. Full Coding Agent (Node.js) The standard setup for building an autonomous backend coding agent with local filesystem access. ```typescript import { createLocalAgent, OpenAIProvider } from 'astorlm' import { createCodingTools } from 'astorlm/tools' const agent = await createLocalAgent({ cwd: process.cwd(), // safe working directory provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }), tools: createCodingTools(), // read, write, edit, bash, bash_spawn, bash_get_output, bash_kill, ls, grep, glob }) agent.on('text', (text) => process.stdout.write(text)) agent.on('tool-start', (t) => console.log(`\n๐Ÿ› ๏ธ [tool: ${t.name}]`, t.input)) await agent.run('Refactor src/utils.ts to use arrow functions.') console.log(agent.getUsage()) // { inputTokens, outputTokens, cacheReadTokens?, cacheCreationTokens? } ``` --- ### ๐Ÿ—ƒ๏ธ 3. Session & History Persistence (FileSessionManager) Persist conversation history on disk to resume the agent's work or branch it at any point. ```typescript import { createLocalAgent, FileSessionManager, OpenAIProvider } from 'astorlm' // 1. On-disk persister (creates a .jsonl history file + .meta.json per session) const sessionManager = new FileSessionManager({ dir: './.astor-sessions' }) // 2. Load or create the persistent session. createLocalAgent is async โ€” it // loads prior history before resolving, so by the time you have `agent` // it's ready to run(). There is no separate initPromise. const agent = await createLocalAgent({ sessionId: 'my-refactor-session', sessionManager, provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }), }) await agent.run('Write an optimized fibonacci function.') ``` #### ๐ŸŒฟ Session Branching Create a child session by copying the messages of an existing session (or truncating up to a given message ID): ```typescript // Either via the session manager directly... const childState = await sessionManager.create({ parentId: 'my-refactor-session', branchFromMessageId: 'optional-message-id-cutoff', // omit to clone full history }) // ...or fork from the live agent: const child = await agent.fork({ branchFromMessageId: 'optional-message-id-cutoff' }) await child.run('Now rewrite it in TypeScript with strict types?') ``` --- ### ๐ŸŽญ 4. Simplified Facade with `AstorAgent` To streamline recurring flows, `AstorAgent` wraps lifecycle management, output subscription, and branching. ```typescript import { AstorAgent, FileSessionManager, OpenAIProvider } from 'astorlm' const agent = new AstorAgent({ provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }), sessionManager: new FileSessionManager({ dir: './.astor-sessions' }), defaultOutputMode: 'verbose', // 'silent' | 'console' | 'verbose' | (event) => void }) const { sessionId, text } = await agent.runTask('Create a test.js script that adds 2 + 2') await agent.runBranchTask({ parentId: sessionId, promptText: 'Change that script so it subtracts instead of adding', outputMode: 'console', }) ``` --- ## ๐Ÿงฌ Subagents (agent-as-tool) Expose a whole child agent to a parent as a single tool. When the parent calls it, `createSubagentTool` spins up an independent session with its own (typically narrower) system prompt and tool set, runs ONE prompt to completion, and returns the child's final text as the `tool_result`. The parent never sees the child's intermediate turns โ€” only the distilled answer. The child inherits the parent's `cwd` / `executor` from the `ToolContext`, and the parent's abort signal propagates (cancelling the parent cancels the child mid-flight). Pure composition over the public API โ€” no loop changes. ```typescript import { createLocalAgent, createSubagentTool, OpenAIProvider } from 'astorlm' import { createReadOnlyTools, createCodingTools } from 'astorlm/tools' const provider = () => new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }) // A focused subagent with a read-only tool surface. const explorer = createSubagentTool({ name: 'repo_explorer', description: 'Delegate repository exploration: list files, read them, summarise. Pass the task in `task`.', provider: provider(), systemPrompt: 'You explore repositories with read-only tools and return a concise summary.', tools: createReadOnlyTools(), maxTurns: 8, }) const orchestrator = await createLocalAgent({ provider: provider(), tools: [explorer, ...createCodingTools()], }) await orchestrator.run('Understand this project: list the root .ts files and summarise each in one line.') ``` > Returns the child's final text only โ€” it does not stream the child's intermediate tokens up to the parent. --- ## ๐Ÿช Control Hooks (`SessionHooks`) Hooks let you intercept the agent loop. Five optional interception points, all can be async: | Hook | When | Can | |---|---|---| | `beforeTurn` | start of each turn | observe `{ turn, messages, ... }` | | `beforeProviderCall` | before `provider.stream` | **mutate** `{ messages, systemPrompt, tools }` sent to the model | | `beforeToolExecution` | before each tool | return `{ authorize, mockResult?, steer?, feedback? }` โ€” permission + mocking + steering in one | | `afterToolExecution` | after each tool | return the final `output` string the model sees (sanitisation / wrapping) | | `afterTurn` | end of each turn | observe `{ turn, lastMessage, ... }` | ```typescript import { createLocalAgent, OpenAIProvider } from 'astorlm' import { createCodingTools } from 'astorlm/tools' const agent = await createLocalAgent({ provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }), tools: createCodingTools(), hooks: { beforeProviderCall: async ({ messages, systemPrompt }) => { return { messages, systemPrompt: `${systemPrompt}\nAlways answer in English.` } }, beforeToolExecution: async ({ toolName, input }) => { if (toolName === 'bash') { const approved = await askUserForPermission((input as any).command) return { authorize: approved, mockResult: approved ? undefined : 'Command canceled by the operator.' } } return { authorize: true } }, afterToolExecution: async ({ toolName, output, durationMs }) => { console.log(`[metric] ${toolName} took ${durationMs}ms`) return output }, }, }) ``` ### ๐ŸŽฏ Steering (redirect without aborting) A `beforeToolExecution` hook returning `{ steer: true, feedback }` cancels the turn's tool calls and feeds the model a `[User Steering Feedback]` note on the next turn โ€” redirecting it without aborting the run. To drive that from *outside* a hook (e.g. a UI button), use `createSteeringController`: ```typescript import { createLocalAgent, createSteeringController, OpenAIProvider } from 'astorlm' import { createCodingTools } from 'astorlm/tools' const controller = createSteeringController() // optionally wraps an existing SessionHooks const agent = await createLocalAgent({ provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }), tools: createCodingTools(), hooks: controller.hooks, }) // From anywhere (button handler, watcher, another process): controller.steer('Stop โ€” do not create files, just list the existing ones.') await agent.run('Create a file BORRAR.txt with "temp".') // The queued feedback is consumed at the next tool boundary. ``` > Steering takes effect at the next tool-call boundary, not mid-token. --- ## ๐Ÿ”Œ MCP Connectivity (Model Context Protocol) Mount external MCP servers (local stdio or remote HTTP). Their tools are adapted to the agent standard automatically and prefixed `__`. ```typescript import { createLocalAgent, mountMcpServer, OpenAIProvider } from 'astorlm' import { createCodingTools } from 'astorlm/tools' const mcpServer = await mountMcpServer({ name: 'local-fs', transport: { type: 'stdio', command: 'npx', args: ['-y', '@modelcontextprotocol/server-filesystem', '/allowed/path'], }, }) const agent = await createLocalAgent({ provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }), tools: [...createCodingTools(), ...mcpServer.tools], }) ``` ### MCP Apps โ€” interactive UI (SEP-1865) A mounted tool can return, alongside the text for the model, a sandboxed UI component (a `ui://` resource, mimeType `text/html;profile=mcp-app`) linked via `_meta.ui.resourceUri`. astorlm surfaces it on a side channel: the text still goes to the model, the UI travels separately (it never pollutes the context). ```typescript const mcp = await mountMcpServer({ name: 'docs', transport: { type: 'stdio', command: 'npx', args: ['tsx', 'server.ts'] }, // Fires when a tool result carries _meta.ui.resourceUri. onToolUi: (ui) => { // ui = { toolName, resourceUri, structuredContent, content } console.log(ui.toolName, ui.resourceUri, ui.structuredContent) }, }) // Read the ui:// template to render it host-side. const view = await mcp.readUiResource('ui://semantic/results') // view.text = component HTML ยท view.mimeType = 'text/html;profile=mcp-app' ``` `onToolUi` receives `{ toolName, resourceUri, structuredContent, content }`. `readUiResource(uri)` / `readResource(uri)` read resources from the mounted server. A full server (semantic search over the embeddings module) plus a host that renders the component is available as `35-mcp-apps-semantic` in the companion examples repository, published separately from this one. --- ## ๐Ÿ“š Skills (loadable knowledge packs) A **skill** is a self-contained piece of instructions the agent can consult โ€” a markdown body plus metadata. The SDK has no opinion about where skills come from: the `SkillSource` interface is the seam (filesystem, HTTP registry, in-memory, database). Three activation modes (`skillMode`): * `'filesystem'` โ€” **the canonical Agent Skills pattern** (Claude Code, OpenAI Codex, Gemini CLI). The system prompt lists each skill's name, description and the absolute path of its `SKILL.md`; the agent reads it with the standard `read` tool when triggered. No meta-tool. Most robust across models; bundled `scripts/`, `references/`, `assets/` are reachable for free via `read`/`bash`. Requires every skill to expose a path (use `createFileSystemSkillSource`) and the session to include a `read` tool. * `'on-demand'` (default) โ€” only `{name, description}` go into the system prompt; the session registers a `load_skill` meta-tool the model calls to materialise a body. Scales to many skills; depends on the model invoking a meta-tool. * `'all'` โ€” every body is concatenated into the system prompt up front. Cheapest at runtime; eats context. Use for a small, always-relevant set. ### From the filesystem Convention: one directory per skill, each with a `SKILL.md` whose YAML frontmatter declares `name` and `description`. The frontmatter `name` is the source of truth and must match the folder name (drift throws). ``` ./.astor-skills/ pptx/SKILL.md refactor/SKILL.md ``` ```markdown --- name: pptx description: Build PowerPoint decks when the user asks for slides. # Optional extra fields โ€” captured into Skill.metadata, ignored by the SDK, # available for your own policy code (filter by tag, gate by version, etc.). version: 1.2.0 tags: documents, presentation allowed-tools: read, write, bash --- # How to build a deck Use pptx-genjs. Prefer one slide per concept; keep titles under 60 chars. ``` #### Frontmatter rules (Agent Skills spec) | Field | Required | Rule | | ------------- | -------- | ---- | | `name` | โœ… | 1โ€“64 chars, regex `^[a-z0-9][a-z0-9-]*$`. Reserved: `anthropic`, `claude`. | | `description` | โœ… | 1โ€“1024 chars. No XML tags (``) โ€” they confuse models that emit native tool-call syntax. | | any other key | โŒ | Captured into `Skill.metadata: Record` verbatim. | ```typescript import { createLocalAgent, createFileSystemSkillSource, OpenAIProvider } from 'astorlm' import { createCodingTools } from 'astorlm/tools' const agent = await createLocalAgent({ provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }), tools: createCodingTools(), skillSources: [createFileSystemSkillSource({ dir: './.astor-skills' })], skillMode: 'filesystem', }) await agent.run('Build a deck about climate change.') ``` ### Hierarchical discovery (user โ†’ project โ†’ repo) Passing several `createFileSystemSkillSource` directly throws on duplicate names (no silent override). `createLayeredSkillSource` resolves precedence internally (last-wins, loudly via `onOverride`) and presents itself as a single source, so the registry's conflict-throw stays intact for genuine clashes. ```typescript import { createLayeredSkillSource } from 'astorlm/core' const layered = createLayeredSkillSource({ layers: [userSkillsDir, projectSkillsDir], // LOW โ†’ HIGH precedence; project wins onOverride: ({ name, winner, loser }) => console.log(`"${name}": layer "${winner}" shadows "${loser}"`), }) // pass in skillSources: [layered] ``` ### Per-skill `allowed-tools` The frontmatter `allowed-tools` field is stored verbatim. Interpret and enforce it with two helpers: ```typescript import { parseAllowedTools, restrictToolsHook } from 'astorlm' const allowed = parseAllowedTools(skill) ?? [] // CSV โ†’ string[] | null const hooks = restrictToolsHook(allowed, { alwaysAllow: ['read', 'load_skill'], // keep these usable regardless denyMessage: (t) => `Tool "${t}" is not in the active skill's allowed-tools.`, }) // pass `hooks` into createLocalAgent({ hooks }) ``` Deciding *which* skill is active (and therefore which allowlist applies) is left to the consumer โ€” combine `parseAllowedTools` with your own logic and merge it into the session hooks. The SDK deliberately does not track an "active skill" in the core. ### From an in-memory bundle ```typescript import { createInMemorySkillSource } from 'astorlm' const source = createInMemorySkillSource({ name: 'shipped-skills', skills: [ { name: 'sql-review', description: 'Review SQL migrations for safety on a live DB.', body: '# SQL review\nCheck for table locks, NOT NULL without default, ...', metadata: { version: '1.0.0', 'allowed-tools': 'read, grep' }, }, ], }) ``` Conflicting names across sources throw at session creation โ€” there is no silent override. --- ## ๐Ÿณ Executors (sandboxing & swappable backends) Bash-family tools (`bash`, `bash_spawn`, `bash_get_output`, `bash_kill`) never talk to `child_process` directly โ€” they delegate to an `Executor`, making the execution backend pluggable. * **`LocalExecutor`** (`astorlm/core`) โ€” runs commands in the host process. Default for `createLocalAgent`. * **`DockerExecutor`** (`astorlm/core`) โ€” runs every command inside a container. Real sandboxing, not a command allowlist. * **Custom** โ€” implement the `Executor` interface (`exec`, `spawn`, `getOutput`, `kill`, `dispose`) and pass it via `executor`. ```typescript import { createLocalAgent, DockerExecutor, OpenAIProvider } from 'astorlm' import { createCodingTools } from 'astorlm/tools' const agent = await createLocalAgent({ provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }), tools: createCodingTools(), executor: new DockerExecutor({ image: 'node:20-alpine' }), }) ``` The agnostic core ships `createNoopExecutor()` as default โ€” it throws a clear error if a bash tool tries to use it without explicit configuration, so the SDK never silently runs commands on the host. --- ## ๐Ÿงช WASM code sandbox (`astorlm/experimental/wasm-runner`) `CodeRunner` is a sibling primitive to `Executor`, not a replacement for it. Where `Executor` runs **shell commands** with the host toolchain (isolated by Docker, or not at all), `CodeRunner` runs a **self-contained code snippet** inside a memory-safe WebAssembly runtime โ€” no filesystem, no network, no host syscalls unless explicitly granted (capability-based, default-deny). The key difference: it needs no daemon and no `child_process`, so it runs anywhere WASM does (Node, Deno, the browser, edge) โ€” exactly where `DockerExecutor` cannot reach. * **`QuickJsCodeRunner`** โ€” JavaScript via QuickJS compiled to WASM (`quickjs-emscripten`, an optional dependency loaded lazily). Per-run fresh context, enforced wall-clock deadline and memory limit, captured `console`, read-only JSON `globals`. * **`createCodeRunnerTool({ runner })`** โ€” wraps a runner as a `run_code` tool. Opt-in: it is *not* part of `createCodingTools()`; register it explicitly. ```typescript import { createLocalAgent, OpenAIProvider } from 'astorlm' import { createReadOnlyTools } from 'astorlm/tools' import { QuickJsCodeRunner, createCodeRunnerTool } from 'astorlm/experimental/wasm-runner' const runner = new QuickJsCodeRunner({ timeoutMs: 3_000 }) const agent = await createLocalAgent({ provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }), tools: [...createReadOnlyTools(), createCodeRunnerTool({ runner })], }) ``` | | `Executor` (Docker) | `CodeRunner` (WASM) | |---|---|---| | Runs | shell commands + toolchain | self-contained code snippets | | Where | Node + Docker daemon only | any runtime (Node/Deno/browser/edge) | | Isolation | OS-level (read-write cwd mount) | capability-empty, default-deny | | Cold start | hundreds of msโ€“s per command | ~ms | > โš ๏ธ Experimental API. Python (`PyodideCodeRunner`) is planned behind the same `CodeRunner` interface. --- ## โฑ๏ธ Background processes (`bash_spawn` / `bash_get_output` / `bash_kill`) Beyond the synchronous `bash` tool, the agent can manage long-running processes: * `bash_spawn { command }` โ†’ returns an opaque `pid`. * `bash_get_output { pid }` โ†’ drains stdout/stderr buffered since the last call, plus status / exit code. * `bash_kill { pid, signal? }` โ†’ terminates the process. This is what lets the agent launch a dev server, inspect logs, and tear it down without blocking the loop. --- ## ๐Ÿงญ Loop patterns `createAgent` / `createLocalAgent` accept `pattern: 'REACT' | 'PLAN_EXECUTE'` (default `'REACT'`). `'PLAN_EXECUTE'` auto-registers `add_plan_item` and `update_plan_item` tools that mutate a `PlanItem[]`. Each turn the loop injects the plan state into the system prompt (same idea as a visible, mutable to-do list). The plan persists in session metadata and survives resume/fork; read it with `agent.getPlan()`. ## ๐Ÿ”„ Goal loops (`runGoalLoop`) A single `agent.run()` ends when the model stops calling tools โ€” which is not the same as the job being *done*. `runGoalLoop` repeats the attempt until a condition you control returns true, giving each iteration a **fresh agent** (and therefore a clean context window) instead of letting one conversation grow unbounded. The stop condition is yours and should be cheap and deterministic โ€” run the test suite and check the exit code, assert a file exists โ€” which is what separates this from "ask the model if it's finished". `maxIterations` (default 10) is a mandatory fuse, so a condition that never holds cannot run forever. ```typescript import { createLocalAgent, runGoalLoop, OpenAIProvider } from 'astorlm' import { createCodingTools } from 'astorlm/tools' const result = await runGoalLoop({ goal: 'Make the test suite pass. Run `pnpm test` to check your work.', createIterationAgent: () => createLocalAgent({ cwd: process.cwd(), provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }), tools: createCodingTools(), }), isDone: async () => (await runTests()).exitCode === 0, // your own check maxIterations: 5, onIteration: ({ iteration, lastText }) => console.log(`#${iteration}: ${lastText.slice(0, 80)}`), }) console.log(result) // { iterations, done, lastText, stopReason: 'done' | 'max_iterations' | 'aborted' } ``` > Iterations share on-disk state (same `cwd`), not conversation history. That is the point: the work accumulates in the repo, the context does not. ## ๐Ÿงฑ Structured output (`generateObject`) When you want data back rather than prose, `generateObject` constrains the answer to a Zod schema and returns a parsed, typed object โ€” with repair attempts if the model emits something invalid. ```typescript import { generateObject, OpenAIProvider } from 'astorlm' import { z } from 'zod' const { object } = await generateObject({ provider: new OpenAIProvider({ model: 'qwen2.5-coder', baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' }), schema: z.object({ severity: z.enum(['low', 'medium', 'high']), files: z.array(z.string()), summary: z.string(), }), prompt: 'Triage the failure described in the log below: ...', }) object.severity // typed as 'low' | 'medium' | 'high' ``` `mode` picks how the constraint is applied: * `'tool'` โ€” a synthetic terminal tool (`provide_final_answer`) whose schema *is* the shape. Runs a real agent, so it works on any endpoint with tool calls and reuses the loop's own repair behaviour (`maxTurns`, default 8). This is the mode that lets you pass `tools`, so the model can gather what it needs before delivering the object. * `'native'` โ€” the provider's `response_format: json_schema`. Output failing validation is retried up to `maxRepairAttempts` (default 2), then throws `GenerateObjectError`. * `'auto'` (default) โ€” `'native'` when no `tools` are passed and the provider is OpenAI-compatible; `'tool'` otherwise. The mode actually used comes back on the result. ## ๐Ÿ“‰ Context optimizer (auto-compaction) When the provider exposes a `contextLimit`, the loop runs a structural optimizer between turns that prunes / dedupes once the conversation crosses a threshold. It's structural (prune / dedupe), not LLM-based summarization. ```typescript const agent = await createLocalAgent({ provider: /* ... */, contextOptimizer: { maxTokens: 200_000, compressThreshold: 0.8, // optimize when usage > 80% of maxTokens keepRecentTurns: 3, // always keep the last N turns verbatim }, }) // Disable entirely: contextOptimizer: false // Implicit default: enabled if provider.contextLimit is set, off otherwise. ``` ## ๐Ÿ” Retry policy for transient provider errors Opt-in retries for transient failures (HTTP 429, 5xx, network timeouts, streams cut before any chunk). Already-streamed events are never duplicated โ€” once any event is emitted on an attempt, the loop will not retry that turn. ```typescript const agent = await createLocalAgent({ provider: /* ... */, retry: { maxAttempts: 3, baseDelayMs: 500, maxDelayMs: 10_000, jitter: true }, }) // Default (omitted): no retries โ€” errors propagate and the session closes with session_end: error. ``` ## ๐Ÿ“Š Token usage tracking ```typescript await agent.run('...') console.log(agent.getUsage()) // { inputTokens, outputTokens, cacheReadTokens?, cacheCreationTokens? } ``` No pricing layer โ€” raw counts only. `cacheReadTokens` / `cacheCreationTokens` are populated when the provider reports them (Anthropic always; OpenAI's `cached_tokens` when applicable; many OpenAI-compat endpoints leave them undefined). ## ๐Ÿซ€ Heartbeat (proactive loop) `agent.startHeartbeat(opts?)` / `agent.stopHeartbeat()` (or pass `heartbeat` to the factory) run a check prompt on an interval. A `localCondition(cwd)` enables the **latent heartbeat**: the check runs locally in TypeScript and only wakes the LLM when it returns `true` โ€” zero token cost until the trigger fires. --- ## ๐Ÿ› ๏ธ Development Commands ```bash pnpm install # install dependencies pnpm build # build the library (dist/ in ESM, CJS, and d.ts) โ€” respects every entrypoint pnpm dev # interactive watch-mode build pnpm test # run the unit test suite (Vitest) pnpm typecheck # run TypeScript type checking without emitting ```