# Codemode The `codemode` tool lets the model write a JavaScript script that calls pi's other tools and runs non-LLM models, such as classifiers and image models. Only the script's output reaches the model, so a script can run calls in parallel and filter large results before the model sees them. To turn it on, see [Enable codemode](cli.md#enable-codemode). ## Scripts The tool input is raw JavaScript source, not JSON and not a markdown code fence. It runs as the body of an async function in a QuickJS sandbox, so top-level `await` and `return` work. The sandbox has no Node APIs, file system, network, or timers; scripts reach the outside world only through tools and `models`. A script may start with an options line: ```js // @options: {"max_output_tokens": 2000, "timeout_ms": 60000} ``` - `max_output_tokens` (default 10000) limits the output. Longer output keeps its start and end, and the full text is written to a temp file whose path is included in the result. - `timeout_ms` is a hard deadline for the whole script. It is unset by default. Image generation can take minutes, so do not set a short deadline for scripts that generate images. The result starts with `Script completed` or `Script failed`, the wall time, and the output. A failed script keeps its partial output, followed by `Script error:` and the error. Tool calls are real: calls made before a failure are not undone. Calls still running when the script ends are cancelled, and unawaited promises are discarded. ## Globals | Global | Purpose | |---|---| | `tools.(args)` | Call a tool. See [Call tools](#call-tools). | | `text(value)` | Add a text item to the output. Strings are added as is, other values as JSON. | | `image(value)` | Add an image to the output: a base64 `data:` URL, an `{ image_url }` object, or an image block `{ type: "image", data, mimeType }` such as those returned by MCP tools and `models.generateImages()`. Remote URLs are not supported. PNG, JPEG, GIF, and WebP are accepted. | | `console.log(...)` | Like `text()`; `info`, `warn`, `error`, and `debug` do the same. | | `return value` | A top-level `return` adds the value like `text()`. | | `exit()` | End the script successfully. | | `store(key, value)` / `load(key)` | Keep small JSON values across `codemode` calls. See [Store values](#store-values). | | `ALL_TOOLS` | Every callable tool as `{ name, description }`, including tools the description does not list. | | `searchTools(query, { limit?, namespace? })` | Rank callable tools by relevance (BM25, default limit 8). Resolves to `{ name, description }[]`. | | `describeTool(name)` | Resolves to a tool's description and TypeScript declaration, or `undefined`. | | `describeNamespace(name)` | Resolves to `{ name, description?, instructions?, tools }` for a namespace such as an MCP server, or `undefined`. | | `models` | List and run non-LLM models. See [Models](#models). | ## Call tools Every tool the session can call is a method of `tools`, named by its identifier: characters that are not valid in a JavaScript identifier become `_`, so the MCP tool `mcp__dev-radius__search` is `tools.mcp__dev_radius__search`. Each method takes one object with the tool's arguments. What a call resolves to depends on the tool: - Tools with an output schema resolve to a structured value. `bash` resolves to `{ output, truncated, full_output_path?, exit_code, wall_time_seconds }`, also for non-zero exit codes. Its `output` is not limited to the 2000 lines or 50KB the model sees: it holds up to 1 MiB, and longer output keeps its first and last 512 KiB around an omission marker, with `truncated` set and the full output in `full_output_path`. - MCP tools resolve to their `CallToolResult`, including `isError` and `structuredContent`. - Other tools, such as `read`, `edit`, and `write`, resolve to their text output. A call that fails, is blocked, or gets invalid arguments rejects with an `Error` that carries the tool's error text. Use `Promise.allSettled()` to keep the results of the calls that succeed. The `codemode` description lists tools with their TypeScript declarations, grouped by namespace (for example one MCP server). Tools with `deferred` exposure, which includes MCP tools with the default `codemode` exposure, are not listed, so the description stays the same while MCP servers connect. Listed declarations share a budget of 3000 estimated tokens (`codemode.inlineBudget` in [settings](settings.md#tools)). Scripts find the other tools with `searchTools()`, `describeTool()`, `describeNamespace()`, or by filtering `ALL_TOOLS`. While `codemode` is active, `codemode.mode` in [settings](settings.md#tools) decides how the other tools are presented. With `on` (default) declared tools stay declared, and their descriptions say how to call them from scripts. With `only` they are hidden from the model and listed in the `codemode` description instead, so the model calls them through scripts. ## Store values `store(key, value)` keeps a JSON value under a string key for later `codemode` calls; storing `undefined` deletes the key. `load(key)` returns the value, or `undefined`. Writes are kept only when the script succeeds: each successful script that stores values appends a `codemode-store` custom entry to the session, so resumed sessions keep the values and each branch sees only the values written on its path. The store is for small state such as IDs, cursors, or summaries. One value may have at most 262144 characters of JSON and all values together at most 1048576. Do not store image data; show images with `image()` or write them to a file with a tool. ## Models `models` reaches the model catalog and runs non-LLM models with the session's credentials: classifiers, which answer typed questions about JSON state, and image models, which generate images. Chat models are listed but cannot be run from scripts. Which classifier and image models exist is described in [Use classifier models](models.md#use-classifier-models) and [Use image models](models.md#use-image-models). ```ts type ModelType = "chat" | "image" | "classifier"; /** A catalog entry. `provider` and `id` identify it; other fields depend on the type. */ interface ModelInfo { type?: ModelType; provider: string; id: string; name: string; api: string; input: ("text" | "image")[]; contextWindow?: number; [key: string]: unknown; } declare const models: { /** Every known model of a type, optionally for one provider. */ getModelsOfType(type: ModelType, provider?: string): Promise; /** Models of a type whose provider has working credentials. */ getAvailableOfType(type: ModelType, provider?: string): Promise; /** One catalog entry, or undefined. */ getModelOfType(type: ModelType, provider: string, id: string): Promise; /** Answer `context.questions` about `context.state`; answers are in `result.answers` by question ID. */ classify(model: ModelInfo, context: ClassifierContext): Promise; /** Generate images from `context.input` text and image blocks; show `result.output` blocks with image(). Can take minutes. */ generateImages(model: ModelInfo, context: ImagesContext): Promise; }; ``` `classify()` and `generateImages()` use only the `provider` and `id` of `model`, so `{ provider, id }` works as well. They do not throw on provider errors: check `stopReason` and `errorMessage`. At most four such calls run at once per script; more calls wait for a free slot, so `Promise.all()` over many items is fine. Their usage is added to the `codemode` tool result and counts toward the session cost. Model IDs differ between providers, for example `typesafe/jev-latest` and `openrouter/typesafe/jev-1.13`. Use `models.getAvailableOfType(type)` to find the IDs that work with the current credentials. ### Classify ```ts interface ClassifierContext { /** The data to classify. */ state: Record; /** Questions by ID. One call answers all of them. */ questions: Record; } type ClassifierQuestion = /** Pick one label. `criteria` maps each label to what it means. */ | { type: "choice"; instructions: string; criteria: Record } /** Score on an ordered scale. `criteria` describes each level, lowest first. */ | { type: "score"; instructions: string; criteria: string[] } /** Yes or no. */ | { type: "bool"; instructions: string; criteria: { true: string; false: string } }; interface ClassifierResult { provider: string; model: string; /** Answers by question ID. */ answers: Record; usage?: ModelUsage; stopReason: "stop" | "error" | "aborted"; errorMessage?: string; } type ClassifierAnswer = | { type: "choice"; choice: string; probabilities: Record; confidence: number } /** `score` is the expected level index, from 0 to `criteria.length - 1`. */ | { type: "score"; score: number; confidence: number } /** Probability of `true`. */ | { type: "bool"; probability: number }; /** Token counts and cost in USD, when the service reports them. */ type ModelUsage = { input: number; output: number; totalTokens: number; cost: { total: number } }; ``` Classify several items by calling `classify()` once per item. This script sorts feedback messages, for example ones a tool returned earlier in the script: ```js const jev = await models.getModelOfType("classifier", "typesafe", "jev-latest"); const results = await Promise.all( messages.map((message) => models.classify(jev, { state: { message }, questions: { sentiment: { type: "choice", instructions: "How does the user feel about the product?", criteria: { positive: "Satisfied or happy", negative: "Unhappy or frustrated", neutral: "Neither" }, }, urgency: { type: "score", instructions: "How urgently does this need a reply?", criteria: ["no reply needed", "reply this week", "reply today"], }, }, }), ), ); return results.map((result, i) => result.stopReason === "stop" ? { message: messages[i], sentiment: result.answers.sentiment.choice, urgency: result.answers.urgency.score } : { message: messages[i], error: result.errorMessage }, ); ``` ### Generate images ```ts interface ImagesContext { /** The prompt as text blocks, plus image blocks to edit or use as references. */ input: (TextBlock | ImageBlock)[]; } interface ImagesResult { provider: string; model: string; /** Generated images, and text blocks for models that also return text. */ output: (TextBlock | ImageBlock)[]; usage?: ModelUsage; stopReason: "stop" | "error" | "aborted"; errorMessage?: string; } type TextBlock = { type: "text"; text: string }; /** `data` is base64. */ type ImageBlock = { type: "image"; data: string; mimeType: string }; ``` Show generated images with `image(block)`. Do not print `data` with `text()`, `console`, or `return`: it is large and the model cannot read it as text. Generated images are not saved to disk; to keep one, write it to a file with a tool. ```js // @options: {"timeout_ms": 300000} const painter = await models.getModelOfType("image", "openrouter", "google/gemini-2.5-flash-image"); const result = await models.generateImages(painter, { input: [{ type: "text", text: "A red fox in the snow, watercolor" }], }); if (result.stopReason !== "stop") return result.errorMessage; for (const block of result.output) { if (block.type === "image") image(block); else text(block.text); } ``` ## Limits - A script's VM has 256 MB of memory. Running out throws `InternalError: out of memory`; filter or aggregate large data instead of accumulating it. - A script that waits on a promise that can never settle (no tool call pending) fails immediately, since there are no timers. - Scripts cannot start other `codemode` scripts.