--- name: vot-serve description: Serve a vLLM preset on the current vllm-on-tap environment - resolve the preset (custom or builtin), start vLLM on the environment's stack as unit vot-, wait until it answers, and print the endpoint with a ready-to-paste request. Use when the user wants a preset running; the argument is the preset name. --- # /vot-serve 1. **Environment.** Read `.vot/config.yaml` and `.vot/environments/.yaml`. None: tell the user to run `/vot-config` and stop. Resolve the stack type per the stack-guide skill (an earlier name is read as the current one); a planned type: say it is not supported in this version, and stop. 2. **Preset.** No preset given: list the available presets per the load-preset skill's *Resolve* and stop. Otherwise resolve, validate and merge it per load-preset's *Resolve*, *Validate* and *Merge*, for the environment's stack type. 3. **Stack.** Read the stack-guide skill's `SKILL.md` and the reference of the environment's stack type. 4. **Checks.** Run the reference's *Serve* checks in their written order, stopping at the first refusal. At the "already served" check, when `vot-` exists: ask whether to destroy it first (then run `/vot-destroy `'s steps and continue) or stop. Never start a second unit. 5. **Start.** Run the reference's *Serve* start command, with tracing per stack-guide when the environment has `otlp_endpoint` (or when the reference's checks resolved one). 6. **Wait.** Run the reference's *Ready when*. On timeout, show the log lines it names and leave the unit for the user to inspect or destroy. 7. **Report.** Print the base URL, the served name, and the vllm-guide skill's *OpenAI API* `curl` filled in (with the reference's `` when it has an *API key* section - never print the key), and what the reference's *Serve* says to tell the user once the unit runs.