openapi: 3.2.0 info: title: LiteLLM auto router API description: 'Proxy Server to call 100+ LLMs in the OpenAI format. **Customize Swagger Docs** 👉 ```LiteLLM Admin Panel on /ui```. Create, Edit Keys with SSO. Having issues? Try ```Fallback Login``` 💸 ```LiteLLM Model Cost Map```. 🔎 ```LiteLLM Model Hub```. See available models on the proxy. **Docs**' version: 1.102.1 tags: - name: auto router paths: /public/complexity_router/scorer_defaults: get: tags: - auto router summary: Get Complexity Scorer Defaults description: Return the complexity router's shipped heuristic scorer defaults, for the dashboard to prefill with. operationId: get_complexity_scorer_defaults_public_complexity_router_scorer_defaults_get responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/ComplexityScorerDefaults' /public/autorouter_presets: get: tags: - auto router summary: Get Public Autorouter Presets description: 'Return the auto-router preset catalog the dashboard''s template picker renders. Resolved once per process, like the model cost map: fetched from ``litellm.autorouter_presets_url`` (override with ``LITELLM_AUTOROUTER_PRESETS_URL``) on the first request, falling back to the catalog bundled with the package on any failure. Set ``LITELLM_LOCAL_AUTOROUTER_PRESETS=True`` to serve the bundled catalog only. A restart picks up a newly published catalog.' operationId: get_public_autorouter_presets_public_autorouter_presets_get responses: '200': description: Successful Response content: application/json: schema: additionalProperties: $ref: '#/components/schemas/AutoRouterPresetRecord' type: object title: Response Get Public Autorouter Presets Public Autorouter Presets Get /auto_router/benchmarks: get: tags: - auto router summary: Get Auto Router Benchmarks description: 'Benchmarks for the auto-router dashboard: session shape, savings against the configured baseline, and prompt-caching behaviour bucketed by what the router did. Reads the LiteLLM_AutoRouterSession rollup, folded once per request at spend-write time, so this endpoint never scans LiteLLM_SpendLogs. A session is in the window when it overlaps it: its last turn is on or after start_date and its first turn is on or before end_date. Overall hit rate is over telemetry-bearing turns; each bucket''s hit rate is over that bucket''s turns. The rollup supplies the measures, never the list. Which routers appear comes from the model registry, so one shows up as soon as it is configured and reads zero until it serves traffic, and `routers_in_scope` counts those too rather than only the routers the window recorded.' operationId: get_auto_router_benchmarks_auto_router_benchmarks_get security: - APIKeyHeader: [] parameters: - name: start_date in: query required: false schema: anyOf: - type: string - type: 'null' description: YYYY-MM-DD UTC, inclusive (defaults to 30 days before end_date) title: Start Date description: YYYY-MM-DD UTC, inclusive (defaults to 30 days before end_date) - name: end_date in: query required: false schema: anyOf: - type: string - type: 'null' description: YYYY-MM-DD UTC, inclusive (defaults to today) title: End Date description: YYYY-MM-DD UTC, inclusive (defaults to today) - name: api_key in: query required: false schema: anyOf: - type: string - type: 'null' description: Filter to one virtual key token hash title: Api Key description: Filter to one virtual key token hash responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/AutoRouterBenchmarksResponse' '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' /auto_router/session: get: tags: - auto router summary: Get Auto Router Session description: 'One auto-routed session, for the key that ran it: the model its last turn was routed to and the session''s spend against the router''s savings baseline. Built for a coding agent''s status line or stop hook, so any virtual key may call it and only ever sees rows written under its own key hash. Reads the LiteLLM_AutoRouterSession rollup, which the asynchronous spend flush fills a moment after each turn; a session with no flushed auto-routed turn yet is a 404. The id is bounded the way the writer bounded it, so an oversized client id still finds its row.' operationId: get_auto_router_session_auto_router_session_get security: - APIKeyHeader: [] parameters: - name: session_id in: query required: true schema: type: string description: The client session id (x-*-session-id header) the turns were sent under title: Session Id description: The client session id (x-*-session-id header) the turns were sent under responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/AutoRouterSessionResponse' '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' /auto_router/shadow_eval/start: post: tags: - auto router summary: Start Shadow Eval description: 'Start a shadow eval: duplicate a sampled slice of one or more targets'' live traffic against a second arm, judge the two responses blind, and stratify win rates by tier, by the model that served the real arm, and by target. A target is a virtual key, a team, or a user. Team and user targets match on the identity every request resolves to at auth time, so they cover JWT-authenticated traffic, which presents no virtual key; a user target samples that user''s traffic across all their teams, whether it arrives on a JWT or a key they own. models narrows every target to requests for those model groups, so a user plus one model samples that user''s traffic on that model across every key they own; it is forward-only, since a reverse job already samples exactly the traffic its own router served. A forward job answers whether the targets should adopt router_name: it samples the requests the router did not serve and duplicates them through it. A reverse job answers whether a target already on the router still gains from it: it samples the requests the router did serve and duplicates them against baseline_model. A target can hold one active job per direction, so both questions can run at once, and a request matching several jobs'' targets (say its key and its team) is sampled by each, separately budgeted. Shadow responses are never served to users. Each target samples until its recorded eval spend, the shadow and judge calls'' own cost, reaches max_budget dollars, the job''s window ends, or the job is stopped, so one target running out of budget does not end sampling for the others; sampling changes propagate to pods within about 10 seconds. Shadow and judge calls bill to the sampled request''s own identity but are excluded from request counts and auto-router adoption metrics.' operationId: start_shadow_eval_auto_router_shadow_eval_start_post requestBody: content: application/json: schema: $ref: '#/components/schemas/StartShadowEvalRequest' required: true responses: '201': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/ShadowEvalJobResponse' '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' security: - APIKeyHeader: [] /auto_router/shadow_eval: get: tags: - auto router summary: List Shadow Eval Jobs description: 'List shadow eval jobs, newest first, each target with its attempt count so status is accurate. Judged counts, spend, and results ride the detail endpoint only.' operationId: list_shadow_eval_jobs_auto_router_shadow_eval_get security: - APIKeyHeader: [] parameters: - name: target_type in: query required: false schema: anyOf: - enum: - key - team - user type: string - type: 'null' description: Kind of target to filter on; requires target_id title: Target Type description: Kind of target to filter on; requires target_id - name: target_id in: query required: false schema: anyOf: - type: string - type: 'null' description: Filter to jobs that shadow this target, alone or alongside others title: Target Id description: Filter to jobs that shadow this target, alone or alongside others - name: limit in: query required: false schema: type: integer maximum: 200 minimum: 1 description: Newest jobs to return default: 50 title: Limit description: Newest jobs to return responses: '200': description: Successful Response content: application/json: schema: type: array items: $ref: '#/components/schemas/ShadowEvalJobResponse' title: Response List Shadow Eval Jobs Auto Router Shadow Eval Get '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' /auto_router/shadow_eval/{job_id}: get: tags: - auto router summary: Get Shadow Eval Job description: One job with derived counts, judge spend, latest error, and stratified results. operationId: get_shadow_eval_job_auto_router_shadow_eval__job_id__get security: - APIKeyHeader: [] parameters: - name: job_id in: path required: true schema: type: string title: Job Id responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/ShadowEvalJobResponse' '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' /auto_router/shadow_eval/{job_id}/stop: post: tags: - auto router summary: Stop Shadow Eval Job description: 'Stop an active shadow eval job, every target it scopes at once. Attempts are kept; sampling halts within ~10s. Targets that already stopped on their own budget keep the stopped_at they earned. The statement is the whole state machine: it claims the job only while a leg still samples inside the window with no stop recorded, so a racing operator, a same-instant budget spend, and a repeat stop all read the same 400 with the status the job actually holds.' operationId: stop_shadow_eval_job_auto_router_shadow_eval__job_id__stop_post security: - APIKeyHeader: [] parameters: - name: job_id in: path required: true schema: type: string title: Job Id responses: '200': description: Successful Response content: application/json: schema: $ref: '#/components/schemas/ShadowEvalJobResponse' '422': description: Validation Error content: application/json: schema: $ref: '#/components/schemas/HTTPValidationError' components: schemas: ShadowEvalJobResponse: properties: job_id: type: string title: Job Id targets: items: $ref: '#/components/schemas/ShadowEvalJobTargetResponse' type: array minItems: 1 title: Targets description: The targets whose traffic this job evaluates, and only theirs, each with its own budget router_names: items: type: string type: array minItems: 1 title: Router Names description: Every auto-router this job runs as a shadow arm. Multi-router jobs sample one slice of traffic and judge every arm against the same real responses models: items: type: string type: array title: Models description: Model groups the sampled traffic is narrowed to; empty means every model the targets use default: [] direction: type: string enum: - forward - reverse title: Direction default: forward baseline_model: anyOf: - type: string - type: 'null' title: Baseline Model judge_model: type: string title: Judge Model shadow_percentage: type: number title: Shadow Percentage created_at: type: string format: date-time title: Created At ends_at: type: string format: date-time title: Ends At stopped_by: anyOf: - type: string - type: 'null' title: Stopped By description: The operator who stopped the job early, recorded by the stop endpoint; 'unknown' backfilled by migration for jobs that displayed stopped when the column arrived; None when the job ended on its own. Its presence is what makes a job read stopped rather than completed judged_count: anyOf: - type: integer - type: 'null' title: Judged Count description: Verdicts recorded; detail endpoint only error_count: anyOf: - type: integer - type: 'null' title: Error Count description: Sampled attempts that errored; detail endpoint only judge_spend: anyOf: - type: number - type: 'null' title: Judge Spend description: Judge cost so far; detail endpoint only last_error: anyOf: - type: string - type: 'null' title: Last Error description: Most recent attempt error; detail endpoint only results: anyOf: - $ref: '#/components/schemas/ShadowEvalResult' - type: 'null' description: Stratified verdicts; detail endpoint only router_name: type: string title: Router Name description: 'The first router, kept for callers that predate router_names; derived so the two fields can never disagree.' readOnly: true status: type: string enum: - running - completed - stopped title: Status description: 'Three recorded facts, no history-guessing: a stop is stopped_by (the migration backfills it for every job that displayed stopped when the column arrived, so the pre-column population is closed), completion is the window passing or every target spending its budget, and anything else is running. The all-targets-stamped fallback covers only stops written by pre-column pods during a rolling deploy.' readOnly: true type: object required: - job_id - targets - router_names - judge_model - shadow_percentage - created_at - ends_at - router_name - status title: ShadowEvalJobResponse description: 'A shadow-eval job over one or more targets, each with its own budget and stop state; status is derived from stopped_by, the targets'' stop and budget state, and ends_at, never stored, so no writer anywhere can produce an inconsistent one. Aggregate fields are populated by the detail endpoint only and stay None on list responses.' StartShadowEvalRequest: properties: api_key_ids: items: type: string type: array maxItems: 100 title: Api Key Ids description: Hashed virtual keys whose traffic will be shadowed. Combined with team_ids and user_ids the job needs at least one target and at most 100, which also bounds every read the job's endpoints make. Each target carries its own max_budget spend budget, so one exhausting its budget leaves the others sampling. default: [] team_ids: items: type: string type: array maxItems: 100 title: Team Ids description: Teams whose traffic will be shadowed, matched on the team every authenticated request resolves to, so a team's JWT-auth and virtual-key traffic are both sampled default: [] user_ids: items: type: string type: array maxItems: 100 title: User Ids description: 'Users whose traffic will be shadowed, matched on the user every authenticated request resolves to across all their teams: JWT requests carrying their subject claim and virtual keys they own' default: [] models: items: type: string type: array maxItems: 100 title: Models description: 'Model groups to narrow the sampled traffic to, matched on the group the caller requested and resolved through model_group_alias, so an alias and its target are one name. Empty samples every model the targets use. This ANDs with the targets: a job over a user and one model samples that user''s requests on that model across every key they own, and none of their other traffic. Forward jobs only: a reverse job samples exactly the traffic its own router served, which no other model group can name' default: [] router_name: anyOf: - type: string - type: 'null' title: Router Name description: 'The auto-router under evaluation, in either direction: the single-router spelling of router_names. Provide exactly one of the two fields' router_names: items: type: string type: array maxItems: 4 title: Router Names description: The auto-routers under evaluation, at most 4. Every sampled request runs through every router listed and each arm is judged independently against the same real response, so routers compare head-to-head on identical traffic. More than one router requires direction 'forward'. After validation this field always carries the full deduplicated set, whichever spelling the caller used default: [] direction: type: string enum: - forward - reverse title: Direction description: 'forward answers ''should this key adopt router_name'': it samples the requests the key did NOT route through the router and duplicates them through it. reverse answers ''is the router still worth it for a key already on it'': it samples the requests the router did serve and duplicates them against baseline_model. The response the caller received is always the real arm' default: forward baseline_model: anyOf: - type: string - type: 'null' title: Baseline Model description: 'Required when direction is reverse and rejected otherwise: the fixed model the router''s own responses are judged against. Must be a plain model rather than another auto-router' shadow_percentage: type: number maximum: 100.0 minimum: 0.1 title: Shadow Percentage description: Percentage of each target's requests to duplicate through the router judge_model: type: string title: Judge Model description: 'Model used to blindly judge real vs. shadow responses. The judge only compares two answers, so a mid-tier model (Claude Sonnet or GPT-4o class) is the sweet spot: small/nano-class models produce unreliable or malformed verdicts, while frontier reasoning models add cost without changing outcomes.' default: anthropic/claude-sonnet-5 duration_days: type: integer maximum: 30.0 minimum: 1.0 title: Duration Days description: How many days the job samples traffic before completing on its own default: 7 max_budget: type: number maximum: 10000.0 minimum: 0.01 title: Max Budget description: Per-target USD budget for the eval's own overhead, the shadow-arm and judge calls, priced with the same figures the spend pipeline bills. EACH scoped target samples until its recorded eval spend reaches this, so a job over N targets spends at most about N times max_budget; in-flight samples can overshoot the cap by one sampling cache window. Every router arm draws from the same per-target budget, so a multi-router job reaches it proportionally sooner default: 10.0 type: object required: - shadow_percentage title: StartShadowEvalRequest description: 'Start duplicating one or more targets'' traffic for blind comparison against an auto-router. A target is a virtual key, a team, or a user; each becomes its own leg with its own budget and stop state. Team and user targets match on the identity every request carries after auth (user_api_key_team_id / user_api_key_user_id), so they cover JWT-authenticated traffic, which presents no virtual key at all.' ComplexityScorerDefaults: properties: tier_boundaries: additionalProperties: type: number type: object title: Tier Boundaries token_thresholds: additionalProperties: type: integer type: object title: Token Thresholds dimension_weights: additionalProperties: type: number type: object title: Dimension Weights type: object required: - tier_boundaries - token_thresholds - dimension_weights title: ComplexityScorerDefaults description: 'The complexity router''s shipped heuristic scorer defaults. The dashboard prefills its Advanced scoring controls from these rather than keeping its own copy, so a recalibration of the defaults cannot leave the form reporting numbers the router no longer uses.' AutoRouterBenchmarkTotals: properties: sessions: type: integer title: Sessions turns: type: integer title: Turns avg_turns_per_session: type: number title: Avg Turns Per Session avg_session_seconds: type: number title: Avg Session Seconds avg_tokens_per_session: type: number title: Avg Tokens Per Session spend: type: number title: Spend description: What the routed traffic actually cost classifier_cost: anyOf: - type: number - type: 'null' title: Classifier Cost description: Recorded LLM classifier cost already included in spend; null when any session turns predate subtotal recording, and zero for an empty window saved_spend: type: number title: Saved Spend description: Signed dollars saved versus each router's savings baseline (derived from its hardest tier, or the configured override), from the same per-request savings record the usage tab reads baseline_spend: type: number title: Baseline Spend description: 'spend plus saved_spend: the estimated single-model cost' saved_pct: type: number title: Saved Pct description: saved_spend over baseline_spend, as a percentage saved_per_session: type: number title: Saved Per Session cache: $ref: '#/components/schemas/AutoRouterCacheStats' type: object required: - sessions - turns - avg_turns_per_session - avg_session_seconds - avg_tokens_per_session - spend - classifier_cost - saved_spend - baseline_spend - saved_pct - saved_per_session - cache title: AutoRouterBenchmarkTotals description: Session-shape and savings aggregates over auto-routed traffic in the window. AutoRouterPresetRecord: properties: label: type: string title: Label description: type: string title: Description complexity_router_config: $ref: '#/components/schemas/AutoRouterPresetConfig' additionalProperties: true type: object required: - label - description - complexity_router_config title: AutoRouterPresetRecord description: One auto-router preset as served to the dashboard's template picker. ShadowEvalJobTargetResponse: properties: target_type: type: string enum: - key - team - user title: Target Type description: What kind of entity this entry scopes target_id: type: string title: Target Id description: The hashed virtual key, team id, or user id whose traffic this entry scopes max_turns: type: integer title: Max Turns description: 'This target''s sample-count ceiling: the whole budget for jobs created before max_budget existed, and the error-loop safety valve otherwise' max_budget: anyOf: - type: number - type: 'null' title: Max Budget description: This target's own USD budget for the eval's shadow and judge spend, independent of its siblings'; None on jobs created before spend budgets existed, which max_turns alone bounds stopped_at: anyOf: - type: string format: date-time - type: 'null' title: Stopped At description: When this target's slot was stamped free, whether its own budget ran out, the window closed, or an operator stopped the job; status is derived, so a spent budget reads completed even while this is still unset attempt_count: anyOf: - type: integer - type: 'null' title: Attempt Count description: This target's sampled attempts so far, judged and errored alike, the same count the sampler budgets against max_turns; populated on list and detail responses. Frozen at stopped_at once the target is stamped, so in-flight attempts landing after a stop never reclassify it spend: anyOf: - type: number - type: 'null' title: Spend description: This target's recorded shadow plus judge spend in USD, the same figure the sampler budgets against max_budget; populated on list and detail responses and frozen at stopped_at exactly like attempt_count verdicts: anyOf: - $ref: '#/components/schemas/ShadowEvalSlice' - type: 'null' description: This target's own judged-verdict slice; detail endpoint only, None until a turn is judged target_alias: anyOf: - type: string - type: 'null' title: Target Alias description: 'Display label resolved from the target''s own row at read time: the key''s alias, the team''s alias, or the user''s email; None when unset or deleted' key_name: anyOf: - type: string - type: 'null' title: Key Name description: Masked display name (sk-...) for key targets, resolved at read time; None for teams and users type: object required: - target_type - target_id - max_turns title: ShadowEvalJobTargetResponse description: One target a job shadows (a key, team, or user), with its own budget and stop state. AutoRouterBenchmarksResponse: properties: start_date: type: string title: Start Date description: Window start day, YYYY-MM-DD UTC, inclusive end_date: type: string title: End Date description: Window end day, YYYY-MM-DD UTC, inclusive routers_in_scope: type: integer title: Routers In Scope description: How many groups this response carries. Every auto-router configured on the proxy counts, whether or not it served anything in the window. To count only the routers that did serve traffic, filter `groups` to the entries whose `sessions` is above zero totals: $ref: '#/components/schemas/AutoRouterBenchmarkTotals' groups: items: $ref: '#/components/schemas/AutoRouterBenchmarkGroup' type: array title: Groups description: 'One entry per auto-router, listed from the model registry rather than from the rollup, so a router appears as soon as it is configured and reads zero until it serves traffic. Semantic auto-routers are absent: they record no routing decision, so no session can ever be attributed to them' type: object required: - start_date - end_date - routers_in_scope - totals - groups title: AutoRouterBenchmarksResponse description: Benchmarks for the auto-router dashboard, aggregated from the per-session rollup. AutoRouterCacheStats: properties: coverage_pct: type: number title: Coverage Pct description: Share of turns that carried cache telemetry hit_rate_pct: type: number title: Hit Rate Pct description: All cache hits over telemetry-bearing turns same_model: $ref: '#/components/schemas/AutoRouterCacheBucket' first_visit: $ref: '#/components/schemas/AutoRouterCacheBucket' return_to_tier: $ref: '#/components/schemas/AutoRouterCacheBucket' unordered_turns: type: integer title: Unordered Turns description: Turns that arrived out of order and were not bucketed return_misses_expired: type: integer title: Return Misses Expired description: Return-to-tier misses where the model's recorded cache TTL had lapsed return_misses_within_ttl: type: integer title: Return Misses Within Ttl description: 'Return-to-tier misses inside the recorded TTL: the prefix changed or the provider evicted the entry early; billing telemetry cannot distinguish the two' return_misses_unknown: type: integer title: Return Misses Unknown description: Return-to-tier misses with no recorded TTL to attribute against ttl_5m_turns: type: integer title: Ttl 5M Turns description: Turns whose cache write used the five-minute TTL ttl_1h_turns: type: integer title: Ttl 1H Turns description: Turns whose cache write used the one-hour TTL type: object required: - coverage_pct - hit_rate_pct - same_model - first_visit - return_to_tier - unordered_turns - return_misses_expired - return_misses_within_ttl - return_misses_unknown - ttl_5m_turns - ttl_1h_turns title: AutoRouterCacheStats description: 'Prompt-caching behaviour of auto-routed turns, bucketed by what the router did. Every in-order turn falls in exactly one bucket: the session stayed on the same model, visited a model for the first time (cold by design), or returned to a model it had already used. Out-of-order turns (cross-pod flush races) are counted but not bucketed.' AutoRouterPresetConfig: properties: tiers: $ref: '#/components/schemas/AutoRouterPresetTiers' additionalProperties: true type: object required: - tiers title: AutoRouterPresetConfig description: 'The complexity_router_config a preset prefills. Only tiers is validated, because every dashboard consumer dereferences it; everything else passes through verbatim with unknown fields kept (extra="allow"), so a catalog published after this proxy shipped still serves its new fields intact.' ShadowEvalResult: properties: by_tier: items: $ref: '#/components/schemas/ShadowEvalSlice' type: array title: By Tier by_current_model: items: $ref: '#/components/schemas/ShadowEvalSlice' type: array title: By Current Model description: 'Sliced by the model that served the real arm: the keys'' incumbent models in forward mode, and in reverse the models the router itself picked' by_router: items: $ref: '#/components/schemas/ShadowEvalSlice' type: array title: By Router description: 'One slice per router arm, grouped on the router name. Every arm of a multi-router job is judged against the same real responses over the same sampled requests, so these slices compare routers head-to-head: like-for-like win rates and spends on identical traffic. Verdicts from before arm stamping existed count toward the job''s own router' default: [] overall_shadow_win_rate_pct: type: number title: Overall Shadow Win Rate Pct overall_tie_rate_pct: type: number title: Overall Tie Rate Pct sampled_real_spend: type: number title: Sampled Real Spend description: USD the real arm billed across all judged turns, cache-served turns excluded. A judged turn is one (request, router arm) verdict, so a multi-router job counts the real response once per arm it was judged against; per-router comparisons read by_router default: 0.0 sampled_shadow_spend: type: number title: Sampled Shadow Spend description: USD the shadow arms billed across the same turns, judge excluded, like for like default: 0.0 not_sampled_count: anyOf: - type: integer - type: 'null' title: Not Sampled Count description: 'Eligible requests the sampling dice skipped, summed over legs: the judged rows stand for judged + this many requests. None for jobs from before the funnel existed' unjudgeable_count: anyOf: - type: integer - type: 'null' title: Unjudgeable Count description: Sampled requests whose shape could not be judged (tool-final turn, empty text) shed_count: anyOf: - type: integer - type: 'null' title: Shed Count description: Sampled requests dropped by the per-pod concurrency cap, so quiet periods are overweighted withheld_count: anyOf: - type: integer - type: 'null' title: Withheld Count description: 'Sampled requests the pipeline declined to spend on: no database to record into, an over-budget key or team, or the eval budget unverifiable or already reached (the in-flight burst as a job crosses max_budget lands here rather than vanishing from coverage)' type: object required: - by_tier - by_current_model - overall_shadow_win_rate_pct - overall_tie_rate_pct title: ShadowEvalResult description: Stratified results of a shadow-eval job's verdicts so far. AutoRouterCacheBucket: properties: turns: type: integer title: Turns description: Turns classified into this bucket hits: type: integer title: Hits description: Turns in this bucket whose response reported cache-read tokens hit_rate_pct: type: number title: Hit Rate Pct description: hits over this bucket's turns, as a percentage type: object required: - turns - hits - hit_rate_pct title: AutoRouterCacheBucket description: One prompt-caching bucket of turns, with how often those turns hit the cache. AutoRouterBenchmarkGroup: properties: sessions: type: integer title: Sessions turns: type: integer title: Turns avg_turns_per_session: type: number title: Avg Turns Per Session avg_session_seconds: type: number title: Avg Session Seconds avg_tokens_per_session: type: number title: Avg Tokens Per Session spend: type: number title: Spend description: What the routed traffic actually cost classifier_cost: anyOf: - type: number - type: 'null' title: Classifier Cost description: Recorded LLM classifier cost already included in spend; null when any session turns predate subtotal recording, and zero for an empty window saved_spend: type: number title: Saved Spend description: Signed dollars saved versus each router's savings baseline (derived from its hardest tier, or the configured override), from the same per-request savings record the usage tab reads baseline_spend: type: number title: Baseline Spend description: 'spend plus saved_spend: the estimated single-model cost' saved_pct: type: number title: Saved Pct description: saved_spend over baseline_spend, as a percentage saved_per_session: type: number title: Saved Per Session cache: $ref: '#/components/schemas/AutoRouterCacheStats' router_name: type: string title: Router Name description: The auto-router alias requests were sent to router_type: type: string title: Router Type description: complexity, adaptive or quality tier_turns: additionalProperties: type: integer type: object title: Tier Turns description: 'Turns per tier, keyed by the tier name the routing decision recorded at request time (never re-derived at read time, since the tier-to-model mapping is mutable config). Tier names are scoped to this group''s router_type and are not comparable across types: a complexity router reports ''SIMPLE''/''MEDIUM''/''COMPLEX''/''REASONING'', a quality router reports its numeric quality tier, and an adaptive router records no tier at all. Turns no tier served (the classifier fell back to default_model) are absent rather than pooled under a sentinel key, so the values may sum to less than turns' type: object required: - sessions - turns - avg_turns_per_session - avg_session_seconds - avg_tokens_per_session - spend - classifier_cost - saved_spend - baseline_spend - saved_pct - saved_per_session - cache - router_name - router_type title: AutoRouterBenchmarkGroup description: One auto-router's slice of the benchmarks. ValidationError: properties: loc: items: anyOf: - type: string - type: integer type: array title: Location msg: type: string title: Message type: type: string title: Error Type input: title: Input ctx: type: object title: Context type: object required: - loc - msg - type title: ValidationError AutoRouterSessionResponse: properties: session_id: type: string title: Session Id router_name: type: string title: Router Name description: The auto-router alias the session's requests were sent to router_type: type: string title: Router Type description: complexity, adaptive or quality turns: type: integer title: Turns description: Auto-routed turns the rollup has recorded for this session so far last_model: type: string title: Last Model description: The deployment model the most recent turn was routed to spend: type: number title: Spend description: What the session's routed traffic actually cost, classifier calls included saved_spend: type: number title: Saved Spend description: Estimated savings against the baseline, net of classifier cost baseline_spend: type: number title: Baseline Spend description: 'spend plus saved_spend: the estimated single-model cost' baseline_model: anyOf: - type: string - type: 'null' title: Baseline Model description: 'The savings baseline most of this session''s turns were priced against, recorded turn by turn, so it still names the counterfactual after the router is reconfigured or removed. None when no turn recorded one: rows from before the baseline was recorded, and adaptive and quality routers, which derive no baseline and so report no savings' baseline_models: additionalProperties: type: integer type: object title: Baseline Models description: Turns priced against each baseline model; more than one entry means the router's baseline changed mid-session and baseline_spend mixes both type: object required: - session_id - router_name - router_type - turns - last_model - spend - saved_spend - baseline_spend - baseline_model - baseline_models title: AutoRouterSessionResponse description: 'One auto-routed session as its own key sees it: what the last turn ran on, and what the session cost against the router''s savings baseline (the priciest model in its hardest tier).' AutoRouterPresetTiers: properties: SIMPLE: items: type: string type: array title: Simple MEDIUM: items: type: string type: array title: Medium COMPLEX: items: type: string type: array title: Complex REASONING: items: type: string type: array title: Reasoning additionalProperties: false type: object required: - SIMPLE - MEDIUM - COMPLEX - REASONING title: AutoRouterPresetTiers description: 'Exactly the four built-in tiers the dashboard''s preset prefill can apply. extra="forbid" on purpose: a tier name this dashboard cannot apply would grey out or crash the picker, so such a catalog is rejected wholesale and the bundled one serves instead.' ShadowEvalSlice: properties: group: type: string title: Group turn_count: type: integer title: Turn Count real_win_rate_pct: type: number title: Real Win Rate Pct description: 'Share of judged turns the real arm won, meaning the response the caller actually received: the key''s own model in forward mode, the router''s pick in reverse' shadow_win_rate_pct: type: number title: Shadow Win Rate Pct description: 'Share of judged turns the shadow arm won, meaning the duplicated response nobody was served: the router''s pick in forward mode, baseline_model in reverse' tie_rate_pct: type: number title: Tie Rate Pct avg_judge_confidence: type: number title: Avg Judge Confidence real_spend: type: number title: Real Spend description: USD the real arm billed on this slice's judged turns, completion plus its own routing classifier when it routed, excluding turns litellm's response cache served for free default: 0.0 shadow_spend: type: number title: Shadow Spend description: USD the shadow arm billed on the same turns, completion plus its own routing classifier, excluding the judge and the same cache-served turns, so the two spends compare like for like default: 0.0 cache_hit_turns: type: integer title: Cache Hit Turns description: 'Judged turns litellm''s response cache served, excluded from both spends: an adopted router would be served by the same cache, so those turns cost the same either way' default: 0 type: object required: - group - turn_count - real_win_rate_pct - shadow_win_rate_pct - tie_rate_pct - avg_judge_confidence title: ShadowEvalSlice description: 'Judge outcomes for one slice of a job''s verdicts: a router tier, one of the models that served the real arm, or one scoped target (embedded on that target''s own entry, so slices never need re-joining to a target by id).' HTTPValidationError: properties: detail: items: $ref: '#/components/schemas/ValidationError' type: array title: Detail type: object title: HTTPValidationError securitySchemes: APIKeyHeader: type: apiKey description: Bearer token in: header name: x-litellm-api-key