+++
title = "Kubernetes Model Serving in 2026: What Changed Since 2024"
description = "An update on OCI model volumes, inference-aware routing, KServe, llm-d, multi-node inference, GPU scheduling, and the remaining gaps."
date = 2026-07-26
draft = false
[taxonomies]
tags = ["kubernetes", "llm-inference", "ai-infrastructure", "gpu", "gateway-api"]
[extra]
selected = true
keywords = "kubernetes, model serving, llm inference, kserve, llm-d, gateway api inference extension, leaderworkerset, dynamic resource allocation, gpu scheduling"
toc = true
static_thumbnail = "/images/social-kubernetes-model-serving-in-2026.png"
+++
In October 2024, [Yuan Tang](https://www.linkedin.com/in/terrytangyuan/) published
[AI/ML Innovation in the Kubernetes Ecosystem](https://terrytangyuan.github.io/2024/10/22/ai-ml-innovation-in-the-kubernetes-ecosystem/).
The article described three important developments: Kubeflow Model Registry, KServe ModelCars, and TrustyAI.
It also pointed toward multi-node serving, inference-aware gateways, speculative decoding, low-rank adaptation
(LoRA) adapters, and APIs designed for generative AI.
Since then, Kubernetes model serving has developed into a stack of specialized control-plane and data-plane
components.
Kubernetes now has better primitives for distributing model files and allocating accelerators. Gateway API has
gained inference-aware extensions. KServe has introduced a separate API for generative inference. `LeaderWorkerSet`
has become a building block for multi-node model servers, while [llm-d](https://github.com/llm-d/llm-d/) coordinates
routing and distributed inference optimizations around engines such as [vLLM](https://vllm.ai/).
These capabilities also introduce more controllers, APIs, compatibility constraints, and failure modes.
{%
Figure 1. A production inference request crosses several independently operated layers.
KServe can manage much of this control plane. Gateway API Inference Extension defines the routing integration. `LeaderWorkerSet` represents groups of pods that must operate together. `llm-d` provides distributed inference and routing components. Engines such as vLLM or SGLang still execute the model. Each component has a separate role despite some overlap. ## 1. OCI image volumes simplify model delivery The 2024 article explained why KServe introduced ModelCars: a model could be packaged into an OCI image and exposed to the model server through a passive sidecar. This reused container-registry distribution and node caching instead of downloading weights from object storage for every replica. At that time, Kubernetes could pull an image but could not directly mount its filesystem as a volume. Kubernetes 1.31 introduced the `image` volume source as alpha. Kubernetes 1.33 promoted it to beta, and [Kubernetes 1.36 made OCI artifact and image volumes stable](https://kubernetes.io/blog/2026/04/22/kubernetes-v1-36-release/). A pod can now mount read-only content directly from an OCI registry: ```yaml apiVersion: v1 kind: Pod metadata: name: model-server spec: containers: - name: server image: example.com/model-server:v1 volumeMounts: - name: model mountPath: /models/llama readOnly: true volumes: - name: model image: reference: registry.example.com/models/llama:2026-07 pullPolicy: IfNotPresent ``` The sidecar pattern is no longer required merely because Kubernetes lacks an OCI-backed volume type. ModelCars can still be useful for KServe integration, compatibility with older clusters, and established tooling, but native image volumes provide a cleaner upstream primitive. OCI packaging does **not** eliminate cold starts. A 100–500 GB model still has to reach the node. Startup behavior still depends on: - registry throughput and rate limits - layer structure and decompression cost - container-runtime support - node disk capacity and eviction policy - whether the model is already cached - how many nodes pull the same model at once KServe also maintains a [Local Model Cache](https://kserve.github.io/website/docs/model-serving/generative-inference/modelcache/localmodel). It can pre-download model artifacts to local NVMe and let several inference pods reuse the warmed copy. Native OCI volumes improve packaging and delivery semantics. Local caching addresses placement and startup latency. They solve related but different problems. ## 2. Inference-aware routing gets a standard API A Kubernetes `Service` assumes that its ready endpoints are roughly interchangeable. That assumption breaks down for LLM serving. Two healthy replicas can have very different costs for the same request: - one replica may already have the prompt prefix in its key-value (KV) cache - one may have a long queue - one may be close to KV-cache exhaustion - only one may have the requested LoRA adapter loaded - replicas may use different accelerators or quantizations - a long prompt may be better assigned to a different pool than a decode-heavy request Round-robin routing ignores all of this. The same problem space described as "LLM Instance Gateway" in the 2024 article is now addressed by the [Gateway API Inference Extension](https://gateway-api-inference-extension.sigs.k8s.io/). By July 2026, the project marks itself generally available and defines `InferencePool` as a specialized backend abstraction for model servers. A gateway delegates endpoint choice to an inference router or endpoint picker, which can use live metrics and model-server capabilities. The request path becomes: ```mermaid flowchart LR accTitle: The inference-aware request path accDescr: An HTTPRoute sends a request to an InferencePool, which consults an endpoint picker before selecting a model-server pod. HTTPRoute["HTTPRoute"] InferencePool["InferencePool"] EndpointPicker["Endpoint picker"] ModelServer["Selected model-server pod"] HTTPRoute --> InferencePool InferencePool --> EndpointPicker EndpointPicker --> ModelServer ```Figure 2. Gateway API delegates model-server selection to an inference-aware endpoint picker.
The components have separate responsibilities: - **Gateway API** handles traffic attachment and routing integration. - **InferencePool** identifies a pool of inference endpoints. - **The endpoint picker** decides which endpoint is best for the request. - **The model server** exposes metrics and capabilities used by the picker. The reference project does not define every scheduling policy. Its documentation points users toward components such as the [llm-d router](https://github.com/llm-d/llm-d-router). Some of the advanced features that were previously developed in the Inference Extension repository have now moved to the llm-d repositories. These features include endpoint selection and model rewrite logic. The extension still owns the Pool API and conformance work. The integration API defines how a gateway connects to an inference pool without requiring every platform to use the same scheduling algorithm. The GA label does not imply that every endpoint-selection policy or gateway implementation is equally mature. ## 3. KServe separates predictive and generative serving KServe's original `InferenceService` API was designed for a broad range of predictive models and serving runtimes. It still fits many workloads well. LLM serving, however, needs concepts that do not map cleanly to a conventional stateless model endpoint: - tensor, data, and expert parallelism - groups of pods forming one logical replica - prefill/decode disaggregation - KV-cache-aware routing - LoRA adapter selection - long-lived streaming responses - autoscaling signals based on queues, tokens, and cache pressure KServe 0.17 presented `LLMInferenceService` as its API for production generative AI serving. KServe 0.18 added multi-node inference without requiring Ray, LeaderWorkerSet-based autoscaling, OpenAI Responses API routing, namespace-scoped model-cache work, and llm-d 0.6 integration. In KServe 0.18, example manifests still use the `serving.kserve.io/v1alpha2` API. Schema evolution therefore belongs in the upgrade plan. The distinction is now explicit: - use **`InferenceService`** for predictive serving and simpler serving patterns - use **`LLMInferenceService`** when you need orchestration designed for generative AI A simplified resource can describe the model, replicas, runtime, and managed routing: ```yaml apiVersion: serving.kserve.io/v1alpha2 kind: LLMInferenceService metadata: name: my-llm spec: model: uri: hf://organization/model name: organization--model replicas: 3 template: containers: - name: main image: vllm/vllm-openai:Figure 3. Disaggregation introduces an explicit KV-state transfer between prefill and decode.
This can reduce interference between long prefills and active decodes and let each phase use a different replica count or hardware shape. But it also adds a KV-transfer path before the first token. The result depends heavily on network bandwidth, latency, topology, transfer libraries, and the distribution of input and output lengths. On a bandwidth-constrained or high-latency inter-node network, disaggregation can move the bottleneck rather than remove it. Use disaggregation when measurements show that it improves the target workload. ## 5. DRA improves accelerator allocation—but depends on drivers For years, most Kubernetes GPU workloads have requested extended resources exposed by Device Plugins: ```yaml resources: limits: nvidia.com/gpu: "8" ``` This works, but it expresses little about the requested devices. It does not naturally represent properties such as GPU model, memory size, interconnect topology, sharing mode, or a reusable claim. [Dynamic Resource Allocation](https://kubernetes.io/docs/concepts/scheduling-eviction/dynamic-resource-allocation/) provides a richer model based on resources such as `DeviceClass`, `ResourceClaim`, and `ResourceClaimTemplate`. The core DRA APIs graduated to stable in Kubernetes 1.34. Kubernetes 1.36 continued work on device health, partitionable devices, consumable capacity, and additional drivers. DRA is Kubernetes's upstream direction for allocating specialized hardware, but a stable API does not make every accelerator stack ready for production. You still need to check: - whether your vendor provides a supported DRA driver - which Kubernetes and driver versions are compatible - how Multi-Instance GPU (MIG), time slicing, or other partitioning modes are represented - whether device health reaches the controllers that perform recovery - how upgrades coexist with existing Device Plugin workloads - whether your managed Kubernetes provider exposes the required features For many current clusters, Device Plugins remain the production default. DRA is useful when its richer selection and sharing semantics solve a concrete placement problem. ## 6. Workload-aware scheduling arrives in alpha Distributed inference may require several pods and devices to become available together. Scheduling them one at a time can leave partially allocated workloads holding scarce GPUs while waiting for the rest of the group. Kubernetes 1.35 introduced native workload-aware scheduling work, including gang-scheduling foundations. Kubernetes 1.36 expanded it with `Workload` and `PodGroup` APIs and atomic scheduling of related pod groups. The feature is **alpha**. It moves group scheduling closer to the upstream scheduler and Job controller and may eventually replace some external scheduling layers. Do not replace an existing production scheduler or queueing system without extensive testing. LeaderWorkerSet and workload-aware scheduling also solve different concerns: - LeaderWorkerSet describes and manages a replicated group of cooperating pods. - Workload-aware scheduling determines how related pods receive resources together. A platform may eventually use both. ## 7. Kubeflow Hub expands model discovery and governance The Model Registry described in 2024 has expanded into [Kubeflow Hub](https://www.kubeflow.org/docs/components/hub/), which combines: - **Model Registry:** metadata, versions, artifacts, lifecycle status, and governance - **Model Catalog:** read-only discovery across configured external catalogs The catalog can federate metadata from sources such as Hugging Face. The registry can integrate with deployment workflows and KServe storage initialization. Kubeflow Hub is primarily a metadata and discovery system. Its architecture documentation explicitly describes the registry as a passive repository rather than a Kubernetes control plane. The catalog does not store model weights. As of July 2026, the registry is still offered as an opt-in alpha component in Kubeflow Community Distribution 1.9 and newer, and its REST API remains versioned as `v1alpha3`. Treat it as an early lifecycle component, not as a dependency for every inference request. Keep large weight distribution in OCI registries, object storage, or a dedicated model-cache layer. Keep approval, lineage, versions, and deployment metadata in the registry. ## 8. AI policy moves toward gateways TrustyAI has continued beyond its early explainability focus, with tools for metrics, evaluation, bias and drift analysis, and language-model testing. It remains a pluggable part of the wider lifecycle rather than the central serving control plane. A newer development is the formation of the [Kubernetes AI Gateway Working Group](https://kubernetes.io/blog/2026/03/09/announcing-ai-gateway-wg/). Its scope includes common patterns such as: - token-based rate limiting - fine-grained access control - payload inspection - routing, caching, and guardrail hooks - secure access to external model providers - regional policy and failover Model-serving engines schedule and execute inference. A gateway can apply tenant policy consistently to self-hosted and external providers. The working group is new. Its proposals are not stable APIs with broad implementation support. ## What I would deploy today There is no single "Kubernetes AI stack." I would choose the smallest architecture that can handle the workload. ### Case 1: One model, one or a few replicas Start with: - a tested model-server image, such as vLLM, SGLang, or TensorRT-LLM - a `Deployment`, `StatefulSet`, or a basic KServe `InferenceService` - a standard `Service` and Gateway API route - model files from object storage, an OCI image volume, or a pre-warmed node cache - Prometheus metrics at the gateway, engine, GPU, and node layers Do not install an inference-aware router until replicas are sufficiently busy for routing quality to matter. ### Case 2: A shared production LLM service Add: - multiple replicas with an explicit latency and throughput SLO - `InferencePool` and a production endpoint picker when queue or cache affinity matters - KServe `LLMInferenceService` when its lifecycle automation is worth the dependencies - request priorities, admission control, quotas, and maximum context/output limits - autoscaling based on queueing, prompt length, output length, concurrency, and request mix—not GPU utilization alone Keep the model server and routing policy independently observable. A healthy gateway does not prove that the engine is healthy, and a healthy engine does not prove that clients receive timely streamed tokens. ### Case 3: Models that require several nodes Add `LeaderWorkerSet` only when one logical replica spans multiple pods or nodes. Before production, test: - startup and restart of the whole worker group - one worker or node failing during generation - collective-communication failures - rack and topology placement - network saturation - rolling updates when old and new model copies cannot fit simultaneously - whether failed replicas release all GPUs promptly Treat network bandwidth and topology as capacity constraints. ### Case 4: Large-scale or heterogeneous inference Evaluate DRA when you need richer accelerator selection, sharing, or health semantics and your vendor's driver is ready. Consider llm-d when prefix affinity, prefill/decode disaggregation, or distributed serving produces a measurable improvement over simpler routing. Do not adopt every available CRD. Every controller adds: - reconciliation behavior - status you must monitor - upgrade compatibility - RBAC and security surface - another place where desired and actual state can diverge ## What is still not solved - **Predictable cold starts:** OCI volumes and local caches improve delivery, but very large models still require substantial disk, network, and initialization time. Autoscaling cannot create warm GPU capacity instantly. - **Safe autoscaling:** Inference demand is measured in prompts, tokens, context lengths, and active sequences, not only requests per second. Replica startup may take many minutes, and scaling down can interrupt streams or destroy useful cache state. - **Portable accelerator management:** DRA provides a stable upstream API, but production behavior still depends on vendor drivers, managed-platform support, partitioning features, and integration with the existing device stack. - **Operational simplicity:** A full deployment may combine KServe, Gateway API, Inference Extension, a gateway, LeaderWorkerSet, KEDA, llm-d components, an engine, a model cache, and accelerator drivers. The combination requires compatibility testing. - **Distributed-inference networking:** Tensor parallelism, expert parallelism, and KV transfer are sensitive to topology and bandwidth. Kubernetes can place and restart pods, but it cannot compensate for insufficient network capacity. - **Stable, interoperable APIs:** Several APIs remain alpha or project-specific and will continue to change. Inference performance still depends on the model engine, kernels, quantization, batching, KV-cache management, accelerator topology, and request characteristics. Kubernetes supplies control, placement, lifecycle, and routing. {%