--- name: llm-d-deploy-stack description: Deploys the llm-d stack on GKE using well-lit paths specification. version: 0.4.0 allowed-tools: kubectl gcloud helm kustomize curl terraform python3 mcp-servers: - kubernetes: reason: "Inspect cluster info, custom compute classes, pod statuses, and secrets." - gcp: reason: "Verify GKE cluster status." --- # Deploy llm-d Stack Skill Follow these instructions to deploy the llm-d benchmarking stack on GKE. ## Terminology & Variables This skill utilizes the following core variables: - `SPEC` (also referred to as `strategy` or `guide`): This represents the name of the `llm-d` well-lit path guide being targeted. It maps directly to GKE overlay folder structures (`llmd-`) - **Allowed Values:** - `optimized-baseline` (Standard optimized baseline configuration). - `precise-prefix-cache-routing` (Includes precise cache routing for multi-turn workloads). - `predicted-latency-routing` (Includes dynamic latency-based routing overlays). - `pd-disaggregation` (Separates prefill and decode onto dedicated model servers. **TPU `v6e` only**, and only with the `qwen/qwen3-32b` model). - **CRITICAL WARNING:** 1. **NEVER run any `teardown-*.sh` scripts** (e.g., `teardown-llmd-optimized-baseline.sh`) when attempting to "fix and rerun" a deployment or remove a workload. These scripts default to a full core platform teardown (`ACP_TEARDOWN_CORE_PLATFORM=true`) and will completely destroy the GKE cluster and all its resources. If a deployment fails, debug in place or re-run the `deploy` scripts. ONLY run teardown script when the user explicitly asks to tear down the stack and confirm with you first. 2. **NEVER Store a secret value or retrieve a value of a secret** in memory or write it to any file. Instead, always use gcloud secrets describe [SECRET_NAME] OR kubectl get kubectl get secret [SECRET_NAME] -o yaml. If a secret does not exist ask the user to create it. ## 1. Prerequisites & Cluster Setup 1. **Ask the user**: "Is there an existing GKE cluster? (yes/no)" 2. **If NO**: - **Ask the user**: "What platform name would you like to use for the new cluster, and what is your Google Cloud project ID?" - Use `sed` to inject these values into `${ACP_REPO_DIR}/platforms/gke/base/_shared_config/platform.auto.tfvars`: ```bash sed -i 's/^platform_name.*/platform_name = ""/g' "${ACP_REPO_DIR}/platforms/gke/base/_shared_config/platform.auto.tfvars" # If platform_default_project_id doesn't exist, append it: grep -q "^platform_default_project_id" "${ACP_REPO_DIR}/platforms/gke/base/_shared_config/platform.auto.tfvars" || echo "platform_default_project_id = \"\"" >> "${ACP_REPO_DIR}/platforms/gke/base/_shared_config/platform.auto.tfvars" sed -i 's/^platform_default_project_id.*/platform_default_project_id = ""/g' "${ACP_REPO_DIR}/platforms/gke/base/_shared_config/platform.auto.tfvars" ``` - Proceed to **Section 2: Configure Model & Accelerator** (to configure variables and create cluster). 3. **If YES**: - Verify connectivity to the target cluster: ```bash kubectl cluster-info ``` - **Ask the user**: "Is the llm-d stack already deployed on this cluster? (yes/no)" - **If YES**: - Skip to **Section 4: Hugging Face Token Setup** (we only need to deploy the model server, which is done in Section 5). - **If NO**: - Proceed to **Section 2: Configure Model & Accelerator** (to configure the deployment variables). ## 2. Configure Model & Accelerator 1. **Ask the user**: "Which model would you like to run?" - **Allowed Models**: - `google/gemma-4-31b-it` - `qwen/qwen3-32b` (default) - `qwen/qwen3-32b-fp8` - `redhatai/gemma-4-31b-it-fp8-block` - _Validation_: If the user inputs anything else, flag it as invalid, display the allowed models, and ask again. 2. **Ask the user**: "Which accelerator would you like to use?" - **Allowed Accelerators**: - `rtx-pro-6000` (default) - `h100` (translates to `nvidia-h100`) - `h200` (translates to `nvidia-h200`) - `v6e` (TPU, translates to `google-tpu-v6e`) - _Validation_: If the user inputs anything else, flag it as invalid, display the allowed accelerators, and ask again. 3. **Update configuration**: - Use `sed` to inject the chosen model and accelerator into `${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/examples/llmd/_shared_config/llmd.auto.tfvars`: ```bash echo "llmd_model_id = \"\"" >> "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/examples/llmd/_shared_config/llmd-shared.auto.tfvars" echo "llmd_accelerator_type = \"\"" >> "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/examples/llmd/_shared_config/llmd-shared.auto.tfvars" ``` ## 3. Select & Deploy Well-Lit Path Guide 1. **Ask the user**: "Which llm-d well-lit path guide would you like to deploy?" - **Allowed Options**: - `optimized-baseline` (corresponds to [llmd-optimized-baseline-vllm-with-hf-model.md](../../docs/platforms/gke/base/use-cases/inference-ref-arch/llmd/well-lit-paths/llmd-optimized-baseline-vllm-with-hf-model.md)) - `precise-prefix-cache-routing` (corresponds to [llmd-precise-prefix-cache-routing-vllm-with-hf-model.md](../../docs/platforms/gke/base/use-cases/inference-ref-arch/llmd/well-lit-paths/llmd-precise-prefix-cache-routing-vllm-with-hf-model.md)) - `predicted-latency-routing` (corresponds to [llmd-predicted-latency-routing-vllm-with-hf-model.md](../../docs/platforms/gke/base/use-cases/inference-ref-arch/llmd/well-lit-paths/llmd-predicted-latency-routing-vllm-with-hf-model.md)) - `pd-disaggregation` (corresponds to [llmd-pd-disaggregation-vllm-with-hf-model.md](../../docs/platforms/gke/base/use-cases/inference-ref-arch/llmd/well-lit-paths/llmd-pd-disaggregation-vllm-with-hf-model.md)) - _Validation_: If the user inputs anything else, flag it as invalid, display the list of allowed options, and ask again. - _Validation_: If the user chose `pd-disaggregation`, the accelerator MUST be `v6e` and the model MUST be `qwen/qwen3-32b`. If either differs, explain that this guide is currently TPU-only in this repository and return to Section 2 to reconfigure. 2. **Deploy the baseline stack**: Run the deployment script corresponding to the chosen guide to create the cluster (if new) and deploy the baseline infra/services: - For `optimized-baseline`: ```bash "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/examples/llmd/deploy-llmd-optimized-baseline.sh" ``` - For `precise-prefix-cache-routing`: ```bash "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/examples/llmd/deploy-llmd-precise-prefix-cache-routing.sh" ``` - For `predicted-latency-routing`: ```bash "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/examples/llmd/deploy-llmd-predicted-latency-routing.sh" ``` - For `pd-disaggregation`: ```bash "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/examples/llmd/deploy-llmd-pd-disaggregation.sh" ``` 3. **Run validation**: - Run the following command to verify the chosen custom compute class exists on the cluster: ```bash kubectl get computeclasses ``` - Ensure that the accelerator type you chose appears in the list. If it does not, warn the user that the accelerator is unsupported or missing its custom compute class. - For `pd-disaggregation`, the required compute class is `tpu-v6e-2x4` (8 chips), not the `tpu-v6e-2x2` used by the other guides. ## 4. Hugging Face Token Setup 1. **Instruct the user** to add their Hugging Face Read Token to Google Secret Manager and as a Kubernetes secret: Provide them with these commands, replacing `` with their actual token. Note that the `source` command must be run in the same shell session as the subsequent commands so the environment variables are preserved: ```bash # Source environment variables source "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/examples/llmd/_shared_config/scripts/set_environment_variables.sh" # Add to Secret Manager HF_TOKEN_READ= echo ${HF_TOKEN_READ} | gcloud secrets versions add ${huggingface_hub_access_token_read_secret_manager_secret_name} --data-file=- --project=${huggingface_secret_manager_project_id} # Add to Kubernetes kubectl -n ${llmd_namespace} create secret generic llm-d-hf-token --from-literal=HF_TOKEN="${HF_TOKEN_READ}" ``` 2. **WAIT**: Stop calling tools and ask the user to confirm once they add the HF token to secret manager and kubernetes secret Do not proceed until the user confirms. ## 5. Deploy Model Download Job & Model Server Once the user confirms the token is configured, proceed with the deployment: 1. **Deploy the model download job**: ```bash # Configure "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/kubernetes-manifests/model-download/configure_huggingface.sh" # Apply kubectl apply --kustomize "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/kubernetes-manifests/model-download/huggingface" ``` 2. **Wait for download to complete**: Monitor the job: ```bash kubectl get job -n ${huggingface_hub_downloader_kubernetes_namespace_name} ``` Wait until the job status shows `Complete`. 3. **Clean up the download job**: Once the download is complete, delete the job to free up GKE resources: ```bash kubectl delete job -n ${huggingface_hub_downloader_kubernetes_namespace_name} ${HF_MODEL_ID_HASH}-hf-model-to-gcs ``` 4. **Deploy the Model Server**: - Configure the model server (run the script for GPU or TPU as indicated by the chosen accelerator): - If GPU: ```bash "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/kubernetes-manifests/online-inference-gpu/llmd-/vllm/configure_vllm.sh" ``` - If TPU: ```bash "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/kubernetes-manifests/online-inference-tpu/llmd-/vllm/configure_vllm.sh" ``` - Deploy using the appropriate overlay directory. Construct the directory path as: `platforms/gke/base/use-cases/inference-ref-arch/kubernetes-manifests/online-inference-[gpu|tpu]/llmd-[spec]/vllm/[prefix]-[suffix]` - `[gpu|tpu]`: Use `tpu` if the accelerator is `v6e`, otherwise `gpu`. - `[spec]`: The well-lit path chosen in Section 3. - `[prefix]`: The accelerator prefix (e.g., `rtx-pro-6000`, `h100`, `h200`, `v6e`). - `[suffix]`: The model name suffix (e.g., `gemma-4-31b-it`, `qwen3-32b`). ```bash kubectl apply --kustomize "${ACP_REPO_DIR}/" ``` ## 6. Verification - Check that all pods are running and services are accessible. (Note: You may need to source the environment variables script first to get `$llmd_namespace`): ```bash source "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/examples/llmd/_shared_config/scripts/set_environment_variables.sh" kubectl get pods -n ${llmd_namespace} kubectl get svc -n ${llmd_namespace} ``` - Verify the model downloader job has completed and the model files are in the GCS bucket: ```bash source "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/examples/llmd/_shared_config/scripts/set_environment_variables.sh" kubectl get jobs -n ${llmd_namespace} gcloud storage ls gs://${huggingface_hub_models_bucket_name}/${llmd_model_id}/ ``` - Ensure the Hugging Face token is securely configured: ```bash source "${ACP_REPO_DIR}/platforms/gke/base/use-cases/inference-ref-arch/examples/llmd/_shared_config/scripts/set_environment_variables.sh" kubectl describe secretProviderClass huggingface-tokens -n ${llmd_namespace} kubectl describe secret llm-d-hf-token -n ${llmd_namespace} ```