# Anyray Helm Chart Deploys the full Anyray stack — gateway, optimizer, proxy (console), and Postgres — to any Kubernetes cluster. No external chart dependencies; all services are self-contained in this chart. The gateway persists content-free traces, spend, and observations to Postgres (`anyray_traces` / `anyray_observations`, auto-created; any stored content is AES-256-GCM encrypted at rest) and reads them in-process — there is no separate observability backend to run. ## Prerequisites - Kubernetes 1.24+ - Helm 3.10+ - A default StorageClass (or set `*.storageClass` in values) - A target namespace that already exists, if you install outside your current namespace - `setup.sh --k8s --connect ` run locally to generate the Secret manifest ## Install from the OCI registry (recommended for ArgoCD / GitOps) The chart is published as an OCI artifact to the same public registry as the images — no git clone, no `setup.sh`: ```bash helm install anyray oci://public.ecr.aws/anyray/anyray \ --version \ --namespace "$ANYRAY_NAMESPACE" \ -f my-values.yaml ``` Browse released versions with `helm show chart oci://public.ecr.aws/anyray/anyray`. Pulls are anonymous — nothing to authenticate. **Your values file must set `postgres.storage`.** The gateway keeps 90 days of trace content by default, so the store grows for three months before the first prune reclaims anything, and the size lands in a StatefulSet `volumeClaimTemplate` that Kubernetes makes immutable. A fresh install below `postgres.minStorageGi` (150Gi) is refused at render time rather than provisioning a volume that wedges later: ```yaml postgres: storage: 150Gi # or your own measured 90-day figure ``` Running a smaller volume deliberately (short `ANYRAY_SPEND_RETENTION_DAYS`, or `ANYRAY_CONTENT_MODE=off` so no content is stored) needs `postgres.acknowledgeSmallVolume: true`. ArgoCD and Flux render with `helm template`, which reports an install, so an existing deployment already below the floor needs that same line to keep syncing; its live volume is untouched. Sizing model and the queries that measure your own rate: https://docs.anyray.ai/get-started/install/choose-your-setup#size-the-datastore-before-you-install ### ArgoCD Point an Application straight at the OCI chart: ```yaml apiVersion: argoproj.io/v1alpha1 kind: Application metadata: name: anyray namespace: argocd spec: project: default source: repoURL: public.ecr.aws/anyray # registry, no chart name chart: anyray targetRevision: 0.4.17 # a released chart version helm: valuesObject: host: anyray.example.com gateway: publicUrl: https://anyray.example.com consolePublicUrl: https://anyray.example.com # image.tag omitted → pinned to the chart's appVersion (see note below) destination: server: https://kubernetes.default.svc namespace: team-ai ``` The chart reads every credential from a Secret named `anyray-secrets` (see [Secrets](#secrets)). Manage it the GitOps way — External Secrets Operator, Sealed Secrets, or SOPS — or apply a plain manifest once. You do **not** need `setup.sh` for a GitOps install; it is only a local convenience that writes that Secret plus a starter values file, and it never runs inside your cluster. The two keys that must be present are `ANYRAY_DEPLOYMENT_TOKEN` (your `adt_…` connect token) and an `ANYRAY_PSEUDONYM_SALT` you generate once and keep in-cluster. > **Image tags are pinned by default.** Each chart version ships a fixed > `appVersion`, and images default to it, so a given `targetRevision` always > deploys the same, auditable build. That is what you want under GitOps, and it > is unchanged from earlier charts: upgrading the chart never moves a running > deployment off its pinned build. > > **Auto-update is one value** (chart 0.6.0+). Set `image.tag: policy-stable` and > the deployment keeps itself current: the channel is promoted on every release, > the chart forces `imagePullPolicy: Always`, and the bundled `autoUpdate` > CronJob rolls the Deployments nightly so the moving tag actually resolves to a > new digest. `setup.sh --k8s` writes that line into the values file it > generates, so a new install is self-updating out of the box. > > ```yaml > image: > tag: policy-stable # the whole opt-in; autoUpdate follows automatically > ``` > > Leave the tag unset and the nightly roll is skipped on its own, because a > restart onto an identical build is pure churn. See > [Automatic image updates](#automatic-image-updates-autoupdate). ## Install from source (setup.sh) ```bash # 1. Choose the namespace (optional, but recommended) export ANYRAY_NAMESPACE="team-ai" # replace with your target namespace # 2. Generate secrets ./setup.sh --k8s --connect adt_XXXX --host --namespace "$ANYRAY_NAMESPACE" # Emits: anyray-secrets.yaml my-values.yaml # 3. Apply the Secret kubectl apply -n "$ANYRAY_NAMESPACE" -f anyray-secrets.yaml # 4. Install the chart helm install anyray ./helm -f my-values.yaml --namespace "$ANYRAY_NAMESPACE" # 5. Wait for pods kubectl rollout status -n "$ANYRAY_NAMESPACE" deployment/anyray-gateway kubectl rollout status -n "$ANYRAY_NAMESPACE" deployment/anyray-proxy # 6. Access # Ingress: console https:/// gateway https:///v1 # NodePort: console http://:30000 gateway http://:30787 ``` Use an existing namespace unless your cluster policy says the installer should create one. If you are allowed to create it, run `kubectl create namespace "$ANYRAY_NAMESPACE"` before applying the Secret. Omit `--namespace` and the `-n` / `--namespace` flags to use your current kubectl and Helm namespace. `setup.sh` never creates a namespace automatically and never assumes `default`. ## Upgrade ```bash git fetch && git reset --keep origin/main # Run before the first upgrade from an older Secret. Safe to repeat. ./setup.sh --k8s --connect adt_XXXX --host --namespace "$ANYRAY_NAMESPACE" kubectl apply -n "$ANYRAY_NAMESPACE" -f anyray-secrets.yaml helm upgrade anyray ./helm -f my-values.yaml --namespace "$ANYRAY_NAMESPACE" ``` Apply the updated Secret before upgrading. The gateway pod will not start without `ANYRAY_DEPLOYMENT_TOKEN` and `ANYRAY_PSEUDONYM_SALT`. When following `policy-stable` the bundled `autoUpdate` CronJob does this for you nightly. To roll immediately instead of waiting for the schedule: ```bash kubectl rollout restart deployment -n "$ANYRAY_NAMESPACE" \ -l app.kubernetes.io/instance=anyray kubectl rollout status deployment -n "$ANYRAY_NAMESPACE" \ -l app.kubernetes.io/instance=anyray --timeout=10m ``` ### Automatic image updates (`autoUpdate`) Added in chart 0.6.0, and armed by a single value. `autoUpdate.enabled` is `true` out of the box but renders **nothing** while `image.tag` is pinned, because a nightly restart onto an identical build is pure churn: setting `image.tag: policy-stable` is what turns the feature on. A CronJob then pins each Deployment to the digest the tag resolves to that night and runs `kubectl rollout restart` against this release's Deployments. The pin is what keeps every replica on one build between rolls. Without it, a pod that starts later (a reschedule, a node drain, an autoscale) pulls whatever `policy-stable` points at by then, and replicas split across builds: proxies on two builds blank the console, optimizers on two builds rewrite a session's prompt cache on every hop between them. The Job's init containers pull each app image at the moving tag, so the kubelet resolves the digests with your own pull secrets and mirror. Every digest resolves before anything is patched, so a failed pull skips the night's roll rather than half-applying it. The Deployments then show images as `public.ecr.aws/anyray/:policy-stable@sha256:…`. A `helm upgrade` renders the plain tag again, which rolls onto the newest build, and the next scheduled run pins it. That pairing is deliberate rather than a quirk. It means upgrading the chart never changes what a running deployment does: an install that never set a tag renders exactly the workloads it did before, with no CronJob. The Postgres StatefulSet is never restarted, and the selector is scoped to this release's instance label, so a second release in the same namespace is untouched. | Value | Default | What it does | | --- | --- | --- | | `autoUpdate.enabled` | `true` | Arms the roll. Renders nothing unless `image.tag` can actually move, so a pinned install is never restarted onto the same build. Set `false` to keep a moving tag without a scheduled roll. | | `autoUpdate.schedule` | `"30 3 * * *"` | Standard cron. Daily, outside working hours. | | `autoUpdate.timeZone` | `""` | IANA name (`Europe/Berlin`). Empty uses the cluster's zone. Needs Kubernetes 1.27+. | | `autoUpdate.image.repository` / `.tag` | `registry.k8s.io/kubectl` / `v1.33.0` | Any kubectl within one minor of your cluster. Mirrored by `global.imageRegistry` like every other image. The upstream image has no `latest` tag, so this is always explicit. | | `autoUpdate.resources` | 10m/32Mi → 100m/128Mi | Per container: the image resolvers, the pin step, and kubectl. Each makes a few API calls at most. | | `autoUpdate.nodeSelector` / `.tolerations` | `{}` / `[]` | Falls back to the global scheduling values. | It needs a `Role` (never a `ClusterRole`) granting `get`, `list` and `patch` on `apps/deployments` in this namespace, and `get` on `pods`: `get`/`list` resolve the label selector, `patch` pins the image and stamps the restart annotation, and the pod read is the Job reading its own resolved digests. Nothing else. **Under ArgoCD or Flux, prefer pinning in Git over this roll.** A Job that patches Deployments competes with the controller that owns them. Leave `image.tag` unset (each chart version deploys its `appVersion`, and the roll renders nothing), and upgrade by raising the chart version: `targetRevision` on ArgoCD, `spec.chart.spec.version` on a Flux `HelmRelease`. What each version changed: https://docs.anyray.ai/changelog (a chart version's `appVersion` names the release). If you do run `policy-stable` under ArgoCD, the pinned image reads as drift from the rendered `:policy-stable`. On that channel the rendered image never changes, so ignore the field rather than let a sync undo the pin: ```yaml ignoreDifferences: - group: apps kind: Deployment jqPathExpressions: - .spec.template.spec.containers[].image syncPolicy: syncOptions: - RespectIgnoreDifferences=true ``` > **Not with `ServerSideApply=true`.** ArgoCD 2.9.2 with server-side apply has > synced this as gateway env entries that kept their names and lost their > values: `ANYRAY_ADMIN_TOKEN`, `ANYRAY_CONTENT_KEY`, `ANYRAY_PSEUDONYM_SALT` > and three more, all empty. Use it with client-side apply only, and after the > first sync check that > `kubectl get deploy anyray-gateway -n "$ANYRAY_NAMESPACE" -o yaml` still shows a > value or `valueFrom` on every entry. A stripped `ANYRAY_ADMIN_TOKEN` shows on > `/admin/health` (and the console Health page) as `secrets.adminToken: env_blank`. **A build that will not start stalls the roll, it does not drop the deployment.** The gateway, optimizer and proxy roll with `maxUnavailable: 0`, so a replacement must pass its readiness probe before an old pod is retired. Turning `optimizer.persistence.enabled` back on is the exception: its `ReadWriteOnce` volume forces `strategy: Recreate` and a real gap. The gateway fails open across that gap, so it costs prompt-cache efficiency rather than availability. **A roll is unconditional.** It cannot tell a release that only needs new bytes from one that needs a new setting from you first, which is the check the gateway makes for itself where an applier exists. Such a release will not start and the roll stalls; the console's **Updates** panel names what to set. ### Rollout behaviour and downtime **The gateway rolls without a gap by default.** Since chart 0.5.0 `gateway.persistence.enabled` defaults to `false`, which is what allows `replicas: 2` and the `RollingUpdate` strategy with `maxUnavailable` held at zero: a replacement is Ready before any old pod is retired. That volume is `ReadWriteOnce`, so one node must release it before a replacement can mount it. Any component still on it uses `strategy: Recreate` (every old pod stops, and only then does a new one start) and is capped at one replica. No setting makes a rolling update possible while a single-attach volume is in play. The optimizer has been off its PVC by default since chart 0.7.0, and its runtime config lives in the shared Postgres, so every replica reads the same config. Running the gateway on `emptyDir` is lossless as of appVersion v1.10.224, and the chart refuses to render an older image without the volume. Everything an operator sets now lives in the shared Postgres and is read identically by every replica: client keys, provider keys, per-user caps, model aliases, team policy, the audit trail, the runtime settings (content-capture mode, heartbeat tier), the default routing config, the end-point fleet config, and the fleetd installers. Spend and trace history were always there. What still resets with the pod is the entitlement-lease cache, which is deliberately node-local and re-fetched. Independent of the strategy, three values shape how termination is handled, and each is overridable per component: | Value | Default | What it does | | --- | --- | --- | | `preStopDrainSeconds` | `5` | Keeps a pod serving after termination starts but before SIGTERM, so requests stop arriving before the process stops listening. On the `Recreate` path this also lengthens the gap, since nothing starts until the old pod is gone — set `0` to trade in-flight requests for a shorter window. | | `terminationGracePeriodSeconds` | `120` | Budget from "terminating" to SIGKILL. Must exceed `preStopDrainSeconds` plus **that component's own** drain of in-flight requests, and the chart charges each one separately: gateway 90000ms and optimizer 15000ms (both `ANYRAY_SHUTDOWN_DRAIN_MS`, different defaults), endpoint-control a hardcoded 35000ms. Too low and streams are cut mid-response, which the developer's tool reports as "Connection closed mid-response". A ceiling, not a delay; the chart refuses to render if the sum no longer fits. | | `minReadySeconds` | `10` | How long a new pod must stay Ready before the rollout counts it available, so a pod that passes one probe and then crashes cannot retire the previous version. | | `progressDeadlineSeconds` | `300` | How long a rollout may make no progress before it is marked failed. Unset, Kubernetes waits 600s, so a rollout whose pods never pass their probes stalls silently for ten minutes. Since `maxUnavailable` is 0 the outgoing pods serve throughout, so a shorter deadline costs no traffic — it is what makes `helm upgrade --atomic` roll back promptly. Must exceed `minReadySeconds`. | | `revisionHistoryLimit` | `3` | Retired ReplicaSets kept per Deployment for rollback (Kubernetes default 10). | ### Probes Each workload has a **readiness** probe (is this pod fit to receive traffic?) and, from appVersion `v1.10.224`, a **liveness** and **startup** probe (should the kubelet restart this pod?). The liveness probe exists for one failure that readiness cannot fix: a *wedged* process — an exhausted connection pool, a hung await — is up, gets pulled from the Service by readiness, and is then never restarted. At `replicas: 1` that is a silent permanent outage. | Value | Default | What it does | | --- | --- | --- | | `livenessProbe.enabled` | `true` | Restarts a wedged pod. Targets `/livez`: the gateway from appVersion `v1.10.224`, the optimizer from `v1.10.432` (earlier optimizer pins keep `/health`). The gateway's readiness route turns 503 while draining; a liveness probe there would restart the pod mid-drain and reset every in-flight stream. The optimizer's newer `/health` turns 503 while warming, so its liveness uses `/livez`. Neither route checks a dependency, so a database blip cannot crashloop the fleet. | | `livenessProbe.failureThreshold` x `periodSeconds` | `6` x `10s` | A full minute of failure before a restart. Slack on purpose: restarting a merely-slow pod is itself an outage. | | `startupProbe.enabled` | `true` | Holds liveness off until the process answers once, so the liveness budget above need not cover a cold start. 360 x 5s = 30 minutes, which must cover a full migration run: the gateway applies the ordered ledger at boot and refuses to serve on an unmigrated schema, so it opens no port until every pending migration has run, and the ledger allows a single statement 30 minutes (`MIGRATION_QUERY_TIMEOUT_MS`). A startup probe only ever delays a *restart*, never traffic, so the budget costs nothing. | | `postgres.livenessProbe.enabled` | `true` | `pg_isready` against the bundled Postgres, with slack thresholds — interrupting crash recovery is worse than waiting it out. | The gateway's probes render **only for an image that serves `/livez`** (appVersion `v1.10.224` or later, or the `policy-stable` channel). Pinned to an earlier tag the chart omits them rather than 404 every check and crashloop the deployment. A `PodDisruptionBudget` (`podDisruptionBudget.maxUnavailable`, default `1`) is created for each Deployment that runs two or more pods, capping how many pods a node drain or autoscaler scale-down may remove at once. It is deliberately not created for a single-replica workload: such a budget could never allow its only pod to be evicted, and `kubectl drain` would block indefinitely on a node upgrade. At the defaults that covers all four workloads, each at two replicas. `podDisruptionBudget.unhealthyPodEvictionPolicy` defaults to `AlwaysAllow`. Kubernetes' own default (`IfHealthyBudget`) refuses to evict pods that are running but not Ready unless the workload is already at full health, which makes a broken workload permanently undrainable — a node upgrade then blocks forever on exactly the pods serving no traffic. `AlwaysAllow` still protects every Ready pod through `maxUnavailable`. Requires Kubernetes 1.26+; older clusters ignore the field. ### Spreading replicas across nodes With no `affinity` set, the chart applies **soft** (preferred) pod anti-affinity per component: replicas separate by `kubernetes.io/hostname` first, then by `topology.kubernetes.io/zone`. Without it, two replicas routinely land on one node, and a single node drain takes out the whole workload the second replica exists to protect. Soft rather than required is deliberate — a required rule on a single-node or capacity-tight cluster leaves pods `Pending` forever. Setting `affinity` (globally or per component) replaces this wholesale. ## Billing Billing is required. The chart always enables content-free usage metering. Create or update the Secret with: ```bash ./setup.sh --k8s --connect adt_XXXX --host --namespace "$ANYRAY_NAMESPACE" ``` `setup.sh` adds the deployment token and a local pseudonym salt to `anyray-secrets.yaml`. The chart requires both keys. The Billing URL and verification key are pinned in the gateway image, and the salt stays in your cluster. Tune the reporting interval with `gateway.metering.intervalMs` / `gateway.metering.graceMs`. ## Exposing services By default all Services are `ClusterIP`. To expose externally: **NodePort (simplest for on-prem / bare metal):** ```yaml proxy: service: type: NodePort nodePort: 30000 gateway: service: type: NodePort nodePort: 30787 ``` Note: the `gateway` and `optimizer` Services use bare (unprefixed) names because nginx inside the proxy image hardcodes those upstream hostnames. Install the chart in its own namespace to avoid name collisions. **LoadBalancer (cloud providers):** Set the proxy and gateway service type in `my-values.yaml`, and scope traffic to your org/VPN: ```yaml proxy: service: type: LoadBalancer loadBalancerSourceRanges: - 203.0.113.0/24 gateway: service: type: LoadBalancer loadBalancerSourceRanges: - 203.0.113.0/24 ``` **Ingress:** Set `ingress.enabled: true` in your values file and fill in `ingress.className` and any cert-manager annotations. See the commented example in `values.yaml`. The `/v1` path carries streaming completions that run for minutes and request bodies over a megabyte, which ingress-nginx's stock 60s `proxy-read-timeout` and 1 MB `proxy-body-size` both cut short. `ingress.streamingDefaults` (default `true`) therefore adds `proxy-read-timeout` and `proxy-send-timeout` `3600`, `proxy-body-size` `32m`, and `proxy-buffering` `off`. Anything you set in `ingress.annotations` wins over these key by key; set `streamingDefaults: false` to drop them entirely. On the Gateway API path, `httpRoute.streamingDefaults` (default `true`) disables the request timeout on the `/v1` and `/connect` rules, because an unset Gateway API request timeout is implementation-specific and some implementations apply a value far below one completion. For TLS / Ingress installs, set the gateway and console public URLs explicitly so links and auth callbacks use the externally reachable scheme and host: ```yaml gateway: publicUrl: https://anyray.example.com consolePublicUrl: https://anyray.example.com ``` If you need multiple Ingress hosts or non-default paths, use `ingress.hosts`: ```yaml ingress: enabled: true className: nginx hosts: - host: anyray.example.com paths: console: / gateway: /v1 admin: /admin ``` ### The end-point agent plane The `endpoint-control` service (device evidence for employee laptops) is the one component whose clients live outside the cluster, so it needs a route of its own. On an Ingress or HTTPRoute install the chart adds three fixed prefixes — `/api/v1/osquery`, `/api/fleet/orbit`, `/api/v1/evidence` — on the same host the console and gateway already use. The agent protocol anchors those paths at the root, so they are not configurable; the gateway serves no `/api/*` route, so nothing collides, and no second hostname or certificate is involved. Everything else the service exposes (`/admin/*`, `/readyz`, `/livez`) stays in-cluster: the gateway reaches its admin plane over the cluster Service. Two consequences worth knowing before you install: - **`LoadBalancer` and `NodePort` installs route none of it.** Nothing external reaches the agent plane, so the chart deliberately leaves `ANYRAY_ENDPOINT_CONTROL_PUBLIC_URL` unset rather than guessing an origin from `host` — a guessed origin would be baked into every installer and MDM profile the service mints and fail only on the laptop. Expose the plane yourself and set `endpoint-control.publicUrl` to the origin that reaches it. - **`endpoint-control.enabled: false`** drops the Deployment, the Service, the ingress rules, and the gateway's pointer at it, leaving the end-point lane dormant. Use it if a device-evidence plane hasn't cleared your security review, or if you pin `image.tag` below `v1.10.246` — the first release that published this image, so older pins have nothing to pull. ## Signing gateway → optimizer requests The optimizer's shared token proves a caller is inside the deployment, not that it is the gateway, and both services read that token from the same Secret. Request signing separates the two: the gateway holds an Ed25519 private key and the optimizer only the public half, so a workload that can read the optimizer's environment still cannot produce a valid request. Worth enabling where other workloads share the namespace — it limits what a compromised sidecar, a co-tenant, or a leaked token can do. ### Generate the keypair `setup.sh` mints the pair for you, but most Kubernetes installs never run it. Generate it yourself and add both halves to the Secret the chart already reads: ```bash openssl genpkey -algorithm ed25519 -out signing.pem openssl pkey -in signing.pem -pubout -out verify.pub kubectl -n "$ANYRAY_NAMESPACE" patch secret anyray-secrets --type merge -p "$(cat <`. Never commit `anyray-secrets.yaml` — it is in `.gitignore` at the install repo root. ## Uninstall ```bash helm uninstall anyray kubectl delete -f anyray-secrets.yaml # To also delete PVC data (destructive): kubectl delete pvc -l app.kubernetes.io/instance=anyray ``` Add `--namespace "$ANYRAY_NAMESPACE"` to `helm uninstall` and `-n "$ANYRAY_NAMESPACE"` to `kubectl delete` commands when you installed into a specific namespace. ## Troubleshooting ### Postgres won't start — `initdb: directory "/var/lib/postgresql/data" exists but is not empty` / `lost+found` A volume mounted at the root of an `ext4` filesystem — which is what most cloud block stores provision, e.g. an **AWS EBS** PVC — always contains a `lost+found` directory, and Postgres's `initdb` refuses any non-empty data directory. The chart avoids this by initializing into a `pgdata` subdirectory of the mount (`PGDATA=/var/lib/postgresql/data/pgdata`). If you fork the Postgres template or mount your own volume, keep `PGDATA` (or a `subPath`) pointed at a subdirectory, never the mount root. ### Mixed-architecture clusters (arm64 + amd64 nodes) Every image the chart ships is a **multi-arch manifest list** (`linux/amd64` + `linux/arm64`) — the gateway, optimizer, and proxy images plus Postgres — so Kubernetes schedules each pod onto any node and pulls the matching architecture automatically. **No `nodeSelector` by architecture is required.** To deliberately pin a workload to a node pool (e.g. keep Postgres on a specific instance type), set `nodeSelector` / `affinity` / `tolerations` / `topologySpreadConstraints` — either globally (every pod) or **per component** on the individual `gateway` / `optimizer` / `proxy` / `postgres` blocks. See [Cluster policy knobs](#cluster-policy-knobs).