--- name: stack-profile description: >- The single stack-definition point — what this team runs today, which stacks it authors versus only supports, the stay-in-lane rule, and the platform boundary. Load before editing code in a support-only stack (Java/JVM), before recommending any runtime, tool, or infrastructure change, and when choosing between observability backends. Triggers: "what's our stack", "should we use X for this", "can we move this to Kubernetes / the cloud", "which backend do I query". This skill bundle changes when the ground shifts. argument-hint: "[the runtime, tool, or infrastructure question]" --- # Stack profile — current facts, not aspirations Phrased as what is true today. When the ground shifts, update this canonical skill bundle and regenerate its projections; no other canonical stack definition should change. Use only `[verified]`, `[sourced]`, and `[unverified]` evidence labels, and keep each claim's state intact in transit; describe an inference in prose and label it `[unverified]` rather than inventing a fourth state. A planned or candidate technology is never a current-stack fact until a human owner records the decision here. ## The business The firm is a stock trading firm; use trading examples (orders, accounts, markets, exchanges), never retail ones. Most incidents trace to dependencies, most often order management, the trading apps, and the quote plant. *[sourced: operator statement 2026-09-22]* ## Runtime On-prem servers + PCF (VMware Tanzu Application Service); this is what runs today. **The team operates PCF through Apps Manager**, not the command line: many SREs do not have the `cf` CLI installed. Skills give first checks as Apps Manager views with the `cf` v8 (CAPI V3) equivalent as a fallback, and the `sre-assistant` agent says when `cf` is absent where it runs rather than pretending to have observed the platform. *[sourced: operator statement 2026-09-02]* **GCP migration is in progress**: GCP is an approved target, arriving (as planned) as reference files inside the obs skills plus the `gcp-ops` triage skill, not as a restructure. The landing runtime is **decision-pending** (Cloud Run is the primary candidate for TAS-shaped apps; GKE only if a workload demands it) — do not present either as decided. [unverified — record the runtime decision here when a human owner accepts it]. **No self-managed Kubernetes**; on-prem stays Kubernetes-free. ## Observability decision The incumbent and additive observability stacks coexist as first-class; no listed backend is retired by team decision. Read the conditional observability reference for the signal inventory, query languages, lifecycle evidence, and GCP additions. During an incident the responder's tools are, in order: Apps Manager for what changed and instance state, **Grafana** for the service's dashboards, panels, and alert state, Splunk for logs beyond the last minutes, and **Wavefront and PCF App Metrics** for application metrics. Grafana's Mimir, Loki, and Tempo backends are the additive stack: GCP workloads and services already instrumented with OpenTelemetry land there. *[sourced: operator statement 2026-09-02; Grafana second, owner 2026-09-22]* ## Read only the conditional stack facts the request needs | If the request involves… | Read first | |---|---| | A broad inventory or cross-domain comparison, including "what's our stack" | All three: [Observability stack](./references/observability-stack.md), [Application and data stack](./references/application-and-data-stack.md), and [Copilot models](./references/copilot-models.md) | | An observability backend, signal, query language, vendor lifecycle, GCP observability choice, or edge/CDN/WAF/RUM product | [Observability stack](./references/observability-stack.md) | | A service language, framework, CI platform, tooling, or authentication design, runner/host assumption, or data-store choice | [Application and data stack](./references/application-and-data-stack.md) | | Selecting or recording the team's current Copilot model and fallbacks | [Copilot models](./references/copilot-models.md) | | A term or role a responder may not know (blast radius, golden signals, release owner, platform team) | [Terms and roles](./references/terms-and-roles.md) | Load every matching row and no others. These references provide current facts; they do not widen the app/ops lane, settle a pending decision, authorize a platform change, or replace target-specific verification. The entrypoint rules remain authoritative after a reference is loaded. ## Incident response A formal on-call rotation is in place. *[sourced: operator statement 2026-08-21]* The team investigates and recommends fixes. An existing bridge or TLC (Techline Chat) is the coordination channel, not a request to open another. *[sourced: operator statement 2026-09-12]* ITO (IT Operations) runs the TLC once one is open: it asks for updates, pages teams, and approves changes, and does not run the investigation. There is no standing incident lead or commander. *[sourced: operator statement 2026-09-22]* `incident-investigation` owns technical advice, the investigation board, and recommendations based on supplied impact and policy. Who sets severity is not recorded. *[unverified — record the owner]* ## Change management Change records live in **both BMC Remedy and Jira**. `production-change-gate` refers to "the formal change record" generically; name whichever system governs the change in hand rather than assuming one. *[sourced: operator statement 2026-08-21]* ## Documentation home **The team-owned GitHub repository is the living source.** Confluence still holds operational documentation and is actively being drained into the repo — one direction, import only. Both therefore exist today, but they are not co-equal homes: a page in Confluence is a source to import, not a destination to write to. What follows for the document lanes: `scribe` authors into the repository (it has no web tools and could not reach Confluence anyway), `runbook` owns the import path, and `skills/runbook/scripts/confluence_to_runbook.py` is live working software, not a migration leftover. Never write new operational documentation into Confluence. *[sourced: operator statement 2026-08-21]* ## Stay in lane Stay in the app/ops lane; hand platform-internal problems to the platform team. GCP managed services are now in-lane **for the migration** (Cloud Run, Cloud Logging/Monitoring/Trace, Secret Manager); do not propose self-managed Kubernetes anywhere, and do not propose GKE while the landing-runtime decision is pending — flag the need instead. On-prem/PCF infra-layer fixes remain out of lane. ## The platform boundary We own our apps up to the platform edge; we do not operate the platform. On PCF: BOSH, Ops Manager, Diego cells, Gorouter, CredHub/UAA, and foundation upgrades belong to the platform team; `pcf-ops` owns which symptoms are platform-side and the evidence to escalate with. On GCP the boundary moves and is **not yet ratified**: the team owns more (service config, revisions, project-scoped observability), while org policy, folder/project structure, shared networking, and IAM beyond project scope sit with the cloud platform owner. Treat that split as [unverified] until recorded here; the `gcp-ops` skill carries the working boundary rules. Akamai delivery and WAF config is team-owned change-managed work (see the `akamai-edge` skill); Akamai the platform — the edge network itself — is Akamai's. ## Model rule No agent sets `model:` today; inheriting the session model remains the default. A Claude agent may use a generation alias (`haiku`, `sonnet`, `opus`, `fable`, or `inherit`) only when its lane's cost or latency profile justifies tiering, and never a full model ID. When a task needs the team's current Copilot picker order or fallback sequence, load the conditional model reference and preserve its verification state; that host inventory does not select a Claude agent alias.