--- name: setting-up-cloudwatch-observability description: >- Sets up CloudWatch observability for the first time - Omni (CloudWatch Application Observability) and classic CloudWatch. Omni: creating a Space or Domain; access grants (who has access, at what level) and access profiles bounding async alerts, integrations, or agents; instrumenting an app or AI agent with the plain ADOT SDK so traces reach Omni (Python/Node/Java/.NET on EC2/ECS/EKS/Lambda), incl. no-image-rebuild and .NET CoreCLR vars; ingesting Azure telemetry via the CloudWatch agent on an Azure VM or AKS; connecting Slack to a Space; whether GitHub or a custom MCP tool server (HTTP/stdio; API key, bearer, OAuth2) can be connected. CloudWatch: onboarding a service to Application Signals - ADOT auto-instrumentation, the amazon-cloudwatch-observability add-on, monitored service, reporting telemetry, ServiceEvents, CI/CD git/deployment metadata, Terraform/manifest edits. For using what is set up - queries, dashboards, alarms, Omni alerts, X-Ray, synthetics, Dynamic Instrumentation - use aws-observability. version: 3 --- # Setting Up CloudWatch Observability > **Scope:** First-time CloudWatch observability setup, for **both** products — a CloudWatch Application Observability (Omni) Space from creation through first traces flowing, **and** onboarding a service to classic CloudWatch **Application Signals**. For using what is already set up (queries, dashboards, alarms, Omni alerts, evaluations, live debugging), route to **aws-observability**. ## Two products, two reference folders This skill owns *setup* for two products that share the CloudWatch name but are separate services. The folder a reference lives in is the signal for which product it belongs to — the same convention `aws-observability` uses: | Folder | Product | What setup means there | |---|---|---| | `references/cloudwatch-omni/` | CloudWatch Omni (Application / Agent Observability) | Domain → Space → grants → telemetry in → **plain ADOT SDK** instrumentation. No add-on, no CloudWatch Agent `application_signals` config, no port 4316. | | `references/cloudwatch/` | Classic CloudWatch | Onboarding a service to **Application Signals**: ADOT *auto*-instrumentation, the `amazon-cloudwatch-observability` EKS add-on, the CloudWatch Agent, monitored service, ServiceEvents, CI/CD git/deployment metadata, port 4316. | The two instrumentation paths are **mutually exclusive on a given workload** — the Omni path explicitly forbids the Application Signals env vars and the add-on, and vice versa. Decide the product first (Routing Rules 1–4 below), then stay inside that folder. Enabling one does not replace the other; an account may run both, on different services. **Works best with** the [AWS MCP server](https://docs.aws.amazon.com/aws-mcp/) — enables running AWS CLI commands directly. All guidance also works with standard AWS CLI access (`aws cloudwatchomni ...` on the Omni path; `aws eks`, `aws iam`, and `aws application-signals` on the Application Signals path). ## Concepts — Omni (`references/cloudwatch-omni/`) These are Omni's resources. **None of them exist on the Application Signals path** — that one has no Domain, Space, grant, or Dataset; its unit is a *monitored service*, and access is plain IAM. Do not ask an Application Signals onboarding customer for a `spaceId`. | Term | What it is | |---|---| | **Domain** | The identity boundary. Carries the authorization provider (IAM or Identity Center) and owns the endpoint URL customers reach Omni through. One per account, or one shared across an AWS Organization. | | **Space** | A workspace holding telemetry, in exactly one account and one Region. Created under a Domain. At most one per account per Region. Region rule: under an IAM-only Domain a Space may sit in a Region other than the Domain's; under an Identity Center Domain the Space must be in the Domain's own Region — Identity Center plus Spaces in several Regions needs the org-scoped Domain. | | **Access grant** | Attaches a principal — person, group, IAM identity, or async workload — to one Space at a permission level. The only way anyone reaches data *through* a Space — it does not restrict the source CloudWatch log groups, which stay readable under their own IAM. | | **Access Profile** | A named boundary for async workloads (alerts, integrations, agents) that act without a person in the loop. It is only a named container: `create-access-profile` takes a Space, a name, and a description and nothing else — no permission, action, or scope input. It does something only once two separate sets of grants exist (what the profile may do; which workloads may assume it) and a workload names it. | | **Dataset** | What queries run against. Telemetry arrives through the CloudWatch OTLP endpoints, or by forwarding what is already in CloudWatch log groups. | Omni setup order from nothing: **Domain → Space → grants → telemetry in → instrumentation.** (The Application Signals equivalent is much shorter and has no prerequisite resources — add-on/agent, IAM, then the per-platform enablement change; see `references/cloudwatch/application-signals-onboarding.md`.) "Telemetry in" (a collector exporting to CloudWatch's OTLP endpoints, plus dataset forwarding for what is already in CloudWatch) is its own step, separate from instrumenting the workloads. Access Profiles are **conditional**, not a step in the sequence — only when async workloads (alerts, integrations, agents) are involved. Whenever you give this sequence, also say that instrumentation or forwarding started before a Space exists appears to succeed while delivering telemetry nowhere the customer can see — a customer who checks Omni first reads a working, empty Space as a failure. For the concept relationships and the full arc, see `references/cloudwatch-omni/app-basics.md`. This is a **routing skill**. Classify the user's setup request and delegate to the correct reference. Omni references live under `references/cloudwatch-omni/`; classic-CloudWatch (Application Signals) references live under `references/cloudwatch/`. | User intent | Reference | |---|---| | **Onboard a service to Application Signals** (auto-instrumentation, the `amazon-cloudwatch-observability` EKS add-on, CloudWatch Agent IAM, monitored service, reporting telemetry, ServiceEvents, the two onboarding tiers) | `references/cloudwatch/application-signals-onboarding.md` | | **Propagate ServiceEvents git/deployment metadata through CI/CD** (the 5 `OTEL_AWS_SERVICE_EVENTS_*` vars, per-provider patterns) | `references/cloudwatch/application-signals-cicd-metadata.md` | | **Per-platform × per-language Application Signals enablement steps** once platform and language are known | The matching `references/cloudwatch/appsignals-guides/-.md` (e.g. `references/cloudwatch/appsignals-guides/eks-python.md`) | | **Turn on Dynamic Instrumentation for a service** at onboarding time (the `OTEL_AWS_DYNAMIC_INSTRUMENTATION_*` vars and their IAM) | `references/cloudwatch/application-signals-onboarding.md` (Step 5d). *Using* it to debug is **aws-observability** | | Understand **what Omni is**, its concepts, or **where to start** | `references/cloudwatch-omni/app-basics.md` | | **Instrument an AI agent** (ADOT, OpenInference, framework detection, trace verification), or **deploy an agent to production** and get traces flowing to CloudWatch — env vars per platform (AgentCore, Lambda, or other platforms such as ECS/EC2/EKS), routing spans to a custom trace log group, IAM permissions needed, ADOT version requirements | `references/cloudwatch-omni/omni-agents-instrumentation/omni-agents-instrumentation.md` (§ Production deployment for the deploy case) | | **Per-framework OpenInference guide** (LangChain, LangGraph, Strands, CrewAI, OpenAI Agents, Vercel AI) once the agent framework is known | `references/cloudwatch-omni/omni-agents-instrumentation/openinference-framework-guide.md`, then the matching `references/cloudwatch-omni/omni-agents-instrumentation/instrument-.md` | | **Instrument an application** (ADOT SDK on EC2/ECS/EKS/Lambda — Python, Node.js, Java, .NET) | `references/cloudwatch-omni/instrumentation/instrumentation.md` | | Emit a **custom application or agent metric** so it is queryable in Omni (why OTLP and not `PutMetricData`/EMF) | `references/cloudwatch-omni/instrumentation/instrumentation.md` (§ Custom metrics) | | Create an **account-scoped Space or Domain** | `references/cloudwatch-omni/spaces-and-domains.md` | | Whether a Space can be in a **different Region from its Domain** | `references/cloudwatch-omni/spaces-and-domains.md` (Prerequisites → Region rules) | | The **AgentCore evaluation role** that `create-space` asks for — what it is, whether to create one | `references/cloudwatch-omni/spaces-and-domains.md` (Step 3 → Also resolve the AgentCore evaluation role) | | A Space that **was created successfully but returns an authorization error** when used | `references/cloudwatch-omni/spaces-and-domains.md` (Troubleshooting → If the Space was created but cannot be used) — a space access role trust-policy problem, not a grant problem | | Create a **Domain shared across an AWS Organization** | `references/cloudwatch-omni/org-domains.md` | | Configure **access grants** for people or IAM identities or alerts | `references/cloudwatch-omni/access-grants.md` | | Bound an **alert, integration, or agent** with an Access Profile | `references/cloudwatch-omni/access-profiles.md` | | Deploy an **OTel Collector** so an instrumented app has somewhere to export to (EC2/ECS/EKS) — it exports to CloudWatch's own per-signal OTLP endpoints | `references/cloudwatch-omni/instrumentation/collector.md` | | Enable **Transaction Search** so traces reach a Space (spans land in `aws/spans` only once it is on — per account, per Region) | `references/cloudwatch-omni/instrumentation/collector.md` (Step 1) | | **Forward telemetry** already in CloudWatch into the Dataset | `references/cloudwatch-omni/data-forwarding-and-centralization.md` | | **Send your application's own telemetry from Azure** (the logs/metrics/traces your service emits, via the CloudWatch agent on an Azure VM/AKS) | `references/cloudwatch-omni/azure-ingestion/azure-ingestion.md` (intent triage), then `references/cloudwatch-omni/azure-ingestion/custom-telemetry.md` (the CloudWatch-agent-on-VM/AKS procedure) | | Connect **Slack** to a Space for the first time | `references/cloudwatch-omni/slack-integration.md` | | Whether a **custom MCP tool server** can be registered (Omni has none) | `references/cloudwatch-omni/custom-mcp-integration.md` | ## Routing Rules 1. If the user is asking **what Omni is**, what a Domain or Space or grant means, **where to start**, or **the end-to-end order of steps** from nothing — rather than asking to perform one setup step — route to `references/cloudwatch-omni/app-basics.md` and answer the sequence from its "Setup order" section (not from a single procedure file such as `collector.md`, which covers one step). It also carries the boundary against **aws-observability**. A question about one concept's rules — a Space's Region relative to its Domain, the roles `create-space` needs — belongs to the procedure file for that concept (`spaces-and-domains.md`), which the routing table names. 2. If the request is about **instrumenting an AI agent project** (adding OTel to a framework such as LangChain, LangGraph, Strands, CrewAI, OpenAI Agents, or Vercel AI), it is Omni — route to `references/cloudwatch-omni/omni-agents-instrumentation/omni-agents-instrumentation.md`. There is no Application Signals path for an agent framework. 3. If the request is about **instrumenting a general application** (a service on EC2, ECS, EKS, or Lambda in Python, Node.js, Java, or .NET — not an AI-agent framework), **decide the product before opening any procedure file** — the two paths are mutually exclusive on a workload: - Names **Application Signals**, ServiceEvents, a "monitored service", the service map, the `amazon-cloudwatch-observability` add-on, the CloudWatch Agent, or port 4316 → the CloudWatch path: `references/cloudwatch/application-signals-onboarding.md`, then the matching `references/cloudwatch/appsignals-guides/-.md`. - Names **Omni**, Application Observability, a Space, a Domain, or a Dataset — or explicitly rules Application Signals out ("plain OTel", "no add-on", "no agent sidecar") → the Omni path: `references/cloudwatch-omni/instrumentation/instrumentation.md`. - Names **neither** → rule 4. 4. If an **instrumentation, ADOT, or collector** request names neither product (neither Omni or a Space, nor Application Signals, ServiceEvents, or the `amazon-cloudwatch-observability` add-on), probe for a Space in the target Region before choosing. `list-spaces` is **account-global**: the `--region` flag only selects the endpoint, and the response lists every Space in the account, each with its own `region`, so filter to the target Region rather than trusting a non-empty list: `aws cloudwatchomni list-spaces --region --query "items[?region=='']"`. A Space in that Region → the Omni path, `references/cloudwatch-omni/instrumentation/instrumentation.md`. An empty filtered list → the customer has not adopted Omni in that Region, so Application Signals is the live product for them: `references/cloudwatch/application-signals-onboarding.md`. **Either way the request stays in this skill** — the probe picks the folder, not the skill. If the probe errors with an unknown service, that is the CLI model, not evidence Omni is absent — fall back to the customer's wording and ask only if still inconclusive. 5. If the request is about **Space/Domain creation or configuration**, route to the matching setup reference. An authorization error that appears the moment a *just-created* Space is used is part of this — it is a space access role trust-policy problem (`create-space` never verifies the role can be assumed), so route to `references/cloudwatch-omni/spaces-and-domains.md` → "If the Space was created but cannot be used", not to the access-grants reference. 6. If the user asks about getting telemetry **to** a destination — telemetry not yet reaching CloudWatch — **deploy an OTel Collector** on EC2/ECS/EKS so an instrumented workload has somewhere to send OTLP, and wire the app's `OTEL_EXPORTER_OTLP_ENDPOINT` to it: `references/cloudwatch-omni/instrumentation/collector.md`. The collector exports straight to CloudWatch's own per-signal OTLP endpoints. If instead the telemetry is **already in CloudWatch log groups** and needs forwarding into the Dataset, route to `references/cloudwatch-omni/data-forwarding-and-centralization.md`. Both the collector's OTLP export and CloudWatch's OTLP endpoints authenticate with **SigV4** (AWS credentials), including OIDC-federated credentials for a workload outside AWS (see `references/cloudwatch-omni/azure-ingestion/custom-telemetry.md`). A purely **token-based** ingestion path for a sender that cannot obtain AWS credentials at all is **not covered by these skills**; do not improvise one, and say so plainly. 7. If the request is about **getting Azure telemetry into CloudWatch**, route to `references/cloudwatch-omni/azure-ingestion/azure-ingestion.md` and decide by intent. If the customer wants the telemetry their own application produces (the logs/metrics/traces from their code), that is supported via the CloudWatch agent on an Azure VM or AKS cluster — follow the reference. If they want telemetry their Azure resources emit on their own (the Azure equivalent of AWS VPC flow logs / Route 53 logs), that is not available — say so and do not attempt a setup. 8. If the user asks to **connect, enable, or authorize Slack** for a Space for the first time (including granting the operator permission to use it), route to `references/cloudwatch-omni/slack-integration.md`. Using Slack after it is connected (posting findings to a channel, mentioning the assistant, or searching Slack) happens in Slack and the console and is not covered by these skills; Slack as an **alert notification target** is covered by **aws-observability**'s Omni alerts reference. 9. **Application Signals, ServiceEvents, and Dynamic Instrumentation are CloudWatch features, and they split by setup versus use — not by skill.** *Enabling* them on a service that does not have them yet is setup and belongs here, under `references/cloudwatch/`: Application Signals onboarding and ServiceEvents in `references/cloudwatch/application-signals-onboarding.md` (with the CI/CD metadata chain in `references/cloudwatch/application-signals-cicd-metadata.md`), and switching Dynamic Instrumentation on at instrumentation time in that same file's Step 5d. They must never appear in an **Omni** instrumentation change — for that they remain out of scope, and rule 3 is how you tell the two paths apart. *Using* them once the service reports — reading the service map, alarming on Application Signals metrics, or placing breakpoints and reading snapshots with Dynamic Instrumentation — is day-to-day work: STOP and route to **aws-observability**. 10. If the user already has a working Space, or a service already reporting to Application Signals, and is asking about **queries/dashboards/alarms/alerts/investigation**, STOP and route to **aws-observability**. 11. If the user asks to **register their own MCP tool server**, connect a custom MCP server, or set up MCP authentication: **CloudWatch Omni does not let customers add a custom MCP server.** There is no console screen for it and the integration type is not enabled for customer use, so the attempt is refused. Say so plainly and stop; do not describe a setup flow, and do not offer a substitute. See `references/cloudwatch-omni/custom-mcp-integration.md`. 12. If the user asks to **connect GitHub**, link a GitHub organization or repositories, or enrich the application map from source code: **CloudWatch Omni has no GitHub integration.** Say so plainly and stop — do not walk through a setup. Do not offer a substitute: the Application Signals GitHub Action and the CloudWatch Logs GitHub audit-log source are separate CloudWatch features, not a way to connect GitHub to Omni, and a custom MCP server is not a GitHub integration. A GitHub-related feature found in public documentation is not evidence of an Omni GitHub integration. 13. If the request names **neither product** — "add observability to my API", "how should I begin monitoring my workload", with no mention of Omni, a Space, a Domain, or of a CloudWatch feature — do not assume either one. Which product a customer wants to *set up* is their intent, not something the account's current state tells you, so ask, framed product versus product: **CloudWatch Omni** (Application/Agent Observability — Domains, Spaces, grants, a Dataset) or **CloudWatch** (log groups, alarms, Log Insights, Application Signals). Both answers are served here, from different reference folders, so this is a question about which product to set up — **not** about which skill to use, and not a reason to hand the request to **aws-observability**. In the same message, explain the boundary so they can choose — it is **setup versus use**: creating or configuring a Domain, Space, grant, or Access Profile, getting telemetry flowing for the first time, instrumenting an application or agent for Omni, and onboarding a service to Application Signals are all first-time setup and belong here; querying telemetry, building dashboards, configuring alarms or Omni alerts, investigating a live problem, and debugging with Dynamic Instrumentation are day-to-day use and belong to **aws-observability**. Do not start walking through Domain/Space creation or Application Signals onboarding until they confirm which product; having asked, stop and wait.