--- name: rocky-config description: Canonical `rocky.toml` authoring reference. Use when writing or reviewing a Rocky pipeline config — covers the 4 pipeline types (replication, transformation, quality, snapshot), adapter variants (duckdb/databricks/snowflake/fivetran), minimal-config defaults, env-var substitution, governance, checks, hooks, and the ${VAR:-default} syntax. --- # rocky.toml authoring reference Rocky reads **one** config file — `rocky.toml` — for everything: adapters, pipelines, governance, state backend, cache. The Rust source of truth is `engine/crates/rocky-core/src/config.rs`. This skill is the canonical authoring reference. ## When to use this skill - Writing a new `rocky.toml` from scratch (POC, example, production) - Migrating an old config to current shape (watch for `[source]` / `[warehouse]` — those are pre-Phase-2 and no longer work) - Deciding which pipeline type to use (replication vs transformation vs quality vs snapshot) - Looking up an adapter's required fields (databricks needs host + http_path, snowflake needs auth variants, etc.) - Adding governance, hooks, state sync, or checks to an existing config ## Top-level structure Two mandatory sections (`[adapter]` + at least one `[pipeline.]`) plus optional global blocks: ```toml [adapter] # Warehouse / source connection # … [pipeline.] # One or more pipelines — discriminated by `type` (default: replication) # … # Optional globals: [state] # Embedded state store backend [cache.schemas] # Schema (DESCRIBE) cache: enabled, ttl_seconds, replicate, trusted_max_age_seconds, strict_sources [cost] # Cost model for `rocky optimize` [hook.] # Lifecycle hooks (one per event) # Governance (tags, grants, workspace bindings) is NOT a top-level table: # it lives under [pipeline..target.governance]. A top-level # [governance] is refused (deny_unknown_fields). ``` ## Env var substitution Every string value supports `${VAR_NAME}` at parse time, with optional default: ```toml token = "${DATABRICKS_TOKEN}" # hard-required namespace = "${ROCKY_NAMESPACE:-default}" # default when unset workspace = "${WORKSPACE_IDS:-}" # default to empty string ``` Substitution happens in `rocky-core/src/config.rs` before serde sees the value. ## Adapter variants An unnamed `[adapter]` with a `type` key auto-wraps as `adapter.default`. Pipeline adapter refs default to `"default"` — so you can omit `adapter = "default"` lines everywhere. Top-level adapter fields are strict (`deny_unknown_fields` — typos are parse errors). Adapter-specific keys Rocky doesn't model go under a nested `[adapter..extra]` table, which passes through untouched for a custom or process adapter to read. ### The `kind` field `kind` declares the role an `[adapter.*]` block plays. Two valid values: `"data"` (warehouse read/write) and `"discovery"` (metadata enumeration). | Adapter type | `kind` rule | |---|---| | `databricks`, `snowflake`, `bigquery`, `postgres`, `redshift`, `clickhouse`, `sqlserver` | Optional — defaults to `"data"`. Setting `"discovery"` is a parse error. | | `fivetran`, `airbyte`, `iceberg`, `manual` | **Required — must be `"discovery"`.** Omitting it is a parse error: these adapters have no data path. | | `duckdb` | Optional — absent means "register both roles" (the common DuckDB case). Setting `"data"` or `"discovery"` narrows to a single role. | Requiring `kind` on discovery-only adapter types is deliberate: a reader should be able to tell from the raw config alone that `[adapter.fivetran]` is metadata-only, without knowing the Rust trait surface. ### DuckDB (no credentials) ```toml [adapter] type = "duckdb" path = "playground.duckdb" # omit for in-memory; required if also used for discovery ``` ### Databricks (PAT first, OAuth M2M fallback) ```toml [adapter] type = "databricks" host = "${DATABRICKS_HOST}" # e.g. dbc-xxxx.cloud.databricks.com (no https://) http_path = "${DATABRICKS_HTTP_PATH}" # /sql/1.0/warehouses/ [adapter.auth] token = "${DATABRICKS_TOKEN}" # PAT (tried first) # client_id = "${DATABRICKS_CLIENT_ID}" # OAuth M2M (fallback) # client_secret = "${DATABRICKS_CLIENT_SECRET}" ``` ### Snowflake (OAuth > RS256 key-pair > password priority) ```toml [adapter] type = "snowflake" account = "${SNOWFLAKE_ACCOUNT}" username = "${SNOWFLAKE_USER}" [adapter.auth] # OAuth (pre-supplied token, highest priority): # token = "${SNOWFLAKE_OAUTH_TOKEN}" # RS256 key-pair JWT (preferred for service principals): private_key_path = "${SNOWFLAKE_KEY_PATH}" # Password (lowest priority): # password = "${SNOWFLAKE_PASSWORD}" ``` ### PostgreSQL / Redshift (password auth, libpq-style `sslmode`) ```toml [adapter] type = "postgres" # or "redshift" (beta) host = "${PGHOST}" # host or host:port (default 5432 / 5439) database = "analytics" # the connected database = the only valid catalog username = "${PGUSER}" password = "${PGPASSWORD}" # timeout_secs = 300 # connect timeout + statement_timeout [adapter.extra] # unknown keys are refused sslmode = "verify-full" # disable | prefer (default) | require | verify-full # sslrootcert = "/etc/ssl/rds-ca.pem" # extra PEM roots under verify-full # port = 6543 # wins over host:port # max_connections = 8 # merge_mode = "on_conflict" # postgres < 15 only; needs a unique index on unique_key # late_binding_views = true # redshift only: views WITH NO SCHEMA BINDING ``` Redshift models can set table attributes in the sidecar `[redshift]` block (`dist_style`, `dist_key`, `sort_key`, `sort_style`); `rocky compile` validates it (E052 / W052) and any other adapter refuses it at SQL generation. IAM auth for Redshift is not built in yet. ### ClickHouse (beta; HTTP interface, user/password, TLS) ```toml [adapter] type = "clickhouse" host = "${CLICKHOUSE_HOST}" # host or host:port (default 8123, 8443 with secure) username = "${CLICKHOUSE_USER}" # default "default" password = "${CLICKHOUSE_PASSWORD}" # database = "default" # session default database for unqualified names # timeout_secs = 300 # per statement; also max_execution_time [adapter.extra] # unknown keys are refused secure = true # HTTPS, certificate always verified # ca_cert = "/etc/ssl/ch-ca.pem" # extra PEM roots (needs secure = true) # port = 9443 # wins over host:port ``` A Rocky schema is a ClickHouse database: set every model's `catalog = ""` (a non-empty catalog is refused). Models can set `[clickhouse]` (`engine` = a parameterless MergeTree-family name, `order_by` = columns, `partition_by` = a column or `fn(column)`); `rocky compile` validates it (E053 / W053). `merge` and `incremental` with `unique_key` are refused (E053, no `MERGE`), as are snapshots (E049), UDFs (E051) and `materialized_view`. ### SQL Server / Azure SQL / Fabric Warehouse (beta) ```toml [adapter] type = "sqlserver" host = "${MSSQL_HOST}" # host, host,port or host:port (default 1433) database = "analytics" # the connected database = the only valid catalog # exactly ONE auth method: username = "${MSSQL_USER}" # SQL auth (not on Fabric) password = "${MSSQL_PASSWORD}" # oauth_token = "${MSSQL_ACCESS_TOKEN}" # Entra ID access token (not refreshed) # client_id = "${AZURE_CLIENT_ID}" # Entra ID service principal, with extra.tenant_id # client_secret = "${AZURE_CLIENT_SECRET}" [adapter.extra] # unknown keys are refused # encrypt = "mandatory" # mandatory (default) | strict (TDS 8.0) | optional # trust_server_certificate = true # local/test servers only # ca_cert = "/etc/ssl/corp-ca.pem" # flavor = "fabric" # Fabric Warehouse renderings # tenant_id = "${AZURE_TENANT_ID}" # port = 1433 # max_connections = 8 ``` Snapshots, `materialized_view`, UDFs (E051) and `regex_match` checks are refused on SQL Server. ### Fivetran (discovery-only) ```toml [adapter.fivetran] type = "fivetran" kind = "discovery" # required: fivetran has no data path destination_id = "${FIVETRAN_DESTINATION_ID}" api_key = "${FIVETRAN_API_KEY}" api_secret = "${FIVETRAN_API_SECRET}" ``` Use this block in `pipeline.*.source.discovery.adapter` to let Rocky query the Fivetran REST API for the list of schemas to sync. Data itself flows through whichever warehouse adapter is referenced by `pipeline.*.source.adapter` (usually Databricks or Snowflake — the destination Fivetran populates). ### Named adapters (multi-adapter configs) ```toml [adapter.warehouse] type = "databricks" # … [adapter.source] type = "fivetran" kind = "discovery" # … [pipeline.raw] [pipeline.raw.source] adapter = "source" # ref by name [pipeline.raw.target] adapter = "warehouse" ``` ## Pipeline types Pipelines are discriminated by `type`. Default is `"replication"`, so a pipeline block with no `type` is a replication pipeline. | `type` | What it does | Canonical use | |---|---|---| | `replication` (default) | Copies tables from source to target with schema-pattern discovery + incremental/full-refresh strategy | Raw/Bronze layer | | `transformation` | Runs user SQL models against the warehouse and materializes them | Silver/Gold layer | | `quality` | Runs data quality checks against existing tables | Standalone QA runs | | `snapshot` | SCD-Type-2 snapshots with history tracking | Slowly-changing dimensions | All four share `execution`, `checks`, `depends_on`, and `target.adapter`. They differ in what `source`/`target` shapes look like. ### Replication pipeline ```toml [pipeline.raw] # type = "replication" # optional, this is the default strategy = "incremental" # or "full_refresh" or "merge" timestamp_column = "_fivetran_synced" # merge_keys = ["id"] # required when strategy = "merge" # merge_keys_fallback = ["id"] # used when merge_keys is unset metadata_columns = [ { name = "_loaded_by", type = "STRING", value = "'rocky'" }, ] [pipeline.raw.source] # adapter = "default" # optional — first adapter by default [pipeline.raw.source.schema_pattern] prefix = "src__" separator = "__" components = ["client", "regions...", "connector"] # "name" = single segment # "name..." = variable-length (1+) # Reserved: `table` and `id` are not allowed as component names # (they're consumed by `--filter` and `[[table_overrides]]`). [pipeline.raw.target] catalog_template = "{client}_warehouse" schema_template = "staging__{regions}__{connector}" # Per-`(connector, table)` overrides on top of the pipeline defaults. # Per-field most-specific-match-wins: each field is supplied by the # most-specific matching rule that sets it; unrelated rules don't # clobber it. See `engine/crates/rocky-core/src/config.rs` :: # `TableOverride` for the full surface. # Skip diagnostic tables on every connector. [[pipeline.raw.table_overrides]] match.table = "_diagnostics_*" # `*`/`?` glob in TOML only — CLI literals enabled = false # `pii_users` on one specific connector needs a composite merge key. [[pipeline.raw.table_overrides]] match.connector = "stripe_main" # matches conn.id OR conn.schema match.table = "pii_users" merge_keys = ["user_id", "tenant_id"] # `audit_log` is append-only — keep it on incremental even when # the pipeline default is `merge`. [[pipeline.raw.table_overrides]] match.connector = "stripe_main" match.table = "audit_log" strategy = "incremental" timestamp_column = "occurred_at" ``` **CLI symmetry:** `--filter table=` filters to one table at a time (no globs at the CLI; shell quoting is too fragile). The connector-level `--filter id=` keys remain unchanged. ### Transformation pipeline ```toml [pipeline.silver] type = "transformation" models = "models/silver/**" # glob, relative to rocky.toml; default "models/**" depends_on = ["raw"] # There is no contracts key: a model's contract is auto-discovered as the # sibling `.contract.toml`, or passed at the CLI via `--contracts `. # A transformation target takes an adapter ref (plus optional governance) # and nothing else — `catalog`/`schema` here are parse errors # (`deny_unknown_fields`). Each model names its own destination in its # sidecar `[target]`: # # # models/silver/dim_customers.toml # [target] # catalog = "analytics" # schema = "marts" # table = "dim_customers" [pipeline.silver.target] adapter = "warehouse" ``` Declare the external tables the models read, with optional dbt-style freshness checked by `rocky freshness` (exit 1 on `error` / `runtime_error`): ```toml [[pipeline.silver.sources]] schema = "raw" # catalog optional (two-part names) table = "orders" [pipeline.silver.sources.freshness] loaded_at_field = "_loaded_at" # DATE / TIMESTAMP column warn_after = "12h" # s | h | d; at least one of the two error_after = "24h" # must be >= warn_after (else E050) filter = "status <> 'test'" # optional WHERE predicate; no `;` ``` ### Quality pipeline ```toml [pipeline.qa] type = "quality" depends_on = ["silver"] # A quality target is an adapter ref only. The tables to check are listed # in `[[pipeline.qa.tables]]` (omit `table` to check the whole schema). [pipeline.qa.target] adapter = "warehouse" [[pipeline.qa.tables]] catalog = "analytics" schema = "marts" # Quality runs execute `row_count`, `custom`, and `assertions` only. Other # check kinds are inert here — `rocky validate` warns (V034). [pipeline.qa.checks] row_count = true ``` ### Snapshot pipeline ```toml [pipeline.dim_history] type = "snapshot" unique_key = ["customer_id"] # required: row identity in the source table updated_at = "updated_at" # required: change-detection column depends_on = ["silver"] [pipeline.dim_history.source] catalog = "analytics" schema = "marts" table = "dim_customers" [pipeline.dim_history.target] catalog = "analytics_history" schema = "snapshots" table = "dim_customers_history" # required: explicit history table ``` ## Minimal-config defaults (omit these) The parser applies sane defaults — keep configs lean by omitting anything that matches the default: | Field | Default | Omit unless | |---|---|---| | `pipeline.type` | `"replication"` | You need transformation/quality/snapshot | | `adapter = "default"` (in pipeline source/target) | First adapter | Multi-adapter config | | `[state]\nbackend = "local"` | local (embedded redb) | Using S3, Valkey, or tiered state sync | | `auto_create_catalogs = false` | false | You want Rocky to CREATE CATALOG | | `auto_create_schemas = false` | false | You want Rocky to CREATE SCHEMA | | Model `name` in sidecar `.toml` | filename stem | You want a different logical name | | Model `target.table` | `name` | Renaming on write | | Directory-level `target` | `models/_defaults.toml` inherited | Overriding per-model | ## Checks Every pipeline type accepts a `[checks]` block, but execution differs: replication runs the full set; quality runs `row_count`, `custom`, and `assertions` only; transformation, snapshot, and load run no pipeline-level checks today (`executed_check_kinds` in `config.rs` is the source of truth). `rocky validate` warns (V034) on a check the pipeline type never executes. ```toml [pipeline..checks] enabled = true row_count = true # source vs target row count column_match = true # source vs target column list freshness = { threshold_seconds = 86400 } # max staleness of newest row null_rate = { columns = ["email"], threshold = 0.05, sample_percent = 10 } [[pipeline..checks.custom]] name = "no_future_dates" sql = "SELECT COUNT(*) FROM {target} WHERE created_at > CURRENT_TIMESTAMP()" threshold = 0 # max failing rows ``` ## Contracts Enforced on `load` pipelines: the load runs into a staging table, validates the landed schema against the contract, and promotes to the target only on pass — a violation drops staging and fails, so non-conforming data never lands. Types compare in Rocky's normalized vocabulary, so one contract ports across warehouses. ```toml [pipeline..contract] required_columns = [ { name = "id", type = "BIGINT", nullable = false }, ] protected_columns = ["id", "email"] # may not be removed from source allowed_type_changes = [ { from = "INT", to = "BIGINT" }, # widening allowlist ] ``` ## Execution & adaptive concurrency ```toml [pipeline..execution] concurrency = 16 fail_fast = false error_rate_abort_pct = 50 table_retries = 1 # adaptive_concurrency is planned but not yet a config field. # The AIMD throttle primitive exists but isn't wired. ``` ## Governance (Databricks Unity Catalog) Governance is configured **per pipeline target**, under `[pipeline..target.governance]`. There is no top-level `[governance]` table; one is refused at load (`deny_unknown_fields` on `RockyConfig`). ```toml [pipeline.bronze.target.governance] auto_create_catalogs = true auto_create_schemas = true [pipeline.bronze.target.governance.tags] managed_by = "rocky" [[pipeline.bronze.target.governance.grants]] principal = "data-readers" permissions = ["BROWSE", "USE CATALOG", "USE SCHEMA", "SELECT"] [pipeline.bronze.target.governance.isolation] enabled = true workspace_ids = "${WORKSPACE_IDS:-}" ``` Permissions handled: BROWSE, USE CATALOG, USE SCHEMA, SELECT, MANAGE, MODIFY. Skipped: OWNERSHIP, ALL PRIVILEGES, CREATE SCHEMA (non-managed). ## State backend ```toml # Embedded redb (default — no config needed) [state] backend = "local" # S3-backed state sync [state] backend = "s3" bucket = "my-rocky-state" prefix = "prod/" region = "us-east-1" # Tiered: local redb + S3 for durability [state] backend = "tiered" # … S3 fields plus local path # State-file namespacing (opt-in, default "none") [state] backend = "local" namespacing = "pipeline" # each pipeline gets its own .rocky-state/.redb ``` `namespacing` (`StateNamespacing`, default `"none"`) controls state-file fan-out. redb allows one writer per file, so running one `rocky run` per pipeline/client against the single global `/.rocky-state.redb` serializes them on one lock. `"pipeline"` gives each pipeline its own `/.rocky-state/.redb`. `"none"` is byte-identical to omitting the key. For per-client fan-out use the per-invocation `--state-namespace ` flag (overrides this config); an explicit `--state-path` disables namespacing for that run. ## Run-skip gate (`[run]` + per-model `[skip]`) ```toml # Opt-in model-skip gate — default OFF (omit the block to keep old behavior) [run] skip_unchanged = true # master switch (also via the --skip-unchanged flag) skip_rowcount_fallback = false # default; allow COUNT(*) when no timestamp column (weaker signal) lag_tolerance_seconds = 0 # default; any MAX(ts) movement forces a rebuild ``` `[run]` (`RunConfig`) tunes the `--skip-unchanged` gate: skip re-materializing a transformation model whose logic **and** every upstream's data both appear unchanged since the last successful build. It is a **best-effort optimization, not a result-equivalence guarantee** — every field defaults to no-skip, and any missing/unreadable/ambiguous input rebuilds (fail-safe). **Not skip-eligible (always rebuild):** - Non-deterministic SQL: `CURRENT_TIMESTAMP` / `NOW()`, `RANDOM()`, `UUID()`, `CURRENT_USER`, `CURRENT_CATALOG`, `ANY_VALUE`, `ARRAY_AGG`, unordered `LIMIT`, or any unknown function. - Models whose lineage isn't provably complete: CTEs, subqueries (`FROM (…)`, `IN (SELECT …)`, `EXISTS`, scalar sub-selects), `PIVOT` / `UNNEST` / nested-join table-factors, and set operations (`UNION` / `INTERSECT` / `EXCEPT`). - `content_addressed` / `time_interval` strategies (`full_refresh` **is** eligible). `--force-rebuild` bypasses the gate entirely. ```toml # Per-model override (in the model's .toml sidecar) name = "fct_orders" [skip] eligible = true # Some(false) = always build; Some(true) = eligible; unset = auto deterministic = true # owner asserts SQL is pure → re-eligible despite the static scan ``` `[skip]` (`SkipConfig` in `rocky-core/src/models.rs`) is a per-model sidecar block. Both fields are `Option` — unset means "trust the automatic rules." `--defer` / `--defer-to ` are runtime-only dev flags (not config): build the `--model`-selected models locally and resolve unbuilt upstream `ref()`s to a prod/defer schema. The rewrite parses model SQL with the Databricks dialect, so `SELECT * EXCEPT(...)`, trailing-comma selects, and `STRUCT(...)` literals can't be rewritten — run those without `--defer`. ## Reuse (`[reuse]`) — EXPERIMENTAL, do not enable in production ```toml [reuse] enabled = false # default; preview-only, NOT live-verified — leave off in production ``` `[reuse]` (`ReuseConfig`) is a **preview** surface scoped to the **Databricks–Iceberg content-addressed write path only** (no DuckDB / Snowflake / BigQuery), and it is **not yet live-verified against a warehouse**. When `enabled = true`, a successful run only *populates* an input-match index + provenance record; the reuse decision path only ever resolves to **BUILD** today (an ONLY-BUILD posture) — a fail-closed verdict is computed but nothing is reused, since the live point-to reuse is not yet wired/verified. It attests an input-logic match + byte-identity of the **recorded** bytes, never that a fresh re-run would reproduce them. Default-off keeps `rocky run` byte- and cost-identical. The per-invocation `--no-reuse` flag forces every model to build. Provenance is auditable per `docs/.../guides/verify-a-run.md`. ## Cache ```toml [cache.schemas] enabled = true # default; false for strict CI (every typecheck hits the warehouse) ttl_seconds = 86400 # default 24h; lower for high-DDL-churn teams replicate = false # default; true to ship the cache through state_sync # trusted_max_age_seconds = 3600 # unset by default; entries younger than this make a missing source column E041 # strict_sources = true # default false; every W041 (seed / old cache entry) becomes E041 ``` `[cache.schemas]` is the only `[cache]` table (`CacheConfig` in `config.rs`): it stores `DESCRIBE TABLE` results in `state.redb` so leaf models typecheck against real warehouse types without a live round-trip on every compile. There is no `[cache] valkey_url` key: `ValkeyCacheConfig` exists as a type but is not wired into `RockyConfig`, and a `[cache]` table with any other key is refused (`deny_unknown_fields`). The Valkey tier is a `[state]` backend, not a cache setting. ## Cost model (for `rocky optimize`) ```toml [cost] storage_cost_per_gb_month = 0.023 compute_cost_per_dbu = 0.40 warehouse_size = "Medium" min_history_runs = 5 ``` ## Hooks ```toml [hook.on_pipeline_start] command = "scripts/notify.sh" timeout_ms = 5000 on_failure = "warn" # default; or "abort" (stop the pipeline) / "ignore" (silent) [hook.on_pipeline_fail] url = "${SLACK_WEBHOOK_URL}" # webhook instead of command template = "default" # built-in preset ``` Events: `on_pipeline_start`, `on_pipeline_complete`, `on_pipeline_fail`, `on_model_start`, `on_model_complete`, `on_model_fail`, `on_check_fail`, `on_drift_detected`. ## Model sidecar files (`.sql` + `.toml`) Model files live under `models/` and use a sidecar pattern: ```toml # models/marts/dim_customers.toml name = "dim_customers" depends_on = ["stg_customers"] [strategy] type = "merge" # e.g. full_refresh, merge, delete_insert, time_interval, view, snapshot; not incremental (E037) unique_key = ["customer_id"] # update_columns = ["name", "email"] # omit for UPDATE SET * [target] catalog = "analytics" schema = "marts" table = "dim_customers" [[sources]] # optional: declare source tables catalog = "analytics" schema = "staging" table = "customers" ``` ```sql -- models/marts/dim_customers.sql -- Pure SQL. No Jinja. No templating. SELECT customer_id, name, email, updated_at FROM {{ analytics.staging.customers }} -- Rocky expands refs at compile time, not via Jinja ``` Directory-level defaults via `models//_defaults.toml`: ```toml [target] catalog = "analytics" schema = "marts" ``` ## Validation SQL identifiers (catalog, schema, table, tenant, region, source names) must match `^[a-zA-Z0-9_]+$`. Rocky rejects anything else. Principal names for GRANT/REVOKE allow the broader `^[a-zA-Z0-9_ \-\.@]+$` pattern and are always wrapped in backticks in generated SQL. ## Time-interval (partition-keyed) materialization For transformation models that need partition-by-partition execution: ```toml # models/fact_events.toml [strategy] type = "time_interval" time_column = "event_date" granularity = "day" # hour | day | month | year lookback = 7 # re-run last 7 partitions batch_size = 4 # partitions per concurrent batch first_partition = "2024-01-01" ``` Model SQL uses `@start_date` and `@end_date` placeholders — Rocky substitutes per partition: ```sql SELECT * FROM raw.events WHERE event_date >= @start_date AND event_date < @end_date ``` CLI: `rocky run --partition KEY` / `--from KEY --to KEY` / `--latest` / `--missing` / `--lookback N` / `--parallel N`. ## Full canonical example See `examples/playground/pocs/00-foundations/01-replication-basics/rocky.toml` for the minimal DuckDB **replication** case (schema-pattern routing), `examples/playground/pocs/00-foundations/00-playground-default/rocky.toml` for the minimal **transformation** case (model DAG), or `engine/examples/multi-layer/rocky.toml` for a full Bronze/Silver/Gold setup. ## What NOT to write (pre-Phase-2 legacy) The following keys **do not work** anymore — they were the pre-Phase-2 config shape and will be rejected by the parser: | ❌ Legacy | ✅ Current | |---|---| | `[source]` top-level | `[pipeline..source]` | | `[warehouse]` | `[adapter]` | | `[replication]` top-level | `[pipeline.]` with `strategy = "incremental"` | | `[checks]` top-level | `[pipeline..checks]` | | `[target]` top-level | `[pipeline..target]` | If you see any of these in a config you're editing, the config is stale — migrate it. ## Reference - `engine/crates/rocky-core/src/config.rs` — Rust source of truth for every field - `engine/AGENTS.md` — "Configuration" section with full annotated example - `editors/vscode/schemas/rocky-config.schema.json` — JSON Schema for IDE autocompletion (autogenerated; keep in sync with `config.rs`) - `examples/playground/AGENTS.md` — POC-specific minimal-config idioms