# KYMA v3 Symbolic Composition Probe — Preregistration Date: 2026-09-29 Status: **frozen before any model is trained.** The freeze is the commit that adds this file and pushes it to the public remote; the code it names (`src/scpn_quantum_control/benchmarks/kyma_v3/`, `scripts/run_kyma_v3_probe.py`) is in the same commit. No substrate or baseline has been trained on this task. The result will be appended below a `RESULT` heading; nothing above it changes after training. 0 QPU. This is a classical oscillator-substrate probe. ## Scientific question Does the gated-coupling oscillator substrate still generalise to a held-out combination of learned operations when the ground truth is produced by a symbolic program rather than by oscillator dynamics? (KYMA Part B assumption A2, MS1 criterion family.) v2 (`docs/campaigns/kyma_v2_composition_probe_2026-07-21.md`, PASS) used an oscillator teacher, so its task lived inside the substrate's hypothesis class. v3 removes that objection: no oscillator, integrator or model is in the label path. Authority: the KYMA v3 probe specification (seven requirements); owner decision A9 (probe GO; a negative is reported with its diagnosis) and owner decision B3 of 2026-09-29 (option (a): evaluate the held-out pair on query `a` only and state that restriction here). ## Ground truth and split Three registers `(a, b, c)` with values in `Z4`. Operations: - `R0`: `(a, b, c) → (b, c, a)`; - `R1`: `a ← (a + b) mod 4`; - `R2`: `b ← (b + 1) mod 4`. A configuration is one operation or an ordered pair applied left to right (3 singles + 9 ordered pairs). An item is (initial state, configuration, queried register); its label is the final value of the queried register. - **Training:** every configuration except the held-out pair, all 64 states, all three queries: 11 × 64 × 3 = **2,112 items**. - **Held-out test:** the ordered pair **`(R0, R1)`**, all 64 states, **query `a` only: 64 items per seed.** The pair's queries `b` and `c` are neither trained nor evaluated. - **Restriction (owner decision B3):** query `a` is the only query of the pair whose answer is not a function of the single-operation answers. For queries `b` and `c` the held-out answer equals the `R0`-alone answer, so they cannot test composition and are excluded. Every query register appears in every trained configuration. ## Teacher-free design checks (recorded before training) Computed by `kyma_v3.task.design_report()` and `kyma_v3.probe.realisability_accuracy()` from the symbolic program and hand-set gates only: | check | value | |---|---| | every configuration is a bijection on `Z4^3` | yes | | label counts per (configuration, query) | exactly 16 per class (all 36) | | held-out states whose answer the single-op answers do not fix: query a / b / c | **1.00** / 0.00 / 0.00 | | minimum Hamming distance of the held-out answer vector (`b + c`) to any trained (configuration, query) answer vector | **48 of 64** | | measured training-marginal chance floor on the test items | **0.25** | | hand-set substrate, all 12 configurations × 64 states × 3 queries | **2,304 / 2,304 exact**; worst phase error 0.026 rad against a 0.785 rad margin | The last row shows the task is realisable inside the substrate class. It is a validity check only: training never sees the hand-set gates. ## Substrate (frozen) Each register is one oscillator; its phase relative to a fixed reference (phase 0) encodes the value, `v ↦ v·π/2`. Two banks of three oscillators. One operation is one **write stage**: the destination bank is reset to the off-lattice phase `π/4`, the source bank receives no coupling (it is held), and the destination integrates dθ_i/dt = Σ_j K[o,i,j] sin(x_j − θ_i + α[o,i,j]) + Σ_p T[o,i,p] sin(x_{p1} + x_{p2} − θ_i + β[o,i,p]) with directed pairwise couplings `K` and lags `α` (source register `j` → target `i`), and triadic couplings `T` and lags `β` over the three source-register pairs `p ∈ {(a,b), (a,c), (b,c)}`. The operation code `o` gates which couplings act. The banks then swap roles; an ordered pair is two write stages with the gates of each operation in program order. Readout: the queried register's final phase, rounded to the nearest lattice value. - Integrator: fixed-step RK4, `dt = 0.05`, **60 steps per stage** (`T = 3`). - Trainable parameters: `K, α, T, β` for 3 operations × 3 targets × 3 sources or pairs = **108**. - Initialisation: every parameter drawn from `N(0, 0.3²)` with the run seed. - Loss: mean `1 − cos(φ_q − label·π/2)` of the queried register's final phase. Only the final queried value is supervised. **Stated limitation.** Two architectural choices were made for realisability, not from any model's performance: the staged schedule (one write stage per operation) and the triadic term (pairwise phase coupling cannot add phases, which `R1` requires). Because the stages are applied in program order, a substrate that learns each operation exactly composes the held-out pair by construction. The probe therefore tests whether gradient descent learns reusable operation gates from end-of-program labels alone, and the diagnostic baselines below test whether staging alone, without oscillator dynamics, does as well. ## Baselines (frozen) Contract baselines, parameter count within ±10 % of the substrate (108): | baseline | architecture | width | parameters | |---|---|---|---| | MLP | one tanh hidden layer over `sin/cos` of the three register phases, one-hot first op, one-hot second op (with "none") and one-hot query (16 inputs) | 5 | 109 (+0.9 %) | | staged GNN | register nodes; per-operation learned 3×3 adjacency and node bias; embedding, then one message-passing round per operation in program order; read at the queried node | 4 | 103 (−4.6 %) | | transformer | one single-head attention layer with residual over tokens `[first op, second op, query, a, b, c]`; learned embeddings and positions; read at the query token | 3 | 100 (−7.4 %) | Diagnostic baselines (reported; they never change the verdict): | diagnostic | purpose | width | parameters | |---|---|---|---| | sequential MLP | one small MLP per operation maps `sin/cos` phases to `sin/cos` phases, applied in program order; label = queried angle rounded to the lattice; loss as the substrate's | 2 | 96 (−11.1 %, outside ±10 %; diagnostic only) | | MLP, large | capacity control | 64 | 1,348 | | staged GNN, large | capacity control | 16 | 703 | | transformer, large | capacity control | 16 | 1,348 | Chance floor: always predicting the most frequent training label (ties to the smallest), measured on the test items (0.25 by the design check). No baseline receives intermediate states or any privileged information. ## Training (frozen) Full-batch Adam on the 2,112 training items for every model; **3,000 epochs** for every model; learning rates fixed in advance without search: substrate 0.05, MLP 0.02, staged GNN 0.01, transformer 0.01, sequential MLP 0.02. Seeds **0, 1, 2, 3, 4**. No early stopping, no model selection and no tuning on the held-out items. Training accuracy is reported for every model and seed. ## Decision procedure (the contract) Primary statistic: held-out accuracy (fraction of the 64 test items classified correctly), per model and seed, reported as mean ± population standard deviation over the five seeds. **PASS** iff both hold: 1. substrate mean ≥ (best contract-baseline mean) + **10 percentage points**; and 2. substrate mean − substrate sd > the measured chance floor. Otherwise **NEGATIVE**, reported with a diagnosis (training accuracy of each operation's single-op items, per-seed results, where composition failed). Design-selection seed: the design checks use no random draws, so no seed was used for design; all five seeds count and the "excluding the design seed" robustness check is vacuous. It is stated here so it is not read as omitted. **Pre-declared attribution rule.** If the staged GNN or the sequential MLP reaches within 10 percentage points of the substrate mean, any Part B wording must attribute the generalisation to staged operator application, which non-oscillator staged models share, and not to oscillator dynamics specifically. If a large-capacity diagnostic reaches within 10 points, the matched-budget margin is stated as budget-dependent. These rules apply whatever the verdict. Exploratory (labelled as such if reported): anything beyond the per-seed accuracies, training accuracies and the attribution flags above. ## Energy For every model: J/task = declared nominal package power of the host CPU × the measured wall-clock seconds per test item of a warm inference pass. This is an **energy proxy**, not a measurement; no power meter is read, and it is not evidence for oscillator-hardware frugality. The declared power and the host are recorded in the artefact. ## Compute and environment Host: ML350 (approved compute host). Before the run: health check, `uptime`, other projects' running units; at most 12 threads; a `systemd-run --user` unit; outputs under `~/`, then copied into the repository. Environment: the pinned `requirements-ci-py312-linux.txt` plus `requirements-ci-jax-py312-linux.txt` (jax 0.10.1, CPU). ## Analysis and data - Analysis: `scripts/run_kyma_v3_probe.py`, which calls `scpn_quantum_control.benchmarks.kyma_v3.probe.run_probe` and applies the contract in `probe.verdict`, frozen at the freeze commit. The artefact records the source commit. - Data: `data/kyma_v3_symbolic_composition/kyma_v3_symbolic_composition.json`; archived with the next Zenodo release. ## Amendments Any change before training is a new commit labelled as an amendment in this file, pushed before the run. No change after training starts. ## Reporting The measured result is reported whichever way it falls, with the attribution flags. The CEO receives this file's path and SHA-256 at freeze, and the result or truthful "in progress" wording by 12 October 2026. Part B quotes no partial or assumed number. Seat: 90ad ## RESULT Appended 2026-09-30 after the run; nothing above this heading has changed since the freeze commit `ddae107dc` (file SHA-256 before this section: `9248cf8a4fe9bfaf4d4807133f86dd9da51fac821fb71f70a3e8932eb232d1a5`). **Verdict under the frozen contract: PASS.** - Run: ML350 (hostname `god-of-the-math`), `systemd-run --user` unit pinned to cores 0–11, JAX CPU; source commit `dbcb0b0778221f85723ab7f60118c27e6369260f`, a descendant of the freeze commit with **zero** changes under `src/scpn_quantum_control/benchmarks/kyma_v3/`, `benchmarks/kyma_v2/` and `scripts/run_kyma_v3_probe.py` (it adds only CI registrations and evidence digests). Started 2026-09-29T19:01:33Z (freeze pushed 17:11Z), finished 2026-09-30T01:25:17Z; 9 h 08 min CPU time. Artefact `data/kyma_v3_symbolic_composition/kyma_v3_symbolic_composition.json`, SHA-256 `dd27ae8953d660626abbd23f23a7ea9c5c13123fa93280bef092982104f90cb8` (identical on the host and in the repository). - Design checks in the artefact match the table above: 2,112 training items, 64 test items, ambiguity by query 1.00 / 0.00 / 0.00, minimum distance 48, uniform label counts, hand-set realisability 1.000. Held-out accuracy on the pair `(R0, R1)`, query `a`, mean ± population SD over seeds 0–4 (training accuracy in brackets): | model | params | held-out | training | |---|---:|---:|---:| | **substrate** | 108 | **1.000 ± 0.000** | 0.9995 ± 0.0009 | | MLP (contract) | 109 | 0.256 ± 0.041 | 0.677 ± 0.012 | | staged GNN (contract) | 103 | 0.441 ± 0.099 | 0.943 ± 0.027 | | transformer (contract) | 100 | 0.253 ± 0.015 | 0.504 ± 0.002 | | sequential MLP (diagnostic) | 96 | 0.247 ± 0.025 | 0.495 ± 0.025 | | MLP, large (diagnostic) | 1,348 | 0.234 ± 0.043 | 1.000 ± 0.000 | | staged GNN, large (diagnostic) | 703 | 0.306 ± 0.021 | 1.000 ± 0.000 | | transformer, large (diagnostic) | 1,348 | 0.250 ± 0.000 | 0.496 ± 0.008 | Every seed of the substrate classified all 64 held-out items correctly (seed 2 missed 5 of 2,112 training items). **Contract.** (1) 1.000 ≥ 0.441 + 0.10: margin over the best contract baseline (staged GNN) **+55.9 percentage points**. (2) 1.000 − 0.000 > chance floor 0.25. Both hold → PASS. **Attribution rule (pre-declared).** Staged GNN within 10 points of the substrate: **no** (0.441). Sequential MLP: **no** (0.247). Large-capacity diagnostics within 10 points: **no** (best `gnn_large` 0.306). The generalisation is therefore not attributed to staged operator application alone, and the matched-budget margin is not stated as budget-dependent within the budgets tested. **What the result supports and what it does not.** As stated in the frozen limitation, the staged schedule makes a substrate that learns each operation exactly compose the held-out pair by construction. The result shows that gradient descent learned the three operation gates from end-of-program labels alone, and that parameter-matched and larger non-oscillator models, including two with the same staging, did not generalise to the pair. It does not separate the oscillator dynamics from the other architectural priors chosen for realisability (phase-lattice value encoding, reset phase, triadic coupling that can add phases); that separation would need an ablation not pre-registered here. **Energy proxy (not a measurement).** Declared nominal package power 190 W × measured warm wall time per test item, on cores shared with other projects' jobs on the host: substrate 1.68 J/item (8.8 ms/item, RK4 simulation on CPU); contract baselines 0.003–0.043 J/item; all baselines ≤ 0.050 J/item. The simulated substrate costs 33–536 times more per item than the baselines on this CPU. This says nothing about oscillator hardware. **Exploratory (labelled; not part of the contract).** The matched-budget contract baselines under-fit the training set (MLP 0.68, transformer 0.50), while the large MLP and large staged GNN fit it exactly and still stay near chance on the held-out pair (0.23, 0.31): the baselines' held-out failure is not explained by capacity alone.