# Claims Ledger Every claim destined for any paper or the thesis gets a row *before* it ships. No row, no claim. Types: `theorem` (proved), `proposition`, `conjecture` (+ supporting experiment), `confirmatory` (pre-registered empirical), `exploratory` (E-track, watermarked), `replication`. | # | claim | type | evidence artifact | script | commit | status | |---|---|---|---|---|---|---| | 1 | The per-neuron functional symmetry group of sine INRs contains ⟨ρ,σ⟩ ≅ D∞ with τ = ρ²; the per-layer group contains D∞ ≀ Sₙ (containment direction only; maximality = PO-2) | theorem (to write, G2) | thesis Ch1 §PO-1 | — | — | scoped (G0-theory-scoping §1) | | 2 | The protocol-suggested features cos 2b·(w⊗u), sin 2b·(w⊗u) are **not** invariant under the PO-1 group; the parity-classified family {(0,0): 1,cos 2b; (1,0): sin 2b·w; (0,1): sin b·u; (1,1): cos b·(w⊗u)} ⊗ matching-parity monomials is | proposition (membership verified by hand; generation/completeness = G2 lemma) | G0-theory-scoping §2 table | tests/test_invariants.py (T6, G1) | — | membership verified on paper; needs T6 + generation proof | | 3 | Per-neuron sine group ⟨τ,ρ,σ⟩ ≅ D∞ with normal form g_{d,j} and composition (d₁⊕d₂, j₂+(−1)^{d₂}j₁); per-layer group ⊇ D∞ ≀ Sₙ; cross-layer actions commute | theorem | ch1-symmetry.tex Thm (PO-1) + Lemma (normal form) | tests T1 | — | written G2 | | 4 | Strata: dead (w=0), invisible (u=0), **parallel-frequency** (w_j = ±w_i, any biases — phasor merge); closed, null, each carrying extra non-group equivalences | proposition | ch1-symmetry.tex Prop (PO-3) | strata audit at G3 | — | written G2; parallel-stratum correction logged | | 5 | Φ = (w⊗w, cos2b, sin2b·w, sinb·u, cosb·(w⊗u)) is invariant and separates D∞-orbits on {w≠0,u≠0} | proposition | ch1-symmetry.tex Prop (PO-4) | T6 | — | written G2 | | 6 | No continuous canonicalizer exists on the generic stratum (τ-winding obstruction; elementary loop proof) | proposition | ch1-symmetry.tex Prop (PO-5) | F7 margin diagnostics | — | written G2 | | 7 | **L=1 identifiability:** on Θ_gen (w≠0, u≠0, no parallel pair), f_θ=f_θ′ ⟹ θ′ = gθ, g ∈ D∞≀Sₙ unique — D∞≀Sₙ is a maximal symmetry group (sense of 2604.23720 Def 2.2), resolving 2409.11697 Rmk 4.5 for sine, L=1 | theorem | ch1-symmetry.tex Thm (PO-2) | S4e falsification hunt (deep case) | — | written G2; proof to be red-teamed at G2 advisor review | | 8 | Deep identifiability | conjecture + roadmap (Bessel-CP reduction; 2 open lemmas; Bessel-Vandermonde ill-conditioning documented) | proof-memos/PO-2-deep-attempt.md | S4e | — | downgraded per timebox, honest attempt recorded | | 9 | Complete invariants factor through function space (conditional on identifiability); informational-equivalence corollary; exact/D∞ counterpart of 2602.01083's approximation results | proposition | ch1-symmetry.tex Prop (PO-6) | S5 Pareto adjudication | — | written G2 | | 10 | Activation classification: tanh ℤ₂; Gaussian ℤ₂ (u untouched); ReLU+PE ℝ₊ continuous; sine D∞; **FINER ℤ₂ (odd; no bias-shift symmetries — zero-gap argument)** | proposition (per row) | ch1-symmetry.tex Prop (PO-7) | S3 | — | written G2; P-A/P-B amended pre-data | | 11 | Microcosm: profiled loss closed form (quadrature-certified 6e−16); zero set = D∞ orbit exactly; 20 minima (1 global, 19 spurious); basin census non-monotone in init range with three regimes — 100% degenerate-ridge capture at ±2, 62% global at ±10, 51% spurious-sidelobe capture at ±20 (ω=7) | theorem (forms) + numeric certification | ch2-fitmap.tex §2; results/microcosm/ (F4) | scripts/02_microcosm_po8.py | — | done G2 | | 12 | Microcosm basin census is **optimizer-robust in its headline and wrong in one sub-claim**: the non-monotone global-capture curve replicates under Adam and plain GD on the *full* (non-profiled) model — global fraction 0.00/0.20/0.56/0.31 (Adam) and 0.00/0.18/0.58/0.33 (GD) at init ranges 2/5/10/20 vs Nelder–Mead's 0.00/0.26/0.62/0.32, sweet spot at init range ≈ ω — but **"degenerate-ridge capture" is a Nelder–Mead/profiled-surface artifact** (0.00 under both gradient methods; those runs are still descending, not in a ridge basin), and the ±20 spurious-sidelobe fraction is optimizer-dependent (0.48 Adam-converged, 0.00 GD-converged) | numeric certification (replication of row 11 under the production optimizer class) | results/microcosm/optimizer_census.json (F4 right panel) | scripts/13_microcosm_optimizers.py | — | done G4; corrects row 11's ridge/sidelobe language | | 13 | At the **frozen corpus setting** (Adam, lr 1e−3, 300 steps) the microcosm fit is nowhere near a critical point: 100% of inits end unconverged at every init range, endpoint ‖∇‖ ≈ 0.5–0.7, and median \|Δw\| ≈ 0.24 *independently of init range* — the fit stays in its initialization neighbourhood. Mechanistic support for PO-9 laziness and hence for the W1−W3 gap | numeric | results/microcosm/optimizer_census.json | scripts/13_microcosm_optimizers.py | — | done G4 | | 14 | **S1 decomposition (MNIST, confirmatory):** information survives the fit (P1 ≈ P0, TOST margin 1.0, p 2.5e−05); the raw-weight gap W1−W3 = 80.43 pts replicates the A1 anchor; optimization noise is **null** (W1−W2 = −0.68); template alignment to the shared init recovers f = 0.628 of the gap, template-free sorting only 0.177, exact L=2 invariants 0.269; group augmentation and K-marginalization both ≈ 0.05 and are statistically indistinguishable from each other (Holm p = .21) | confirmatory (pre-registered S1, hash 8c029cf43f01a94c + addendum 01) | results/ladder/mnist/S1_analysis.json, F9_waterfall.png | scripts/11_ladder.py, 14_ladder_analysis.py | — | done G4; H-S1-3, H-S1-4c, H-S1-5 missed their registered intervals | | 15 | **Exact L=2 phase-invariant encoding**: for two-layer sine networks the layer-2 Gram G = W₂ᵀW₂ is invariant under the whole layer-2 group and picks up ε_iε_l under layer 1, so (sin b_i sin b_l)G, (cos b_i w_i · cos b_l w_l)G and (sin2b_i w_i · sin2b_l w_l) are sign-cancelling and their sorted spectra are fully invariant — closing the L≥2 gap that PO-4 left open (OPEN_PROBLEMS #4) for the L=2 case | proposition (invariance proved by parity bookkeeping; numerically certified to 3e−7) | ch3 §3.6 (to write); tests/test_t10_deep_invariants.py | src/sirengap/canon/deep_invariants.py | — | done G4 | | 16 | **Weight-space brittleness (X1):** a decoder trained on shared-init weights is at chance (10.7) on random-init weights of the same images, and 13.2 in reverse — the two protocols are not merely differently-scaled versions of one representation | confirmatory (S1 rung X1) | results/ladder/mnist/X1.json | scripts/11_ladder.py | — | done G4 | (rows appended as work completes) | 17 | **Template alignment recovers ~60% of the perception gap regardless of template** (exploratory support for row 14's f(W5)): f = 0.628 aligning to the corpus's shared init θ₀, 0.640 and 0.601 to *unrelated* random inits, 0.552 and 0.508 to fitted INRs. The gain is a consistent frame, not privileged knowledge of the initialization; fitted networks are worse templates than untrained ones | exploratory (watermarked; the confirmatory number is row 14's registered θ₀ rung) | results/ladder/mnist/EXPLORATORY_w5_template_sensitivity.json | scripts/15_w5_template_sensitivity.py | — | done G4; retires the template caveat in DEFENSE row 15 | | 18 | **The S1 decomposition replicates on FashionMNIST:** f(W4) 0.170, f(W5) 0.664, f(W6) 0.032, f(W7) 0.034, f(W9) −0.008, W8 at chance, W1−W2 = −0.70 — every qualitative claim of row 14 survives the dataset change; the one difference is W10, which does markedly better on FMNIST (0.428 vs 0.269) | replication (same frozen apparatus and analysis; MNIST-calibrated *magnitude* intervals do not transfer and are not scored) | results/ladder/fashionmnist/S1_analysis.json, F9_waterfall.png | scripts/11_ladder.py, 14_ladder_analysis.py | — | done G4 | | 19 | **The S1 decomposition on CIFAR-10 (confirmatory, own registration):** against 17 intervals frozen before any cell existed, **14 hit (82% against nominal 80%)**. P0 55.81, P1 56.23, W1 44.29, W2 45.19, W3 12.64; gap W1−W3 = 31.64; f(W4) 0.108, f(W5) **0.324**, f(W6) 0.128, f(W7) 0.101, f(W9) 0.001, f(W10) **0.534**; W8 10.65 (chance), X1 10.4 forward / 12.0 reverse. The three misses are H-C1-8 and H-C1-17 (both the crossover, row 20) and H-C1-9 (augmentation recovers 0.128 against a registered ceiling of 0.12) | confirmatory (pre-registered S1-CIFAR, hash f7906fc6904c7c81) | results/ladder/cifar10/S1_analysis.json, F9_waterfall.png | scripts/11_ladder.py, 14_ladder_analysis.py, 24_score_cifar_arm.py | — | done G5 | | 20 | **The two-thirds recovery is not a law, and the two exact methods cross over.** f(W5) = 0.628 / 0.664 / **0.324** and f(W10) = 0.269 / 0.428 / **0.534** on MNIST / FashionMNIST / CIFAR-10: template alignment halves on natural RGB images while the exact L=2 invariant encoding nearly doubles and overtakes it. The direction of W10's rise was registered in advance (P-C1-B, P = 0.60, from the encoding's algebra: c = 3 gives each neuron more D∞-visible outgoing structure) and resolved correctly; the companion call that the grayscale ordering would persist (P-C1-A, P = 0.65) resolved incorrectly | confirmatory for the levels (rows registered per dataset); the crossover itself is the joint reading of three arms | results/ladder/{mnist,fashionmnist,cifar10}/S1_analysis.json | scripts/11_ladder.py, 14_ladder_analysis.py | — | done G5 | | 21 | **The fit-length confound is dead, and the channel mechanism explains ~29% of W10's rise.** (a) Median relative travel ‖θ_T − θ₀‖/‖θ₀‖ = 0.186 / 0.191 / 0.187 and layer-1 direction cosine to init 0.998 / 0.999 / 0.999 across the three datasets, despite CIFAR being fitted for 1000 steps against 300 — the extra steps bought no extra displacement, so f(W5)'s drop is not a fit-length artifact. (b) Restricting the W10 encoder to one output channel (D 384 → 320, the grayscale dimension) costs only 0.077 of 0.534, i.e. ~29% of the 0.265 rise from MNIST; the *averaged* control has the same D and strictly more information yet does worse (0.425), so the effect is neither dimension nor information but per-channel sign structure | (a) numeric, no new fitting; (b) exploratory (watermarked) | results/fit_travel.json; results/ladder/cifar10/EXPLORATORY_w10_channel_ablation.json | scripts/23_fit_travel.py, 25_w10_channel_ablation.py | — | done G5 | | 22 | **Correction to rows 14/16 and DEFENSE row 8: a recovery fraction is a certified *lower* bound on the symmetry share, not a decomposition.** For any exact reframing c(θ) ∈ Gθ, f_{c(θ)} = f_θ, so no information about the signal is created and every point gained is attributable to the change of orbit representative — but c need not be *canonical*, so a better exact reframing can only raise f. The complement is therefore an *upper* bound on the non-symmetry share, not a measurement of it. Evidence that the bound moves: c_sort → c_align took the MNIST "residual" from 82% to 37%, and CIFAR's best exact method is W10 (0.534), not c_align (0.324). W10 needs a further caveat — it is a *nonlinear* invariant feature map, so part of its gain may be feature engineering rather than group removal | proposition (elementary) + the program's own history as evidence of looseness | paper §6 (Prop. 4); DEFENSE rows 8, 15 | — | — | done G5; corrects the gloss in rows 14/16 rather than editing them | | 23 | **Exact orbit minimisation over G is tractable at L≥2**: given the other layers fixed, the optimal element of one layer's D∞ ≀ Sₙ is exactly computable — the per-neuron D∞ cost depends on the winding j only through its *parity*, so four (d, parity) cases give the exact minimum over the whole infinite group, and the permutation is then a Hungarian assignment on those per-pair minima; layers are swept by coordinate descent from restarts. Planted group elements are undone to <1e−6 relative at L=1,2,3, with c=3 outputs, at windings up to 12 | proposition (construction) + numeric certification | tests/test_t12_refine.py (T12, 13 cases) | src/sirengap/canon/refine.py | — | done S4e | | 24 | **S4e: Conjecture 6.5 (deep identifiability) survives, with positive evidence and no counterexample.** At w=2, 1 of 128 independent students fitted to a teacher's *exact outputs* reached R_f = 5.9e−08 and aligned to R_θ = 1.2e−07 — float32 epsilon, max per-coordinate relative disagreement 2.8e−06, i.e. recovery to 6–7 significant figures, and 2.6e−07 of the unrelated-network scale (0.468). No counterexample at any width under the amended criterion | confirmatory (pre-registered S4e, hash aa5426a4245bd22f; 7/9 intervals hit) | results/s4e/s4e.json, candidate_verification.json | scripts/26_s4e_identifiability.py, 27_score_s4e.py, 29_s4e_verify_candidate.py | — | done S4e | | 25 | **…but identifiability at L=2 has no empirical content at production width.** κ = R_θ/R_f falls 0.042→0.0055 from w=2 to w=32, so local recovery is *well* conditioned (the forward map is expansive) — the *opposite* of the naive reading of the proof memo's Bessel–Vandermonde collapse, which governs the global system rather than the local Jacobian. What collapses is the basin's **volume**: started inside it, 78–91% of runs return at w≤8, 44% at w=16, **0% at w=32**, where the optimiser walks from R_θ=1e−5 out to 1.3e−1 while the function barely improves; a control at 5× the step count is identical, so this is not under-training. Independent restarts find the basin only at w=2 | confirmatory + one exploratory budget control | results/s4e/s4e.json, budget_control.json | scripts/26, 28_s4e_budget_control.sh | — | done S4e; retires S4e as the answer to DEFENSE row 15 and hands the question to the two open lemmas | | 26 | **Same-image repeated fits are no closer, modulo the whole group, than fits of different images.** At production width (w=32, MNIST P-random-K) two independent fits of one image sit at R_θ = 0.279 after *optimal* alignment over G; two fits of different images sit at 0.280, difference −0.001. This is the W1−W3 gap of row 14 measured in parameter space rather than through a decoder | confirmatory (S4e production arm; P-S4e-8/9 both HIT) | results/s4e/s4e.json | scripts/26_s4e_identifiability.py | — | done S4e; their R_f is 0.73, so the arm characterises the corpus and does not test the conjecture | | 27 | **A pre-registered falsification criterion can be under-specified in a way only data reveals — recorded as a methods finding.** S4e §4 (R_f < 1e−5 AND R_θ > 20κR_f) *fired* on the exact-recovery case of row 24, because (a) it is ratio-only with no absolute floor, so as R_f → machine epsilon the threshold falls below any float32 aligner's resolution, and (b) κ was measured on *random* directions while a minimiser's residual lies in the flattest directions of the loss, where R_f is least sensitive to R_θ, so R_θ/R_f > κ is expected for any converged minimiser (2.10 vs 0.042). Post-hoc amendment A1 adds R_θ > 1e−3; frozen §4 left unedited; P-S4e-C still scored against the criterion as written (Brier 0.7225) | methods (post-hoc, declared) | docs/prereg/S4e.md amendment A1; results/s4e/candidate_verification.json | scripts/29_s4e_verify_candidate.py | — | done S4e; third calibration failure mode, added to the template | | 28 | **The cross-dataset behaviour is driven by image statistics, not by output-channel count — and the paper's own mechanism conjecture is withdrawn.** Luminance CIFAR-10 (identical images, geometry, architecture and 1000-step budget; c = 3 → 1) gives f(W5) = **0.324** against **0.324** at c = 3, identical to three decimals, and the crossover does not reverse (f(W10) = 0.493 > f(W5)). The registered §9 conjecture — that c_align's collapse is about outgoing structure growing with c — is therefore **wrong and withdrawn**, per the pre-committed wording of the registration | confirmatory (pre-registered S1-gray, hash b84b660829aa6d40; 9/10 intervals hit, both probability calls resolved against the conjecture, Brier 0.1225 / 0.2025) | results/ladder/cifar10gray/S1_analysis.json, verdict.json | scripts/30_cifar_gray_corpus.sh, 31_cifar_gray_ladder.sh, 32_score_gray_arm.py | — | done; partial ladder (8 rungs, 2 protocols) as S1 §6 permits | | 29 | **Higher render fidelity does not buy alignment recovery.** Dropping to c = 1 makes the fit over-parameterised (1185 params to 1024 targets) and lifts median render PSNR from 40.1 dB to 59.8 dB, yet f(W5) is unchanged at 0.324 — and this corpus is fitted *more* accurately than MNIST (39.2 dB) while aligning *far worse* (0.324 vs 0.628). Kills the "the fit is simply easier" reading that the registration named as owed | confirmatory (same arm; the PSNR rise is a declared, uncontrollable consequence of the intervention) | results/inrbench/cifar10gray_P-shared-det_test_gate.json | scripts/04_quality_gate.py | — | done | | 30 | **Three candidate causes eliminated; no replacement mechanism offered.** Fit length (row 21a: travel 0.186/0.191/0.187 despite 300 vs 1000 steps), output-channel count (row 28), and render fidelity (row 29) are all ruled out for the cross-dataset behaviour of W5 and W10. What remains is image statistics. The program deliberately declines to supply a second mechanism on the same evidence immediately after the first was falsified, and instead names the experiment that would identify one | methods position | paper §9; docs/LAB_NOTEBOOK.md | — | — | done | | 31 | **A capacity-matched equivariant reader does not substitute for frame choice.** W11a — bipartite permutation-equivariant message passing on raw weights, the coverage the DWSNets/NFN/GMN family has for sine networks — recovers f = **0.265** on MNIST P-random against c_align's 0.628, i.e. about the same as an exact invariant encoding with lossy pooling (0.269) and less than half of an exact change of frame. Reader sized by rule to the frozen decoder's 1,873,162 params (+0.4%), so it does not lose for being smaller | confirmatory (pre-registered S1-W11, hash e3bbc081a5810956; 5/5 intervals hit, P-W11-A resolved as registered, Brier 0.0625) | results/ladder/mnist/W11.json, W11_verdict.json | scripts/33_w11_equivariant.py, 34_score_w11.py | — | done; bounded to one equivariant construction at one capacity on one corpus, stated as such | | 32 | **The exact L=2 invariant encoding's shortfall is 72% pooling, 28% incompleteness — OPEN_PROBLEMS #4 closed.** Feeding W10's *own* invariants (per-neuron even scalars; the sign-cancelling matrices A, B, C) to a permutation-equivariant reader with **learned** pooling instead of sorted eigenvalue spectra gives f = **0.526** against W10's 0.269. Of the 0.359 gap to c_align's 0.628, **0.257 (72%) is attributable to the pooling and 0.102 (28%) to the invariants' incompleteness** | confirmatory (same registration; H-W11-5 registered +0.17 [0.02, 0.34], observed +0.257 — HIT) | results/ladder/mnist/W11.json | scripts/33_w11_equivariant.py | — | done; the practically useful form is a G-invariant equivariant reader, which needs no template and comes within 0.10 of alignment | | 33 | **S5, the FLOPs-matched adjudication: weight access is dominated by function access on BOTH axes.** Classifying an INR by querying it at K=64 *learned* coordinates reaches **95.34%** at **1.59 MFLOP/INR** (analytic accounting); the best weight-access rung on the same corpus, c_align, reaches 64.41% at 5.45 MFLOP — **30.9 points worse at 3.4× the compute**. At K=256 function access reaches 98.23%, above the real-pixel MLP's 97.97% | confirmatory (pre-registered S5, hash 80bdc96ce9497c3d; 5/8 intervals, P-S5-A resolved TRUE, Brier 0.0225) | results/s5/pareto_mnist.json, verdict.json | scripts/35_s5_pareto.py, 36_score_s5.py; src/sirengap/eval/{flops,probes}.py | — | done; the three interval misses are one shape error — the accuracy curve is far more sigmoidal in K than registered | | 34 | **Amortization does not rescue weight access, and the reason is general.** Over T downstream tasks weight access costs 1.70 + 3.74·T MFLOP against function-query's 1.59·T; the lines never cross because the weight reader's *per-task* cost on a 1185-dim input already exceeds function-query's *entire* per-task cost on a 64-dim one. Mechanism: reading P parameters into a decoder of first-layer width W costs ≈2PW while querying K points costs ≈2KcW + K·siren, so function access is cheaper whenever **Kc ≪ P** — a condition on probes needed, not on INR size. This closes the one escape PO-6's corollary left open | confirmatory (P-S5-B registered at 0.30, resolved FALSE, Brier 0.09) + analytic | results/s5/verdict.json | src/sirengap/eval/flops.py | — | done | | 35 | **The perception gap is invisible from the function side.** Between P-random and P-shared-det, function-query accuracy moves **5.4 points** — a fit-quality effect (median PSNR 37.5 vs 39.2 dB) — where weight access moves **80.4**. The entire object this program decomposes is an artifact of choosing to read parameters | confirmatory (H-S5-7 registered 4.0 [0, 10], observed 5.43 — HIT) | results/s5/pareto_mnist.json | scripts/35_s5_pareto.py | — | done | | 36 | **Correction to rows 14/16/22 and 28–32, after external review: the recovery fraction is NOT a certified lower bound on the symmetry share.** The argument "an exact reframing creates no function-level information, therefore its gain measures nuisance removal" is invalid. Counterexample: for any orbit-invariant binary y(θ), the map applying τ_k at the first neuron with k = M·y(θ) is orbit-valued and exactly function-preserving, yet a linear probe on the first bias then predicts y — the group's degrees of freedom were used as a channel. Two further gaps: the decoder is retrained per rung (so what is fixed is the *algorithm*), and the contrast is between corpora fitted from different *initializations*, which intervenes on the fit map rather than on the group. **f is renamed the algorithm-relative recoverable fraction, f ∈ ℝ (not [0,1] — W8/W9 are negative).** The causal estimand is measured separately by S6 | correction (methods) | paper §8; docs/prereg/S6.md | — | — | done; row 22's weakened claim was itself still too strong | | 37 | **Correction: the CIFAR-10 headline must report 0.324, not 0.534.** The paper's own §8 states W10 is a *nonlinear invariant encoding*, not an orbit-valued reframing, and that part of its gain may be feature engineering; the gap table nonetheless took a max over "exact methods" including W10. That is an internal inconsistency. The strongest **reframing** result on CIFAR-10 is c_align at **0.324**. Reframings (W4, W5) and encodings (W10, W11b) are now reported in separate families throughout, and "the gap is mostly, but not entirely, symmetry" is withdrawn — an incomplete canonicalizer cannot identify the residual | correction | paper §8, Table 2 | scripts/22_paper_tables.py | — | done | | 38 | **Correction: the 72%/28% pooling-vs-incompleteness split (row 32) is withdrawn.** W10 and W11b differ in pooling *and* reader architecture, parameter count (0.99M vs 1.85M), relational capacity and optimisation geometry, so the movement 0.269 → 0.526 cannot be assigned to pooling alone. The supported statement is that replacing fixed spectral pooling with a learned equivariant reader over the same invariant objects substantially improves accuracy. The factorial ablation that would license a split (matched-capacity MLP on the same objects; richer fixed invariant statistics; the same learned pooling over non-invariant controls) is future work | correction | paper §12 | — | — | done | | 39 | **Correction: Propositions PO-6 and its corollary are proved only under Theorem PO-2 (L=1).** At L=2, where every corpus lives, they are conditional on the deep-identifiability conjecture. The S5 FLOPs comparison at L=2 is therefore an empirical result, not a consequence of the proved theorem, and is presented as such (paper Remark, §3) | correction (scope) | paper §3 Remark | — | — | done | | 40 | **Positioning correction: the two function-preserving SIREN transformations are prior work.** Shamsian et al. (2402.04081) use neuron negation and integer-π bias shifts as weight-space augmentations, and Tran et al. (2409.11697) incorporate sign symmetry for sine/tanh into a monomial-matrix framework. This was recorded in our own close-read memo on 2026-07-17 and was not foregrounded in the paper. The novelty claim is now stated as: the *group* they generate (D∞ ≀ Sₙ), the affine nature of the phase component that places it outside monomial actions, and generic maximality at L=1 — not the discovery of the transformations | positioning correction | paper §1; docs/THINKING/close-reads/2402.04081-ws-augmentations.md | — | — | done | | 41 | **c_align and c_sort are canonicalizers on the production corpora, not alignment heuristics.** The review asked whether c(gθ) = c(θ) holds for arbitrary tested g on *fitted* networks, whose first-layer directions come within 3e−4 rad of the parallel stratum so the sort keys nearly tie. Measured on 512 held-out INRs per dataset under 4 random group elements each (windings \|j\|≤3, non-trivial permutations): residual relative to each INR's own scale has median 1.2e−07 and max 1.5e−07 on MNIST, 1.3e−07 / 1.5e−07 on CIFAR-10, with **0% of INRs above the 1e−4 tolerance**. That is fp32 round-off. PO-5 still applies — these maps must be discontinuous somewhere — but not on the sampled orbits | audit (numeric, no new fitting) | results/audits/canon_equivariance_{mnist,cifar10}.json | scripts/42_canon_equivariance_audit.py | — | done; answers review question 1 | | 42 | **A component-by-component provenance ledger now governs every novelty claim.** Six theory components, six method components and five experimental components are each classified prior work / specialization / consequence / extension / ours, with the closest prior work named and the difference stated. Two rows rest on a null keyword search (eight targeted arXiv queries, 2026-08-03) and say so, since a null search is weak evidence. The ledger has already forced five corrections (rows 36–40) and is reproduced as a table in the paper's related-work section rather than left to a closing paragraph | positioning (process) | docs/PROVENANCE.md; docs/lit_snapshots/G8-novelty-scan.txt; paper Table 3 | scripts/00_lit_scan.sh | — | done | | 43 | **A matched non-invariant control exists for the W10 encoding (apparatus).** W10c emits the same monomials in (w,u) at the same trigonometric orders as W10, pooled by the same eigenvalue spectra under the same ‖w‖² sort key, at the same dimension (320 / 384), decoded by the same frozen apparatus — with only the parity class of each trigonometric factor swapped. It is therefore **still exactly permutation-invariant** (relative move < 1e−5) and **not** D∞-invariant (relative move > 1e−2), so the difference between the two rungs isolates the D∞ component specifically rather than confounding it with permutation handling | proposition (construction) + numeric certification | tests/test_t15_deep_control.py (12 cases) | src/sirengap/canon/deep_control.py; docs/prereg/S7.md | — | apparatus done; the decoded result is the S7 row | | 44 | **S6: the group alone reproduces nearly the whole gap — but the gap is NOT reducible to the group.** Randomizing g ~ μ_B over the untouched `P-shared-det` corpus (same networks, same functions, functional gap ≤ 8.7e−06) costs **Δ_sym = 79.07 pts** at B=0 and **79.09** at B=10 — flat in B — against an 80.4-pt observed gap. Recovery of Δ_sym vs of the observed gap: c_align **0.865 vs 0.628**, invariants **0.724 vs 0.269**, c_sort **0.573 vs 0.177**, W11a 0.631, W11b 0.886. Every treatment does markedly better against synthetic scatter. Applying the intervention to `P-random` costs **+0.13 / −0.50 pts**: an independently-initialised corpus is already group-saturated | confirmatory for arms (ii)–(iv) (pre-registered S6, hash 4826774c2e2c0d9d, 3/6 intervals); arm (i) declared exposed and reported as measurement | results/s6/orbit_mnist_{perm,noperm,equivariant,prandom}.json, verdict.json | scripts/37_orbit_intervention.py, 43_score_s6.py | — | done; replaces the invalid lower-bound reading of row 36 | | 45 | **The decisive triple: an exactly G-invariant reader still loses 28.6 points, and the group cannot explain it.** W11b scores **84.81%** on `P-shared-det`, **85.39%** on the *same corpus after the group is randomized* at B=3 — a 0.59-pt difference, i.e. seed noise, so its invariance is **measured** and not merely asserted (H-S6-5 HIT) — and **56.24%** on `P-random`. A reader empirically blind to D∞ ≀ Sₙ therefore loses 28.6 pts between the two corpora. That loss is not group scatter, and it is not lost signal: function-query accuracy moves only 5.4 pts between the same corpora (row 35). **"The gap is not reducible to parameter symmetry" is therefore supported again — by an intervention, not by a canonicalizer's residual.** What the residual *is* (genuinely different orbits, per row 26's R_θ = 0.279 vs 0.280; or an incomplete invariant family that reads more from a shared chart) an incomplete invariant cannot decide, and no decomposition is claimed | confirmatory + measurement | results/s6/orbit_mnist_equivariant.json; results/ladder/mnist/{W11,W11_shareddet}.json | scripts/37, 33_w11_equivariant.py | — | done; reinstates, in scoped form, the language row 37 withdrew — **magnitude superseded by row 52**, which re-runs this triple with W12 and gets 7.8 pts | | 46 | **Within the group, reflection dominates and winding is nearly free — the reverse of what was registered.** With the permutation fixed to identity, Δ_sym is **62.90** (B=0) rising only to **64.08** (B=10). Of the 79 points the full group costs, ~**63 are per-neuron sign flips**, ~**15** is relabelling 32 neurons, and ~**1** is windings up to \|j\|=10. H-S6-1 (10 [2,30]) and H-S6-3 (64 [45,76]) both miss badly and **P-S6-A resolves FALSE** (Brier 0.64). Consequence for positioning: D∞ beats Sₙ four-to-one, which is the empirical case for treating the sine group as more than permutations — but *within* D∞ it is σ, the generator monomial-matrix frameworks already cover, that carries almost all of it. **The affine phase component is mathematically necessary for PO-2 and empirically minor as a source of scatter**, and the paper now states both separately | confirmatory (S6 arm (ii)) | results/s6/orbit_mnist_noperm.json | scripts/37, 43 | — | done | | 47 | **S7: most of the invariant encoding's CIFAR-10 gain is invariance, but a substantial minority is not.** The matched control (row 43) gives f(W10c) = **0.125 / 0.216** on MNIST / CIFAR-10 against W10's 0.269 / 0.534, so **0.318 of the CIFAR-10 recovery is attributable to quotienting D∞** and 0.216 to the same nonlinearity without it. **3/3 intervals hit** (0.14 [0.02,0.28], 0.22 [0.05,0.42], and the difference 0.31 [0.11,0.48] against an observed 0.318), and the pre-committed falsifier — which would have voided every symmetry reading of the CIFAR-10 encoding result — **did not fire**. But **P-S7-B resolves FALSE**: W10c (0.216) beats c_sort (0.108), so the nonlinearity contributes on its own and **the W10-vs-W4 comparison is not a clean symmetry comparison**; the clean one is W10 vs W10c, and that is what the paper now quotes | confirmatory (pre-registered S7, hash 682f760e526f4850, 3/3 intervals; Brier 0.0625 / 0.3025) | results/s7/control.json; results/ladder/{mnist,cifar10}/W10c.json | scripts/11_ladder.py, 41_score_s7.py | — | done; answers the review's Priority 3 | | 48 | **A reader that quotients D∞ ≀ Sₙ on the RAW parameters recovers 92% of the perception gap, and the paper's practical claim is withdrawn.** W12 puts the bias in phasor coordinates, which reduces the infinite winding to a **parity** and leaves a finite Z₂×Z₂ acting by signs; every layer preserves that grading (linear within a character, products add characters, odd nonlinearity off the neutral block, only neutral channels reach the head), and W² couples the layers with character (1,1) on the layer-1 side and (1,0) on the layer-2 side, leaving exactly two legal message channels per direction. **f(W12) = 0.917** (87.64% against W1's 94.36%) at 1,874,898 params — **+0.651** over W11a, **+0.390** over W11b, **+0.288** over c_align, all within 1.5% of the frozen decoder's capacity. The phasor route was **proposed by an external reviewer**, credited at first mention and in PROVENANCE row M7; ours is the two-layer realization | confirmatory (pre-registered S9, hash 1c74280c55a1a3f0, **0/3 intervals** — all missed high; P-S9-A/B/C all resolved TRUE, Brier 0.0225 / 0.3025 / 0.5625) | results/ladder/mnist/W12.json, results/s9/verdict.json | scripts/47_w12_phasor.py, 50_score_s9.py; src/sirengap/models/phasor.py; T16 | — | done | | 49 | **Correction, and it reverses row 31: frame choice does NOT beat reader architecture.** Row 31 read W11a's 0.265 against c_align's 0.628 as showing that within weight space the orbit representative matters more than the reader. W12 (row 48) recovers 0.917 at the same capacity, so the correct reading of row 31 is the narrower one: **permutation equivariance alone is not enough**. That is the empirical counterpart of row 46 — relabelling carries ~15 of the group's 79 points and reflection/phase carry ~64 — and W11a quotients only the relabelling. S9 §4 pre-committed to this withdrawal if P-S9-C resolved true; it did | correction (pre-committed) | docs/prereg/S9.md §4; results/s9/verdict.json | scripts/50_score_s9.py | — | done; row 31's *measurement* stands, its interpretation does not | | 50 | **Invariance verified on the corpus it is decoded from, not only on synthetic parameters.** W12's logits move 1.5e−06 to 3.3e−06 on 512 fitted MNIST INRs under group elements with windings to \|j\|=40 — unbounded windings are meaningful here because the phasor quotients the integer translation *exactly* — and the neutral blocks move 3e−08 to 7e−07 under the D∞ part alone. S9's void condition (relative logit move > 1e−4) does not fire | audit (numeric, no new fitting) | results/audits/w12_invariance_mnist.json | scripts/52_w12_invariance_audit.py | — | done; the first version of the neutral-block check compared blocks elementwise after a permutation and reported ~1.0, which was the check being wrong, not the code | | 51 | **The layer-level grading is NOT what does the work — the symmetry-adapted coordinates are.** The matched control W12u (row 43's argument applied to W12) removes the grading and keeps everything else: same skeleton, same block shapes, same phasor features, capacity re-solved by the same rule (width 122, 1,888,212 params). It is genuinely non-invariant — logits move **0.25** relative under the group on fitted INRs against W12's 3e−06 — and still reaches **f = 0.858**, beating c_align (0.628) and W11b (0.526). So the grading is worth **+0.059** of W12's 0.917. Mechanism: both readers receive the same exactly-invariant neutral block from the shared feature map (‖w‖², cos 2b, ‖u‖², Gram diagonals; moves 3e−08 under D∞), and what the grading adds is the *guarantee* that nothing outside it reaches the head. **What is not separated** is the coordinates from the architecture: the step W11a 0.265 → W12u 0.858 conflates them, and the third arm that would separate them (the same graded skeleton reading raw bias) was not run | exploratory (not registered; S7's design reapplied) + audit | results/ladder/mnist/W12u.json; results/audits/w12u_invariance_mnist.json | scripts/47_w12_phasor.py --ungraded, 52_w12_invariance_audit.py --ungraded | — | done; qualifies row 48 without touching the withdrawal in row 49, which both readers license — **and is itself qualified by row 58**: with the third arm run, the coordinates carry about half the step, not the bulk of it | | 52 | **The residual that survives removing the group is reader-relative, and row 45's 28.6 points is withdrawn as a measurement of it.** Re-running row 45's triple with W12 in place of W11b: W12 scores **95.46%** [95.03, 95.99] on `P-shared-det` and **87.64%** on `P-random`, a loss of **7.82 pts** (9.7% of the 80.4-pt gap) where W11b lost 28.56 (35.5%) — a factor of 3.7 from changing the reader, not the corpora. Normalized against each reader's own shared-init ceiling rather than W1's, W11b carries 0.597 of its gap across and W12 carries 0.904. The middle leg of the triple is redundant for W12: its invariance is exact by construction and audited at a 3.3e−06 logit move to |j|=40 (row 50), so evaluating it on the group-randomized corpus can only reproduce the same accuracy to float noise. **Consequences:** (a) every such figure is an *upper* bound on the non-symmetry share that a better invariant reader can lower — the same algorithm-relativity Prop. 4 established for recovery fractions (row 22); (b) the qualitative claim survives (the intervals do not overlap, so the loss is ≥6.7 pts and cannot be group scatter), but the 'not lost signal' leg now clears the 5.4-pt function-side movement between the same corpora (row 35) by only **2.4 pts** where it cleared it by 23 with W11b. 'The gap is not reducible to parameter symmetry' is retained only in the scoped form *some* residual survives, of unknown size | exploratory (not registered; row 45's registered design re-run with a better reader) | results/ladder/mnist/W12_shareddet.json vs W12.json, W11_shareddet.json; results/audits/w12_invariance_mnist.json | scripts/47_w12_phasor.py --protocol P-shared-det | — | done; supersedes the magnitude in row 45 and the corresponding text in the abstract, §9, §15, §18 and the conclusion | | 53 | **W12 exceeds the raw-weight ceiling on the corpus with no nuisance, so part of its gain is not group removal.** On `P-shared-det`, where every INR is fitted from the same initialization and there is nothing to quotient, W12 reaches 95.46% against W1's 94.36% — **f = 1.014 [1.008, 1.021]**, strictly above one at matched capacity (1,874,898 params, width 186, 5 seeds). The phasor coordinates are therefore a better way to read parameters *per se*, which independently corroborates row 51's finding from the other side and caps how much of W12's 0.917 on `P-random` any invariant construction may claim as symmetry work. It also means f, defined against W1, is not bounded by 1 for readers stronger than W1 | exploratory (not registered) | results/ladder/mnist/W12_shareddet.json | scripts/47_w12_phasor.py --protocol P-shared-det | — | done | | 54 | **Methods finding: an off-protocol re-run silently overwrote a registered artifact and double-scored the ledger.** The S8/S9 master chain re-priced the S5 frontier by calling `35_s5_pareto.py` at its default `--seeds 3`, but S5 §6 registers **n=5** for the headline K sweep (the stopping rule permits 3 only for K=256, and only if the sweep exceeds 3 h, which it did not). The 3-seed run overwrote `results/s5/pareto_mnist.json` and `36_score_s5.py` appended a **second copy of all 11 S5 rows** to PREDICTION_OUTCOMES.csv, which would have double-counted 11 calls in the calibration audit. No verdict changed (all 11 resolve identically at n=3 and n=5; K=64 95.34 vs 95.40, K=256 98.23 vs 98.16), so nothing in the paper moves. The seed-count guard added for S8 in `48_s8_sweep.py` was never extended to the S5 path, and the scorer had no idempotence check. Both are now enforced: the pareto artifact is reverted to the n=5 run, the duplicate rows are removed, `35_s5_pareto.py` refuses off-protocol seed counts and `36_score_s5.py` refuses to append rows that are already in the ledger | methods (post-hoc, declared) | docs/PREDICTION_OUTCOMES.csv; results/s5/pareto_mnist.json | scripts/35_s5_pareto.py, 36_score_s5.py, 53_resume_s8_decodes.sh | — | done; fourth entry in the off-protocol-execution family after row 27 | | 55 | **S8: the convergence sweep did not reach stationarity, so it cannot answer the question it was built to answer — executed as pre-committed.** Across {300, 1000, 3000, 10000} steps on both protocols: the gap moves 77.64 → 75.94, f(c_align) 0.502 → 0.489 → 0.470 → **0.459**, f(invariants) flat at 0.249–0.258. **P-S8-C fails**: the median relative endpoint gradient norm ratio between 300 and 10000 steps is **1.17** on P-random and **0.91** on P-shared-det, against the registered ≥10. S8 §4 pre-committed that if P-S8-C failed we say the sweep did not reach stationarity and do not read the accuracy numbers as though it had; that is executed verbatim in paper §13, the README and here. The **falsifier did not fire** (f = 0.459 ≫ 0.15), so no ladder claim is rescoped. Per the decline branch, f(c_align)'s **−0.043** over a 33× budget is reported as a real but small budget dependence and every ladder number is labelled as measured at the frozen 300-step config. Scores **6/8** intervals (H-S8-6 misses high at 6.3e−03 vs [3e−05, 3e−03]; H-S8-7 misses low at 0.197 vs [0.20, 0.60]) with P-S8-A and P-S8-B true, P-S8-C false | confirmatory (pre-registered S8, hash ab228c6deb526eed, 6/8 intervals; Brier 0.09 / 0.04 / 0.36) | results/s8/sweep.json | scripts/48_s8_sweep.py, 03_generate_inrbench.py, 53_resume_s8_decodes.sh | — | done; both interval misses are the stationarity diagnostics, i.e. the same finding read through the registration | | 56 | **The fits never leave θ₀'s neighbourhood at any budget, and the reason stationarity is unreachable is the optimizer.** Relative parameter travel from θ₀ is **0.186, 0.194, 0.194, 0.197** across a 33× budget — saturated by 1000 steps. So alignment to θ₀ keeps working not because the fits are under-trained but because they never leave θ₀'s neighbourhood, which measures directly the lazy regime row 10's mechanism infers from displacement. Fit quality is **not monotone** in budget (PSNR 37.4 → 64.6 → 62.4 → 58.4 dB; gradient norm 7.4e−03 → 2.8e−04 → 4.6e−03 → 6.3e−03, on both protocols), which the constant-learning-rate Adam fitter explains: Adam's step is a ratio of moments and does not shrink with the gradient, so past the end of descent the iterate diffuses in a band whose width is set by the learning rate. The 1000-step corpus catches the end of descent; larger budgets sample the band. **More steps cannot buy stationarity under this fitter** — a decaying schedule or a per-INR stationarity stopping rule is the repair, and is not run | exploratory (mechanism) + confirmatory for the travel figures | results/s8/sweep.json; src/sirengap/fitting/batched.py | scripts/48_s8_sweep.py | — | done; registered in docs/prereg/S8-addendum-02.md between the 3000- and 10000-step decodes and resolved 5/5 at mean Brier 0.054 | | 57 | **Methods finding: monitoring a long job exposed a registered quantity, because the generator and the scorer share one log.** `03_generate_inrbench.py` prints a per-shard median PSNR into `results/s8/run_master.log`, the file the decode also writes to. Checking that the resumed chain was alive therefore meant reading fitted values (57.6–59.3 dB) of a quantity that a call in S8-addendum-02 then predicted a [58, 66] dB band for. The call is **struck out and declared, not scored**; the other five calls in that addendum are clean, because gradient norm, travel and every f are computed only at decode time and appear nowhere in the log. Caught in the same session that created it, which is the first time in this family (cf. rows 27, 54 and S9-addendum-01). Generalizable fix, recorded as owed: separate the generator's progress log from the results log — not done here because changing the path mid-chain would have broken the resume | methods (post-hoc, declared) | docs/prereg/S8-addendum-02.md | — | — | done | | 58 | **The third arm closes the attribution: coordinates and architecture contribute comparably, and the pre-committed branch is 'split'.** W12b keeps W12's graded skeleton, block structure and edge coupling exactly, re-solves capacity by the same rule (width 186, **1,873,782** params, within 0.06% of W12), and changes only the feature map — the bias enters **raw** instead of phasor-lifted. **f(W12b) = 0.602** (62.34%). So the W11a → W12 step (0.265 → 0.917) splits as **+0.337 architecture** (W11a → W12b) and **+0.315 coordinates** (W12b → W12), with the layer-level grading's **+0.059** sitting inside the second. S10 §4 pre-committed the reading rule *before* the run: ≥0.75 withdraw row 51's claim, ≤0.55 keep it, in between report a split and do not pick the closer side. **0.602 is in between**, so row 51 is qualified, not vindicated: the coordinates do not 'carry the win', they carry about half of it. Two registered scope conditions: (a) the arm is handicapped by the grading it retains — a non-character feature reaches the head only through the bilinear rounds' even products — so +0.337 bounds the architecture's contribution *within the graded skeleton*, not in general; (b) it is emphatically non-invariant, and the audit shows the failure **growing with winding** (relative logit move 6.2, 1.7e2, 2.4e4 at |j|≤3, 10, 40 against W12's 3e−06), because a raw bias grows linearly in πj where its phasor does not. Its feature-level neutral block is exactly fixed under D∞ (0.00), so invariance dies precisely where the theory says: in the bilinear rounds | confirmatory (pre-registered S10, **3/3 intervals**, points 0.60/0.26/0.34 against observed 0.602/0.256/0.337; P-S10-A/B true, P-S10-C false; Brier 0.0225 / 0.0100 / 0.1225) | results/ladder/mnist/W12b.json, results/s10/verdict.json, results/audits/w12b_invariance_mnist.json | scripts/47_w12_phasor.py --raw-bias, 52_w12_invariance_audit.py --raw-bias, 56_score_s10.py; T18 | — | done; closes the limitation row 51 left open and qualifies row 51's reading | | 59 | **External review: the claims that outran the evidence are corrected, and three are withdrawn.** A referee report on the manuscript identified overclaims, all now fixed in the paper. **(a)** The title's "An Exact Decomposition for Periodic-Activation Networks" is withdrawn on both counts — no exact decomposition of the *natural* gap is established, and every proof uses sin's oddness and π-antiperiodicity, so it is sine-specific; the subtitle is now "Evidence from Sine Networks". **(b)** The orbit intervention is recast as **sufficiency**, not causal attribution: it shows scatter within the group can reproduce ~79 of 80 points, not that the natural gap is symmetry-mediated, since μ_B is a chosen distribution and equality of two effect magnitudes is not mediation. "The symmetry story is right about the cause" is deleted. **(c)** The 7.8-point residual is **no longer called an upper bound** on non-symmetry: no such estimand was defined, no monotonicity theorem supports it, and the unconstrained infimum over invariant readers is vacuous (a constant classifier gives 0). It is reported as the strongest invariant reader's shared-versus-random difference. **(d)** New notation: f (eq. 1) is valid only at fixed reader and training algorithm, so architecture comparisons now use a **reference-normalized score s(A)** (eq. 2) with the frozen-MLP anchors, and the two are never mixed. **(e)** W12's invariance is now **Proposition (prop:w12) with a proof** over the four layer types, answering the referee's question about the nonlinearity: GELU acts only on the neutral block, tanh (odd) on the other three, which is what preserves character covariance. Also corrected: the §7 capacity argument that §8 refutes (withdrawn in place), the "group-saturated" reading of a null (floor effect not excluded), orbit distances relabelled as best-found by the registered solver rather than global minima, "Adam on the mean" corrected to the sum and qualified by ε and float32, W1 no longer called a ceiling, W10's exact invariance scoped to the no-tie stratum, "destroys no class information" scoped to the render classifier, the 103× compute factor attributed to the strongest reader rather than all of weight space, and the vector-output step of Theorem 5 written out (a coordinatewise argument alone would appear to permit inconsistent signs across output coordinates) | correction (external review) | paper §1, §5, §7, §9, §12, §15, §18, App. A | — | — | done; the experiments the review asks for are registered as S11 rather than deferred | | 60 | **The 2×2 is complete and the two ingredients are additive: the interaction is +0.013.** W12ub (ungraded skeleton, raw bias — neither ingredient) reaches **f = 0.5565** at 1,885,284 params. The square: ungraded {raw 0.5565, phasor 0.8579}, graded {raw 0.6020, phasor 0.9165}. Phasor lift **+0.301** ungraded and **+0.315** graded; grading **+0.046** on raw bias and **+0.059** on phasor; **interaction +0.0131**. So the W11a→W12 step decomposes additively as **+0.291 skeleton, +0.301 phasor lift, +0.059 grading**, summing to 0.9165 — W12's exact value. This **supersedes row 58's +0.337/+0.315 split**, which measured 'architecture' along W11a→W12b and so folded the grading into it; with the fourth cell present, the skeleton's own contribution is +0.291 measured with both other ingredients off. Answers the external review's objection (§3.11) that the attribution was not factorial | confirmatory (pre-registered S11, 2/2 intervals so far: H-S11-1 point 0.55 vs observed 0.5565, H-S11-2 interaction point 0.00 vs observed 0.0131; P-S11-C resolved FALSE at Brier 0.09) | results/ladder/mnist/W12ub.json, results/s11/verdict.json | scripts/47_w12_phasor.py --ungraded --raw-bias, 58_score_s11.py | — | done; W12 on FashionMNIST, luminance CIFAR-10 and RGB CIFAR-10 still running under the same registration | | 61 | **W12 was designed on MNIST and holds on three corpora it never saw: it beats both families everywhere, including where they trade places.** Run unchanged (same capacity rule, optimizer, schedule, seed policy), s(W12) = **0.828** FashionMNIST, **0.826** luminance CIFAR-10, **0.965** RGB CIFAR-10, against 0.917 MNIST. Best exact reframing on the same corpora: 0.664 / 0.324 / 0.324; best invariant encoding: 0.428 / 0.493 / 0.534. **W12 is above both families on all four**, which is the sharp test, since the CIFAR corpora are exactly where reframing and encoding reverse (row 30). The crossover is therefore a property of *those two families*, not of weight-space reading in general. Invariance re-audited per corpus on its own fitted networks (max relative logit move 3.7e−06 at |j|≤40). Per S11 §3's pre-commitment, the manuscript's MNIST-only scoping is **lifted for this claim** and the corresponding limitation is discharged | confirmatory (pre-registered S11, **4/5 intervals**; P-S11-A true at Brier 0.04, P-S11-B true at 0.2025, P-S11-C false at 0.09) | results/ladder/{fashionmnist,cifar10gray,cifar10}/W12.json, results/s11/verdict.json, results/audits/w12_invariance_*.json | scripts/57_s11_chain.sh, 58_score_s11.py, 22_paper_tables.py | — | done; the single miss is H-S11-5, registered 0.60 [0.30, 0.85] against 0.965 observed — the fifth interval lost to under-predicting our own construction | | 62 | **Disclosure correction: the "external reviewer" was an AI system, and the paper now says so.** The acknowledgment described the source of the phasor route as "an anonymous external reviewer of an earlier draft", which reads as a human referee. It was a large language model acting as a reviewer. Since M7 is an *idea* contribution and not a request to run a control — and it is the idea the paper's strongest empirical result rests on — describing it as a human referee misstated the provenance of the central construction. Corrected in the paper (new "Use of AI systems" disclosure), PROVENANCE (M7, E3, E4, E6 and the header note) and the README. The disclosure also covers what the acknowledgment never did: language models were used to write and revise code in this repository and to draft and edit manuscript prose, including text that survives in the final version. The responsibility statement is unchanged — every theorem, proof, design, threshold and number was checked against artifacts by the author | correction (provenance/integrity) | paper Use-of-AI-systems paragraph; docs/PROVENANCE.md | — | — | done; raised as audit item I-8, unresolvable from the repository alone and confirmed by the author | | 63 | **S12: the converged-fit ladder was attempted and FAILED its own pre-registered validity gate. Nothing was decoded.** A cosine-decayed schedule with a per-INR stationarity stop at relative gradient norm 1e−4, three independently seeded replications of both protocols, 6 corpora. The fitter worked: median render PSNR **58 → 68–71 dB**, median relative gradient norm down **~50×** against S8's 10000-step arm to **1.03–1.36e−4** on five of six corpora. It still missed: the registered threshold was **1e−4** and 6/6 corpora sat above it (worst 7.34e−4 on P-shared-det r1, which also broke the ±2 dB render-matching condition at −2.16 dB). Per S12 §5 the gate blocked decoding and **no number from these corpora enters the paper**. The threshold was NOT loosened and the corpora were not re-run at a friendlier bar; the gate ran once, cold, from code committed before the first corpus existed. **P-S12-B** (0.70 that all conditions would be met first time) resolved FALSE at Brier **0.49**, the program's worst-scored call. Two defects recorded rather than quietly patched: (a) `stopped_at` is computed by the fitter but never written to shard metadata, so the artifacts cannot say how many INRs tripped the stop; (b) two shards took 6.2 h and 3.2 h against a 145 s baseline because **swap filled** (4.2/5.1 GB) after 11 h of continuous MPS allocation — not thermals — and a restart restored 100 s/shard. The watchdog detects dead processes, not degraded ones | confirmatory-null (pre-registered S12; gate failed, the failure is the result) | results/s12/gate.json, results/s12/verdict.json | scripts/59_s12_converged.sh, 60_score_s12.py, 61_s12_decode.sh | — | done; the paper's limitation now reports the attempt and how close it came rather than claiming it was never tried | | 64 | **W12 on the standard MNIST-INR benchmark: 93.08 ± 0.44, third of seven, ahead of DWSNets/NFN/NG-GNN and 3.5 points behind ScaleGMN.** The objection that W12 had only been compared to readers we built is closed. Run in its frozen configuration with no tuning against the benchmark, five seeds, on the corpus of Navon et al. that DWSNets, NFN, NG-GNN and ScaleGMN all report on and whose INRs have exactly our architecture. Leaderboard: ScaleGMN-B 96.59, ScaleGMN 96.57, **W12 93.08**, NG-GNN 91.40, DWSNets 85.71, NFN_HNP 79.11, NFN_NP 78.50. Per S13 §5's middle branch, the paper reports W12 as **competitive but not leading**, with the gap quoted, in the abstract and §15. **Two findings beyond the ranking.** (a) That benchmark is **independently initialized** — median pairwise relative parameter distance 1.419 vs our P-random 1.40 and shared-init 0.19 — so the published leaderboard is set in the hard regime and those methods' architectures, not their corpora, do the work. This paper emphasises corpus artifacts, and it is worth saying the field's main benchmark is not one. (b) W12 scores **93.1 there vs 87.6 on our P-random** at the same architecture and regime, so part of our ladder's difficulty is our fitting protocol rather than the problem — the same caution S12 raises from the other side. **Caveat recorded:** the runner's recovery_fraction (0.984) normalizes against OUR anchors and is meaningless on this corpus; absolute accuracy is the only comparable quantity | confirmatory (pre-registered S13, **3/3 intervals**; P-S13-A true 0.2025, P-S13-B true at Brier **0.64**, P-S13-C false 0.0049) | results/ladder/mnist/W12_dwsbench.json, results/s13/verdict.json | scripts/62_import_dws_benchmark.py, 47_w12_phasor.py, 63_bench_table.py | — | done; P-S13-B is the sixth call lost to under-predicting our own construction | | 65 | **The leaderboard row is complete: W12 places 3rd, 3rd and 2nd of seven, and beats ScaleGMN on CIFAR-10.** Run unchanged and frozen per dataset on the three standard INR-classification corpora: **MNIST 93.08 ± 0.44, FashionMNIST 75.15 ± 0.31, CIFAR-10 38.16 ± 0.22** (five seeds each). Above DWSNets, both NFN variants and NG-GNN on **all three**; above **ScaleGMN** (36.43) on CIFAR-10 by 1.7; below ScaleGMN-B everywhere (−3.5, −5.6, −0.66). The CIFAR gap to ScaleGMN-B is ~3 SD of our seed spread, so it is a real gap and not noise. Import provenance: FashionMNIST uses the **authors' own splits.json** (55k/5k/10k, ScaleGMN's exact protocol); MNIST and CIFAR-10 ship no split file, so their test halves are the authors' and validation is carved deterministically from training, declared wherever quoted. CIFAR's train/test division was **verified** rather than assumed (ids <50000 give exactly 5000/class, ≥50000 exactly 1000/class = CIFAR's own split). The 950k augmented INRs in that archive belong to ScaleGMN's separate Augmented CIFAR-10 column and were excluded. **Caveat repeated per cell:** the runner's recovery_fraction (0.984 / 0.889 / 0.806) normalizes against OUR anchors and is meaningless on these corpora; only absolute accuracy is comparable | confirmatory (S13 registered for MNIST; the two further cells are the same frozen configuration run on more corpora) | results/ladder/{mnist,fashionmnist,cifar10}/W12_dwsbench.json | scripts/62_import_dws_benchmark.py, 64_import_inr_benchmark2.py, 65_import_cifar_inr_benchmark.py, 63_bench_table.py | — | done; supersedes row 64's single-cell result | | 66 | **The channel-count crossover is NOT evidence about the grading: the (1,1) block's width contributes nothing.** Three arms on `cifar10`, frozen W12, n=5, only the output weights entering the (1,1) block changed. **A** (u full, width 4, full colour) f=0.9645; **B** (u→channel mean, width 2, less colour) f=0.7698; **C** (that mean padded back to 3 channels, width 4, same information as B) f=0.7557. **Width effect (C−B, information fixed) = −0.014** — zero, and slightly negative. **Information effect (A−C, width fixed) = +0.209** — fifteen times larger. Padding the block back to full width bought nothing, so the mechanism story (that the (1,1) block is the only one whose width grows with c, and that this is why W12 suits RGB) is **false**. All of the channel-count advantage is that RGB INRs carry more information for any reader. Per S14 §5, pre-committed: the crossover is not evidence about our grading, and **W12's CIFAR-10 win over ScaleGMN needs an explanation we do not have**. The three-arm design is what makes this conclusive — a single collapse shows −0.195 and would have supported the false story; the pad arm cost one hour and killed it. **Loose thread for the follow-up:** arm B lands 0.056 *below* the luminance corpus (0.826), so mangling u inside the reader hurts more than genuinely fitting grayscale images does, which means the luminance corpus recovers something through a route the collapse destroys | confirmatory (pre-registered S14, **3/4 intervals**; P-S14-A true 0.09, P-S14-B false 0.1225, P-S14-C false **0.36**) | results/s14/verdict.json, results/ladder/cifar10/W12_umean{,pad}.json | scripts/47_w12_phasor.py --u-mode, 68_s14_u_ablation.sh; T20 | — | done; H-S14-4 missed high (0.07 registered, 0.209 observed) — the seventh effect on our own construction we have under-predicted | | 67 | **S12 attempt 2: doubling the budget made stationarity ~3× WORSE, at identical fit quality. Second failure, reported not escalated.** Same 1e-4 bar, cosine schedule, per-INR stop, step cap 6000 → 12000, staged so replication 0 alone decided whether the rest were fitted. Median relative endpoint gradient norm **1.19e−4 → 3.66e−4** (P-shared-det) and **1.31e−4 → 3.79e−4** (P-random), while median render PSNR was **unchanged to the decimal** (69.3 and 70.2 dB in both attempts). Twice the budget bought fits sitting further from a stationary point that reconstruct the image exactly as well. Stopped at replication 0 per S12-addendum-01 rather than fitting two more or searching for a third schedule; the threshold was never loosened. **This strengthens rather than weakens the S8 conclusion**: constant-step Adam diffuses (S8), cosine-annealed Adam at 6k lands just above the bar (attempt 1), and at 12k lands further from it (attempt 2). Three schedules agree the obstruction is optimizer geometry on these objectives, not step count. Machine note: two of the three earlier throughput collapses were swap exhaustion, and one apparent stall was my own watchdog killing a healthy resuming job — see row 63 and scripts/69_stall_watchdog.sh | confirmatory-null (pre-registered S12 + addendum 01; gate failed, nothing decoded) | results/s12/gate_attempt2_stage.json | scripts/67_s12_attempt2.sh, 60_score_s12.py --tag s12b --reps 0 | — | done; the converged-fit question stays open and the paper says the remainder is not a budget problem |