# CARVE: Contrastive Adapter Rotation for Verified Erasure **Replaces:** SPECTRAL-UNBIND Phases 1–3 (attribution, classification, PSN/CCD surgery) **Keeps:** Phase 0 (snapshot), Phase 4 (Maat-style repair), Phase 5 (GATE verification), the entire evaluation protocol, and the probe-building infrastructure. --- ## 1. The Core Insight (Why CARVE Fixes What PSN Broke) SPECTRAL-UNBIND's fatal assumption: capabilities align with the **singular vectors of ΔW**. They don't — SVD directions maximize update variance, not capability separation. So the MIXED bucket is large, and scaling mixed σₖ (PSN) damages retain exactly as much as it removes forget. CARVE's move: **don't accept the SVD basis. Find the basis that separates.** The adapter's update lives in an r-dimensional inner space (r = LoRA rank). Any invertible r×r transform of that space leaves ΔW's function unchanged — it's a gauge freedom: ``` ΔW = B A = (B R⁻¹)(R A) for any invertible R ∈ ℝ^{r×r} ``` So we are free to *rotate* the inner space into whatever basis is most useful **before** cutting. The most useful basis is the one where forget-relevance and retain-relevance are maximally separated — and finding it is a classical generalized eigenvalue problem (the same math as Fisher discriminant analysis and contrastive PCA), solvable in closed form on r×r matrices. r is 8–64. This costs microseconds. One sentence: **Maat cuts along whatever basis SVD happens to give; CARVE first rotates to the capability-discriminating basis, then cuts.** --- ## 2. Notation Per LoRA layer l (dropping subscript l): - B ∈ ℝ^{d_out×r}, A ∈ ℝ^{r×d_in}, ΔW = s·BA with s = lora_alpha/r - Inner space: ℝ^r, the space "between" A and B - D_f, D_r: forget / retain probe sets (as in original docs) - G(x) = ∂L/∂ΔW ∈ ℝ^{d_out×d_in}: loss gradient w.r.t. the update, for input x --- ## 3. Phase C1: Contrastive Second-Moment Statistics (replaces CRA scalars) Instead of scalar per-component scores, build **matrices** that capture the full geometry of how each capability uses the inner space. Pull the gradient into the inner space. For direction w ∈ ℝ^r (unit vector), the rank-1 sub-update "through w" is s·(Bw)(wᵀA). Its first-order effect on the loss for input x: ``` g(x, w) = s · (Bw)ᵀ G(x) (wᵀA)ᵀ = s · wᵀ [Bᵀ G(x) Aᵀ] w ``` Define the **inner-space sensitivity matrix** per example: ``` M(x) = sym( Bᵀ G(x) Aᵀ ) ∈ ℝ^{r×r} (sym(X) = (X+Xᵀ)/2) ``` Accumulate second moments over each probe set: ``` S_f = E_{x∈D_f} [ M(x) M(x)ᵀ ] + ε I (forget sensitivity structure) S_r = E_{x∈D_r} [ M(x) M(x)ᵀ ] + ε I (retain sensitivity structure) ``` Both are r×r, symmetric PSD. They encode not just *how much* each inner direction matters to each capability (diagonal), but *how directions co-vary* (off-diagonal) — which is precisely the entanglement information PSN threw away. **Practical alternative (gradient-free):** replace M(x) with activation-based statistics — inner activations z(x) = A·h(x) ∈ ℝ^r collected on D_f vs D_r, S_f = Cov(z|D_f), S_r = Cov(z|D_r). This is exactly the LEACE/cPCA setup transplanted to the adapter bottleneck, is robust on quantized models, and needs forward passes only. Recommended default; use gradient version as ablation. --- ## 4. Phase C2: The Contrastive Rotation (closed form — the novel core) Find inner directions maximally forget-relevant *relative to* retain-relevance. Solve the **generalized eigenvalue problem**: ``` S_f w = λ S_r w ``` Equivalently: whiten by retain (S_r = LLᵀ Cholesky; S̃_f = L⁻¹ S_f L⁻ᵀ), eigendecompose S̃_f = Q Λ Qᵀ, map back W = L⁻ᵀ Q. Columns w₁…w_r ordered by λ₁ ≥ … ≥ λ_r. Interpretation of λⱼ = (wⱼᵀS_f wⱼ)/(wⱼᵀS_r wⱼ): - **λⱼ ≫ 1:** direction carries forget signal with little retain cost → **cut** - **λⱼ ≈ 1:** genuinely mixed → this is the *irreducible* entanglement; cutting always costs retain. Decide by Pareto preference, not pretend-demixing. - **λⱼ ≪ 1:** retain-dominant → **keep untouched** This replaces SPECTRAL-UNBIND's three-way classification with a **principled, continuous spectrum** where the threshold has a direct meaning (forget-per-unit-retain-damage ratio). And crucially: because the basis was *optimized for separation*, the mixed region is as small as linear algebra allows — provably no rotation separates better (Courant–Fischer). **Diagnostic for free:** the eigenvalue spectrum {λⱼ} IS the entanglement measurement. Flat spectrum (all λ≈1) → linearly inseparable, surgery will fail, retrain modularly — you know *before* cutting. This makes the honest "when will this fail" story quantitative. --- ## 5. Phase C3: The Cut — Oblique Projection (replaces PSN/CCD) Select the cut set J = {j : λⱼ ≥ λ_thresh} (default λ_thresh ≈ 2–4, sweep for Pareto curve). Build the oblique (S_r-orthogonal) projector removing those directions from the inner space: ``` P = I_r − W_J (W_Jᵀ S_r W_J)⁻¹ W_Jᵀ S_r where W_J = [wⱼ]_{j∈J} ``` (Using the S_r inner product means the removal does least damage as measured by retain sensitivity — the LEACE "least-squares" property, here in the inner space.) Apply as a gauge-respecting edit: ``` A' = P A (or equivalently B' = B P', symmetric variant) ΔW' = s · B A' ``` No SVD refactoring needed, no re-scaling pitfalls, no rank bookkeeping — A is edited in place, B untouched, PEFT format preserved. (Optionally also *negate* rather than zero: A' = (I − (1+α)·proj)A, Maat-style, as an ablation.) **Soft variant:** instead of hard cut, shrink each direction by factor 1/(1+γλⱼ) — gives a smooth Pareto knob γ instead of a threshold. --- ## 6. Phase C4: Sequential Scrubbing (fixes the cross-layer flaw) Edit layers **front-to-back**. After editing layer l, recompute statistics for layer l+1 **on the already-edited model** (LEACE's concept-scrubbing procedure, transplanted). Cost: one extra forward-stats pass per edited layer group; in practice batch by blocks of 4–8 layers. Then run Phase 4 repair (Maat hybrid objective) and Phase 5 GATE verification exactly as the original docs specify. --- ## 7. Full Pseudocode ``` Algorithm CARVE Input: adapter {A_l, B_l}, base W0, D_f, D_r, λ_thresh, block_size Output: edited adapter snapshot ← copy(adapter) for each block of layers (front to back): # statistics on the CURRENT (partially edited) model for x in D_f ∪ D_r: record inner stats (activation z_l(x) = A_l h_l(x), or gradient M_l(x)) for each layer l in block: S_f, S_r ← second moments (+ εI) (λ, W) ← generalized_eig(S_f, S_r) # r×r, closed form J ← {j : λ_j ≥ λ_thresh} if J empty: continue P ← I − W_J (W_Jᵀ S_r W_J)⁻¹ W_Jᵀ S_r A_l ← P A_l repair(adapter, D_r, D_f) # Maat Phase-3 objective (unchanged) if not GATE_verify(holdouts): rollback(snapshot) return adapter ``` Complexity: dominated by |D_f|+|D_r| forward (or forward+backward) passes per block — same order as SPECTRAL-UNBIND, minus its per-component inner optimization loops. Everything else is r×r linear algebra. --- ## 8. What You Can Prove (Paper Ammunition) 1. **Optimal-basis lemma:** among all rank-|J| cuts of the inner space, the generalized-eigenvector cut maximizes removed forget-sensitivity per unit retain-sensitivity (Courant–Fischer on the S_r-whitened pencil). PSN/Maat's SVD-basis cut is strictly suboptimal whenever S_f, S_r don't commute with ΔWᵀΔW — i.e., whenever there's real entanglement. 2. **Linear-guarding corollary (LEACE analog):** after the cut, no linear functional of the adapter's inner representation restricted to span(W_J) carries forget signal; with the activation-statistics variant this inherits LEACE's guarantee at the adapter bottleneck. 3. **Failure certificate:** if max λⱼ < λ_thresh, no linear surgery in this adapter achieves the target forget/retain ratio — output "retrain modularly" with a number attached. Turning "we can't" into a certificate is itself a contribution. Honest limits (state them in the paper): guarantees are linear and local (first/second-order); nonlinear recovery and relearning attacks can restore behavior — measure with the relearning protocol; guarantees degrade if D_f/D_r are unrepresentative. --- ## 9. CARVE vs SPECTRAL-UNBIND, Side by Side | | SPECTRAL-UNBIND | CARVE | |---|---|---| | Basis for surgery | Fixed SVD of ΔW | Learned contrastive rotation (generalized eigvecs) | | Attribution | Scalar per SVD component | Full r×r second-moment geometry | | Mixed directions | PSN scaling (broken) / CCD (ill-posed) | Provably-optimal linear cut + honest irreducibility certificate | | Cross-layer | Independent (wrong) | Sequential scrubbing | | Closed form | Partially | Fully (stats + eig + projector) | | Failure prediction | Post-hoc entanglement index | Pre-cut eigenvalue spectrum certificate | | PEFT compatibility | SVD refactor with scaling pitfalls | In-place edit of A | | Theory | Assumption sketch | Optimality lemma + guarding corollary | --- ## 10. Implementation Delta (vs original doc 06) Reuse ~70% of the implementation spec. Changes: - `attribution.py` → collect z = A·h (hook on lora_A output) per probe; accumulate S_f, S_r streaming (LEACE-fitter style, O(r²) memory — trivial) - `classify.py` → replace with `rotation.py`: `scipy.linalg.eigh(S_f, S_r)` (or Cholesky-whiten + `torch.linalg.eigh`) - `surgery.py` → replace PSN/CCD with the oblique projector; delete SVD refactor path - `pipeline.py` → add block-sequential scrubbing loop - Everything else (probe builder, repair, verify, CLI, synthetic benchmark) unchanged Name for the system paper: **CARVE** standalone, or **GATE-CARVE** inside the Founder OS story.