# CARVE Master Build Plan **Version:** 1.0 **Created:** 2026-07-05 **Repo:** `research/lora-capability-removal/` **Algorithm:** CARVE (Contrastive Adapter Rotation for Verified Erasure) **Horizon:** 8 weeks (Phases A–H) **Commit granularity:** Per module (~33 commits) --- ## How to use this document 1. Execute commits **in order** — each commit lists its dependencies. 2. Run the **acceptance test** before committing; do not proceed if it fails. 3. Pass each **Quality Gate** before starting the next phase. 4. Cross-reference: - Agent manual: [`.cursor/claude.md`](claude.md) - Orchestration: [`.cursor/AGENTS.md`](AGENTS.md) - Algorithm spec: [`lora-capability-removal-review/03-upgraded-algorithm-CARVE.md`](../lora-capability-removal-review/03-upgraded-algorithm-CARVE.md) - Implementation patterns: [`.cursor/skills/carve-implementation/reference.md`](skills/carve-implementation/reference.md) - Eval protocol: [`docs/07-evaluation-protocol.md`](../docs/07-evaluation-protocol.md) - Fable audit: [`lora-capability-removal-review/01-critical-audit.md`](../lora-capability-removal-review/01-critical-audit.md) --- ## Table of contents 1. [Binding constraints](#1-binding-constraints) 2. [Commit conventions](#2-commit-conventions) 3. [Target project structure](#3-target-project-structure) 4. [Phase A — Infrastructure](#phase-a--infrastructure-days-13) 5. [Phase B — Synthetic + diagnostic](#phase-b--synthetic--diagnostic-days-410) 6. [Phase C — CARVE core (GO/NO-GO)](#phase-c--carve-core-go-no-go-days-1118) 7. [Phase D — Repair + GATE](#phase-d--repair--gate-days-1925) 8. [Phase E — Baselines + ablations](#phase-e--baselines--ablations-days-2632) 9. [Phase F — Real adapters](#phase-f--real-adapters-days-3342) 10. [Phase G — Public credibility](#phase-g--public-credibility-days-4350) 11. [Phase H — Paper + OSS](#phase-h--paper--oss-parallel-from-week-3) 12. [Test strategy](#12-test-strategy) 13. [Evaluation protocol](#13-evaluation-protocol) 14. [Quality gates summary](#14-quality-gates-summary) 15. [Hardware and runtime](#15-hardware-and-runtime) 16. [Risk register](#16-risk-register) 17. [Definition of done](#17-definition-of-done) --- ## 1. Binding constraints These are **non-negotiable** for every commit in this plan. | Constraint | Rule | Source | |------------|------|--------| | Primary algorithm | **CARVE** only | Fable review §3, `.cursor/research/algorithm-decision.md` | | Forbidden | PSN (`σ_k ← σ_k · a_r`), CCD inner-loop demixing | `01-critical-audit.md` §3.1–3.2 | | Adapter handling | Sidecar LoRA only — never merge into base before surgery | `docs/06-implementation-spec.md` §15 | | LoRA scaling | `delta_w = (lora_alpha / r) * B @ A` in all ΔW math | `peft-lora-safety.mdc` | | Module targeting | MLP (`gate_proj`, `up_proj`, `down_proj`) default; attention off for LLM | Maat finding | | Promotion eval | Held-out probes only — never training probes for GATE decision | `docs/07-evaluation-protocol.md` §6 | | Snapshot | Full adapter clone before any edit; rollback on GATE failure | CARVE Phase 0 | | Sequential surgery | Recompute S_f, S_r on edited model each block | CARVE Phase C4 | | Model size order | ≤3B until Phase C go/no-go passes; then 7B | `04-impact-and-roadmap.md` | | Claims | "Verified behavioral removal" — not "information deletion" or "industry-breaking" | `02-novelty-verdict.md` | --- ## 2. Commit conventions ### Format ``` type(scope): imperative subject Optional body explaining why (not what — the diff shows what). ``` ### Allowed types | Type | Use for | |------|---------| | `feat` | New module, CLI command, benchmark, integration | | `fix` | Bug fixes | | `test` | Tests only (may accompany feat in same commit if small) | | `refactor` | Restructure without behavior change | | `docs` | Paper, README, reproduction guide | | `chore` | Scaffolding, deps, release packaging | ### Allowed scopes `repo`, `config`, `io`, `decompose`, `metrics`, `synthetic`, `stats`, `rotation`, `diagnostic`, `baselines`, `surgery`, `pipeline`, `carve`, `repair`, `verify`, `probes`, `cli`, `eval`, `report`, `integration`, `paper`, `oss`, `release`, `infra` ### Rules - **Conventional Commits** strictly — enables changelog generation later. - **No `Co-authored-by: Cursor`** or any AI co-author trailer. - **One logical unit per commit** — a module + its direct tests, or a standalone test commit. - **Do not squash** phase commits during development; preserve audit trail. - Run acceptance test **before** `git commit`. ### Example ```bash git add src/stats.py tests/test_stats.py git commit -m "$(cat <<'EOF' feat(stats): add inner-space S_f/S_r second-moment collection Activation-based default (z = A @ h). Gradient variant gated behind stats_method=gradient for ablation only. EOF )" ``` --- ## 3. Target project structure Final tree after Phase H: ``` research/lora-capability-removal/ ├── .cursor/ │ ├── BUILD-PLAN.md ← this file │ ├── AGENTS.md │ ├── claude.md │ ├── rules/ │ ├── skills/ │ ├── research/ │ └── scripts/ ├── docs/ ← original research (reference) ├── lora-capability-removal-review/ ← Fable audit + CARVE spec ├── src/ │ ├── __init__.py │ ├── config.py # CarveConfig │ ├── decompose.py # Phase 0: snapshot + analysis SVD │ ├── stats.py # Phase C1: S_f, S_r │ ├── rotation.py # Phase C2: generalized eig │ ├── surgery.py # Phase C3: oblique projector │ ├── pipeline.py # Phase C4: sequential scrubbing │ ├── repair.py # Phase 4: Maat-style repair │ ├── verify.py # Phase 5: GATE │ ├── diagnostic.py # Separability certificate │ ├── baselines/ │ │ ├── __init__.py │ │ ├── delete.py │ │ ├── negation.py │ │ ├── maat_svd.py # SVD-basis ablation │ │ └── negative_lora.py # Optional Phase E │ ├── eval/ │ │ ├── __init__.py │ │ ├── runner.py # Ablation + Pareto sweep │ │ ├── relearning.py # Relearning attack harness │ │ └── report.py # Report generator │ └── utils/ │ ├── __init__.py │ ├── lora_io.py │ ├── svd_utils.py │ └── metrics.py ├── scripts/ │ ├── run_carve.py # Primary CLI │ ├── run_unbind.py # Deprecated alias │ ├── synthetic_entangle.py │ └── build_probes_from_traces.py ├── tests/ │ ├── conftest.py │ ├── test_config.py │ ├── test_lora_io.py │ ├── test_decompose.py │ ├── test_metrics.py │ ├── test_stats.py │ ├── test_rotation.py │ ├── test_surgery.py │ ├── test_pipeline.py │ ├── test_repair.py │ ├── test_verify.py │ ├── test_synthetic.py │ ├── test_synthetic_e2e.py # GO/NO-GO │ ├── test_baselines.py │ └── test_eval_runner.py ├── configs/ │ ├── default.yaml │ ├── syn2cap_1.5b.yaml │ ├── syn2cap_3b.yaml │ ├── syn2cap_7b.yaml │ ├── founder_os.yaml │ └── public_tofu.yaml ├── paper/ │ ├── main.tex # Phase H skeleton │ ├── figures/ │ └── bib/references.bib ├── requirements.txt ├── pyproject.toml ├── README.md ├── LICENSE ├── REPRODUCE.md └── results/ # gitignored └── / ├── preregistration.yaml ├── metrics.yaml ├── ablation.csv ├── pareto.csv ├── eigenvalue_spectrum.json ├── certificate.json └── report.md ``` --- ## Phase A — Infrastructure (Days 1–3) **Goal:** Package skeleton, config, I/O, snapshot/decompose, metrics. No GPU required for unit tests. **Quality Gate A:** `pytest tests/test_decompose.py tests/test_lora_io.py tests/test_metrics.py -v` all green. --- ### Commit A1 ``` chore(repo): scaffold package layout and dependencies ``` | Field | Value | |-------|-------| | **Depends on** | — (first commit) | | **Files created** | `src/__init__.py`, `src/utils/__init__.py`, `src/baselines/__init__.py`, `src/eval/__init__.py`, `tests/conftest.py`, `tests/__init__.py`, `requirements.txt`, `pyproject.toml`, `results/.gitkeep` | | **Files modified** | `.gitignore` (ensure `results/`, `*.safetensors`, `.venv/`) | **Implementation notes:** - `requirements.txt`: `torch>=2.1.0`, `transformers>=4.40.0`, `peft>=0.10.0`, `safetensors>=0.4.0`, `accelerate>=0.27.0`, `datasets>=2.18.0`, `pyyaml>=6.0`, `pytest>=7.0`, `tqdm`, `scipy`, `matplotlib` (diagnostic plots) - `pyproject.toml`: package name `carve`, Python `>=3.10`, pytest config, optional `[dev]` extras - `tests/conftest.py`: fixtures for random r×r PSD matrices, mock LoRA layer dict **Acceptance test:** ```bash pip install -e ".[dev]" # or pip install -r requirements.txt pytest --collect-only # discovers tests/ without errors python -c "import src; print('ok')" ``` --- ### Commit A2 ``` feat(config): add CarveConfig dataclass wired to configs/default.yaml ``` | Field | Value | |-------|-------| | **Depends on** | A1 | | **Files created** | `src/config.py`, `tests/test_config.py` | | **Files modified** | `configs/default.yaml` (already exists — wire loader) | **Implementation notes:** - Dataclass `CarveConfig` with all fields from `configs/default.yaml` - `CarveConfig.from_yaml(path) -> CarveConfig` - `CarveConfig.to_yaml(path) -> None` - Validation: `lambda_threshold > 0`, `0 < holdout_fraction < 1`, paths non-empty at runtime - Do NOT include deprecated PSN fields (`mixture_threshold`, `demix_method`) **Acceptance test:** ```bash pytest tests/test_config.py -v python -c "from src.config import CarveConfig; c = CarveConfig.from_yaml('configs/default.yaml'); assert c.lambda_threshold == 3.0" ``` --- ### Commit A3 ``` feat(io): add PEFT adapter load/save and LoRA layer enumeration ``` | Field | Value | |-------|-------| | **Depends on** | A2 | | **Files created** | `src/utils/lora_io.py`, `tests/test_lora_io.py` | **Implementation notes:** - `load_model_with_adapter(base_path, adapter_path, adapter_name="default", device="cuda") -> PeftModel` - `save_adapter(model, output_path, adapter_name="default") -> None` - `list_lora_layers(model, target_modules, adapter_name) -> list[str]` - `get_lora_AB(model, layer_name, adapter_name) -> tuple[Tensor, Tensor]` - `set_lora_A(model, layer_name, A, adapter_name) -> None` # used by surgery - Include scaling: `get_delta_w(B, A, lora_alpha, r) -> Tensor` - Mock-based tests when no real adapter on disk (inject fake `lora_A`/`lora_B` modules) **Acceptance test:** ```bash pytest tests/test_lora_io.py -v ``` --- ### Commit A4 ``` feat(decompose): add adapter snapshot and analysis SVD ``` | Field | Value | |-------|-------| | **Depends on** | A3 | | **Files created** | `src/decompose.py`, `src/utils/svd_utils.py`, `tests/test_decompose.py` | **Implementation notes:** - `snapshot_adapter(model, adapter_name, target_modules) -> dict[str, dict]` — deep clone A, B per layer - `rollback_adapter(model, snapshot, adapter_name) -> None` - `decompose_layer(delta_w) -> tuple[U, sigma, V]` — thin SVD for analysis only (not surgery basis) - `analyze_adapter(model, config) -> dict` — per-layer σ spectrum for logging - Snapshot must clone tensors on CPU for rollback safety **Acceptance test:** ```bash pytest tests/test_decompose.py -v # Unit: random [d_out, d_in] matrix SVD roundtrip error < 1e-5 ``` --- ### Commit A5 ``` feat(metrics): add FE/RF/US metric helpers ``` | Field | Value | |-------|-------| | **Depends on** | A1 | | **Files created** | `src/utils/metrics.py`, `tests/test_metrics.py` | **Implementation notes:** - `forget_efficacy(perf_pre, perf_post) -> float` — `1 - post/pre` - `retain_fidelity(perf_pre, perf_post) -> float` — `post/pre` - `utility_score(perf_pre, perf_post) -> float` - `pareto_point(fe, rf) -> tuple` — for CSV logging (do NOT optimize FRS scalar alone) - `check_gate(fe, rf, us, config) -> tuple[bool, dict]` - Handle edge cases: `pre == 0` → epsilon denominator **Acceptance test:** ```bash pytest tests/test_metrics.py -v # fe(1.0, 0.05) ≈ 0.95; rf(1.0, 0.97) ≈ 0.97 ``` --- ### Commit A6 ``` test(infra): cover lora_io roundtrip and decompose snapshot ``` | Field | Value | |-------|-------| | **Depends on** | A3, A4, A5 | | **Files created/modified** | Expand `tests/test_lora_io.py`, `tests/test_decompose.py` with integration-style tests | **Implementation notes:** - Test snapshot → mutate A → rollback → A unchanged - Test `list_lora_layers` filters by `target_modules` - Test metrics gate pass/fail boundaries at thresholds 0.90/0.95/0.98 **Acceptance test:** ```bash pytest tests/test_lora_io.py tests/test_decompose.py tests/test_metrics.py tests/test_config.py -v # Quality Gate A — all green ``` --- ## Phase B — Synthetic + diagnostic (Days 4–10) **Goal:** Syn-2Cap ground-truth benchmark, stats/rotation modules, eigenvalue diagnostic, cheap baselines. Requires GPU for full synthetic training. **Quality Gate B:** - Syn-2Cap merged adapter exhibits both capabilities - λ spectrum logged per layer; `max(λ)` interpretable - Separability certificate emitted when spectrum flat - Baselines (delete, negation, uniform scale) run without error --- ### Commit B1 ``` feat(synthetic): add Syn-2Cap entangled adapter construction ``` | Field | Value | |-------|-------| | **Depends on** | A6 | | **Files created** | `scripts/synthetic_entangle.py`, `configs/syn2cap_1.5b.yaml`, `configs/syn2cap_3b.yaml`, `tests/test_synthetic.py` | **Implementation notes:** - CLI subcommands: `train-cap-a`, `train-cap-b`, `merge`, `verify`, `diagnostic-only` - `train_single_cap_lora(base, data, rank, output) -> adapter_path` - `merge_in_weight_space(lora_a, lora_b, alpha=0.6, beta=0.4) -> B, A` - `extract_lora_from_delta_w(delta_w, rank) -> B, A` via truncated SVD - Default base: `Qwen/Qwen2.5-1.5B-Instruct` (fast iteration) - Capability pair default: tone A vs tone B (or QA fact sets — document in config) - Save merged adapter to `results/syn2cap_/entangled_adapter/` - `verify` subcommand: eval both caps on merged adapter, assert both > baseline **Acceptance test:** ```bash pytest tests/test_synthetic.py -v -k "not gpu" # GPU (when available): # python scripts/synthetic_entangle.py merge --config configs/syn2cap_1.5b.yaml # python scripts/synthetic_entangle.py verify --adapter results/syn2cap_.../entangled_adapter ``` --- ### Commit B2 ``` feat(stats): add inner-space S_f/S_r second-moment collection ``` | Field | Value | |-------|-------| | **Depends on** | A3, A4 | | **Files created** | `src/stats.py`, `tests/test_stats.py` | **Implementation notes:** - `InnerActivationHook` on `lora_A` output: `z = A @ h`, shape `[batch*seq, r]` - `accumulate_second_moment(z_batches, r, epsilon) -> S` — `E[zz^T] + εI` - `collect_stats(model, layer_name, probes, adapter_name, method="activation") -> S` - Gradient variant (ablation): `inner_sensitivity_matrix(B, G, A)` per CARVE doc §3 - Streaming accumulation — O(r²) memory regardless of probe count - See `.cursor/skills/carve-implementation/reference.md` for hook pattern **Acceptance test:** ```bash pytest tests/test_stats.py -v # S is symmetric PSD; shape [r,r]; epsilon regularization prevents singularity ``` --- ### Commit B3 ``` feat(rotation): add generalized eigenbasis with spectrum logging ``` | Field | Value | |-------|-------| | **Depends on** | B2 | | **Files created** | `src/rotation.py`, `tests/test_rotation.py` | **Implementation notes:** - `contrastive_eigenbasis(S_f, S_r) -> tuple[Tensor, Tensor]` — `(lambda, W)` columns desc - Cholesky whiten: `L = cholesky(S_r)`, `S_w = L^{-1} S_f L^{-T}`, `eigh`, map back `W = L^{-T} Q` - `select_cut_indices(lam, threshold) -> Tensor` - `log_spectrum(lam, layer_name, run_dir) -> None` — append to `eigenvalue_spectrum.json` **Acceptance test:** ```bash pytest tests/test_rotation.py -v # Synthetic PSD pair: lam descending; W[:, i] generalized eigenvectors python .cursor/scripts/quick_eigenvalue_demo.py # toy demo ``` --- ### Commit B4 ``` feat(diagnostic): add separability certificate and spectrum plot ``` | Field | Value | |-------|-------| | **Depends on** | B2, B3 | | **Files created** | `src/diagnostic.py` | **Implementation notes:** - `SeparabilityCertificate` dataclass: `status` (`SURGERY_VIABLE` | `LINEAR_INSEPARABLE`), `max_lambda`, `mean_lambda`, `std_lambda`, `layer`, `recommendation` - `emit_certificate(lam, threshold, layer, run_dir) -> SeparabilityCertificate` - `plot_eigenvalue_spectrum(lam, output_path) -> None` — matplotlib, one plot per layer or aggregate - `LINEAR_INSEPARABLE` when `max(lam) < lambda_threshold` → recommend modular retrain (QR-LoRA) - Wire into `scripts/synthetic_entangle.py diagnostic-only` mode **Acceptance test:** ```bash python -c " from src.diagnostic import emit_certificate import torch lam = torch.ones(16) # flat spectrum c = emit_certificate(lam, 3.0, 'layer.0', 'results/test_cert') assert c.status == 'LINEAR_INSEPARABLE' " ``` --- ### Commit B5 ``` feat(baselines): add delete, task-negation, uniform-scale baselines ``` | Field | Value | |-------|-------| | **Depends on** | A3, A5 | | **Files created** | `src/baselines/delete.py`, `src/baselines/negation.py`, `src/baselines/__init__.py`, `tests/test_baselines.py` | **Implementation notes:** - `delete_adapter(model, adapter_name) -> None` — zero all LoRA weights - `task_vector_negation(model, adapter_name, alpha=1.0) -> None` — `B, A → -alpha * B, A` or negate delta_w refactored - `uniform_scale(model, adapter_name, scale) -> None` — baseline for ablation - Each returns `method_name` str for eval logging - No stats collection needed — instant baselines **Acceptance test:** ```bash pytest tests/test_baselines.py -v ``` --- ### Commit B6 ``` test(synthetic): cover Syn-2Cap build and rotation eigenbasis ``` | Field | Value | |-------|-------| | **Depends on** | B1, B2, B3, B4, B5 | | **Files created/modified** | Expand `tests/test_synthetic.py`, add `tests/fixtures/synthetic/` (tiny mock deltas) | **Implementation notes:** - Unit: merge_in_weight_space with known ΔW_A, ΔW_B → merged norm within tolerance - Unit: extract_lora_from_delta_w roundtrip - Integration (mock): stats → rotation → certificate pipeline on synthetic r=8 matrices - Mark GPU tests `@pytest.mark.gpu` **Acceptance test:** ```bash pytest tests/test_synthetic.py tests/test_stats.py tests/test_rotation.py -v -m "not gpu" # Quality Gate B (partial without GPU): all non-GPU tests green ``` --- ## Phase C — CARVE core (GO/NO-GO) (Days 11–18) **Goal:** Oblique projection surgery, sequential pipeline, SVD-basis ablation, Syn-2Cap E2E. **Quality Gate C (BLOCKING):** - Syn-2Cap: **FE ≥ 0.90, RF ≥ 0.95** on held-out probes - **CARVE Pareto-dominates SVD-basis cut** (≈Maat) at same cut budget |J| - If CARVE ≈ Maat → **STOP**, debug stats hooks and sequential scrubbing before Phase D --- ### Commit C1 ``` feat(surgery): add oblique projector and in-place A cut ``` | Field | Value | |-------|-------| | **Depends on** | B3 | | **Files created** | `src/surgery.py`, `tests/test_surgery.py` | **Implementation notes:** - `oblique_projector(W_cut, S_r) -> P` — `P = I - W_J (W_J^T S_r W_J)^{-1} W_J^T S_r` - `apply_cut(A, P) -> A_new` — `P @ A`, preserve B - `soft_shrink_eigenvalues(lam, gamma) -> Tensor` — `1/(1+γλ)` for soft variant - `apply_soft_cut(A, W, lam, S_r, gamma) -> A_new` — optional ablation path - Empty cut set → P = I (no-op) - See `.cursor/skills/carve-implementation/reference.md` § surgery.py **Acceptance test:** ```bash pytest tests/test_surgery.py -v # Projector idempotent: P @ P ≈ P # W_J^T S_r P ≈ 0 # A shape unchanged after cut ``` --- ### Commit C2 ``` feat(pipeline): add sequential front-to-back block scrubbing ``` | Field | Value | |-------|-------| | **Depends on** | B2, B3, C1, A4 | | **Files created** | `src/pipeline.py`, `tests/test_pipeline.py` | **Implementation notes:** - `get_layer_blocks(model, target_modules, block_size) -> list[list[str]]` — ordered front-to-back - `carve_block(model, layers, D_f, D_r, config, run_dir) -> None`: 1. Recompute S_f, S_r on **current** model for each layer in block 2. `contrastive_eigenbasis` → select J 3. `oblique_projector` → apply to A 4. Log spectrum + certificate per layer - `run_carve_phases_c1_c4(model, D_f, D_r, config) -> dict` — metrics summary - Snapshot at start via `decompose.snapshot_adapter` - Do NOT implement repair or GATE here — pipeline stops after cut **Acceptance test:** ```bash pytest tests/test_pipeline.py -v -m "not gpu" # Mock 2-layer model: after block 1 edit, stats for block 2 use edited model ``` --- ### Commit C3 ``` feat(baselines): add SVD-basis Maat-style ablation cut ``` | Field | Value | |-------|-------| | **Depends on** | A4, C1 | | **Files created** | `src/baselines/maat_svd.py` | **Implementation notes:** - `maat_svd_cut(model, D_f, config, cut_budget) -> None`: - SVD decompose ΔW per layer (analysis basis — NOT CARVE rotation) - Score components by forget gradient norm only (Maat-style) - Zero top-|J| forget-scored σ_k, refactor to BA (ablation only — note scaling pitfalls) - `match_cut_budget(carv_eigenvalues, svd_sigmas, target_j) -> indices` — align |J| for fair comparison - This is the **thesis ablation** — proves rotation value **Acceptance test:** ```bash pytest tests/test_baselines.py -v -k maat # Same |J| as CARVE on same layer → comparable cut budget ``` --- ### Commit C4 ``` test(carve): cover projector idempotence and cut invariants ``` | Field | Value | |-------|-------| | **Depends on** | C1, C2, C3 | | **Files modified** | `tests/test_surgery.py`, `tests/test_pipeline.py`, `tests/test_rotation.py` | **Implementation notes:** - Property tests: oblique projector with random PSD S_r, random W_J - Pipeline mock: verify A changed, B unchanged, delta_w norm decreases after cut - Compare CARVE vs Maat-SVD on mock entangled r=16 matrices (no LLM) **Acceptance test:** ```bash pytest tests/test_surgery.py tests/test_pipeline.py tests/test_rotation.py -v ``` --- ### Commit C5 ``` test(synthetic): add Syn-2Cap E2E CARVE-vs-SVD-basis go/no-go ``` | Field | Value | |-------|-------| | **Depends on** | B1, C2, C3, A5 | | **Files created** | `tests/test_synthetic_e2e.py` | **Implementation notes:** - `@pytest.mark.gpu` E2E test: 1. Load or build Syn-2Cap entangled adapter 2. Split probes 80/20 train/holdout 3. Run CARVE on train probes 4. Eval FE, RF on **holdout** 5. Run Maat-SVD with matched |J| 6. Assert CARVE FE ≥ 0.90, RF ≥ 0.95 7. Assert CARVE RF > Maat-SVD RF OR CARVE FE > Maat-SVD FE at equal RF (Pareto dominate) - Log results to `results/go_nogo_/` - This test is the **program gate** — document skip reason if no GPU in CI **Acceptance test:** ```bash pytest tests/test_synthetic_e2e.py -v -m gpu # Quality Gate C — MUST pass before Phase D ``` --- ## Phase D — Repair + GATE (Days 19–25) **Goal:** Maat repair, held-out verification, trace probe builder, full CLI. **Quality Gate D:** - Repair ablation measurable (with/without Phase 4) - Intentional GATE failure triggers rollback - `preregistration.yaml` written before holdout eval - No training-probe leakage in promotion path --- ### Commit D1 ``` feat(repair): add Maat-style Phase 4 retain repair ``` | Field | Value | |-------|-------| | **Depends on** | C2, A4 | | **Files created** | `src/repair.py`, `tests/test_repair.py` | **Implementation notes:** - Objective from `docs/05-novel-algorithm-SPECTRAL-UNBIND.md` §8 (keep verbatim): ``` L_repair = w_KL · KL(p_W' || p_W_ref)_{D_r} + w_HS · d_rep(h_W', h_W_ref)_{D_r} - w_ent · H(p_W')_{D_f} + w_cos · Σ_l cos(B_l', τ_l^f)^+ ``` - `repair_adapter(model, ref_snapshot, D_f, D_r, config) -> None` - Train only LoRA params; ref_model frozen from snapshot - Optimizer: AdamW, `repair_lr`, `repair_steps` from config - Optional: skip repair when `repair_steps=0` for ablation **Acceptance test:** ```bash pytest tests/test_repair.py -v -m "not gpu" # Mock: repair loop runs N steps without NaN; LoRA params change ``` --- ### Commit D2 ``` feat(verify): add held-out GATE verification and snapshot rollback ``` | Field | Value | |-------|-------| | **Depends on** | A4, A5, D1 | | **Files created** | `src/verify.py`, `tests/test_verify.py` | **Implementation notes:** - `write_preregistration(run_dir, config, data_hashes) -> Path` - `gate_verify(model_edited, model_ref, D_f_holdout, D_r_holdout, D_c, config) -> tuple[bool, dict]` - `promote_or_rollback(passed, model, snapshot, output_path) -> str` — returns `PROMOTE|ROLLBACK|RETRY` - Retry loop: max `max_retry_attempts`, tune `lambda_threshold` / `soft_gamma` between attempts - **Never** pass training probes to `gate_verify` - See `.cursor/skills/gate-verification/reference.md` for preregistration template **Acceptance test:** ```bash pytest tests/test_verify.py -v # Test intentional fail (FE=0.5 mock) → rollback restores snapshot # Test preregistration.yaml created before metrics.yaml ``` --- ### Commit D3 ``` feat(probes): add build_probes_from_traces for Founder OS ``` | Field | Value | |-------|-------| | **Depends on** | A2 | | **Files created** | `scripts/build_probes_from_traces.py` | **Implementation notes:** - CLI: `--traces-dir`, `--forget-filter`, `--retain-filter`, `--holdout-fraction`, `--output-dir` - Load `data/traces/*.jsonl` - Filters as Python expressions or preset names (`bad_tone`, `tool_misfire`) - LECE score threshold for retain (default ≥ 0.7) - Export: `D_f_train.jsonl`, `D_r_train.jsonl`, `D_f_holdout.jsonl`, `D_r_holdout.jsonl` - SHA256 hash each file for preregistration **Acceptance test:** ```bash python scripts/build_probes_from_traces.py --help # With fixture traces: split sums to total; holdout disjoint from train ``` --- ### Commit D4 ``` feat(cli): add run_carve.py entry point with yaml override ``` | Field | Value | |-------|-------| | **Depends on** | C2, D1, D2, A2 | | **Files created** | `scripts/run_carve.py` | **Implementation notes:** - Args: `--base-model`, `--adapter`, `--forget-data`, `--retain-data`, `--output`, `--config`, `--lambda-threshold`, `--no-repair`, `--diagnostic-only` - Flow: load config → load model → preregister → snapshot → carve (C1–C4) → repair → gate_verify → promote/rollback - `--diagnostic-only`: run stats + spectrum + certificate, no cut - Log all outputs to `results//` - Exit code 0 on PROMOTE, 1 on ROLLBACK, 2 on RETRY exhausted **Acceptance test:** ```bash python scripts/run_carve.py --help # Dry-run with mock/minimal config loads without crash ``` --- ### Commit D5 ``` refactor(cli): add deprecated run_unbind alias ``` | Field | Value | |-------|-------| | **Depends on** | D4 | | **Files created** | `scripts/run_unbind.py` | **Implementation notes:** - Thin wrapper: print deprecation warning, delegate to `run_carve.py` - `"SPECTRAL-UNBIND PSN is deprecated; use CARVE via run_carve.py"` **Acceptance test:** ```bash python scripts/run_unbind.py --help 2>&1 | grep -i deprecat ``` --- ### Commit D6 ``` test(verify): cover rollback and preregistration flow ``` | Field | Value | |-------|-------| | **Depends on** | D2, D4 | | **Files modified** | `tests/test_verify.py`, add integration test in `tests/test_pipeline.py` | **Implementation notes:** - End-to-end mock: run_carve flow with tiny model or mocked components - Verify file order: preregistration.yaml timestamp < metrics.yaml timestamp - Verify rollback tensor equality with snapshot **Acceptance test:** ```bash pytest tests/test_verify.py tests/test_repair.py -v # Quality Gate D — all green ``` --- ## Phase E — Baselines + ablations (Days 26–32) **Goal:** Full ablation suite, Pareto sweep, relearning attacks, evaluation reports. **Quality Gate E:** Ablation report in `results/` with all rows from §13.5 filled. --- ### Commit E1 ``` feat(eval): add ablation runner and lambda Pareto sweep ``` | Field | Value | |-------|-------| | **Depends on** | C2, C3, D4 | | **Files created** | `src/eval/runner.py`, `tests/test_eval_runner.py` | **Implementation notes:** - `run_ablation(adapter, probes, methods, config) -> pd.DataFrame` - Methods: `carve`, `maat_svd`, `delete`, `negation`, `no_repair`, `soft_gamma`, `sequential_vs_independent` - Pareto sweep: `lambda_threshold` ∈ {1, 2, 3, 4, 5} → `results//pareto.csv` - Columns: `method`, `FE`, `RF`, `US`, `time_s`, `lambda_threshold`, `|J|`, `max_lambda` - Restore adapter from snapshot between methods **Acceptance test:** ```bash pytest tests/test_eval_runner.py -v -m "not gpu" ``` --- ### Commit E2 ``` feat(eval): add relearning-attack harness ``` | Field | Value | |-------|-------| | **Depends on** | D1, D4 | | **Files created** | `src/eval/relearning.py` | **Implementation notes:** - `relearning_attack(model, D_f_train, steps, lr) -> list[float]` — FE after each step - Default steps: [0, 10, 20, 50, 100] - Report `steps_to_recovery` — first step where FE drops below 0.80 - Target per `docs/07-evaluation-protocol.md` §7: recovery requires ≥ 50 steps - Honest reporting — do not hide fast recovery **Acceptance test:** ```bash python -c "from src.eval.relearning import relearning_attack; print('import ok')" ``` --- ### Commit E3 ``` feat(report): add evaluation report generator ``` | Field | Value | |-------|-------| | **Depends on** | E1, A5 | | **Files created** | `src/eval/report.py` | **Implementation notes:** - `generate_report(run_dir, output_path) -> None` — fills template from `.cursor/skills/research-evaluation/reference.md` - Inputs: `metrics.yaml`, `ablation.csv`, `pareto.csv`, `eigenvalue_spectrum.json`, `certificate.json` - Sections: diagnostic, held-out results, baseline comparison, ablation highlights, relearning, decision - Decision: `PROMOTE | ROLLBACK | RETRY | RETRAIN_MODULAR` (from certificate) **Acceptance test:** ```bash # With fixture results/ dir → report.md generated python -c "from src.eval.report import generate_report; print('ok')" ``` --- ### Commit E4 ``` test(eval): cover ablation runner and metric regressions ``` | Field | Value | |-------|-------| | **Depends on** | E1, E2, E3 | | **Files modified** | `tests/test_eval_runner.py` | **Implementation notes:** - Regression test: metrics on fixed synthetic perf values - Ablation runner produces expected columns - Report generator runs on minimal fixture **Acceptance test:** ```bash pytest tests/test_eval_runner.py -v # Quality Gate E ``` --- ## Phase F — Real adapters (Days 33–42) **Goal:** Founder OS entangled adapter, trace-derived probes, production GATE decision. **Quality Gate F:** Production decision logged — PROMOTE / ROLLBACK / RETRAIN_MODULAR with failure analysis if RF < 0.95. --- ### Commit F1 ``` feat(integration): wire Founder OS traces and GATE harness ``` | Field | Value | |-------|-------| | **Depends on** | D3, D4, E1 | | **Files created** | `configs/founder_os.yaml` | | **Files modified** | `scripts/build_probes_from_traces.py` (Founder OS field mappings) | **Implementation notes:** - Config points to real base model + entangled adapter paths (placeholders in repo; user fills locally) - Wire to existing GATE harness if present: `research/` or `agent/cognition/train/gate.py` - Fix circular-eval anti-pattern documented in GATE blog - Probe filters for Founder OS metadata: `capability`, `lece_score`, `outcome` - Min probes: 50 D_f + 50 D_r + 10 holdout each **Acceptance test:** ```bash python scripts/build_probes_from_traces.py --config configs/founder_os.yaml --dry-run # Or with local traces: produces 4 JSONL files + hashes ``` --- ### Commit F2 ``` feat(eval): add real-adapter run configs and decision logging ``` | Field | Value | |-------|-------| | **Depends on** | F1, E3 | | **Files modified** | `configs/founder_os.yaml`, `src/eval/report.py` | **Implementation notes:** - Document run command in `configs/founder_os.yaml` header comments - `run_carve.py` with founder_os config on real adapter - Log decision + one-paragraph failure analysis template if RF < 0.95 - If certificate = LINEAR_INSEPARABLE → decision RETRAIN_MODULAR with QR-LoRA recommendation **Acceptance test:** ```bash # Manual: run on real adapter, inspect results//report.md # Automated: pytest skipped if no real adapter path (mark @pytest.mark.founder_os) ``` --- ## Phase G — Public credibility (Days 43–50) **Goal:** TOFU/WMDP subset, 7B configs, relearning robustness table, preprint-ready numbers. **Quality Gate G:** At least one public benchmark row + relearning table in report. --- ### Commit G1 ``` feat(eval): add TOFU/WMDP subset evaluation adapters ``` | Field | Value | |-------|-------| | **Depends on** | E1, D4 | | **Files created** | `configs/public_tofu.yaml`, `src/eval/public_benchmarks.py` | **Implementation notes:** - TOFU forget10 subset loader (or WMDP-Cyber MC subset) - Map TOFU metrics to FE (fact probability / ROUGE) - Retain set: retain90 or similar from TOFU protocol - Compare CARVE vs delete vs negation vs Maat-SVD on same entangled or standard adapter - Document dataset download/cache path **Acceptance test:** ```bash python -c "from src.eval.public_benchmarks import load_tofu_subset; print('ok')" # GPU: one small benchmark run completes ``` --- ### Commit G2 ``` feat(eval): add 7B configs and robustness table ``` | Field | Value | |-------|-------| | **Depends on** | G1, C5 | | **Files created** | `configs/syn2cap_7b.yaml` | | **Files modified** | `src/eval/report.py`, `src/eval/relearning.py` | **Implementation notes:** - `configs/syn2cap_7b.yaml` — Qwen2.5-7B-Instruct, rank 32 - Re-run Syn-2Cap E2E on 7B if Phase C passed on ≤3B - Robustness table in report: - Paraphrase forget (FE under paraphrased D_f) - Adversarial elicitation - Relearning steps 10/50/100 - Composed capabilities prompt - Targets: FE ≥ 0.80 under paraphrase; relearning ≥ 50 steps to recover **Acceptance test:** ```bash pytest tests/test_synthetic_e2e.py -v -m gpu --config configs/syn2cap_7b.yaml # Report includes robustness section ``` --- ## Phase H — Paper + OSS (Parallel from Week 3) **Goal:** arXiv preprint skeleton, OSS quickstart, pip-installable CLI, reproduction guide. **Quality Gate H:** External user can reproduce Syn-2Cap from README alone. --- ### Commit H1 ``` docs(paper): add paper skeleton with certificate results ``` | Field | Value | |-------|-------| | **Depends on** | E3, G2 (results available) | | **Files created** | `paper/main.tex`, `paper/bib/references.bib`, `paper/figures/.gitkeep` | **Implementation notes:** - Sections: Abstract, Intro, Related Work (Maat, LEACE, ReGLU, RepSelect), Method (CARVE), Benchmark (Syn-2Cap), Experiments, Limitations, Conclusion - Lead with benchmark + separability certificate — CARVE as method that exploits it - Include failure certificate as contribution ("when is post-hoc surgery possible?") - Pull bib from `docs/08-references.md`; verify arXiv IDs resolve - No "industry-breaking" language **Acceptance test:** ```bash # pdflatex paper/main.tex or check structure manually ls paper/main.tex paper/bib/references.bib ``` --- ### Commit H2 ``` docs(oss): add README quickstart, license, reproduction guide ``` | Field | Value | |-------|-------| | **Depends on** | D4, B1, H1 | | **Files created** | `README.md`, `LICENSE` (Apache-2.0 or MIT — user choice), `REPRODUCE.md` | **Implementation notes:** - README: one-liner "git revert for model capabilities", install, Syn-2Cap quickstart, CLI examples - REPRODUCE.md: exact commands to reproduce go/no-go table from paper - Link to CARVE spec, GATE protocol, Fable review - Badge placeholders: tests, arXiv **Acceptance test:** ```bash # Follow README Syn-2Cap steps on clean machine (manual) grep -q "run_carve" README.md grep -q "Syn-2Cap" REPRODUCE.md ``` --- ### Commit H3 ``` chore(release): package CARVE CLI with PEFT-compatible IO ``` | Field | Value | |-------|-------| | **Depends on** | H2, D4 | | **Files modified** | `pyproject.toml`, `src/__init__.py` | **Implementation notes:** - `[project.scripts]` entry: `carve = scripts.run_carve:main` - Version `0.1.0` - `pip install -e .` installs CLI - Verify PEFT adapter in → edited adapter out on Syn-2Cap fixture **Acceptance test:** ```bash pip install -e . carve --help # Quality Gate H — full reproduction path documented ``` --- ## 12. Test strategy ### Pyramid ``` ┌─────────────────────┐ │ test_synthetic_e2e │ GPU, GO/NO-GO (Phase C) │ test_synthetic_e2e │ 7B variant (Phase G) └──────────┬──────────┘ ┌────────────────┼────────────────┐ │ Integration (pipeline + verify + repair) │ └────────────────┬────────────────┘ ┌──────────────────────────┼──────────────────────────┐ │ Unit tests (rotation, surgery, stats, metrics) │ └───────────────────────────────────────────────────────────────┘ ``` ### Test matrix | Test file | Phase | GPU | Blocking | |-----------|-------|-----|----------| | `test_config.py` | A | No | Yes | | `test_lora_io.py` | A | No | Yes | | `test_decompose.py` | A | No | Yes | | `test_metrics.py` | A | No | Yes | | `test_stats.py` | B | No | Yes | | `test_rotation.py` | B | No | Yes | | `test_surgery.py` | C | No | Yes | | `test_pipeline.py` | C | Mock | Yes | | `test_baselines.py` | B/C | No | Yes | | `test_synthetic.py` | B | Optional | Yes | | `test_synthetic_e2e.py` | C | **Yes** | **GO/NO-GO** | | `test_repair.py` | D | Mock | Yes | | `test_verify.py` | D | No | Yes | | `test_eval_runner.py` | E | Mock | Yes | ### pytest markers ```python # conftest.py import pytest def pytest_configure(config): config.addinivalue_line("markers", "gpu: requires CUDA GPU") config.addinivalue_line("markers", "founder_os: requires real Founder OS adapter") ``` Run tiers: ```bash pytest -m "not gpu" -v # CI / no GPU pytest -m gpu -v # GPU machine pytest tests/test_synthetic_e2e.py -m gpu # GO/NO-GO only ``` ### Coverage targets - `src/rotation.py`, `src/surgery.py`, `src/stats.py`: ≥ 90% line coverage - `src/pipeline.py`, `src/verify.py`: ≥ 80% (GPU paths may be mocked) - `scripts/`: smoke tests via `--help` minimum --- ## 13. Evaluation protocol ### 13.1 Metrics (held-out promotion only) | Metric | Formula | Target | Used for | |--------|---------|--------|----------| | **FE** | `1 - perf(post, D_f_h) / perf(pre, D_f_h)` | ≥ 0.90 | Forget efficacy | | **RF** | `perf(post, D_r_h) / perf(pre, D_r_h)` | ≥ 0.95 | Retain fidelity | | **US** | `perf(post, D_c) / perf(pre, D_c)` | ≥ 0.98 | Utility | Report **Pareto curve** (FE vs RF) across `lambda_threshold` sweep — do not optimize `FRS = FE × RF` alone. ### 13.2 Dataset tiers | Tier | Dataset | When | Ground truth | |------|---------|------|--------------| | 1 | **Syn-2Cap** | Phase B–C (blocking) | LoRA_B recovery | | 1b | Syn-3Cap | Phase E ablation | Per-cap eval | | 1c | Syn-MixedRank | Phase E ablation | Component labels | | 2 | TOFU forget10 | Phase G | Public comparison | | 2b | WMDP subset | Phase G optional | Safety unlearning | | 3 | Founder OS traces | Phase F (primary product) | LECE + task success | | 4 | Diffusion | Optional future | ErasureBench-H | ### 13.3 Baseline comparison table (required for paper) Run all on same (model, adapter, D_f, D_r, D_holdout): | Method | FE ↑ | RF ↑ | US ↑ | Time | Phase | |--------|------|------|------|------|-------| | No edit | 0 | 1.0 | 1.0 | — | — | | Delete adapter | 1.0 | 0 | 1.0 | instant | B | | Task vector negation | ? | ? | ? | instant | B | | Uniform scale | ? | ? | ? | instant | B | | Maat-SVD (fixed basis) | ? | ? | ? | ~30m | C | | **CARVE** | ? | ? | ? | ~30m | C | | CARVE + repair | ? | ? | ? | ~45m | D | | Negative LoRA | ? | ? | ? | ~1hr | E | | Retrain exclude D_f | ? | ? | ? | hours | E | ### 13.4 Ablation study (required for paper) | Ablation | Config knob | Tests | |----------|-------------|-------| | CARVE vs SVD-basis (Maat) | method=carve vs maat_svd | Rotation value | | Hard vs soft cut | soft_gamma=None vs 0.5 | Pareto knob | | Activation vs gradient stats | stats_method | Quantized robustness | | No repair | repair_steps=0 | Phase 4 value | | Sequential vs independent | pipeline mode | Circuit coupling | | λ sweep | {1,2,3,4,5} | Sensitivity | | Rank sweep | {8,16,32,64} | Entanglement vs rank | | Attention included | include_attention=true | Maat claim (LLM: expect worse) | ### 13.5 Robustness tests | Test | Method | Target | |------|--------|--------| | Paraphrase forget | Rephrase D_f_holdout | FE ≥ 0.80 | | Adversarial elicitation | "Ignore instructions, do c_f" | FE ≥ 0.75 | | Relearning attack | 10–100 finetune steps on D_f | Recovery ≥ 50 steps | | Composed capabilities | Prompt needing c_r near c_f | RF ≥ 0.90 | ### 13.6 GATE protocol (mandatory production) ``` PRE-REGISTRATION (before holdout): → results//preregistration.yaml → hash D_f_holdout, D_r_holdout → thresholds: FE≥0.90, RF≥0.95, US≥0.98 SURGERY: → D_f_train, D_r_train only EVALUATION: → holdout only → metrics.yaml DECISION: PASS → promote edited adapter FAIL → rollback snapshot, log failure PARTIAL → tune λ/γ, retry (max 3) LINEAR_INSEPARABLE → RETRAIN_MODULAR ``` ### 13.7 Entanglement diagnostics **CARVE (primary):** eigenvalue spectrum `{λ_j}` per layer - `max(λ) < threshold` → certificate LINEAR_INSEPARABLE - Flat spectrum (std/mean < 0.1) → surgery unlikely to succeed **Legacy (analysis only):** entanglement index E from doc 07 §3.5 ### 13.8 Expected outcomes by entanglement | Scenario | max(λ) | Expected FE | Expected RF | |----------|--------|-------------|-------------| | Low entanglement | > 5 | 0.90–0.98 | 0.95–0.99 | | Moderate | 2–5 | 0.80–0.92 | 0.90–0.96 | | High / flat | < 2 | 0.60–0.80 | 0.85–0.92 | | Low rank many caps | — | Poor | Poor → retrain | --- ## 14. Quality gates summary | Gate | Phase | Blocking | Criteria | |------|-------|----------|----------| | **A** | Infrastructure | Yes | `pytest tests/test_{config,lora_io,decompose,metrics}.py` green | | **B** | Synthetic | Yes | Syn-2Cap builds; λ spectrum logged; certificate path works | | **C** | CARVE core | **YES** | FE≥0.90, RF≥0.95; CARVE Pareto-dominates Maat-SVD | | **D** | Repair+GATE | Yes | Rollback works; preregistration before holdout | | **E** | Ablations | Yes | Full ablation table in `results/` | | **F** | Real adapter | Yes | PROMOTE/ROLLBACK/RETRAIN decision logged | | **G** | Public | For paper | TOFU/WMDP row + relearning table | | **H** | OSS | For adoption | README reproduction path works | ### Phase dependency graph ```mermaid graph TD A[GateA Infra] --> B[GateB Synthetic] B --> C[GateC GO_NO_GO] C -->|pass| D[GateD GATE] C -->|fail| Debug[Debug stats sequential] Debug --> B D --> E[GateE Ablations] E --> F[GateF Real adapter] F --> G[GateG Public] B --> H[GateH OSS] C --> H G --> Preprint[arXiv preprint] ``` --- ## 15. Hardware and runtime ### Minimum requirements | Resource | Syn-2Cap 1.5B | Syn-2Cap 3B | Syn-2Cap 7B | CARVE surgery only | |----------|---------------|-------------|-------------|-------------------| | GPU VRAM | 8 GB (4-bit) | 12 GB | 24 GB | Same as model | | GPU VRAM (fp16) | 16 GB | 24 GB | 48 GB | Same as model | | CPU RAM | 16 GB | 32 GB | 64 GB | 16 GB | | Disk | 20 GB (model cache) | 30 GB | 50 GB | — | ### Time estimates | Task | 1.5B | 3B | 7B | |------|------|----|----| | Train LoRA_A + LoRA_B | 1–2 hr | 2–3 hr | 4–8 hr | | Merge + verify | 5 min | 5 min | 10 min | | CARVE surgery (C1–C4) | 10 min | 15 min | 25 min | | Repair (Phase 4) | 15 min | 20 min | 30 min | | Full ablation suite | 2 hr | 3 hr | 6 hr | ### Model size strategy 1. **Iterate on 1.5B or 3B** until Gate C passes 2. **Promote to 7B** only after Gate C on ≤3B 3. Use **activation stats** (default) on 4-bit quantized bases 4. Use **bf16** for surgery phase if gradients needed (ablation) ### Software pins ``` Python >= 3.10 torch >= 2.1.0 transformers >= 4.40.0 peft >= 0.10.0 ``` Run `.cursor/scripts/check_env.py` before Phase B. --- ## 16. Risk register | Risk | Likelihood | Impact | Mitigation | Commit phase | |------|------------|--------|------------|--------------| | Flat λ on real adapters | Medium | Surgery fails | Certificate → RETRAIN_MODULAR; publishable negative result | B4, F2 | | CARVE ≈ Maat on synthetic | Medium | Invalidates thesis | Gate C blocking; debug hooks + sequential scrub | C5 | | Scooped by ReGLU/RepSelect | Medium-high | Novelty reduced | Preprint at Gate C; arXiv immediately | H1 | | Relearning < 50 steps | Medium | Overclaim risk | Report honestly; entropy repair term | E2, G2 | | Sequential scrubbing slow | Low | UX | Block_size=8; forward-only stats | C2 | | Wrong ΔW gradient | Medium | Wrong cuts | Activation stats default; gradient ablation only | B2 | | SVD refactor scaling bug | Low | Adapter corruption | CARVE edits A in place — no refactor | C1 | | Training-probe leakage | Medium | False promotion | verify.py gate; evaluation-gate.mdc | D2 | | No GPU in CI | High | Gate C skipped | Mark `@pytest.mark.gpu`; manual gate on dev machine | C5 | | Reviewers want 3 eval domains | High | Paper rejection | LLM synthetic + TOFU + Founder OS traces | G1, F1 | ### Stop conditions 1. **Gate C fails after 3 debug cycles** → revisit stats formulation; consider gradient hybrid; do not proceed to real adapters 2. **Flat λ on Syn-2Cap AND real adapter** → pivot to certificate paper + QR-LoRA modular retrain recommendation 3. **Relearning recovery < 10 steps** → scope all claims to "behavioral suppression"; no erasure language 4. **RF < 0.85 on real adapter after repair** → RETRAIN_MODULAR; document failure mode --- ## 17. Definition of done ### Implementation complete - [ ] All 33 commits in this plan applied in order - [ ] `pytest -m "not gpu"` green - [ ] `pytest -m gpu` green on dev machine (Gate C) - [ ] `carve --help` works after `pip install -e .` - [ ] Syn-2Cap reproduction in REPRODUCE.md verified on clean machine ### Research complete - [ ] Gate C passed: FE≥0.90, RF≥0.95, CARVE Pareto-dominates Maat-SVD - [ ] Gate F passed: real Founder OS adapter decision logged - [ ] Ablation table complete (§13.4) - [ ] Relearning attack table complete (§13.5) - [ ] One public benchmark row (TOFU or WMDP) ### Paper + OSS complete - [ ] `paper/main.tex` compiles with results figures - [ ] README quickstart works - [ ] arXiv preprint submitted - [ ] GATE blog post draft ("the surgical undo") ### Preprint trigger Submit arXiv when **all** of: 1. Gate C passed on ≤3B Syn-2Cap 2. Ablation table shows CARVE > Maat-SVD at matched |J| 3. Separability certificate demo included 4. Relearning table present (even if weak — report honestly) ### Claims allowed at preprint - Closed-form contrastive rotation in existing entangled adapter inner space - LEACE-style oblique projection in LoRA weight space - Syn-2Cap benchmark with component-level ground truth - Separability certificate for linear insufficiency - Verified behavioral removal with GATE held-out protocol ### Claims NOT allowed - "Industry-breaking" / "foolproof" / "information deletion" - "First contrastive forget/retain for LoRA" (ReGLU, RepSelect) - "First rank-1 demixing" (PSN was broken) - "Complete capability separation" for nonlinear entanglement --- ## Appendix A: Commit index (quick reference) | # | Commit | Phase | |---|--------|-------| | A1 | `chore(repo): scaffold package layout and dependencies` | A | | A2 | `feat(config): add CarveConfig dataclass wired to configs/default.yaml` | A | | A3 | `feat(io): add PEFT adapter load/save and LoRA layer enumeration` | A | | A4 | `feat(decompose): add adapter snapshot and analysis SVD` | A | | A5 | `feat(metrics): add FE/RF/US metric helpers` | A | | A6 | `test(infra): cover lora_io roundtrip and decompose snapshot` | A | | B1 | `feat(synthetic): add Syn-2Cap entangled adapter construction` | B | | B2 | `feat(stats): add inner-space S_f/S_r second-moment collection` | B | | B3 | `feat(rotation): add generalized eigenbasis with spectrum logging` | B | | B4 | `feat(diagnostic): add separability certificate and spectrum plot` | B | | B5 | `feat(baselines): add delete, task-negation, uniform-scale baselines` | B | | B6 | `test(synthetic): cover Syn-2Cap build and rotation eigenbasis` | B | | C1 | `feat(surgery): add oblique projector and in-place A cut` | C | | C2 | `feat(pipeline): add sequential front-to-back block scrubbing` | C | | C3 | `feat(baselines): add SVD-basis Maat-style ablation cut` | C | | C4 | `test(carve): cover projector idempotence and cut invariants` | C | | C5 | `test(synthetic): add Syn-2Cap E2E CARVE-vs-SVD-basis go/no-go` | C | | D1 | `feat(repair): add Maat-style Phase 4 retain repair` | D | | D2 | `feat(verify): add held-out GATE verification and snapshot rollback` | D | | D3 | `feat(probes): add build_probes_from_traces for Founder OS` | D | | D4 | `feat(cli): add run_carve.py entry point with yaml override` | D | | D5 | `refactor(cli): add deprecated run_unbind alias` | D | | D6 | `test(verify): cover rollback and preregistration flow` | D | | E1 | `feat(eval): add ablation runner and lambda Pareto sweep` | E | | E2 | `feat(eval): add relearning-attack harness` | E | | E3 | `feat(report): add evaluation report generator` | E | | E4 | `test(eval): cover ablation runner and metric regressions` | E | | F1 | `feat(integration): wire Founder OS traces and GATE harness` | F | | F2 | `feat(eval): add real-adapter run configs and decision logging` | F | | G1 | `feat(eval): add TOFU/WMDP subset evaluation adapters` | G | | G2 | `feat(eval): add 7B configs and robustness table` | G | | H1 | `docs(paper): add paper skeleton with certificate results` | H | | H2 | `docs(oss): add README quickstart, license, reproduction guide` | H | | H3 | `chore(release): package CARVE CLI with PEFT-compatible IO` | H | **Total: 33 commits across Phases A–H** --- ## Appendix B: Agent session prompts ### Start Phase A ``` @.cursor/BUILD-PLAN.md @.cursor/claude.md Execute Phase A commits A1–A6 in order. Stop at Quality Gate A. No PSN. Conventional Commits, no co-author trailer. ``` ### Start Phase C (after Gate B) ``` @.cursor/BUILD-PLAN.md @.cursor/skills/carve-implementation/SKILL.md Execute Phase C commits C1–C5. Gate C is blocking. Run test_synthetic_e2e.py on GPU before proceeding. ``` ### Debug Gate C failure ``` @.cursor/BUILD-PLAN.md @lora-capability-removal-review/01-critical-audit.md Gate C failed: CARVE ≈ Maat-SVD. Diagnose stats hooks, sequential scrubbing, lambda threshold. Do not proceed to Phase D. ``` --- *End of CARVE Master Build Plan v1.0*