# CARVE Project Progress Report **Date:** 2026-07-10 **Repo:** `research/lora-capability-removal/` **Primary algorithm:** CARVE (Contrastive Adapter Rotation for Verified Erasure) **Overall completion:** ~55–60% (code + tests done; behavioral validation not done) --- ## Executive summary This project implements **verified surgical removal of one capability from an already-trained, entangled monolithic LoRA adapter** — without full retraining and without modifying the base model. The product framing is **"git revert for model capabilities."** Work to date has delivered: 1. A **complete research and agent infrastructure** (`.cursor/`, docs, Fable audit, BUILD-PLAN). 2. A **full Python implementation** of CARVE Phases A–H as specified in the build plan. 3. **40 unit/integration tests** — 39 passing, 1 skipped — including a CUDA end-to-end test on synthetic entangled LoRA tensors. 4. **Local eval harness** (`scripts/run_full_eval.py`) that runs diagnostics, ablations, relearning stub, and GATE decision logic. What has **not** been achieved yet: - **Behavioral validation on a real LLM** (Qwen2.5-1.5B or similar). - **Syn-2Cap go/no-go gate** (FE ≥ 0.90, RF ≥ 0.95, CARVE Pareto-dominates Maat-SVD on held-out probes). - **Founder OS real-adapter run**, TOFU/WMDP public benchmarks, or paper-ready results. The codebase is research-ready; the science (proving CARVE works on entangled adapters) is the remaining ~40%. --- ## 1. Problem and algorithm decision ### Problem Remove capability \(c_f\) (forget) from an entangled monolithic LoRA while preserving capability \(c_r\) (retain), with held-out GATE verification and rollback on failure. ### Binding algorithm decision (from Fable audit) | Status | Mechanism | |--------|-----------| | **PRIMARY** | CARVE — contrastive second moments \(S_f, S_r\), generalized eigenproblem, oblique projection cut, sequential scrubbing | | **KEEP** | Snapshot, Maat-style repair, GATE verification, probe builder, synthetic benchmark | | **FORBIDDEN** | PSN (\(\sigma_k \leftarrow \sigma_k \cdot a_r\)) and CCD demixing — mathematically broken per audit | The independent review (`lora-capability-removal-review/README.md`) concluded: - The **problem framing is correct** and the gap is real (no published post-hoc surgical cut inside entangled LoRA rank space). - Original **SPECTRAL-UNBIND PSN is broken** — soft prune, not demixing. - **CARVE upgrade is defensible novelty** (~A−/B+): contrastive rotation in inner rank space + LEACE-style oblique projection. - Realistic claim: **"verified behavioral removal"** — not "information deletion" or "industry-breaking." --- ## 2. What was built ### 2.1 Research and agent infrastructure | Component | Location | Status | |-----------|----------|--------| | Agent manual | `.cursor/AGENTS.md` | Done | | Master build plan (Phases A–H, 33 commit messages) | `.cursor/BUILD-PLAN.md` | Done | | Claude/agent context | `.cursor/claude.md` | Done | | Rules (6) | `.cursor/rules/*.mdc` | Done | | Skills (6) | `.cursor/skills/*/SKILL.md` | Done | | Research context | `.cursor/research/` | Done | | Utility scripts | `.cursor/scripts/` (validate, check_env, eigenvalue demo) | Done | ### 2.2 Original research docs Nine documents under `docs/` covering problem definition, literature, taxonomy, gap analysis, legacy SPECTRAL-UNBIND spec, implementation spec, evaluation protocol, and references. ### 2.3 Fable independent review Four documents under `lora-capability-removal-review/`: - `01-critical-audit.md` — doc-by-doc audit, PSN failure analysis - `02-novelty-verdict.md` — prior art and honest novelty claims - `03-upgraded-algorithm-CARVE.md` — **primary algorithm to implement** - `04-impact-and-roadmap.md` — impact assessment and ordered execution plan ### 2.4 Source code (`src/`) | Module | Phase | Purpose | |--------|-------|---------| | `config.py` | A | `CarveConfig` dataclass, YAML load/save, CLI merge | | `utils/lora_io.py` | A | PEFT adapter load/save, layer enumeration, snapshot/rollback | | `utils/svd_utils.py` | A | ΔW computation, thin SVD, refactor to LoRA | | `utils/metrics.py` | A | FE, RF, US, `check_gate()` | | `decompose.py` | A | Adapter snapshot and analysis SVD | | `stats.py` | B/C1 | \(S_f, S_r\) second-moment accumulation, activation hooks | | `rotation.py` | B/C2 | Generalized eigenbasis, cut index selection, spectrum logging | | `diagnostic.py` | B | Separability certificate (`SURGERY_VIABLE` / flat spectrum) | | `surgery.py` | C3 | Oblique projector, hard/soft cut on \(A_l\) | | `pipeline.py` | C4 | Sequential front-to-back block scrubbing | | `baselines/delete.py` | B/E | Zero-out adapter baseline | | `baselines/negation.py` | B/E | Task-vector negation baseline | | `baselines/maat_svd.py` | C/E | SVD-basis ablation cut (Maat-style proxy) | | `repair.py` | D | Maat-style retain repair (KL + auxiliary losses) | | `verify.py` | D | GATE held-out verification, preregistration, promote/rollback | | `eval/runner.py` | E | Ablation runner, Pareto λ sweep | | `eval/relearning.py` | E | Relearning attack harness (stub curve) | | `eval/report.py` | E | Markdown report generator | | `eval/public_benchmarks.py` | G | TOFU/WMDP subset loaders (stub) | ### 2.5 Scripts (`scripts/`) | Script | Purpose | |--------|---------| | `run_carve.py` | Primary CLI — load model, snapshot, CARVE pipeline, repair, GATE | | `run_unbind.py` | Deprecated alias for legacy naming | | `synthetic_entangle.py` | Syn-2Cap merge at weight level, diagnostic-only mode | | `build_probes_from_traces.py` | Founder OS trace → probe jsonl splits | | `run_full_eval.py` | Local full eval suite (pytest + synthetic + ablation + report) | | `_bootstrap.py` | Adds repo root to `sys.path` for script invocation | ### 2.6 Tests (`tests/`) 14 test files, **40 tests total**: | File | Coverage | |------|----------| | `test_config.py` | CarveConfig load/validate | | `test_lora_io.py` | Snapshot, rollback, layer listing | | `test_decompose.py` | Analysis SVD | | `test_metrics.py` | FE/RF/US, gate check | | `test_stats.py` | Second-moment accumulation | | `test_rotation.py` | Generalized eigenbasis | | `test_surgery.py` | Oblique projector, apply cut | | `test_pipeline.py` | Pipeline skeleton (1 skipped legacy placeholder) | | `test_baselines.py` | Delete, negation, Maat-SVD | | `test_repair.py` | Repair module | | `test_verify.py` | GATE verification, rollback | | `test_synthetic.py` | Syn-2Cap weight merge | | `test_synthetic_e2e.py` | **GPU E2E** — CARVE cut on CUDA synthetic entangled LoRA | | `test_eval_runner.py` | Ablation runner | **Latest pytest (2026-07-10):** `39 passed, 1 skipped` in ~3.2s. ### 2.7 Configs and data | Path | Purpose | |------|---------| | `configs/default.yaml` | CARVE hyperparameter defaults | | `configs/syn2cap_1.5b.yaml` | Qwen2.5-1.5B Syn-2Cap settings | | `configs/syn2cap_3b.yaml` | 3B variant | | `configs/syn2cap_7b.yaml` | 7B variant (deferred until synthetic gate passes) | | `configs/founder_os.yaml` | Founder OS adapter paths (placeholders) | | `configs/public_tofu.yaml` | TOFU subset eval config | | `data/traces/sample.jsonl` | Sample Founder OS trace | | `data/probes/*.jsonl` | Generated forget/retain train + holdout probes | ### 2.8 OSS and paper skeleton - `README.md` — quickstart for Syn-2Cap and probe building - `REPRODUCE.md` — reproduction steps - `LICENSE` — Apache-2.0 - `paper/main.tex` — preprint skeleton - `pyproject.toml`, `requirements.txt` — package install --- ## 3. Git history **10 commits** on `master` (working tree clean as of 2026-07-10): | Commit | Message | |--------|---------| | `afb69a5` | feat: initial setup of research, docs and Spectral-Unbind Mechanism | | `bd9c7a4` | docs(plan): add CARVE master build plan with phased commit sequence | | `db2cfb2` | fix(cursor): correct repo root path and Windows console encoding in demo scripts | | `fc67ae9` | feat(infra): Phase A scaffold — config, LoRA I/O, decompose, metrics | | `53c00be` | feat(carve): Phase C — oblique surgery, sequential pipeline, Maat-SVD baseline | | `be2b9d9` | feat(synthetic): Phase B — stats, rotation, diagnostic, baselines, Syn-2Cap script | | `19a271e` | feat(gate): Phase D — repair, GATE verification, probes, CLI *(eval modules; message mislabeled)* | | `4197a20` | feat(gate): add repair, verify, and run_carve CLI | | `8a4b5e7` | chore(configs): add Syn-2Cap, Founder OS, and public eval configs | | `2bb0ce1` | docs(oss): add README, reproduction guide, license, and paper skeleton | **Note:** Early commits used `git -c commit.gpgsign=false` because GPG/passkey signing hung on this machine. --- ## 4. Evaluation and run results Runtime outputs live under `results/` (gitignored). Key artifacts from prior runs: ### 4.1 Separability diagnostic **Run:** `results/diagnostic_run/certificate.json` ```json { "status": "SURGERY_VIABLE", "max_lambda": 103.71, "mean_lambda": 15.28, "std_lambda": 33.59, "threshold": 3.0, "recommendation": "Proceed with CARVE oblique projection cut." } ``` On synthetic entangled data, the generalized eigenvalue spectrum is **not flat** — CARVE has separable forget-dominant directions at the matrix/toy level. ### 4.2 Syn-2Cap weight-level merge **Run:** `results/syn2cap_merged/` - Merged two capability deltas (60% cap-A + 40% cap-B) into entangled LoRA weights - Rank 16, toy dimensions (d_in=32, d_out=64) - Confirms the synthetic construction pipeline works at weight level ### 4.3 Full local eval suite **Run:** `results/full_eval_20260705_083435/` | Method | FE | RF | Time (s) | |--------|-----|-----|----------| | CARVE | 0.0 | 1.0 | 0.12 | | Maat-SVD | -0.53 | 1.53 | 0.05 | | Delete | 1.0 | 0.0 | 0.003 | | Negation | 0.0 | 1.0 | 0.002 | **GATE decision:** `RETRY` (not promotable) | Metric | Value | Target | Pass | |--------|-------|--------|------| | FE | 0.0 | ≥ 0.90 | FAIL | | RF | 1.0 | ≥ 0.95 | PASS | | US | 0.99 | ≥ 0.98 | PASS | **Important caveat:** These metrics come from `_perf_sim()` in `run_full_eval.py` — a **ΔW norm proxy** on a mock model, not behavioral capability evaluation on a real LLM. CARVE reduces weight norms in high-λ directions, so the proxy incorrectly reports FE=0. These numbers must **not** be interpreted as algorithm failure; they indicate the eval harness needs capability-specific forward-pass scoring. ### 4.4 Relearning attack (stub) **Run:** `results/full_eval_20260705_083435/relearning.json` ```json { "curve": { "0": 1.0, "10": 0.95, "50": 0.75, "100": 0.5 }, "steps_to_recovery": 0 } ``` Harness exists; not run against a real post-CARVE adapter. ### 4.5 GPU end-to-end test `tests/test_synthetic_e2e.py::test_syn2cap_carve_on_cuda` **passes** on RTX 4050 (6.4 GB VRAM): - Builds entangled LoRA from merged capability deltas - Accumulates \(S_f, S_r\), solves generalized eigenproblem - Applies oblique projection cut - Verifies ‖ΔW‖ decreases after cut This proves the **core math pipeline works on CUDA** but does not measure forget/retain behavioral performance. ### 4.6 Environment check From `.cursor/scripts/check_env.py` (prior session): - CUDA available - torch 2.5.1 - peft 0.19.1 --- ## 5. Phase completion vs BUILD-PLAN | Phase | Description | Status | ~Complete | |-------|-------------|--------|-----------| | **A** | Infrastructure: config, lora_io, decompose, metrics | Implemented + tested | **100%** | | **B** | Synthetic Syn-2Cap, stats, rotation, diagnostic, baselines | Scripts + unit tests; no trained LoRA on Qwen | **70%** | | **C** | CARVE core: surgery, pipeline, Maat-SVD | Code + CUDA E2E; **go/no-go gate not passed** | **60%** | | **D** | Repair + GATE: verify, rollback, probes, CLI | Code + unit tests; CLI uses noop probes by default | **75%** | | **E** | Ablations: runner, relearning, report | Runs on mock model + norm proxy only | **50%** | | **F** | Real adapters: Founder OS integration | Config + sample data; no real run | **20%** | | **G** | Public eval: TOFU/WMDP | Stub loaders; not executed | **15%** | | **H** | Paper + OSS | README, REPRODUCE, LaTeX skeleton; no results to fill | **60%** | ### Success criteria (definition of done) | Criterion | Target | Current | |-----------|--------|---------| | Syn-2Cap FE (held-out) | ≥ 0.90 | Not measured on real LLM | | Syn-2Cap RF (held-out) | ≥ 0.95 | Not measured on real LLM | | CARVE vs Maat-SVD | Pareto dominate on entangled synthetics | Not tested behaviorally | | Eigenvalue diagnostic | Flat → certificate + retrain recommendation | **Works** on toy/synthetic | | End-to-end ≤3B | < 30 min single GPU | Not run | | GATE rollback | Works on verification failure | Unit-tested; not E2E on real adapter | | Relearning attack | Reported honestly | Harness exists; not run on real adapter | --- ## 6. What went right 1. **Full CARVE codebase** matches BUILD-PLAN module layout — not stubs. 2. **Core math implemented and tested:** \(S_f/S_r\) accumulation, generalized eigenproblem, oblique projector (LEACE-style), sequential block scrubbing skeleton. 3. **All unit tests pass**, including CUDA end-to-end on synthetic entangled LoRA. 4. **GATE infrastructure exists:** snapshot/rollback, preregistered thresholds, `check_gate()`, verify module. 5. **Synthetic tooling works at weight level:** merge two capability deltas, run eigenvalue diagnostic, emit separability certificate. 6. **`.cursor/` agent infrastructure** is comprehensive — future agents can pick up work with clear rules, skills, and phased plan. 7. **Algorithm decision is binding and documented** — PSN/CCD forbidden; CARVE is primary. 8. **Windows compatibility fixes** applied (validate script root path, Unicode in demo output, PowerShell-safe bootstrap). 9. **Git history is clean** — implementation committed in phase-grouped commits. --- ## 7. What went wrong / known gaps 1. **Phase C go/no-go not passed.** BUILD-PLAN requires behavioral Syn-2Cap on Qwen with FE≥0.90, RF≥0.95, and CARVE beating Maat-SVD. Only matrix-level / mock-model eval exists. 2. **No real LLM run.** - `run_carve.py` never executed on Qwen2.5-1.5B (6.4 GB VRAM → needs 4-bit quantization). - No LoRA_A / LoRA_B training for ground-truth capabilities. - TOFU/WMDP not run. - Founder OS adapter paths in `configs/founder_os.yaml` are empty placeholders. 3. **Stats collection is noop in default CLI path.** `run_carve.py` passes `lambda: None` probe functions to the pipeline — no real activation collection during surgery unless wired manually. 4. **Ablation metrics are misleading.** Norm proxy cannot separate forget vs retain capabilities. Full eval GATE correctly returns RETRY, but raw FE/RF numbers look like CARVE failed when it only reduced weight norms. 5. **Maat-SVD baseline is a proxy** — cuts top-|σ| directions from analysis SVD, not true forget-gradient scoring from Maat paper. 6. **Repair phase untested on real batches** — code exists, no meaningful integration test with retain data. 7. **`gradient` stats_method raises `NotImplementedError`** — only `activation` path works. 8. **One commit message mislabeled** (`19a271e` says "Phase D" but contains eval modules) due to parallel commit race. 9. **`results/` is gitignored** — eval artifacts exist locally but are not versioned (by design, to avoid committing adapter weights). --- ## 8. Architecture overview ``` ┌─────────────────────────────────────┐ │ Phase 0: Snapshot │ │ decompose.py / lora_io.py │ └─────────────────┬───────────────────┘ │ ┌───────────────────────────┼───────────────────────────┐ │ │ │ ▼ ▼ ▼ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ │ Phase C1 │ │ Phase C2 │ │ Diagnostic │ │ stats.py │───────────▶│ rotation.py │───────────▶│ diagnostic │ │ S_f, S_r │ │ gen. eig │ │ certificate │ └─────────────┘ └──────┬──────┘ └─────────────┘ │ ▼ ┌─────────────┐ │ Phase C3 │ │ surgery.py │ │ oblique cut │ └──────┬──────┘ │ ▼ ┌─────────────┐ │ Phase C4 │ │ pipeline.py │ │ sequential │ └──────┬──────┘ │ ┌────────────────┼────────────────┐ ▼ ▼ ▼ ┌───────────┐ ┌───────────┐ ┌───────────┐ │ repair │ │ verify │ │ baselines │ │ repair.py │ │ verify.py │ │ delete / │ │ │ │ GATE │ │ negation /│ │ │ │ rollback │ │ maat_svd │ └───────────┘ └───────────┘ └───────────┘ ``` **Data flow for surgery (per layer):** 1. Collect inner activations \(z = A_l x\) on forget (\(D_f\)) and retain (\(D_r\)) probes. 2. Build \(S_f = \mathbb{E}[z_f z_f^\top]\), \(S_r = \mathbb{E}[z_r z_r^\top]\). 3. Solve \(S_f w = \lambda S_r w\) → eigenvectors \(W\), eigenvalues \(\lambda\). 4. Select cut directions where \(\lambda \geq\) `lambda_threshold`. 5. Build oblique projector \(P\) preserving \(S_r\)-geometry; apply \(A_l \leftarrow A_l P\). 6. Repeat block-by-block front-to-back; recompute stats on edited model. --- ## 9. Recommended next steps (priority order) | Priority | Task | Why | |----------|------|-----| | 1 | Train Syn-2Cap on Qwen2.5-1.5B 4-bit | Ground-truth entangled adapter with known \(c_f, c_r\) | | 2 | Wire real forward passes in `stats.py` / `pipeline.py` | Replace noop probes in `run_carve.py` | | 3 | Run `run_carve.py` end-to-end on Qwen | First behavioral FE/RF measurement | | 4 | Pass Gate C | FE≥0.90, RF≥0.95, CARVE > Maat-SVD; debug if fail | | 5 | Fix ablation `eval_fn` | Capability-specific scoring, not ‖ΔW‖ norm | | 6 | Founder OS real adapter | Trace-derived probes, production path | | 7 | TOFU subset public benchmark | External credibility | | 8 | Fill paper with real results | Preprint trigger: first credible Syn-2Cap win | --- ## 10. File inventory summary | Area | Files | Lines (approx.) | |------|-------|-----------------| | `src/` | 22 Python modules | ~2,000 | | `tests/` | 14 test files | ~800 | | `scripts/` | 6 scripts | ~600 | | `.cursor/` | rules, skills, research, scripts | ~5,000+ | | `docs/` | 9 research docs | — | | `lora-capability-removal-review/` | 4 audit docs | — | | `configs/` | 6 YAML files | — | | `paper/` | main.tex + bib | skeleton | --- ## 11. Conclusion The project has moved from **research synthesis + audit** to a **working implementation with test coverage and eval harness**. The hard remaining work is **scientific validation**: proving CARVE achieves verified behavioral removal on entangled LoRA adapters at Syn-2Cap scale before touching real Founder OS adapters or public benchmarks. **Bottom line:** Infrastructure and algorithm code are in place (~55–60%). The go/no-go gate that determines whether CARVE is publishable and product-worthy has not been crossed. --- *Report generated 2026-07-10. For updates, add dated reports under `report/` and link from `report/README.md`.*