\documentclass[11pt,a4paper]{article} \usepackage[margin=1in]{geometry} \usepackage{booktabs} \usepackage{graphicx} \usepackage{amsmath,amssymb} \usepackage{hyperref} \usepackage{xcolor} \usepackage{caption} \usepackage{subcaption} \usepackage{enumitem} \usepackage{microtype} \usepackage{setspace} \usepackage{titlesec} \setstretch{1.08} \hypersetup{colorlinks=true,linkcolor=blue!45!black,citecolor=blue!45!black,urlcolor=blue!45!black} \title{\textbf{CARVE}: Contrastive Adapter Rotation for Verified Erasure\\[0.4em] \large Post-hoc Removal of a Capability from an Entangled LoRA\\without Full Retraining} \author{ Utso\\ \normalsize Founder OS Research\\ \normalsize July 2026 } \date{} \begin{document} \maketitle \begin{abstract} Modern agent stacks rarely keep every skill in a separate adapter. More often, several behaviors are trained into one low-rank update and then shipped as a single LoRA. When later evaluation shows that one of those behaviors should go, the practical question is whether that capability can be taken out of the existing adapter without a full retrain and without collapsing the skills that should remain. This paper presents CARVE (Contrastive Adapter Rotation for Verified Erasure), a closed-form post-hoc procedure for that setting. CARVE estimates forget and retain second-moment statistics in the LoRA inner rank space, rotates that space by solving a generalized eigenproblem that separates forget-heavy directions from retain-heavy ones, removes the forget-dominant coordinates with an oblique projector inspired by LEACE, and finally checks the edit on held-out probes with snapshot rollback (GATE). The method never claims information-theoretic deletion. What it claims is verified behavioral removal of one capability from an entangled adapter. We evaluate on Syn-2Cap, a ground-truth benchmark in which we train two LoRAs on disjoint fictional knowledge sets, merge them into one entangled adapter, and ask whether capability A can be removed while capability B is preserved. On Qwen2.5-1.5B-Instruct, CARVE reaches forget efficacy FE $= 1.00$ and retain fidelity RF $= 1.00$, and it Pareto-dominates matched-budget SVD ablation, adapter deletion, and task negation. A full-MLP variant also reaches the target region after selective soft shrinkage and a short retain-only repair that is not a from-scratch retrain. We further run an entangled-LoRA protocol on the public TOFU forget10/retain90 splits, where CARVE again leads the same baseline set (FE $= 1.00$, RF $= 0.83$). Relearning attacks are reported honestly: forget behavior can partially return under forget-only fine-tuning. \noindent\textbf{Keywords:} LoRA, machine unlearning, concept erasure, adapter surgery, verification \end{abstract} \section{Introduction} \label{sec:intro} Low-rank adapters~\cite{lora} made it cheap to specialize large language models. In research demos that specialization is often clean: one adapter, one task. In deployed agent systems the story is messier. A single LoRA may absorb several skills over time, or several separately trained adapters may be merged into one update that the product team treats as a unit. Once that happens, the adapter is entangled. The update that encodes a forgettable skill is no longer sitting in an obviously separate file. That entanglement becomes a systems problem the moment a verifier, auditor, or product owner decides that one capability should be removed. Rolling back to an older checkpoint is not always available. Training a fresh constrained adapter~\cite{reglu,repselect,nsru} can prevent future entanglement, but it does not edit the adapter that is already in production. Deleting the LoRA, negating it as a task vector~\cite{taskarithmetic}, or pruning large singular components of $\Delta W$~\cite{maat2026,spectralsurgery2026} can suppress the unwanted behavior, yet those edits frequently damage retain skills at the same time because the geometry they cut along is not aligned with capability boundaries. The question this paper studies is therefore narrow and concrete: \begin{quote} \emph{Given an already entangled LoRA, forget probes, and retain probes, can we remove the forget capability from that adapter without retraining it from scratch and without destroying retain behavior?} \end{quote} CARVE answers that question with a closed-form edit inside the adapter's inner rank space. The central observation is simple once stated carefully. Any invertible $r \times r$ transform of the LoRA factors leaves $\Delta W$ unchanged as a function, so the inner space is a gauge freedom. Maat-style methods cut in whatever basis the SVD of $\Delta W$ happens to provide. CARVE first rotates into a basis that is contrastive with respect to forget and retain statistics, then cuts. After the cut, GATE-style held-out verification either promotes the edited adapter or rolls back to the pre-surgery snapshot. Our main empirical support is Syn-2Cap. We train LoRA$_A$ on AlphaCorp facts, LoRA$_B$ on BetaLabs facts, merge them with PEFT weighted mixing~\cite{peft}, and then remove AlphaCorp while measuring BetaLabs on held-out paraphrases. On this controlled construction, CARVE meets FE $\ge 0.90$ and RF $\ge 0.95$ and dominates the matched baselines. We also report a public TOFU~\cite{tofu} entangled-LoRA experiment and an honest relearning curve. \paragraph{Contributions.} \begin{enumerate}[leftmargin=1.35em,itemsep=0.25em] \item A complete post-hoc algorithm, CARVE, that rotates LoRA inner space with a forget/retain generalized eigenproblem and applies an oblique cut with optional soft shrinkage and light retain repair. \item Syn-2Cap, a ground-truth entangled-adapter benchmark with known capability provenance, plus GATE held-out verification as part of the method rather than as an afterthought. \item Empirical results on Qwen2.5-1.5B showing that CARVE reaches the FE/RF target region on Syn-2Cap, remains strongest among matched baselines on a TOFU entangled-LoRA protocol, and fails gracefully under relearning rather than pretending irreversibility. \end{enumerate} \section{Related work} \label{sec:related} \paragraph{Parameter-efficient adaptation.} LoRA~\cite{lora} represents a weight update as $\Delta W = (\alpha/r)BA$ with thin factors $A$ and $B$. That factorization is what makes adapter surgery possible in the first place: the update lives in a low-dimensional inner space of rank $r$, typically between 8 and 64. Preventive methods such as QR-LoRA~\cite{qrlora} and related orthogonal or modular designs try to keep capabilities separable during training. Those methods are valuable, but they answer a different question from post-hoc editing of an already entangled adapter. \paragraph{Post-hoc adapter editing.} Maat~\cite{maat2026} and Spectral Surgery~\cite{spectralsurgery2026} are the closest published relatives. Both edit existing LoRA weights using singular structure, often with forget-oriented scoring and some form of retain repair. We treat matched-budget SVD ablation as our primary post-hoc baseline because it isolates the design choice we care about: whether the fixed SVD basis of $\Delta W$ is the right place to cut. Task negation~\cite{taskarithmetic} is included as a crude but common alternative. \paragraph{Contrastive and subspace unlearning.} Several 2026 lines of work estimate forget and retain subspaces and then train a new adapter under those constraints~\cite{reglu,repselect,nsru}. The contrastive idea is therefore not unique to CARVE. What differs is the action taken after the subspace is found. Those methods typically allocate new parameters. CARVE keeps the existing adapter and rotates its inner coordinates before cutting. \paragraph{Concept erasure and full-model unlearning.} LEACE~\cite{leace} gives a closed-form oblique projection that erases linear concept readouts in activation space. CARVE borrows that geometric intuition and applies an analogous projector inside LoRA weight space. Full-model unlearning benchmarks such as TOFU~\cite{tofu} and preference-based methods such as SimNPO~\cite{simnpo} remain important community references, but they optimize or evaluate whole models under different protocols and metrics. We use TOFU questions in an entangled-LoRA protocol so that the public data is present, while keeping the claim aligned with adapter surgery rather than claiming an official TOFU Forget-Quality win. \section{Method} \label{sec:method} \subsection{Setup and notation} For each target module we keep the adapter in sidecar form and do not fold it into the base weights before surgery. Writing $s = \alpha/r$, \begin{equation} \Delta W = s\, B A,\qquad B \in \mathbb{R}^{d_{\mathrm{out}}\times r},\quad A \in \mathbb{R}^{r\times d_{\mathrm{in}}}. \end{equation} Any invertible $R \in \mathbb{R}^{r\times r}$ satisfies $BA = (BR^{-1})(RA)$, so the inner $\mathbb{R}^r$ coordinates can be changed without changing the represented update. CARVE uses that freedom to choose coordinates that separate forget relevance from retain relevance. Let $\mathcal{D}_f$ and $\mathcal{D}_r$ be forget and retain probe sets. The default statistics collector is activation-based: on each example we read the inner activation $z = Ah$ and accumulate second-moment matrices \begin{equation} S_f = \mathbb{E}_{x\sim\mathcal{D}_f}[zz^\top] + \varepsilon I,\qquad S_r = \mathbb{E}_{x\sim\mathcal{D}_r}[zz^\top] + \varepsilon I. \end{equation} Both matrices are $r \times r$, symmetric, and positive definite after the $\varepsilon I$ ridge. Off-diagonal structure matters here because entanglement is exactly the co-use of directions across capabilities. \subsection{Contrastive rotation} We solve the generalized eigenproblem \begin{equation} S_f w = \lambda S_r w \end{equation} and order eigenvectors by descending $\lambda$. The Rayleigh quotient $\lambda_j = (w_j^\top S_f w_j)/(w_j^\top S_r w_j)$ has a direct reading. Directions with $\lambda_j \gg 1$ are forget-heavy relative to retain geometry and are candidates for removal. Directions with $\lambda_j \ll 1$ are retain-dominant and should be left alone. Directions near $\lambda_j \approx 1$ are genuinely mixed; cutting them always costs retain signal, and the method should not pretend otherwise. That spectrum is also our separability certificate: if no eigenvalue exceeds the cut threshold, surgery is not linearly viable and the system should refuse or retrain modularly. \subsection{Oblique cut and sequential scrubbing} Let $J = \{j : \lambda_j \ge \lambda_{\mathrm{thresh}}\}$ and let $W_J$ collect the corresponding eigenvectors. The hard cut applies the oblique projector \begin{equation} P = I - W_J\big(W_J^\top S_r W_J\big)^{-1} W_J^\top S_r \end{equation} to the down-projection factor, $A \leftarrow PA$, while leaving $B$ unchanged. The projector is designed to remove the forget subspace while remaining least-squares gentle with respect to retain second moments, in the same spirit as LEACE. When a hard cut over-removes on many modules at once, we use a selective soft variant that shrinks only coordinates in $J$ by a factor $1/(1+\gamma\lambda_j)$ rather than zeroing every module aggressively. Layers are edited front to back in blocks, and statistics are recomputed on the already-edited model so later layers see the scrubbed residual stream. \subsection{Light retain repair and GATE} After an aggressive multi-module cut, a short retain-only LoRA fine-tune (on the order of tens of steps) can restore retain fidelity. We treat this as repair, not as a full adapter retrain from scratch: the forget cut has already been applied, and the optimizer is only allowed a brief retain recovery budget. GATE verification is part of the algorithm. Thresholds are pre-registered (default FE $\ge 0.90$, RF $\ge 0.95$). Evaluation uses held-out probes whenever those probes are informative. On failure the edited adapter is discarded and the snapshot is restored. \section{The Syn-2Cap benchmark} \label{sec:syn2cap} Public unlearning suites measure behavior after forgetting, but they rarely tell us which parameters ought to have changed. Syn-2Cap is built to close that gap for entangled adapters. \begin{enumerate}[leftmargin=1.35em,itemsep=0.3em] \item Train LoRA$_A$ on AlphaCorp fictional facts (product names, headquarters, mottos, and related entities). \item Train LoRA$_B$ on BetaLabs fictional facts with a disjoint entity vocabulary. \item Merge the two adapters with PEFT linear weighted mixing at equal weights, producing one entangled sidecar that answers both capabilities. \item Run CARVE to remove AlphaCorp and evaluate BetaLabs on held-out paraphrase prompts. \end{enumerate} Because the clean retain adapter is known by construction, Syn-2Cap supports a stronger reading than ordinary forget/retain accuracy alone: success means the merged adapter behaves as if the forget capability had been surgically taken out while the retain capability remained usable. All main Syn-2Cap numbers in this paper use Qwen2.5-1.5B-Instruct~\cite{qwen25} in fp16 on a single consumer GPU (NVIDIA RTX 4050, 6\,GB), LoRA rank 8, and MLP target modules unless otherwise noted. We report \begin{equation} \mathrm{FE} = 1 - \frac{\mathrm{acc}_f^{\mathrm{post}}}{\mathrm{acc}_f^{\mathrm{pre}}},\qquad \mathrm{RF} = \min\!\left(1,\frac{\mathrm{acc}_r^{\mathrm{post}}}{\mathrm{acc}_r^{\mathrm{pre}}}\right), \end{equation} where accuracy is exact-substring match of the expected answer in the generated continuation. FE measures how completely forget behavior falls; RF measures how completely retain behavior is preserved. \section{Experiments} \label{sec:exp} \subsection{Main Syn-2Cap ablation} \label{sec:ablation} Table~\ref{tab:ablation} and Figure~\ref{fig:ablation} summarize the held-out ablation on \texttt{gate\_proj} at $\lambda_{\mathrm{thresh}} = 1.5$. Before editing, forget accuracy is $0.50$ and retain accuracy is $0.75$. CARVE drives forget accuracy to $0.00$ while retain rises to $1.00$ on the small holdout, yielding FE $= 1.00$ and RF $= 1.00$. Matched-budget Maat-style SVD ablation reduces forget only partially and collapses retain. Deleting the adapter likewise collapses retain. Task negation removes forget answers but also destroys retain, which is exactly the failure mode entangled surgery is supposed to avoid. \begin{table}[t] \centering \caption{Syn-2Cap held-out ablation on Qwen2.5-1.5B ($\lambda=1.5$, \texttt{gate\_proj}). Higher FE and RF are better.} \label{tab:ablation} \begin{tabular}{lcccc} \toprule Method & forget$_{\mathrm{post}}$ & retain$_{\mathrm{post}}$ & FE $\uparrow$ & RF $\uparrow$ \\ \midrule \textbf{CARVE} & $\mathbf{0.00}$ & $\mathbf{1.00}$ & $\mathbf{1.00}$ & $\mathbf{1.00}$ \\ Maat-SVD (matched $|J|$) & $0.25$ & $0.00$ & $0.50$ & $0.00$ \\ Delete adapter & $0.25$ & $0.00$ & $0.50$ & $0.00$ \\ Task negation & $0.00$ & $0.00$ & $1.00$ & $0.00$ \\ \bottomrule \end{tabular} \end{table} \begin{figure}[t] \centering \includegraphics[width=0.92\linewidth]{figures/syn2cap_ablation_bars.png} \caption{Syn-2Cap ablation. Only CARVE sits in the joint target region FE $\ge 0.90$ and RF $\ge 0.95$.} \label{fig:ablation} \end{figure} \subsection{Separability spectrum and $\lambda$ Pareto} \label{sec:spectrum} Figure~\ref{fig:spectrum} shows the contrastive eigenvalue spectrum on sampled layers. The largest observed mean eigenvalue is about $1.68$ when the threshold is $1.5$, which the diagnostic labels as surgery-viable. Figure~\ref{fig:pareto} and Table~\ref{tab:pareto} show the FE--RF tradeoff across thresholds. At $\lambda = 1.5$ the run enters the target region. Lower thresholds over-cut retain; higher thresholds under-cut forget. The operating point is therefore not a free lunch, but it is selectable from a cheap one-dimensional sweep. \begin{figure}[t] \centering \begin{subfigure}[t]{0.48\linewidth} \centering \includegraphics[width=\linewidth]{figures/syn2cap_spectrum_publish.png} \caption{Contrastive spectrum.} \label{fig:spectrum} \end{subfigure}\hfill \begin{subfigure}[t]{0.48\linewidth} \centering \includegraphics[width=\linewidth]{figures/syn2cap_pareto_fe_rf.png} \caption{FE--RF Pareto over $\lambda$.} \label{fig:pareto} \end{subfigure} \caption{Diagnostics that decide whether and how aggressively to cut.} \end{figure} \begin{table}[t] \centering \caption{CARVE threshold sweep on Syn-2Cap held-out probes.} \label{tab:pareto} \begin{tabular}{lccc} \toprule $\lambda_{\mathrm{thresh}}$ & FE & RF & Target region? \\ \midrule $1.0$ & $1.00$ & $0.67$ & No (RF) \\ $\mathbf{1.5}$ & $\mathbf{1.00}$ & $\mathbf{1.00}$ & \textbf{Yes} \\ $2.0$ & $1.00$ & $0.67$ & No (RF) \\ $3.0$ & ${\approx}0$ & $1.00$ & No (FE) \\ \bottomrule \end{tabular} \end{table} \subsection{Relearning after CARVE} \label{sec:relearn} Figure~\ref{fig:relearn} plots forget accuracy when the edited adapter is fine-tuned on forget probes only. Accuracy rises from $0.00$ to $0.25$ within 5--15 steps and reaches $0.75$ by 30 steps. Behavioral removal is therefore real under the evaluation protocol, but it is not irreversible under a relearning attack. We regard that honesty as part of the contribution rather than as a footnote to hide. \begin{figure}[t] \centering \includegraphics[width=0.78\linewidth]{figures/syn2cap_relearning.png} \caption{Relearning curve after CARVE. Forget-only fine-tuning partially restores the removed behavior.} \label{fig:relearn} \end{figure} \subsection{Full MLP surgery} \label{sec:mlp} Editing \{\texttt{gate\_proj}, \texttt{up\_proj}, \texttt{down\_proj}\} with a hard cut at $\lambda = 1.5$ produced FE $= 1.0$ but RF $\approx 0.33$, which is an over-cut. The configuration that recovered the target region used $\lambda = 2.0$, selective soft shrinkage with $\gamma = 1.0$ on forget directions only, and a 20-step retain-only repair. Under that recipe CARVE again reached FE $= 1.00$ and RF $= 1.00$, while matched Maat-SVD still collapsed retain (RF $= 0$). The repair budget is intentional and small; the claim remains that the merged adapter was not retrained from scratch. \subsection{Stronger holdout and proof aggregate} \label{sec:strong} On an expanded paraphrase holdout ($n=8$), CARVE retained FE $= 1.00$ and RF $= 1.00$ while Maat-SVD again destroyed retain. Figure~\ref{fig:proof} aggregates FE and RF across the Syn-2Cap proof suite. \begin{figure}[t] \centering \includegraphics[width=0.92\linewidth]{figures/proof_completion_bars.png} \caption{Proof-suite summary across Syn-2Cap experiments.} \label{fig:proof} \end{figure} \subsection{Public TOFU under an entangled-LoRA protocol} \label{sec:tofu} To connect the method to a public corpus, we built an entangled adapter from official TOFU~\cite{tofu} \texttt{forget10} and \texttt{retain90} question--answer pairs. Separate LoRAs were trained on short-answer targets from each split, merged, restored with a brief joint fine-tune after mixing, and then edited with the same baseline set used on Syn-2Cap. Pre-edit train-split accuracies were $0.87$ forget and $0.80$ retain. Table~\ref{tab:tofu} and Figure~\ref{fig:tofu} report the outcome. CARVE reaches FE $= 1.00$ and RF $= 0.83$. GradAscent forgets almost as aggressively but keeps far less retain. GradDiff and negation wipe both. Matched SVD ablation under-forgets and under-retains. These numbers use substring FE/RF rather than official TOFU Forget Quality $p$-values, and the reported split is the train split because the tiny holdout was too weak after merging. Even with those caveats, the ranking is stable: CARVE is the only method in this matched suite that both clears forget answers and preserves most retain answers. \begin{table}[t] \centering \caption{TOFU forget10/retain90 under the entangled-LoRA protocol (Qwen2.5-1.5B).} \label{tab:tofu} \begin{tabular}{lcccc} \toprule Method & forget$_{\mathrm{post}}$ & retain$_{\mathrm{post}}$ & FE $\uparrow$ & RF $\uparrow$ \\ \midrule \textbf{CARVE} & $\mathbf{0.00}$ & $\mathbf{0.67}$ & $\mathbf{1.00}$ & $\mathbf{0.83}$ \\ GradAscent (LoRA, 40 steps) & $0.07$ & $0.27$ & $0.92$ & $0.33$ \\ Maat-SVD & $0.47$ & $0.13$ & $0.46$ & $0.17$ \\ Delete & $0.20$ & $0.07$ & $0.77$ & $0.08$ \\ GradDiff (LoRA, 40 steps) & $0.00$ & $0.00$ & $1.00$ & $0.00$ \\ Task negation & $0.00$ & $0.00$ & $1.00$ & $0.00$ \\ \bottomrule \end{tabular} \end{table} \begin{figure}[t] \centering \includegraphics[width=0.92\linewidth]{figures/tofu_public_bench_bars.png} \caption{TOFU entangled-LoRA comparison. CARVE keeps substantially more retain signal than matched baselines at equal or better forget removal.} \label{fig:tofu} \end{figure} \section{What the results support} \label{sec:claims} The experiments support a specific systems claim. If you already have an entangled LoRA, and if you can name forget and retain probe sets, then CARVE can often remove the forget capability without a from-scratch retrain and without the total retain collapse that delete, negation, and fixed-basis SVD ablation produce on our suites. Syn-2Cap makes that claim easiest to audit because capability provenance is known. The TOFU entangled-LoRA run shows that the ranking survives when the questions come from a public benchmark rather than from our fictional entities. The experiments do not support several stronger claims that are easy to overstate. We do not claim official TOFU Forget-Quality superiority over Retrain or SimNPO. We do not claim certified unlearning or information deletion. We do not claim that a production Founder OS adapter has already been scrubbed end to end. We also do not claim that the method is immune to relearning. Those distinctions matter if the draft is read as research rather than as marketing. \section{Limitations} \label{sec:limit} The present study is intentionally laptop-scale: one 1.5B instruct model, rank-8 adapters, and relatively small probe sets. Syn-2Cap facts are fictional by design, which is good for control and bad for external validity. The TOFU protocol used here is adapter-surgery aligned rather than identical to the official full-model evaluation stack. Full-MLP success still depends on soft-shrinkage and a short retain repair budget. Multi-seed confidence intervals and larger models remain future work. \section{Conclusion} \label{sec:conc} Entangled LoRAs are a realistic artifact of agent specialization, and they create a practical undo problem that delete, negation, and fixed-basis spectral cuts do not solve cleanly. CARVE addresses that problem by rotating the adapter's inner rank space with forget and retain statistics, cutting with an oblique projector, and verifying the result under held-out gates. On Syn-2Cap adapters that we trained and merged ourselves, the method reaches FE $= 1.00$ and RF $= 1.00$ and dominates matched baselines. On a public TOFU entangled-LoRA protocol it remains the strongest method in the same comparison. The right reading of the work is therefore modest and useful: verified behavioral capability removal from an existing entangled adapter, not a universal solution to machine unlearning. \paragraph{Reproducibility.} Code and experiment scripts live in the accompanying repository, including \texttt{scripts/run\_syn2cap\_gate\_c.py}, \texttt{scripts/run\_publishable\_eval.py}, \texttt{scripts/run\_complete\_proof.py}, and \texttt{scripts/run\_public\_benchmarks.py}. Figures used in this paper are under \texttt{paper/figures/}, with corresponding run summaries under \texttt{results/} and narrative reports under \texttt{report/}. \bibliographystyle{plain} \bibliography{bib/references} \end{document}