Skip to content

The Code Only Goes One Way: Why Agents Won't Delete

18 min read

The Code Only Goes One Way: Why Agents Won't Delete

· 18 min read
An editorial illustration on warm cream paper in black ink line work, drawn as an engraved instrument plate: a large ratchet wheel seen flat on, its rim cut into two dozen asymmetric saw teeth, with a spring-loaded pawl dropped into the gap between two of them and resting hard against a steep tooth face. A long curved arrow sweeps clockwise around the outside of the wheel unobstructed and is labelled ADD. A second arrow approaching counter-clockwise stops dead against the pawl with a small collision mark where the two meet, and is labelled DELETE. A smaller dashed arrow branches off the blocked path just short of the pawl, loops back around and rejoins the clockwise direction, labelled GUARD-AND-GO. A thin leader line labels the pawl arm itself. A small all-caps serif title reading THE CODE ONLY GOES ONE WAY sits in the upper left, and one thin ink-blue line runs along the bottom margin above the caption DELETION NEEDS A PROOF THE MODEL CANNOT RUN.

On 30 July 2026, five researchers published a paper with a title that does most of the work: To Add Is Machine, To Delete Is Human. They measured something nobody had isolated before, which they call deletion avoidance, the systematic tendency of a language model to keep code that the edit it was asked to make requires it to remove.

The numbers are specific and they are worse than folklore. Across the five leading models on the official SWE-bench Verified leaderboard, restricted to tasks that all five models solve, deletion recall against the developer’s own patch reaches at most 71.7%. The models reach the right file for over 92% of required deletions and cut the exact line in under 52% of cases.

Read those two figures next to each other. The model finds the code. It then does not remove it. Finding and removing are separate behaviours and only one of them is reliable.

Finding the code versus removing it

Across the five leading SWE-bench Verified models, on tasks all five solve: the correct file is reached for over 92% of required deletions, but the exact line is cut in under 52% of cases. On 34 Verified tasks retrofitted with tests that fail if the targeted code remains, four frontier models fall from a combined 63.2% pass rate to 41.9%.

FINDING IT IS NOT REMOVING ITFrontier models on SWE-bench Verified tasks requiring a deletionLocates the right file92%Cuts the exact line52%Passes the original tests63.2%Passes tests that check removal41.9%The code survived because nothing ever tested that it was gone.Source: Ebrahimi et al., arXiv 2607.28887, July 2026. Lower pair is 34 retrofitted tasks, four models.

What happens instead has a name too. In 29.0% of passing patches the model wraps the targeted code in a guard or a fallback rather than deleting it, a pattern the authors call Guard-and-Go. The patch passes. The code stays. Something that used to be one path is now two, and the dead one is protected by a condition that will never be false.

I have spent a year writing about what agents do inside my own sessions. This result was specific enough to check against my own data, so I did.

Key Takeaways

  • Deletion avoidance is measured, not folklore. Across the five leading models on the SWE-bench Verified leaderboard, deletion recall against the developer’s own patch tops out at 71.7%, and the models reach the right file for over 92% of required deletions while cutting the exact line under 52% of the time.
  • My own sessions say the same thing from the other end. 73.7% of 2,115 agent edits to code files made the file longer and 10.4% made it shorter, a line ratio of 2.33 to 1, and only 390 of 66,206 shell commands touch a delete at all.
  • My review pass does not compensate. Across 9,635 of my own commits in 14 repositories the ratio is 3.15 to 1, 8.4% of commits remove more than they add, and 1.0% only remove.
  • Instruction does not fix it. My CLAUDE.md has ordered deletion for months and produced every number in this post, and the paper’s own ablation barely moves until you supply the exact lines to cut, which was the hard part to begin with.

Does this show up outside a benchmark?

It shows up at industry scale first.

GitClear’s June 2026 report, The Maintainability Gap, analysed 623 million code changes from 2023 through 2026 and tracked eight maintainability signals. They moved the wrong way together.

The one that matters here is moved code, which is GitClear’s proxy for refactoring: relocating a block rather than writing a second copy of it. It was 21% of changed lines in 2022. It is 3.8% year to date in 2026, a 70% decline. Over the same window copy and paste went from 9.4% to 15.7%, block duplication rose 81%, cross-file function calls fell 35%, and updates to code older than a year fell 74%, from 1.7% to 0.46%.

GitClear’s summary of the reversal is the cleanest sentence in the report. In 2022 developers were roughly twice as likely to refactor as to duplicate. They are now about five times more likely to duplicate than to refactor.

Refactoring is the single behaviour that stops a codebase accumulating, and it lost 17 points of share in four years. But it is industry telemetry across many teams and many tools, and it cannot tell you whether the agent or the human is the one declining to refactor. My transcripts can.

What do my own agent edits actually do?

I took a frozen snapshot of 1,207 Claude Code session transcripts on 15 September 2026, excluding the session I was working in, and counted every structured edit by direction: did the replacement have more lines than the text it replaced, fewer, or the same?

Restricted to code files, there are 2,115 such edits.

Direction of agent edits, code versus prose

Code files, 2,115 edits: 1,558 grew the file (73.7%), 338 left it the same length (16.0%), 219 shrank it (10.4%). Lines added 42,351 against 18,153 removed, a ratio of 2.33 to 1. Prose files, 680 edits: 440 grew (64.7%), 224 same (32.9%), 16 shrank (2.4%), a ratio of 3.36 to 1.

WHICH WAY DOES AN AGENT EDIT GO?Replacement longer, same length, or shorter than what it replacedCode73.7%n=2,115Prose64.7%n=680Grew the fileSame lengthShrank the fileCode edits added 42,351 lines and removed 18,153. Ratio 2.33 to 1.Source: 1,207 Claude Code session transcripts, snapshot 2026-09-15

73.7% of edits to code files made the file longer. 10.4% made it shorter. The rest were same-length substitutions, a renamed variable or a changed condition.

The line counts say the same thing with less rounding. Those edits added 42,351 lines and removed 18,153, a ratio of 2.33 to 1. That is before counting whole-file writes, which are pure addition by definition. There are 750 of them in the same snapshot, 391 to code files, contributing 124,088 lines in total with nothing on the other side of the ledger.

Then there is the shell. Across 66,206 Bash commands in those sessions, 390 involve rm or git rm. That is 0.6%. My agent runs about 170 shell commands for every one that deletes a file.

The prose row is a useful control. I write blog posts with the same agent in the same harness, and prose edits shrink the file 2.4% of the time against code’s 10.4%. Both ratchet upward, so some of this is just what “edit” means when the task is “add a section.” But code is where the compounding cost lives, and code is where the paper says the mechanism bites.

Does anyone delete it afterwards?

This is the part I expected to rescue the number. The agent adds, I review, I cut. That is the deal.

I went through the git history of every repository I own, since 1 January 2025, restricted to commits authored by me, with generated files excluded: lockfiles, build output, vendored clones, generated schemas, migration snapshots, images, and this blog’s own post directory, which is pure addition and would flatter the result badly.

That leaves 9,635 commits across 14 repositories. They added 2,331,305 lines and removed 741,240, a ratio of 3.15 to 1. 807 of those commits, 8.4%, removed more than they added. 95 of them, 1.0%, only removed.

So the review pass does not rescue it. One commit in a hundred is a deletion. The ratio after my review is worse than the ratio inside the agent’s own edits, which is the part worth sitting with: I have not hand-written production code in six months, so review is the only step in this pipeline where subtraction could happen, and it is where the ratio gets worse rather than better. Accepting is easy and cutting is work, and the gate that was supposed to compensate turns out to be the one doing the least.

Why won’t the model just delete it?

Because deleting requires a proof that the model cannot construct.

To add a function you need to know what it should do. To delete one you need to know that nothing depends on it, and that is a question about the whole repository, not the file in the context window. The SWE Atlas benchmark puts a number on the gap. Its refactoring split is 70 expert-authored tasks with an average of 18 tests each, graded not only on regression tests but on rubrics that explicitly penalise over-deletion and broken interfaces. Pass@1 for the best agent tested is 48.57%, and even top models miss call sites in 30% to 40% of trials.

The authors describe the failures as structural work left half-done: obsolete definitions, helpers and imports that survive a refactor because the model never established they were safe to remove.

That is the same finding as the deletion paper from a different direction. The model reaches the right file for over 92% of required deletions. It knows where the code is. What it does not have is the call graph, so it cannot answer the only question that licenses removal, and under that uncertainty Guard-and-Go is not laziness. It is the correct play. Wrapping the block in a condition is strictly safer than deleting it if you genuinely do not know who calls it. The model is optimising for the thing we graded it on, which is not breaking the tests.

I have made this argument before about what agents need that a filesystem cannot give them, and measured that giving an agent a real symbol graph beats grep on answer quality. Deletion is the sharpest case for it, because it is the one operation where a wrong answer is silent and expensive and the right answer requires global knowledge.

The tests do not help either, which is the part that should bother you. The deletion paper retrofitted 34 SWE-bench Verified tasks with tests that fail if the targeted code is still present. Four frontier models fell from a combined 63.2% pass rate to 41.9%. The code was surviving because nothing ever checked whether it was gone. That is the same structural hole I found in merged PRs that nobody actually reviewed: the gate exists, it reports green, and it was never measuring the thing you cared about.

Can you just tell it to delete?

I assumed the fix was instruction, because that is the cheap fix and I had already applied it. My own CLAUDE.md has carried this line for months:

Delete dead code rather than comment it out; git has it.

Every number in this post was produced by an agent reading that instruction on every single turn. It does not bite.

The paper ran the controlled version of that experiment so I do not have to. They ablated one model under four cumulative prompts, escalating in specificity. Success barely moved until they supplied the exact lines to remove. That finally fixed incomplete deletion, and still only reached 80.5%, because once the model was told to delete it started deleting past the span it was given or adding code in compensation.

So the ladder is: asking nicely does nothing, and telling it precisely what to cut works only if you already know what to cut, which was the hard part.

This is a sharper version of something I wrote about whether your CLAUDE.md earns its tokens. A context file can reliably change behaviour that the model is capable of and merely not defaulting to. It cannot install a capability. Deletion avoidance is not a default the model picked, it is a proof it cannot run, and no amount of imperative mood in a markdown file supplies the missing call graph.

The one genuinely hopeful result in the paper is the last one. A pilot study teaching deletion during post-training reduced deletion avoidance and improved broader code-editing performance, which says the behaviour is undertrained rather than out of reach. That is a fix in the labs’ hands, not yours, and not this quarter.

What should you actually do?

Four things, in order of how much they cost you.

Measure the ratchet. The whole diagnosis in this post is two commands over data you already have. Lines added over lines removed, and the share of commits that are net negative, per quarter, with generated files excluded. If your ratio is climbing and your net-negative share is falling, the thing described here is happening to you. This belongs on the same board as the cost and failure signals worth instrumenting, and it is cheaper to collect than either.

Make deletion its own task. Deletion loses to addition when they compete inside one prompt, because the addition is what the tests grade. Give the agent a task whose entire success condition is removal, the way CanItDelete does, and give it the reference list up front rather than asking it to find one. That is the only intervention in the paper that moved the number materially.

Grade removal, not just behaviour. Retrofitting removal-checking tests cost four frontier models 21 points of pass rate, which means the pass rate was measuring something else. When you ask an agent to replace a mechanism, assert that the old one is gone. A test that only checks the new path is green whether or not you now ship two.

Watch for Guard-and-Go in review. This is the cheapest and it is nearly free once you know the shape. A new conditional wrapping an old block, a fallback branch that the new code path can never reach, a flag defaulted so one side is dead. 29.0% of passing patches did this. It is not a bug, it reads as caution, and it is how one code path quietly becomes two.

None of this is exotic. It is the same discipline as everything else in the agent era: the model does the work, and you own the measurement that tells you what kind of work it did.

FAQ

Do AI coding agents avoid deleting code?

Yes, measurably. A July 2026 study of the five leading models on the SWE-bench Verified leaderboard found deletion recall against the developer’s patch reaches at most 71.7%, and models cut the exact required line in under 52% of cases despite locating the correct file over 92% of the time. In my own 1,207 sessions, 73.7% of agent edits to code files made the file longer and 10.4% made it shorter.

What is Guard-and-Go?

Guard-and-Go is the pattern where a model wraps code it was supposed to delete in a conditional or fallback instead of removing it. It appeared in 29.0% of passing patches in the 2026 deletion-avoidance study. The patch passes its tests because the tests do not check that the old code is gone, and the codebase silently gains a second code path that nothing reaches.

Why do coding agents refuse to delete code?

Deleting safely requires knowing that nothing else calls the code, which is a question about the entire repository rather than the file in context. Agents miss call sites in 30% to 40% of refactoring trials. Given that uncertainty, guarding the code is genuinely safer than removing it, so the behaviour is a rational response to missing information rather than carelessness.

Will telling the agent to delete dead code fix it?

Largely no. The 2026 study escalated prompt specificity across four levels and found success barely moved until the exact lines to delete were supplied, which only reached 80.5% because the model then over-deleted or added compensating code. My own CLAUDE.md has instructed deletion for months and every measurement in this post was produced under that instruction.

How do I measure whether this is happening in my codebase?

Compute lines added divided by lines removed from git history, excluding generated files, lockfiles and build output, and track the share of commits that remove more than they add. Mine are 3.15 to 1 and 8.4% across 9,635 commits. Compare repositories of similar age and stage rather than quarters, because a young repository legitimately adds more than it removes and will flatter or alarm you for the wrong reason.

Methodology

Transcripts. 1,207 Claude Code session transcript files, copied to a frozen snapshot on 15 September 2026, excluding the session used to write this post. Every tool_use block was parsed from the JSONL. Edit direction compares the newline count of new_string against old_string for each structured edit call; files were classed as code or prose by extension. Totals: 2,888 edits overall, 2,115 to code files, 680 to prose. Whole-file writes (750, of which 391 to code) and shell commands (66,206) were counted separately. Shell deletions matched rm - or git rm at the start of a command or after a shell separator, which will miss deletions inside scripts. In-place sed and perl rewrites (313) are counted but not direction-scored.

Git history. All repositories under my projects directory with a git directory, since 1 January 2025, merges excluded, restricted to commits authored by me, which excludes one cloned repository I do not contribute to. Excluded paths: node_modules, dist, build, out, .next, coverage, vendor, __pycache__, vendored clone directories, generated and __generated__ directories, migration metadata, all common lockfiles, binary and image extensions, generated schema files, and this blog’s src/content/blog directory. Result: 9,635 commits across 14 repositories, 2,331,305 lines added, 741,240 removed. A commit counts as net negative when its deletions exceed its additions after exclusions, and as pure deletion when it has deletions and no additions.

Limitations. Two, stated plainly. The transcript counts have a known blind spot: I run a lot of work in a mode that routes file edits through shell commands rather than the structured edit tool, so 313 of the 66,206 Bash calls are in-place sed or perl rewrites that are counted but never direction-scored, roughly a tenth of the structured edit volume. And this is a snapshot rather than a trend. I deliberately do not claim the ratio has changed over time, because the corpus cannot support that claim: most of these repositories were created in 2026, only one has any meaningful pre-2026 history, and comparing quarters therefore compares different codebases at different stages of life rather than one codebase over time.

External sources. Ebrahimi, Hasan, Bhatia, Rajbahadur and Hassan, To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing, arXiv 2607.28887, submitted 30 July 2026, retrieved 15 September 2026. SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution, arXiv 2605.08366, retrieved 15 September 2026. GitClear, The Maintainability Gap: 2026 AI Code Quality Research, June 2026, retrieved 15 September 2026.

Share this post

If it was useful, pass it along.

What the link looks like when shared.
X LinkedIn Bluesky

Search posts, projects, resume, and site pages.

Jump to

  1. Home Engineering notes from the agent era
  2. Resume Work history, skills, and contact
  3. Projects Selected work and experiments
  4. About Who I am and how I work
  5. Contact Email, LinkedIn, and GitHub