Agent Memory, Revisited: What Held Up and What Didn't
Between April and June 2026 this blog published four posts on agent memory, skills, and context files. The measurement arrived after them. Two of the claims do not survive it.
Here is what landed in the meantime. A practitioner running agents at scale reports “zero performance benefit on SWE tasks when agents have search access to their previous transcript sessions, provided they have access to other forms of context” (12 Grams of Carbon, 2026-07-02). A controlled minimal-pair study across 660 trials found that cleaner code leaves an agent’s pass rate unchanged while cutting tokens 7 to 8% and file revisitations 34% (arXiv:2605.20049, May 2026).
So this post is an audit. Thirteen specific claims, each with the new evidence and a verdict. Six hold, five weaken, two fail. Plus the metric nobody publishes, and what my own memory directory turned out to contain when I finally counted it. The method here was always measurement, which is what makes revising the output cheap.
Key Takeaways
- Thirteen claims from four posts on this blog, checked against measurement published after them: six hold, five weaken, two fail.
- The claim that fails hardest is the vendor memory benchmark. Swapping only the embedding model in an identical pipeline moves accuracy 6.2pp (p=0.004), so one variable flips the conclusion (arXiv:2606.29914, Jun 2026).
- Agent-authored memory is the weak link, not the pipeline. Self-memory scores 42% against 47% for basic retrieval (arXiv:2606.29914), and self-generated skills add nothing on average (SkillsBench v1, Feb 2026).
- Structured artifacts win on a different axis than expected. Cleaner code leaves pass rates unchanged while cutting tokens 7-8% and file revisitations 34% across 660 trials (arXiv:2605.20049, May 2026).
- My own auto-memory corpus: 51 files, 10 repositories, six months, zero deletions. Not one of them was a fact that had no better home.
What actually got measured since June?
Four things landed, and they do not agree in the way a roundup would suggest. Transcript recall shows no reported benefit. Skills show a real but domain-lopsided one. Rule-file content turns out to be largely interchangeable. And clean code moves cost without moving outcomes. Each comes with a caveat that belongs attached to the number.
Transcript recall. theahura at 12 Grams of Carbon reports zero benefit from agent search over prior sessions. He names the mechanism intent drift. “Every line of code, every existing bit of memory, every token is treated as an expression of intent, even if that code or that memory was generated from a random decision made by some previous agent session” (12 Grams of Carbon, 2026). The caveat matters. No sample size, task count, or harness is disclosed anywhere in the post. This is a reported finding, not a measured one, and I use “reports” for it throughout.
Code shape. SonarSource ran 33 tasks across six matched repository pairs built in both directions, with hidden tests, for 660 trials on Claude Code. The abstract is blunt: “code cleanliness does not change the agent’s pass rate. However, it substantially alters the agent’s operational footprint” (arXiv:2605.20049, Trivedi and Schmitt, May 2026). Caveat: both authors work at SonarSource, which sells the static analyzer that defines the independent variable.
Rule files. Guardrails Beat Guidance analysed 679 rule files containing 25,532 rules across more than 5,000 Claude Code runs on Opus 4.6. Its headline is uncomfortable: “Random rules improve a coding agent’s task performance as much as expert-curated ones (both +13.8pp on a discriminative subset of SWE-bench Verified)” (arXiv:2604.11088, Apr 2026, rev. May 2026). Caveat: that subset is 58 tasks screened to 30-70% baseline pass rates, not the full 500.
The inversion. Databricks built a benchmark on their own multi-million-line codebase, saw scores that looked too good, and inspected the traces. The “correct” implementation was still recoverable in the Git history of the worktree, so they sealed it: “we cut the working copy off from the repository entirely” (Databricks, Jul 2026). Hold that one. Git history is memory that worked so well it broke an eval, and it is the best argument in this post that the substrate matters more than the retrieval layer.
The same write-up found a harness they call Pi sent about 3x less context per turn at equal quality. That is the finding the harness is the cost line was built around.
Which claims held, and which failed?
Thirteen claims, three verdicts, and the distribution is the headline. The claims that held were mostly arithmetic and mostly about cost. The claims that failed were the two carrying a number borrowed from somebody else’s methodology.
The same ledger in full, with the evidence the chart has no room for:
| Claim | Post | Verdict | What the new evidence says |
|---|---|---|---|
| Append every event, summarise selectively, replay almost never | Memory architecture | Holds | Transcript recall shows no reported benefit, and unfiltered context gives “limited or negative benefits” |
| Skill metadata is ~22x cheaper than skill bodies | Skills | Holds | The vendor now publishes the tier costs, and the arithmetic was right |
| Prune skills whose description will not fire | Skills | Holds, wrong reason | 39 of 49 skills give zero pass-rate gain, but the failure mode is staleness, not vague wording |
| A useful CLAUDE.md is under ~200 lines | Context engineering | Holds | Anthropic’s docs now state the same number and the same reason |
| Context files should carry minimal requirements only | CLAUDE.md audit | Holds | Every individually beneficial rule turns out to be a prohibition |
| Never ship an LLM-generated context file as-is | CLAUDE.md audit | Holds | Self-memory 42% against retrieval 47%, and self-generated skills add nothing on average |
| Procedural memory compounds, semantic memory does not | Memory architecture | Weakens | Compounding is real and smallest in software engineering, at +4.5pp |
| The four-tier consolidation pipeline | Memory architecture | Weakens | The pipeline shape is fine; the unreviewed distiller is the weak link |
| Progressive disclosure is architecture, not optimisation | Skills | Weakens | It “buys context, not intelligence,” and the outcome gain is +4.1% |
| Progressive disclosure cuts catalog cost 94.7% | Context engineering | Weakens | Tier-4 source, and the outcome the figure implied is unsupported |
| A context file earns its tokens through leverage | CLAUDE.md audit | Weakens | Random rule files match curated ones at +13.8pp, so leverage cannot be the mechanism |
| Mem0 on LoCoMo backs the compact-wins framing | Memory architecture | Fails | One variable flips the conclusion, and LoCoMo never tested coding work |
| Pruning correlates with ~40% fewer bad-suggestion sessions | Context engineering | Fails | The correlation was about having a file; rule count is measurably not the lever |
Two of the six holds are now confirmed by the vendor itself. Context engineering in practice guessed that “a useful CLAUDE.md is under ~200 lines. Past that, signal drops fast.” Anthropic’s documentation now says it in almost the same words: “target under 200 lines per CLAUDE.md file. Longer files consume more context and reduce adherence” (Claude Code memory docs). And the 22x load-cost arithmetic holds because the tier costs are now published. Metadata is “~100 tokens per Skill” always loaded, instructions are “Under 5k tokens” on trigger, and resources cost “None until accessed” (Agent Skills overview).
The widest evidence base sits behind the least glamorous claim. Never ship an LLM-generated CLAUDE.md as-is now has three independent confirmations. SkillsBench v1 found that “Self-generated Skills provide no benefit on average, showing that models cannot reliably author the procedural knowledge they benefit from consuming” (arXiv:2602.12670v1). MemDelta found agent self-memory at 42% against 47% for basic retrieval (arXiv:2606.29914). And theahura accepts under 20% of the skill updates his own agents propose. Generalise the rule: never ship agent-authored persistent context as-is, on any surface.
One hold survives with the right action and the wrong reason. Pruning skills is correct, but the dominant failure is not vague descriptions. SWE-Skills-Bench found three skills that degraded performance up to -10% “due to version-mismatched guidance conflicting with project context” (arXiv:2603.15401, Mar 2026). Staleness, not vagueness. Prune by age and drift, not only by how the description reads.
The pattern across all thirteen is worth naming. Claims about cost survived. Claims about outcomes did not. That is the same two-ledgers split the CLAUDE.md audit drew between success and efficiency, applied to itself. The post that named the distinction is the one whose claims aged best.
Where did the two failures come from?
Both failures have one shape. A number measured on a different task was allowed to carry an architectural recommendation.
Failure one, the vendor benchmark. Agent memory architecture argued that “the benchmarking evidence backs the compact-wins framing,” citing Mem0 at 91.6 on LoCoMo under 7,000 tokens, with 91% lower p95 latency. MemDelta re-ran memory evaluation on LongMemEval-S changing one variable at a time, and the baseline choice decides the result. “Swapping only the embedding model in an identical pipeline shifts accuracy by +6.2pp at n = 500 (p = 0.004), and Mem0 beats MiniLM-RAG by +11pp but loses to cloud-RAG by 1.2pp, so one variable flips the conclusion” (arXiv:2606.29914, Jun 2026). On 2 of 6 question types, “Mem0 matches cloud RAG (72.7% vs. 73.9%, p = 1.0) at 50x the cost.”
There is a quieter problem underneath the loud one. LoCoMo measures recall over long personal conversations. The claim was about coding agents. The benchmark never tested the job.
Failure two, the borrowed correlation. Context engineering in practice claimed pruning a CLAUDE.md to what had been load-bearing in the last 30 days “correlates with ~40% fewer ‘bad suggestion’ sessions,” sourced to Redmonk. That survey correlation is about having a context file. It was attached to an action about pruning one, which the survey never tested. Guardrails Beat Guidance measures count directly and finds pass rates stable across rule counts from 0 to 50, with 50 rules slightly ahead of zero (arXiv:2604.11088). Pruning may still be good hygiene. The causal claim was never in the citation.
The generalisable lesson is one line. A memory-vendor benchmark and a developer survey are both fine sources for what they measured, and neither was measuring an agent completing a coding task. For any load-bearing claim, check that the benchmark’s task is your reader’s task. That is the argument for evaluating models yourself, turned back on the citations rather than the models.
Structured artifacts win, but not for the reason I assumed
Docs, commit messages, PR metadata, and the shape of the code do beat transcript recall. The reason is not that they carry more structure. It is that they are short, human-reviewed before they persist, and shaped as constraints rather than as intent.
The prescription comes from the same practitioner post. theahura’s replacement for transcript memory is to “emphasize good commit messages, good pr messages, and comprehensive documentation. Every code change comes with extensive metadata that is committed alongside the code.” Worth saying plainly: that is a memory system. It just puts the curation cost on humans and the storage cost on git.
Clean code did not make the agent smarter. It made the agent’s search shorter. That is what your codebase as the agent’s operating environment argued observationally, now with a controlled minimal-pair design behind it. It is the same shape as the project graph: give the agent a better substrate and it stops re-deriving.
Then there is the inversion. Databricks had to seal git history because agents were recovering correct implementations from it. Nobody prompts an agent to search git log. The substrate is already indexed, already reviewed, and already attached to the code it describes.
Now the honest complication, because it breaks the tidy version of this thesis. Guardrails Beat Guidance found the gains “largely content-independent: random, shuffled, mismatched-domain, and unconverted-format rule files all match curated rules, pointing to a context priming mechanism” (arXiv:2604.11088). If structure carried the knowledge, that result would be impossible.
So restate the thesis to survive it. Artifacts beat transcripts on three properties that have nothing to do with structure. They are short, bounded by the diff or the page. They pass a human review gate before they persist. And they are constraint-shaped, which is the polarity the same paper found actually helps: “every individually beneficial rule is a negative constraint (‘do not refactor unrelated code’), while every individually harmful one is a positive directive (‘follow code style’).” Transcripts are long, unreviewed, and intent-shaped. That is theahura’s intent drift restated from the other side.
Which memory metric does nobody publish?
If your setup proposes its own memory or skill updates, the acceptance rate on those proposals is the number that tells you whether the loop is working. Exactly one practitioner publishes it. It sits under 20%, which means roughly four in five proposed updates would have made the system worse.
The quote is worth its length: theahura’s bots “propose a set of changes to our built in nori skillsets, tagging the team in slack.” Then, “These are all default rejected. In order to accept a change, you have to go in and actually look at the diff and make sure it fits the intent. We accept less than 20% of these.” Scope it honestly. That is one company’s internal skill system, and no denominator is published.
I set out to publish a second figure and could not. Rejections leave no trace. Claude Code’s auto memory writes a file when it decides something is worth remembering. A memory you would have declined is just a file you delete later, indistinguishable from one that went stale. The surface does not instrument the decision that matters. That is a finding in itself, and it is the first thing I would change about the feature.
Put the published research alongside the practitioner number and they agree. SWE-Skills-Bench found “39 of 49 skills yield zero pass-rate improvement, and the average gain is only +1.2%,” with “Token overhead varies from modest savings to a 451% increase while pass rates remain unchanged” (arXiv:2603.15401). Only seven skills produced meaningful gains. An acceptance rate under 20% is not pessimism. It is roughly the hit rate on public skills when someone measures it.
So instrument acceptance before you instrument anything else. Above 50%, either your proposer is unusually good or your review is not real. Near zero, turn the proposer off and keep the diagnostic. Whatever survives review belongs on the promotion ladder, not back in the memory directory.
What 51 auto-written memories actually contained
Fifty-one files, ten repositories, six months, zero deletions. That is what Claude Code auto memory produced on my machine while I was not looking. Classifying every entry says more about the feature than an acceptance rate would have.
Our finding: Across 10 repositories and roughly six months (oldest file 2026-02-08, newest 2026-08-14), Claude Code’s auto memory has written 51 topic files on this machine. Zero have been deleted. Every memory path that appears as a write in my 2,406 session transcripts still exists on disk, and no session ever removed one. Classifying all 51 by hand from their frontmatter description and body: 22 are ephemeral status snapshots (“R008 benchmark status,” “branch is complete but unmerged”), 11 are documentation gaps that belong in a committed doc, 10 are machine-local harness facts (Colima env vars, a toolchain switch, a phantom-lint quirk in worktrees), 6 are constraints that belong in CLAUDE.md as prohibitions, and 2 are empty or undescribed. A keyword sweep for DONE, LANDED, shipped, status, pending, and handover flags 19 of the 49 described files, so the status-snapshot share is reproducible without trusting my judgement. Not one of the 51 was a fact with no better home.
Two details make that corpus more interesting than the counts alone. One memory exists only because the memory system produced a false claim. Its description reads: “The remember/auto-memory summaries have twice claimed unbuilt work shipped; verify ‘was X built’ against git, never against the index.” The agent wrote a memory to warn itself about its own memory, and the correct authority turned out to be git.
The other detail is a mechanism, not an accident. Anthropic’s docs state that the 200-line and 25KB budget “applies only to MEMORY.md,” and that Claude Code “excludes the files in the memory directory” from the retention sweep that deletes old session transcripts (Claude Code memory docs). Topic files are unbudgeted and unswept. Append-only is the default behaviour of the design, which is exactly what theahura reports at much larger scale: “The agents are also terrible at actually removing context, which is a critical capability for maintaining long term memory. I mean, across literal thousands of sessions, I’ve never seen it happen even once.”
What is agent memory actually good for?
Do not delete the memory write path. Redirect it. A memory the agent felt it had to write is a fact the harness failed to supply, which makes the memory queue a bug tracker for your setup.
The reframe comes from a commenter on the Hacker News thread about the transcript post, which drew 180 points and 83 comments. The first clause is the important one: “I don’t do anything with full session transcripts, but I find value when Claude writes a memory. I don’t actually want Claude to have those memories, but they often point to gaps in my harness. I’ll occasionally sweep the memories, pick out what should go into the harness or CLAUDE.md, then delete the memories” (chickensong, HN 48776232, 2026-07-03).
My own corpus is what happens when you skip that last step for six months. Here is the loop, in four parts.
- Leave auto memory writing on. It ships on by default and stores to
~/.claude/projects/<project>/memory/with aMEMORY.mdindex (Claude Code memory docs). - Read the queue weekly and classify each entry. Harness gap, meaning the agent should not have needed to learn it. Doc gap, meaning it belongs in a committed file. Constraint, meaning it belongs in CLAUDE.md as a prohibition. Or noise.
- Fix the gap at its source, then delete the memory. The fix is the artifact. The memory was the symptom.
- Track the queue’s size over time. Shrinking means the harness is absorbing the lessons. Flat means you are re-learning the same thing every week.
Two people in that thread report memory working, and they deserve an answer rather than silence. shepherdjerred disagrees outright and has Claude and Codex keep session logs, prompted from an AGENTS.md, with a public repo to check. That is worth taking seriously, and it also fits the thesis rather than breaking it. A log a human prompted, shaped, and can read is an artifact. An unreviewed transcript is not. The review gate again.
There is one class where memory genuinely earns its slot, and my classification found it: machine-local facts. Ten of my 51 files describe things the repository cannot hold, such as which Docker runtime the test suite needs on this laptop and which lint errors are phantoms in a worktree. Those cannot be committed, because they are not true for the next person. They are also, precisely, harness gaps. Fix the harness and even those disappear.
This is the eviction destination for context that fails all four tests in the CLAUDE.md audit and fails the skill test too. It becomes a signal, not a stored fact. It also shortens the list in session handoff and the attention budget. Less needs to survive a cut than that post assumed. Most of what an agent wants to carry across is a note about your setup, not about your problem.
What changes now, and what will change this again
Four edits to the standing recommendations, and one reason to expect this audit to need its own audit.
- Do not build transcript search. Build commit-message and PR-description discipline instead, and make the code navigable. The substrate is already reviewed and already indexed.
- Write persistent instructions as prohibitions, not directives. Polarity is the one content property that measurably separates helpful rules from harmful ones (arXiv:2604.11088).
- Keep progressive disclosure one level deep, and stop expecting accuracy from it. A second routing level “never helps and sometimes breaks accuracy outright, so one level is enough” (arXiv:2607.17598, Jul 2026), which converges with Anthropic’s own authoring guidance.
- Instrument the acceptance rate on anything the agent proposes about itself. If your tooling makes that impossible, as mine did, treat that as the gap to close first.
Now the churn caveat, because it is the honest ending. Scaffolding expires, and the sharpest evidence is first-party. Anthropic “removed over 80% of Claude Code’s system prompt for models like Claude Opus 5 and Claude Fable 5 with no measurable loss on our coding evaluations” (Anthropic, Jul 2026). The caveat: “our coding evaluations” rests on a suite that is not published. The direction is still unambiguous. Scaffolding a 2026 model needed, a 2027 model may not.
My own analysis, unattributed because nobody has measured it: skill and prompt content tuned against one model’s behaviour decays when the model changes. Almost nobody runs a regression suite over their skill catalog on upgrade. The churn is documented even if the decay is not, in the retired model IDs and tokenizer changes covered in the deprecation treadmill. There is a second-order version too. Paul Bakaus notes that “if everybody uses the same skill to do frontend design work or something like that, everything ends up looking the same” (Latent Space, Jul 2026). Shared skills converge on shared output.
Which brings this back to engineering that outlasts the paradigm. The durable asset is not the recommendation. It is the willingness to re-run the check.
FAQ
Does AI agent memory actually improve agent performance?
For raw session transcripts, there is no published evidence that it does. One practitioner reports zero benefit on software-engineering tasks when other context is available (12 Grams of Carbon, 2026), with no methodology disclosed. Controlled work finds agent-written self-memory scoring 42% against 47% for basic retrieval (arXiv:2606.29914, 2026). Curated, human-reviewed context is a different matter and does help.
Do self-generated agent skills work as well as curated ones?
No. SkillsBench v1 found curated skills raised average pass rate 16.2 percentage points, while “self-generated Skills provide no benefit on average, showing that models cannot reliably author the procedural knowledge they benefit from consuming” (arXiv:2602.12670v1, Feb 2026). One practitioner accepts under 20% of the skill updates his agents propose.
Does clean code make coding agents better, or just cheaper?
Cheaper. Across 660 trials on matched repository pairs, cleaner code left the pass rate unchanged while cutting tokens 7 to 8% and file revisitations 34% (arXiv:2605.20049, 2026). Note that both authors work at SonarSource, which sells the analyzer defining “clean.”
Should I turn on Claude Code auto memory?
Turn on the writing, not the trusting. Auto memory is on by default and loads the first 200 lines or 25KB of MEMORY.md every session (Claude Code memory docs). The higher-value use is to read the queue as a list of gaps in your harness and fix each one at its source. In my own directory, 51 files accumulated over six months and none of them was a fact with no better home.
Pick one number this week
Six claims hold, five weaken, two fail. Both failures came from importing a number measured on a different task. Transcript recall shows no reported benefit and a named mechanism for why. Structured artifacts win on brevity, review, and polarity rather than on structure. And the acceptance rate on self-proposed updates is the metric to instrument, right after you make it possible to compute.
The one-line version: cleaner and shorter input does not raise the ceiling, it lowers the cost. Every result in this post is a variant of that.
So pick one number this week. Open your memory directory, take the last twenty things your agent decided to remember about itself, and count how many you would actually accept today. If the count is low, you do not have a memory problem. You have a harness with gaps, and a queue that has been listing them for you the whole time.
Sources
- theahura, 12 Grams of Carbon, Agentics: memorizing session transcripts isn’t useful, 2026-07-02, retrieved 2026-08-15.
- Hacker News, Memorizing session transcripts isn’t useful, 2026-07-03, 180 points, 83 comments, retrieved 2026-08-15.
- Kuan Wang, MemDelta: controlled baselines and hidden confounds in agent memory evaluation (arXiv:2606.29914), 2026-06-29, retrieved 2026-08-15.
- Priyansh Trivedi and Olivier Schmitt (SonarSource), Does code cleanliness affect coding agents? A controlled minimal-pair study (arXiv:2605.20049), 2026-05-19, retrieved 2026-08-15.
- Zhang et al., Guardrails beat guidance: a large-scale study of rules, skills, and persistent configuration for coding agents (arXiv:2604.11088), 2026-04-13, rev. 2026-05-28, retrieved 2026-08-15.
- Li et al., SkillsBench: benchmarking how well agent skills work across diverse tasks, v1 (arXiv:2602.12670v1), 2026-02-13, retrieved 2026-08-15. Cited as v1 because later revisions drop the self-generated result from the abstract.
- Han et al., SWE-Skills-Bench: do agent skills actually help in real-world software engineering? (arXiv:2603.15401), 2026-03-16, retrieved 2026-08-15.
- He et al., Is progressive disclosure all you need for long-context agents? (arXiv:2607.17598), 2026-07-20, retrieved 2026-08-15.
- Chen et al., SkillJuror: measuring how agent skill organization changes runtime behavior (arXiv:2606.11543), 2026-06-10, retrieved 2026-08-15.
- Zhu et al., SWE Context Bench: a benchmark for context learning in coding (arXiv:2602.08316), v3 2026-05-06, retrieved 2026-08-15.
- Databricks, Benchmarking coding agents on a multi-million-line codebase, 2026-07-08, retrieved 2026-08-15.
- Anthropic, The new rules of context engineering for Claude 5 generation models, 2026-07-24, retrieved 2026-08-15.
- Anthropic, How Claude remembers your project (Claude Code memory docs), living document, retrieved 2026-08-15.
- Anthropic, Agent Skills overview, living document, retrieved 2026-08-15.
- Latent Space, Skill engineering and design, with Paul Bakaus, 2026-07-02, retrieved 2026-08-15.
- Redmonk, 10 things developers want from their agentic IDEs in 2025, 2025-12-22, retrieved 2026-08-15. Cited here only as the source of a correlation this post retracts.
- First-party data: 51 Claude Code auto-memory topic files across 10 repository directories and 2,406 session transcripts on the author’s machine, counted and classified 2026-08-15.
If it was useful, pass it along.