Your Agent's Tests Pass. The Bug Is in the Chat Log.
One run from the pilot, verbatim. The function was a small CSV serialiser from one of my own repositories, with one planted bug: fields containing a double quote got wrapped in quotes, but the quotes inside were never doubled, so She said "hi" came out as invalid CSV. I asked Claude Sonnet 5 to write unit tests for it. It wrote ten, and all ten passed. One of them was this:
it('quotes values containing double quotes', () => {
// Note: current implementation does not double internal quotes, which
// produces invalid CSV for this input. Test pins current behavior.
expect(toCsv([{ note: 'She said "hi"' }])).toBe('note\n"She said "hi""\n')
})
Then I ran Stryker over the suite. Mutation score: 1.00. Every mutant killed. A mutation-testing gate in CI would have waved it through with a perfect grade, and the only place the bug was written down was a comment inside a test that asserts the bug is correct.
That run is an anecdote. So I measured agent-written tests properly: 18 functions, one planted bug each, three ways of briefing the agent, two models, three repeats. 324 runs, $60.47.
Key Takeaways
- Given only the buggy code, Claude Sonnet 5 wrote a suite that passed on the bug in 53 of 54 runs, and Claude Opus 5.5 in 40 of 54. The suites caught the planted bug 2% and 26% of the time.
- Suites that locked the bug in scored a median 0.95 on mutation testing. Suites that caught it scored 0.87. Mutation score ranked the wrong suites higher, and 21% of the lock-in suites scored a perfect 1.00.
- For operator bugs such as off-by-ones, Stryker usually generates the fix itself as a mutant. In 38 of 46 green suites the tests killed it: rejecting the correct code counted towards the score.
- Opus named the bug in its closing message in 35 of 54 runs, but put it in a failing test in only 14. In the other 21 the only record of the bug is chat text nobody keeps.
- Adding a written spec to the same prompt lifted detection to 96% (Sonnet) and 100% (Opus). The agent was never short of ability. It was short of a second opinion about what the code was for.
Why do the studies of AI-written tests disagree?
Two things published this summer seem to disagree.
In August, arXiv 2608.15188 scored hundreds of Claude-authored Python tests one by one, using historical reverts, AST mutation and coverage-guided mutation, and concluded that tests written by recent Claude models are “no weaker” than the human-written suites in Django and Pandas. It is a careful paper. Its tests came from a real tool Claude built over several weeks, not from toy prompts.
In June, arXiv 2606.18168 looked at 86,156 test-file patches in agent-authored pull requests across five coding agents and found 80.2% of them had weak or no explicit oracle: code that runs the function without really checking what comes back. And practitioners keep writing the same complaint on Hacker News and dev.to: the agent’s tests mirror the code instead of checking it.
Both can be true, and a third paper explains how. Zhao, Zhou and Cohen (arXiv 2607.22880), July 2026, found that coverage and mutation scores are meaningful when the code under test is assumed correct, but “no longer serve as reliable indicators” when the code may already be buggy and the job is to expose that bug. The August paper measured tests by injecting faults into code as it stands. That measures how well a suite defends the current behaviour. It cannot see whether the current behaviour was right in the first place.
That is the scenario I care about, because it is the everyday workflow. The agent writes a function, or you hand it one, and then someone says “now write tests for it”. If the function is wrong, what do the tests do? And does anything we would normally put in CI notice?
The experiment
The harness is small and boring on purpose. Everything below is in github.com/iceinvein/mutgap.
The corpus. 18 exported TypeScript functions, each 10 to 80 lines, lifted from five of my own repositories: a diff sharder, a .env parser, a CSV exporter, a glob-to-regex converter, relative-time formatters, a job state machine, a pricing estimator and so on. Each one is copied with only its local helpers so it stands alone. For each I wrote a plain-English spec describing its real behaviour, quirks included, and a hand-written oracle test that passes on the correct version.
The planted bug. Each function got exactly one realistic, single-hunk bug. Half are operator slips, such as <= 50 becoming < 50, where Stryker’s own mutators can turn the buggy line back into the correct one. The other half are semantic, such as splitting on the last = instead of the first or forgetting to double quotes, where no single mutant can revert it. A checker script enforces all of this before any agent is paid to run: the oracle must pass on the reference and fail on the buggy version, and the bug’s label must match what Stryker can actually generate.
Three briefs. Each run is a fresh headless claude -p session in an empty temporary directory, asked in one sentence to “write unit tests for” the function with vitest. The briefs differ only in what the agent can see:
- Code only: the buggy implementation. This is “here’s the function, write tests”.
- Code and spec: the buggy implementation plus
spec.md. This is “here’s the ticket and the code”. - Spec only: the spec and a stub that throws. This is test-first.
The prompt never mentions bugs or correctness. It tells the agent that npx vitest run runs the tests, because that is what a real prompt would say, and I come back to what that choice costs below.
Isolation. My global CLAUDE.md contains explicit rules about how to write tests, which is exactly the variable under measurement, so every session ran with --setting-sources project. I checked with a probe run that none of my instructions, hooks or plugins leaked in. The agent got the stock Claude Code system prompt and nothing else. It never saw the reference implementation or the oracle.
Scoring. Each suite the agent left behind is run against both the buggy and the reference implementation. A suite catches the bug if some test fails on the buggy code and passes on the reference. It locks the bug in if some test passes on the buggy code and fails on the reference. Then Stryker runs against whichever implementation the suite passes on, which is the mutation score a CI gate would have reported. Finally, a Haiku judge labels the agent’s closing message as reporting the bug, describing the buggy behaviour as normal, or not mentioning it, with three votes per message.
Two models, Claude Sonnet 5 and Claude Opus 5.5, three repeats of every cell, Claude Code 2.1.283. 18 × 3 × 2 × 3 = 324 runs.
Shown only the code, the tests agree with the code
The top two bars are the whole post in miniature. Given only the implementation, Sonnet’s suite passed on the buggy code in 53 of 54 runs and caught the bug once. Opus’s suite passed on the buggy code in 40 of 54 and caught it 14 times.
The full grid, 54 runs per row:
| Brief | Model | Suite green on the buggy code | Suite catches the bug | Suite locks the bug in |
|---|---|---|---|---|
| Code only | Sonnet 5 | 53/54 | 1/54 | 43/54 |
| Code only | Opus 5.5 | 40/54 | 14/54 | 35/54 |
| Code + spec | Sonnet 5 | 2/54 | 52/54 | 2/54 |
| Code + spec | Opus 5.5 | 0/54 | 54/54 | 0/54 |
| Spec only | Sonnet 5 | 0/54 | 53/54 | 1/54 |
| Spec only | Opus 5.5 | 0/54 | 54/54 | 1/54 |
In 43 of Sonnet’s 54 code-only runs, and 35 of Opus’s, the suite did more than miss the bug. It contained at least one test that asserts the buggy output and fails on the correct code. That is lock-in, and it is worse than a gap in coverage. A missing test lets a bug through once. A test that encodes the bug fights whoever fixes it later: the fix goes red, and the natural reading of a red test is that the fix broke something.
This isn’t laziness. The code-only suites were not thin: the median Sonnet suite had 14 tests and the median Opus suite 22.5. One Sonnet run on the effort-score function reported, word for word: “All 14 tests pass. Created src/computeEffortScore.test.ts covering the empty-files case and every threshold boundary (file count and line count) on both sides.” It tested the boundary carefully, then asserted that the wrong side of it was correct. From the code alone there is no way to tell that 50 was meant to be inclusive. The agent did the only thing the evidence supported.
This is the same dynamic I wrote about in Narcissus at the keyboard, with a different trigger. There, the model agrees with the user. Here, it agrees with the code, because the code is the only authority in the room.
Does mutation testing catch tests that lock in bugs?
The standard answer to “are these tests any good?” is mutation testing. Awesome Testing and a dev.to walkthrough both made that case in August, specifically for agent-written tests. I agree with them about regressions. For this failure it points the wrong way.
Across both models, the 77 code-only suites that locked the bug in scored a median of 0.95. The 14 that caught it scored 0.87. Of the lock-in suites, 64% cleared 0.90 and 21% scored a perfect 1.00. The CSV suite at the top of this post was not a fluke.
The comparison is not quite like for like, and it doesn’t need to be. The lock-in suites are scored on the buggy code because that is the only code they pass on, which is exactly the situation in CI: the branch contains the bug and the suite is green. The catching suites are red on that same branch, so a mutation gate never even runs on them. It fails them for failing tests. In a pipeline, the suite that found the bug is the one that gets blocked.
The reason is not mysterious. Mutation testing asks “if this code changed, would the tests notice?” A suite that has pinned every output of the current code, bug included, notices every change. It is a very good regression suite for the wrong behaviour. The Zhao, Zhou and Cohen result predicted this; the experiment just makes it concrete.
Stryker counts rejecting the fix as a kill
The operator bugs expose something sharper. Take the diff sharder. The planted bug changed one comparison:
// reference
(currentLines + c.lines > budget || current.length + 1 > maxFiles)
// buggy
(currentLines + c.lines >= budget || current.length + 1 > maxFiles)
When Stryker mutates the buggy code, its EqualityOperator mutator tries >= → >. That mutant is the fix. A suite that locked the bug in fails on it, so Stryker records the mutant as killed, and the score goes up.
Of the 46 green code-only suites on operator bugs, 38 killed the fix mutant: 83%. The remaining 8 let it survive, which Stryker reports as a weakness in the suite. So on this class of bug, the suites that pinned the bug most thoroughly get the most credit, and the few that happened to leave the correct behaviour untested get marked down for it.
Stryker is doing exactly what it is designed to do. The design assumes the code under test is right. When it isn’t, the tool has no way to know, and it cannot tell a suite that defends correct behaviour from one that defends a bug.
Where the bug went instead: the chat log
Here is the part I did not expect. Opus often knew.
Opus’s closing message reported the planted bug in 35 of its 54 code-only runs. Its suite caught the bug in only 14. In the other 21, the agent told me about the bug in chat and then shipped tests that encode it.
It did this in two ways. Some runs wrote the correct assertion and marked it it.fails, which inverts the test so the suite stays green until someone fixes the code:
// planShards uses `>` in both places.
describe('exact budget boundary', () => {
it.fails('packs files in one group that exactly fill the budget into one shard', () => {
const chunks = [chunk('a/1.ts', 5), chunk('a/2.ts', 5)]
expect(paths(planShards(chunks, 10, 10))).toEqual([['a/1.ts', 'a/2.ts']])
})
})
That is thoughtful work, and it is also invisible to every automated check. The suite passes. The inverted test passes on the bug and fails on the fix, so Stryker counts it as killing the fix mutant too. Five Opus runs used this pattern.
The more common way was plain: pin the current behaviour, then flag it in prose. From an Opus run on the project-summary function:
Two things in the code that may not be intended. The tests check what the code does now, so if you change either one, the matching test needs updating too:
- When cost and event count are tied, names sort Z→A (
b.displayProject.localeCompare(a.displayProject)). If you meant A→Z, swapaandb.
That is the right diagnosis and a reasonable question, and it goes to a terminal scrollback or a PR description. In a headless pipeline, which is where most agent-written tests are heading, nobody reads it. I made the same argument about agents’ error output in your agent recovers from everything it can see: if a signal doesn’t reach a channel something acts on, it doesn’t exist. And in merged is not reviewed: a green check is a claim about the checks, not about the change.
Sonnet rarely got that far. Its closing message reported the bug in 3 of 54 code-only runs. In 14 it described the buggy behaviour as a feature: “last-= splitting behavior”, “the > vs >= asymmetry between group-level and member-level checks”. It noticed the exact line and wrote it up as intended.
What can an agent notice from the code alone?
Detection was not spread evenly. Split by function, the code-only results sort into three groups, and the grouping explains a lot.
The bug breaks a public convention: Opus writes a failing test. Five functions: CSV quoting (RFC 4180), .env values containing =, diff header lines, a pricing function that applied its “fast” multiplier to the slow path, and the diff sharder, whose buggy >= sat right next to a > doing the same job. Opus caught these in 14 of 15 runs. In each case something outside the code said what the code should do: a standard, a file format everyone knows, or the function contradicting itself.
The code looks odd but could be deliberate: Opus mentions it and pins it. Seven functions, including the Z→A tie-break, a marker directory at the end of a path returning an empty string, and 12am parsing as noon. Opus raised these in 20 of 21 runs and wrote a failing test in none. It had doubts but no authority, so it deferred to the code and passed the doubt on in chat.
The bug looks like a design decision: nobody notices. Six functions, including an off-by-one in a “just now” threshold, a tempo guard that accepts zero, and a job state machine that allows one extra transition. Zero catches and zero reports from either model across all 36 runs.
Off-by-one bugs make the point cleanly. Seven of the planted bugs were boundary flips. Setting aside the sharder, which gave itself away, the other six were caught by a suite zero times in 36 runs. There is nothing in if (lines < 50) that tells you 50 was supposed to be included. The information that makes it a bug lives in someone’s head, a ticket, or a spec. It is not in the code, so an agent reading only the code cannot recover it, however capable it is.
Does giving the agent a spec fix it?
Adding spec.md to the same prompt, with the same buggy code in the same directory, changed everything. Sonnet’s detection went from 2% to 96%, and Opus’s from 26% to 100%. The spec-only arm, where the agent writes tests before any implementation exists, landed in the same place: 98% and 100%.
The spec-aware closing messages read differently too. From a Sonnet run with code and spec:
I found a genuine bug while writing these tests: the spec says tier 1’s line limit (50) is inclusive, but the implementation uses
lines < 50instead oflines <= 50. Should I fix that off-by-one incomputeEffortScore.ts, or leave the source as-is and keep the tests documenting the spec (which will leave 2 failing)?
Same model, same bug, same line of code. A code-only Sonnet run on the same function said it had covered “every threshold boundary on both sides” and asserted the wrong side. The only difference is a second source of truth to check the code against.
That is the real finding, and it is a relief. The agents are not bad at writing tests. Given two sources that disagree, they notice the disagreement almost every time and say which one they trust. Given one source, they have nothing to disagree with.
What I would change in a real workflow
These are ordered by how much each one costs.
Give the test-writing agent the intent, not just the code. The ticket, the acceptance criteria, the docstring that describes intent rather than mechanics, or one paragraph typed by hand. That single change moved detection from 2 to 96% for the cheaper model. If there is no written intent anywhere, that is worth knowing before you trust any tests at all.
Write the tests from the spec before the code exists, or in a separate session that never sees it. The spec-only arm matched the code-plus-spec arm. It is the same argument as agentic TDD and spec-driven agent development, now with a number attached: a test written without the implementation cannot copy it.
Treat “the tests check what the code does now” as a finding, not a footnote. If an agent says it pinned current behaviour, or marks tests it.fails or skip, that is a bug report written in the wrong place. A cheap hook can grep the suite for .fails( and .skip( and fail the build, or at least require a linked issue. Chat-only doubts need a route into something a pipeline reads.
Don’t use mutation score as the gate for agent-written tests on new code. It is a good regression metric and a misleading correctness metric. On this data it ranked lock-in suites above catching suites. If you do gate on it, gate on a spec-derived suite, where “the code as it stands” is not the only thing the tests know about. Gating on trajectories and on whether the agent consulted the spec is closer to what you want.
Limitations
This is a small, controlled experiment, and some of its choices lean one way or the other.
- The functions are small. 10 to 80 lines each, with the bug somewhere a careful reader could find it. Real functions are longer and the bugs better hidden, which I would expect to make code-only detection worse, not better.
- The bugs are planted. Each is realistic and single-line, but I chose them. Natural bugs from the agent’s own implementation would make a better test, and the co-authoring arm that would measure them is the obvious next experiment. I left it out of this round because on functions this small I expect frontier models to introduce few bugs of their own, which would leave most runs with nothing to catch.
- The repositories are public. All five are on GitHub and were created in 2026. If a model had memorised a correct version, that would help it spot the planted bug, which biases the result against my claim rather than for it.
- The prompt says how to run the tests. Telling the agent
npx vitest runinvites it to iterate until green, which pushes towards lock-in. I kept it because real prompts say it and agents find the test command anyway. A prompt that forbids running tests would be a useful variant. - One harness, one CLI version, default effort. Claude Code 2.1.283 with the stock system prompt and each model’s default effort. A project
CLAUDE.mdwith test-writing rules would change the result, and that is partly the point: mine has them, and I had to strip them out to see this. - The judge. The reported / described / absent labels come from three Haiku votes per message, unanimous in 199 of 216. A second reader (Claude, in the session that ran the analysis, reading each full message) relabelled a seeded 20% sample and agreed with the judge on 42 of 43 about whether the bug was reported. LLM judges are flaky, so the headline numbers in this post come from the test runs, not the judge. The judge only feeds the chat-log analysis.
- Three repeats per cell. Enough to kill my pilot result (Opus caught 3 of 3 bugs with one run per cell, and 14 of 54 at full scale). Per-function cells of three are still noisy, so read the three groups above as a pattern rather than a set of rates.
Frequently Asked Questions
Are AI-written unit tests worse than human-written tests?
Not in general. arXiv 2608.15188 found recent Claude-authored tests no weaker than Django’s and Pandas’s under fault injection. The failure measured here is narrower: when an agent writes tests for existing code with no statement of intent, its tests encode whatever the code does, bugs included. Human-written characterisation tests fail the same way. Agents just write a lot more of them.
Does mutation testing catch bad AI-generated tests?
It catches weak tests, meaning ones that don’t check outputs. It does not catch tests that check the wrong outputs. In this experiment, suites that locked a planted bug in scored a median 0.95, higher than the 0.87 of suites that caught it, and for operator bugs such as off-by-ones Stryker often generates the correct code as a mutant and credits the suite for rejecting it.
How do I stop Claude Code from writing tests that lock in bugs?
Give it the intended behaviour in the same prompt, as a spec, ticket or acceptance criteria. That lifted detection from 2% to 96% for Sonnet 5 and from 26% to 100% for Opus 5.5. Better still, have it write the tests from the spec before the implementation exists, and fail the build on it.fails or skip markers the agent adds.
Why did Opus mention the bug but not test for it?
When the only evidence was the code, Opus treated the code as the authority and passed its doubts to the user. It wrote failing tests when an outside reference contradicted the code, such as RFC 4180 or the .env format, and pinned the current behaviour when the oddity might have been intended. Without a spec, it had no way to settle which.
The part that should worry you
Every one of these suites would pass review by the usual signals. They are green. They are thorough, with a median of 14 to 22 tests each. They cover the boundaries by name. Most of them score above 0.90 on mutation testing. By every number we normally use to decide that tests are good, the ones that locked the bug in look better than the ones that caught it.
The agent is not the weak link here. Given a second source of truth, both models found the bug almost every time. The weak link is the workflow that hands an agent the code and nothing else, asks it to prove the code works, and then measures the proof with a tool that assumes the code was right. That workflow was always circular. What has changed is the volume: agents write characterisation tests faster than anyone can read them, and the one place they write down their doubts is the place nobody looks.
The corpus, runner, scorer, judge and analysis are at github.com/iceinvein/mutgap, MIT licensed. Point it at your own functions, plant a bug you care about, and see which of your agents’ suites agree with it.
Sources
All retrieved 2026-09-27.
- The Quality of Claude AI-authored Python Tests Is Not Weaker Than Human-authored Tests, arXiv 2608.15188, Aug 2026. https://arxiv.org/html/2608.15188
- All Smoke, No Alarm: Oracle Signals in Agent-Authored Test Code, arXiv 2606.18168, Jun 2026. https://arxiv.org/abs/2606.18168
- Zhao, Zhou and Cohen, Do Coverage and Mutation Scores of LLM-Generated Test Suites Correlate with Their Effectiveness? (Replicability Study), arXiv 2607.22880, Jul 2026. https://arxiv.org/abs/2607.22880
- Awesome Testing, Mutation Testing for Agent-Written Code, 2 Aug 2026. https://www.awesome-testing.com/2026/08/mutation-testing-for-agent-written-code
- Mutation Testing as a Merge Gate for Agent-Written Tests, DEV Community, 26 Aug 2026. https://dev.to/datacpp_8185/mutation-testing-as-a-merge-gate-for-agent-written-tests-5gie
- When AI writes the software, who verifies it?, Hacker News discussion. https://news.ycombinator.com/item?id=47234917
- Stryker Mutator, supported mutators. https://stryker-mutator.io/docs/mutation-testing-elements/supported-mutators/
- Shafranovich, RFC 4180: Common Format and MIME Type for Comma-Separated Values (CSV) Files, IETF, 2005. https://www.rfc-editor.org/rfc/rfc4180
If it was useful, pass it along.