Is Your Code Context Good? The Benchmarks Can't Tell You.
I maintain a code-intelligence MCP server. This post is about failing to prove it improves an agent’s code context.
On 6 November 2025, Cursor published a benchmark for semantic code search. They had trained their own embedding model on agent session traces, scored it on an internal set called Cursor Context Bench, and reported 12.5% higher accuracy on average, ranging from 6.5% to 23.5% depending on the model. In January they scaled the index across organisations, cutting p99 time-to-first-query from 4.03 hours to 21 seconds. In March they shipped Instant Grep, a sparse n-gram index that took a 16.8 second regex search down to 13 milliseconds.
In August 2026 they deleted the semantic search path. No changelog entry. The feature simply stopped existing.
Kevin Neilson, on the Cursor forum on 5 August, explained it in two sentences:
We changed how agents search the codebase. Models got very good at grep / indexed search, so the older dedicated semantic search path was no longer helping in a meaningful way.
That is the only major claim in this entire debate whose incentive runs against the person making it. Augment sells an index. Sourcegraph sells code search. Milvus and Chroma sell vector databases. Their pro-retrieval findings are real, and they are also aligned with the till. Cursor sells a seat, and deleting a marketed differentiator costs them something. When the evidence cuts against the speaker, it is worth more.
So I tried to settle it on my own repositories. Two codebases I own, an eval harness, three models, and a budget. The harness is the same shape as the one in a test suite for your Claude Code setup. What follows is the measurement, and then the considerably more interesting story of how the measurement kept lying to me.
Key Takeaways
- Cursor published a 12.5% accuracy gain for semantic code search in November 2025 (Cursor), then removed the feature in August 2026 with no changelog entry (Cursor forum).
- Across 32 candidate tasks, 214 agent runs and $85 on two repositories I own, zero tasks discriminate between grep and a code-intelligence index on Claude Opus 5.
- Three earlier results looked publishable and all three dissolved: a 54% cost saving that reversed sign on the second repository, a recall delta exactly the size of its own run-to-run noise, and an answer key pointing at the wrong file.
- The strongest paper against this position reports +39.6pp from a structural index, but that is the within-harness ablation. Against an independent grep agent the gap is +9.2pp at p=0.080, which does not clear significance (arXiv:2606.22417, June 2026).
What is code context, actually?
Code context is not the tokens in the window. It is the retrieval step that decides which tokens get there, and it is worth separating into three layers that the public argument keeps collapsing into one.
The haystack is the shape of the repository: how many files, how they are organised, whether the thing you are looking for shares vocabulary with the thing you would search for. The retrieval mechanism is what finds candidates: text search, a lexical index, a structural index over call graphs and types, or semantic similarity over embeddings. The stopping rule is what the agent does once it has something plausible, which turns out to matter more than most of this discussion admits.
Anthropic gave the dominant pattern a name in September 2025: just-in-time context loading. Rather than pre-processing a corpus, the agent holds lightweight identifiers, file paths and stored queries, and pulls content at runtime through tools. Every major coding agent now works this way. The argument is only about what sits behind the tool.
There is a hard floor under all of it. Chroma’s context rot work measured 18 frontier models and found every one degrading as input length grew, even on trivial copy-and-retrieve tasks. That is a property of transformer attention, not a capability gap that scale closes. Retrieval does not go away because windows got bigger. It gets more important, because feeding the window badly now costs more.
For where each piece of context should live in practice, I wrote that up separately in context engineering in practice. For why the repository’s own shape is the variable nobody controls for, see your codebase is the agent’s operating environment.
The industry reversed itself twice, and almost nobody noticed
The story everyone tells is “RAG is dead, agents just grep now.” The actual sequence is stranger, and the strangeness is the point.
Anthropic pulled embeddings and a local vector database out of Claude Code in mid-2025. Boris Cherny’s summary was that plain agentic search “outperformed everything. By a lot.” Windsurf, Cline, Devin and Amp followed. Sourcegraph had already dropped embeddings from Cody in favour of BM25F over its code graph.
Then Cursor went the other way, hard, and published numbers for it. Then Cursor went back.
Here is the part the “RAG is dead” framing gets wrong. The industry did not reject indexing. It rejected fuzzy indexing. Cursor did not go back to raw ripgrep in March; they built a harder index, just a lexical one. Sourcegraph kept SCIP. Nobody dropped structure. Two vendors dropped guessing, twice, and the thing that survived both rounds is deterministic lookup with a freshness guarantee.
Vicent Martí, the Cursor engineer who wrote up Instant Grep, put the freshness requirement in terms that make the distinction concrete: if an agent searches for text it just wrote and doesn’t find it, it “goes into a wild goose chase.” Embeddings degrade gracefully when code changes underneath them. Exact-match search does not, which is exactly why the exact-match index is the one that needs to be right.
One thing to note before moving on, because it affects what you’ll read elsewhere. A good deal of the secondary coverage still asserts that Cursor uses semantic search, quoting the November 2025 post as though it were current. It isn’t. Cursor’s own agent-tools documentation now lists two search capabilities: Instant Grep, and an Explore subagent that runs in its own context window. Semantic search is not among them.
Why is the counter-evidence all older than your model?
Augment’s case is the serious one, and it is not a vendor blog post with a vibe. Auggie tops SWE-Bench Pro at 51.80% using Claude Opus 4.5, against Cursor at 50.21%, Claude Code at 49.75% and OpenAI Codex at 46.47%. Same underlying model, roughly 15 to 17 more problems solved out of 731. In February 2026 they unbundled the Context Engine as an MCP server and measured third-party agents with and without it across 300 Elasticsearch pull requests with three prompts each, 900 attempts: Claude Code improved 80%, Cursor with Opus 4.5 improved 71%.
Sourcegraph’s June 2026 result is sharper still. Claude Sonnet 4.6 with their MCP server scored 0.698 at $1.02 per quality point. Claude Fable 5, a more capable model working from a plain local checkout, scored 0.568 at $1.83. The cheaper model with retrieval beat the frontier model without it, and the gap concentrated on cross-repository symbol discovery rather than spreading evenly.
Now look at what every one of those numbers was measured on.
Amazon’s “Keyword search is all you need” paper at AAAI 2026, the most-cited support for the grep position, measured agentic keyword search at 94.5% of RAG faithfulness with no vector store at all. It ran on Claude 3 Sonnet at 200K context. That is not a slightly older model. Augment’s figures are Opus 4.5. Chroma’s context rot work is the GPT-4.1 and Claude 4 era.
Opus 5, Fable 5.1, GPT-5.6 and GPT-6 Astra all shipped after the most recent number in that list. Nobody has published a current-generation measurement, and the one decision that actually reversed the field, Cursor’s removal, shipped with no published number at all.
That is the gap I went after.
What happens when you measure it yourself?
The design is deliberately boring. Localisation tasks: describe a behaviour, ask which source files implement it, score recall against a hand-written expected set. Ground truth comes from real commits in the repository’s own history, which gives an answer that a human already committed to rather than one I invented. Prompts are written by hand from the diffs so they describe behaviour without naming files.
Two arms, differing in exactly one thing: whether code intelligence exists on the machine. The grep arm runs with a PATH where the code-intel binary is absent, which both removes the tool and stands down the hook that would otherwise deny text search. Everything else is held constant: same prompt, same tools, same effort setting, same repository, same model.
One confounder, stated up front because it affects how you should read the cost numbers. In the code-intel arm, my own PreToolUse hook denies grep and redirects to the index. The agent is pushed rather than choosing. That measures a configuration, which is what I actually run, but it is not a neutral A/B between two freely-selected tools. I built a third arm later to settle it, and the answer turned out to be more interesting than the cost figures it was meant to correct.
The first repository was a Rust and TypeScript desktop application, roughly 3,500 files. The result looked clean. Code-intel was 54% cheaper on Sonnet and 28% cheaper on Opus, with fewer turns, at identical recall.
Then I ran the second repository, a TypeScript monorepo where the agent searches about 13,500 files and the answers live in roughly 1,800 of first-party code. Code-intel came out 15 to 17% dearer, using about 2.5 more turns per task.
Same tool. Same harness. Opposite sign.
If I had stopped after the first repository, I would have published a cost saving. It would not have replicated, and nobody reading it would have known.
Replication killed the only delta I had
At one run per cell, the accuracy numbers looked like a small real effect: Sonnet +0.07 recall for code-intel, Opus +0.00. Small, plausible, the sort of thing that goes in a table.
So I ran each cell three times. On the only task in the set that wasn’t already saturated, the within-cell spread was 0.33: one file out of three, appearing and disappearing between identical runs.
My original finding, the one that would have gone in the post, was run one of the Sonnet code-intel cell. Runs two and three came back 0.67. It was a coin flip, and I would have published it.
This is the same failure shape I wrote about in LLM-as-judge is the new flaky test, arriving from a completely different direction. A measurement that varies as much within a condition as between conditions is not a measurement. Three of four cells flipped on the same single file.
Then the tasks turned out to be the problem
The moment the whole thing came apart was small. I ran one command out of curiosity, searching for the name of the one shared helper the task was built around:
rg -l <the shared symbol>
It returned exactly the three files in my ground truth. Exactly. No extras, nothing missing. The same held for the other task I had built: one search for the symbol name produced the complete answer.
Both arms scored the same because when one grep answers the question, the retrieval layer is irrelevant. I had spent real money comparing two ways of finding something that a single command finds perfectly.
Then it got worse. On the one task that wasn’t trivially greppable, 19 of 24 runs missed the same file, and every one of them substituted the same alternative. My prompt described a deduplication step happening before a set of items reached a menu. My ground truth named the file where the dedupe call sits, because that is what the commit touched. The runs named the menu component itself.
They were reading my prompt correctly. I was marking them wrong.
That was the third ground-truth defect. The first was a file whose only change in the source commit sat inside a Rust mod smoke_tests block, while my own prompt instructed the agent to exclude test files: six runs penalised for following instructions. The second was a prompt so ambiguous that every arm and every model scored exactly 0.50 on it.
Three defective task sets, in a row, built by someone with full repository access who read every diff before writing every prompt.
How do you tell a real retrieval task from a broken one?
At that point the useful artefact stopped being the number and became the screening. Every rule below exists because I violated it and paid to find out. All eight are implemented in ctxeval, along with the runner and the scorer.
| Rule | Rejects | What it cost to learn |
|---|---|---|
| R1 | Expected file absent at HEAD | Caught pre-emptively; a rename silently breaks a task |
| R2 | Expected file’s change is test-only | 6 runs marked wrong for obeying “exclude test files” |
| R3 | One literal grep reproduces the expected set | An entire 5-task set, about $20 |
| R4 | One word of the asker’s vocabulary finds the answer | The word “assembles”: 14 hits, 100% of an answer |
| R5 | Repeated runs disagree with each other | 19 of 24 runs named the same defensible alternative |
| R6 | Both arms saturate at 1.00 | 8 tasks across three rounds |
| R6b | Arms score identically at any level | A task both arms answered 0.67 on, every run |
| R7 | Runs agree on an answer that isn’t yours | All four runs converged on a different, coherent answer |
Two of these are worth dwelling on.
R4 caught a word I picked specifically to be safe. I was writing prompts in plain language precisely to avoid the code’s jargon, and I chose “assembles” to describe building a video. It appears in 14 files in that repository, and those 14 include every file in the answer. Hand-written prompts leak vocabulary in ways that are invisible to the person writing them. You cannot introspect your way out of this; you have to test it.
R7 is the one I would most want other people to steal. R5 checks that repeated runs agree with each other. That is necessary and it is not sufficient, because runs can agree perfectly on an answer that is not the recorded one. One of my tasks had 75% inter-run agreement at 0.00 recall: every run confidently converged on the body-extraction OCR path while my ground truth named the sidebar-redaction path. Both exist in that codebase. My prompt never said which. Inter-run agreement and correctness are different axes, and only the pair of them together tells you whether you have a hard task or a wrong answer key.
One design point matters more than it looks. The saturation rule has to be symmetric. The tempting version is “reject tasks the grep arm already solves,” which would leave only tasks grep fails, and the index would then win by construction. That is selecting on the dependent variable, and it is precisely the error this whole post accuses the field of. Rejecting only when both arms saturate keeps the cull neutral.
The strongest result against me, read properly
In June 2026 a team published Code Isn’t Memory: A Structural Codebase Index Inside a Coding Agent. It is the most careful work in this space and it points the opposite way from my result. Claude Opus 4.7, SWE-PolyBench Verified plus SWE-bench Pro, 91 instances across Go, Java and Python, three seeds per configuration, significance testing, and public code and data.
Their index is a hybrid: embeddings for similarity, a call graph for reachability, BM25 over identifiers for exact match. Their headline numbers are large. Localisation accuracy at 5 went from 44.3% to 84.5%, a 39.6 point gain at p below 0.0001. Resolve rate went from 41.9% to 50.4%, 7.9 points at p = 0.003. Cost per solved task fell 21%. Turns dropped from 36.0 to 28.3.
Read the arms carefully, though.
The headline figures are the within-harness ablation: their own agent, with its own index switched off. That measures architecture coupling, and it measures it convincingly. An agent designed around an index degrades sharply when you take the index away.
The cross-harness comparison, their agent against an independent grep-based agent, gives +9.2 points on localisation at p = 0.080 and +5.1 points on resolve at p = 0.087. Neither clears 0.05.
Those are different questions. “Does my agent need the index it was built around” is answered emphatically yes. “Should I add an index to an agent that doesn’t have one” is answered by the cross-harness arm, and that arm does not separate. The number that travels in summaries is the first one.
I want to be scrupulous here, because this paper is better than my eval in several respects and I am not going to pretend otherwise. They measured task completion; I measured localisation. Their index is purpose-built and hybrid; I tested one shipped product. They ran Opus 4.7; I ran Opus 5. Their tasks come from curated benchmarks with known-good ground truth; mine came from commit history and three of my sets were broken. On task resolution, with a hybrid structural index, inside a harness designed for it, there is a real and statistically separated gain, and my null does not touch that claim.
What my result does sit beside, comfortably, is their non-significant cross-harness arm.
What happens if you just let the agent choose?
The cost numbers above have a thumb on the scale, and it is mine. In the index arm a hook denies text search and redirects to the index, so the agent is pushed rather than choosing. Measuring what it does unpushed needs the index available and the hook inert at the same time, which sounds contradictory until you look at what the hook actually checks.
It checks whether a binary named code-intel is on PATH. So expose the identical CLI under a different name, ci-search, and keep code-intel off the path. The hook finds nothing to enforce and stands down. The agent has full code intelligence, described in a system prompt that names both options and says in as many words that neither is preferred.
Fifteen runs, the same five tasks, the same model as the other two arms.
The agent called the index zero times.
| Tool | Calls across 15 free-choice runs |
|---|---|
| Grep | 89 |
| Read | 12 |
grep via Bash | 3 |
| Glob | 2 |
ci-search | 0 |
A null that convenient deserves suspicion, so I checked it. All fifteen transcripts contain the prompt offering both tools. The string ci-search appears in every one. The CLI returns correct results from that exact PATH with exit code 0. It was present, working, and described. It was never once invoked.
That closes the confound. Every cost and turn difference between my grep and index arms was the hook’s doing rather than the tool’s. Left alone, the agent behaves like the grep arm, because behaviourally it is the grep arm.
It also asks a better question than the one I started with. The index was not weighed and found wanting. It was not reached for. Retrieval quality is what this entire debate argues about, and it never got as far as mattering, because an index nobody calls is worth nothing however good its answers are. Whether that is the model’s judgement, my prompt’s framing, or simply that grep is the reflex a model has seen ten million times in training, fifteen runs cannot tell you. But “is the retrieval better” turns out to sit downstream of “will anything actually call it.”
What I now think, and what it does to my earlier post
Here is the finding at the width the evidence actually supports. Across 32 candidate tasks and three screening rounds on repositories I own, I could not construct a single question where Claude Opus 5 localises code better with a code-intelligence index than with grep alone. On the two tasks that ever discriminated for Sonnet 5, at −0.40 and +0.20, Opus flattened both to 0.00 and −0.07.
That is not “indexes are useless.” It is a statement about how hard the measurement is, and about where the headroom went. It is also the second time I have argued that you have to evaluate this yourself rather than trust a leaderboard.
It also forces me to revisit The Project Graph, and behind it the case I made for local code intelligence, where I reported code-intelligence winning on 40 questions across two large repositories, with three LLM judges scoring it 7.12 against 6.30 for default tooling, and citing sources in 50% of answers against CodeGraph’s 32%. I am not retracting that post. I am saying three things about it.
It used LLM-judge scoring on open-ended questions, which is a more forgiving instrument than file-level recall against a fixed answer key, and one whose own reliability I have since written about. It ran on an older model generation. And it reported no variance and applied no admission screening, which means I cannot now tell you how much of that 0.82 point gap would survive three repeats. I would not publish it today in that form. The honest summary is that the effect it measured is one the current frontier model appears to close, at least for localisation.
If you are deciding what to do on Monday: keep lexical and structural retrieval, because it is cheap, deterministic, fresh, and nothing in any of this argues against it. Stop expecting semantic retrieval to show up in your numbers. Check whether your agent is calling the index you installed at all, because mine never did until a hook made it. And if you are going to measure, report the variance or do not report at all. That last one is not a style preference. It is the difference between my first three results and my fourth.
The limits, stated rather than buried: one repository for the vetted work, localisation only rather than task completion, Haiku never run on the second repository, and a free-choice result resting on fifteen runs with a single prompt wording, which is enough to show the index went untouched and nowhere near enough to say why.
Frequently Asked Questions
Is grep enough for AI coding agents in 2026?
For localisation on the repositories I tested, yes. Grep plus Opus 5 found the right files on tasks built specifically to defeat text search. Amazon’s AAAI 2026 paper measured agentic keyword search at 94.5% of RAG faithfulness with no vector store, though on Claude 3 Sonnet, three model generations ago.
Why did Cursor remove semantic search?
Cursor engineer Kevin Neilson stated on 5 August 2026 that models had become good enough at grep and indexed search that the dedicated semantic path “was no longer helping in a meaningful way.” No changelog entry accompanied the removal, despite the feature shipping with a published 12.5% benchmark in November 2025.
Does this mean code-intelligence tooling is useless?
No, and this eval cannot support that claim. The strongest contrary evidence is arXiv 2606.22417, where a hybrid structural index on Opus 4.7 lifted resolve rate from 41.9% to 50.4% at p = 0.003 and 21% lower cost per solved task. Note that is the ablation arm; the cross-harness comparison gives +5.1 points at p = 0.087.
How do I tell whether my own code context is good?
Write localisation tasks from your own commit history, then screen them with ctxeval. Reject any task where one rg -l <symbol> reproduces the expected file set, and any where independent runs disagree with each other. Run every cell at least three times and report the spread. Most candidates will fail screening, and that is the finding rather than a setback.
The part that should worry you
I had full access to both repositories. I read every diff before writing every prompt. I verified every expected file existed. I still shipped three defective task sets in a row, and each one produced a number I would have published if I had stopped at that round: a 54% cost saving that reversed sign on the next repository, a recall delta identical in size to its own noise, and a “code intelligence wins” result that was my answer key pointing at the wrong file.
Constructing a working code-context benchmark turned out to be much harder than running one. If it is this hard for someone measuring their own repositories with no product to sell, then a vendor publishing a headline delta without its admission criteria or its variance has most likely measured the thin, ill-posed band where ambiguity dominates the signal. That is not an accusation of dishonesty. It is a mechanism, and it explains how Cursor’s 12.5% could be honestly measured in November 2025 and genuinely gone by August 2026.
The instrument is the contribution here, not the number. The number has a shelf life measured in model releases, and so will yours.
The eight rules, the runner and the scorer are at github.com/iceinvein/ctxeval, MIT licensed. It ships with five hand-written tasks against a public repository, four of which the rules reject. Clone it, point it at your own repo, and the first thing it will tell you is how many of your tasks were never going to measure anything.
Sources
All retrieved 2026-09-19.
- Cursor, Improving agent with semantic search, 6 Nov 2025. https://cursor.com/blog/semsearch
- Cursor, Securely indexing large codebases, 27 Jan 2026. https://cursor.com/blog/secure-codebase-indexing
- Cursor, Fast regex search: indexing text for agent tools, 24 Mar 2026. https://cursor.com/blog/fast-regex-search
- Cursor community forum, Codebase indexing not working in v3.12 onwards, Kevin Neilson, 5 Aug 2026. https://forum.cursor.com/t/codebase-indexing-not-working-in-v-3-12-onwards/167527
- Cursor, Agent search tools documentation. https://cursor.com/docs/agent/tools/search
- Augment Code, Auggie tops SWE-Bench Pro, 4 Feb 2026. https://www.augmentcode.com/blog/auggie-tops-swe-bench-pro
- Augment Code, Context Engine MCP is now available for any AI coding agent, 6 Feb 2026. https://www.augmentcode.com/blog/context-engine-mcp-now-live
- Sourcegraph, Sourcegraph MCP server and a cheaper model beat a Mythos-class model alone, 16 Jun 2026. https://sourcegraph.com/blog/sourcegraph-mcp-and-a-cheaper-model-beat-a-mythos-class-model-alone
- Subramanian et al., Keyword search is all you need: achieving RAG-level performance without vector databases using agentic tool use, AAAI 2026. https://www.amazon.science/publications/keyword-search-is-all-you-need-achieving-rag-level-performance-without-vector-databases-using-agentic-tool-use
- Hong, Troynikov and Huber, Context Rot: how increasing input tokens impacts LLM performance, Chroma, Jul 2025. https://www.trychroma.com/research/context-rot
- Zhang et al., CORE-Bench: a comprehensive benchmark for code retrieval in the era of agentic coding, arXiv 2606.11864, Jun 2026. https://arxiv.org/abs/2606.11864
- Code Isn’t Memory: a structural codebase index inside a coding agent, arXiv 2606.22417, Jun 2026. https://arxiv.org/html/2606.22417
- ContextBench: a benchmark for context retrieval in coding agents, arXiv 2602.05892, Feb 2026. https://arxiv.org/abs/2602.05892
- Groshin, I benchmarked code retrieval for AI coding agents on 60 tasks, 29 Apr 2026. https://dev.to/nike-17/i-benchmarked-code-retrieval-for-ai-coding-agents-on-60-tasks-f9h
- Anthropic, Effective context engineering for AI agents, Sep 2025. https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
If it was useful, pass it along.