SP//LOG

What happens after you press Enter in an AI IDE

Between pressing Enter and files changing on disk sit fourteen distinct systems. Most of them are not the model.

A patent-style drawing of a mechanical pipeline. A finger presses a key and the motion travels through punched tape, a bank of typebars, a branching tree, a patch panel and a stack of cards before arriving at a printed page.
Published
Reading
34 min
Difficulty
Advanced
Status
Stable

You type a sentence into a chat box. Something like "fix the race condition in the upload handler." You press Enter. Four seconds later, files change on disk.

Most explanations of that gap are useless. They say "the AI reads your code and writes a fix," which is the technical equivalent of explaining a car as "you press the pedal and it goes." What actually happens is a chain of maybe a dozen distinct systems, most of which are not AI at all, most of which were invented before the model was, and several of which exist purely to compensate for the fact that language models are bad at things that sound easy.

This is that chain, from the lowest layer to the highest. Every stage here is a real system with a real name, real failure modes, and real design arguments still being had in public.

Stage 0: Before the AI exists, there is a text buffer

Nothing in the pipeline works if the editor cannot answer one question fast: what is in this file right now?

That sounds trivial. It is not. Your file is not a string in memory. A naive editor storing a 4 MB file as one contiguous string would have to copy the entire tail of that string every time you type a character in the middle. Editors solve this with data structures built for insertion in the middle: a piece table (the original text plus an append-only buffer of your additions, with an ordered list of "pieces" describing which spans to read in what order) or a rope (a balanced tree of small string chunks, so edits touch one leaf and rebalance a few nodes).

VS Code, which is the substrate under Cursor, Windsurf, and most of the fork ecosystem, uses a piece table variant. This is the layer that determines whether your editor stays at 60fps while an agent streams 800 lines into a file.

Why does this matter to the AI story? Two reasons.

First, everything downstream needs a stable coordinate system. Line 42, column 8. Byte offset 1,204. A parse tree node spanning bytes 900 to 1,150. When an agent applies an edit, every one of those coordinates shifts. The buffer layer is what makes it possible to translate a position through an edit instead of recomputing the world.

Second, and this is the theme of the entire article: the model never sees your file. It sees a snapshot, serialized to text, at some moment in the past, possibly already stale by the time the response comes back. Almost every weird AI IDE bug you have ever hit is a staleness bug in disguise.

Stage 1: The tokenizer, or why your code costs more than your prose

The model does not read characters. It reads integers. The tokenizer is the deterministic function that turns your text into those integers, and it is the single most underrated component in the stack.

How BPE actually works

Byte Pair Encoding starts with a vocabulary of raw bytes and repeatedly merges the most frequent adjacent pair into a new symbol. Do that tens of thousands of times over a training corpus and you get a merge table: an ordered list of pair merges. Encoding a new string means applying those merges in order until nothing more can merge.

The properties that fall out of this are worth stating precisely, because people get them wrong constantly:

  • It is lossless and reversible. Tokens decode back to the exact original bytes.
  • It works on arbitrary input, including text the tokenizer never saw, because the floor of the vocabulary is raw bytes. There is no out-of-vocabulary failure.
  • It compresses, at roughly 4 bytes per token on average English prose.

That "on average" is doing enormous work. Average English prose. Not your code.

Why code tokenizes badly

The merge table was learned from a corpus. Whatever was frequent in that corpus got a short encoding. Whatever was rare did not. So:

  • Common English words are one token. understanding is likely one token.
  • Common code idioms are cheap. def , function, return, import usually land on a single token each.
  • Your identifiers are not. calculateUserSubscriptionRenewalDate is not in the merge table, so it fragments into five or six pieces. Camel case, snake case, and long descriptive names, the exact things good style guides tell you to use, are tokenizer-hostile.
  • Whitespace is expensive. Deeply indented code pays for its indentation. Some tokenizers add explicit multi-space tokens to blunt this, which is a direct admission that Python cost too much.
  • Numbers are chaotic. Digit sequences chunk unpredictably. 1024 might be one token, 1025 might be two, split in a place that makes no arithmetic sense.

That last one is not trivia. Hold onto it. It shows up again in Stage 9 and explains an entire architectural decision made by every serious AI IDE.

The budget nobody shows you

Everything downstream is a budget allocation problem denominated in tokens. A 200K context window is not 200K "units of code understanding." It is 200K integers, and your repository converts to them at a worse exchange rate than this sentence does. A 500-line TypeScript file with real identifiers can easily be 6K to 8K tokens. Twenty such files and you have spent most of a large window before the model has done any thinking.

Tokenization is also where the first silent quality loss happens. When a context builder needs to trim, it trims by token count, and token count does not respect semantic boundaries. The truncation lands wherever the budget ran out.

Stage 2: Structure, part one, tree-sitter

Now the editor needs to know what the text means structurally. Not semantically, not yet. Structurally. Where does this function begin and end. Is this token a type name or a variable.

The classic answer, a compiler front end, is unusable here for one reason: it assumes the input is valid. Your file, while you are typing, is essentially never valid. It is broken about 95% of the time you are actively working in it.

Tree-sitter was built for exactly this environment, and it makes two design choices that matter.

Incremental parsing

It does not reparse the file on each keystroke. It reparses the affected region and reuses the rest of the previous syntax tree. Change one character inside a function body and the parse work is proportional to that edit, not to the file. This is what makes syntax-aware behavior viable at typing latency, and at repository scale it is why a push only requires reparsing and reindexing the changed portions of affected files rather than the whole repo.

Error recovery that produces a usable tree anyway

Tree-sitter parses with a GLR (Generalized LR) algorithm, which can pursue multiple parse possibilities in parallel rather than committing to one and dying. When it hits invalid input, it tries several recovery strategies at once, for example ignoring the offending token, or skipping ahead to a position that matches the current token, and evaluates which one produces the best result a few tokens later. What it could not fit gets wrapped in explicit error nodes and left in the tree.

That is the crucial part. You still get a tree. A half-typed file yields a mostly-correct structure with the broken region localized and labeled, instead of a parse failure. Every feature that depends on structure keeps working while you type.

For the AI pipeline specifically, tree-sitter is what makes syntax-aware chunking possible. When Cursor indexes a file, it runs tree-sitter to build an AST, walks the tree depth-first, and merges sibling nodes until each chunk fits under the embedding model's size limit. The result is chunks whose boundaries are functions and classes rather than arbitrary character counts. That single choice, splitting on structure instead of on length, is worth more retrieval quality than most of the fancier things layered on top of it.

Stage 3: Structure, part two, LSP

Tree-sitter tells you shape. It does not tell you truth. It cannot tell you that user.name is a type error because user is User | null. That requires a real type system, name resolution, and a project-wide view.

That is the Language Server Protocol, and its whole reason for existing is combinatorial. Before LSP, N editors times M languages meant N times M integrations. LSP collapses that to N plus M by defining one JSON-RPC protocol between an editor (client) and a language server (a separate process that actually understands the language).

The lifecycle:

  1. Client spawns the server process (or opens a socket) and sends initialize, declaring its own capabilities.
  2. Server responds advertising its capabilities. This is a negotiation, not a fixed contract. That is why some features work in one language and not another.
  3. Client sends initialized.
  4. Client sends textDocument/didOpen with the file contents, then a stream of textDocument/didChange deltas as you type.
  5. Server maintains its own model of the document, and pushes textDocument/publishDiagnostics whenever its opinion of your errors changes.
  6. Client sends requests on demand: textDocument/hover, textDocument/definition, textDocument/completion, textDocument/references.

Note the direction of arrow 5. Diagnostics are pushed, not polled. The server decides when it has an opinion. This is asynchronous, and it means "does this code have errors" is a question with no synchronous answer. You have to send a change, then wait an unspecified amount of time, then see what arrives. Remember this too. It becomes the central engineering problem of Stage 11.

Both tree-sitter and LSP run in a modern AI IDE, and they are not redundant. Tree-sitter is fast, local, always available, tolerant of broken input, and shallow. LSP is slow to start, project-wide, needs valid-ish input to be useful, and deep. Highlighting and chunking use the first. Correctness checking uses the second.

Stage 4: The context builder

Here is where the AI part actually starts, and it is not a model. It is a string concatenation problem with a budget constraint, and it is where most of the quality difference between competing products lives.

The model receives one flat sequence. Every "smart" thing an IDE does reduces to deciding what goes into that sequence and in what order. A typical assembly:

  1. System prompt. Identity, tool definitions, formatting rules, safety rules, and the edit format specification. Often thousands of tokens before you have said anything.
  2. Environment block. OS, shell, working directory, git branch, today's date. Models have no clock and no filesystem. If you do not tell them, they guess, and they guess wrong.
  3. Project instructions. CLAUDE.md, .cursorrules, AGENTS.md. Persistent user-authored rules.
  4. Retrieved context. The output of Stage 5.
  5. Open editor state. Current file, cursor position, current selection, sometimes recently viewed files. This is a cheap and shockingly strong signal. The file you are looking at is usually the file you mean.
  6. Conversation history. Previous turns, previous tool calls, and previous tool results. In a long agent session this dominates everything else.
  7. Your message.

Ordering is not cosmetic

Two forces push on the order.

Attention. Models attend unevenly across a long context. Material at the very beginning and very end is used more reliably than material buried in the middle. So the actual instruction goes last, near the generation point.

Caching. This is the bigger one, and it is a systems constraint leaking upward into prompt design. Prompt caching works on prefixes. If the first 20K tokens of this request byte-match the first 20K of the last one, the server can reuse the computed attention state instead of recomputing it. That is a large cost and latency saving, and it is worth real money at agent-loop volumes.

So the layout above is not aesthetic. It is: rarely changes, changes per session, changes per turn. That ordering is a cache-hit-rate optimization wearing a prompt-engineering costume.

Stage 5: Retrieval, and the argument the industry just had

Your repo is 400K lines. The window fits maybe 3K. Which lines go in?

Two answers, and the fight between them was the most interesting thing to happen in this space in the last two years.

The RAG approach, and mechanically it is well understood.

Chunk with tree-sitter, on syntactic boundaries.

Embed each chunk with a code-specialized model. Generic text embedders underperform here because code similarity is not prose similarity. voyage-code-3 reported beating OpenAI-v3-large by roughly 13.8% and CodeSage-large by roughly 16.8% averaged across 32 code retrieval datasets, and it supports 2048, 1024, 512, and 256 dimensions plus int8 and binary quantization, which matters enormously when you are storing tens of millions of vectors.

Index into a vector store. Exact nearest-neighbor search is linear in the number of vectors, so production systems use approximate search, and HNSW (a navigable small-world graph you greedily descend through layers) is the default. You trade a few percent of recall for orders of magnitude of speed.

Retrieve and rerank. Bi-encoder retrieval embeds query and document independently, which is what makes precomputation possible and also what makes it imprecise. So the standard shape is two-stage: cheap bi-encoder pulls maybe 100 candidates, then a cross-encoder reranker scores query and candidate jointly, which is far more accurate and far too slow to run over the whole corpus. Reported gains for the rerank stage sit in the 10 to 30 percent range on top-10 quality across most benchmarks, which makes it the cheapest quality lift in the entire retrieval stack.

Keep it fresh. This is the operationally hard part. Re-embedding a monorepo on every save is not viable, so Cursor builds a Merkle tree over the repo: hash every file, hash the hashes up the directory structure, and to sync, compare roots. If the roots match, nothing changed and you are done in one comparison. If they differ, walk only the branches whose hashes differ. You get an efficient diff of "what actually changed" without transferring or rehashing the whole tree, and reindexing runs incrementally on a timer.

There is a privacy design consequence worth noting: the server can hold embeddings and obfuscated metadata without holding plaintext source, which is the only reason cloud indexing of proprietary code is sellable at all.

Answer two: just grep

Then Anthropic, in May 2025, removed vector search from Claude Code entirely. No embedding pipeline, no local vector database, no chunking heuristics. Replaced with tool calls: grep, glob, read file. The model searches your codebase the way you would, iteratively, by running searches and reading results and running better searches.

Boris Cherny's account is that early Claude Code did use RAG, and they moved to agentic search because it outperformed everything else by a lot.

The reasons given, in rough order of weight:

  • Staleness. An index is a claim about the past. Grep runs against the present. In a repo where an agent is editing files in a loop, the index is wrong by construction.
  • Precision. Programming is exact. If you need handleUploadError, you want the 3 places that literally contain it, not the 20 chunks that are semantically about upload error handling.
  • No infrastructure. No embedding service, no vector DB, no sync daemon, no cache invalidation, no cross-machine consistency. Deleting a subsystem is a real engineering win.
  • Privacy. Nothing leaves the machine except what the model asked to see.
  • Composability. The model can chain. Grep for the symbol, read the file, grep for its callers, read those. That is a search strategy, not a similarity lookup, and strategies can recover from bad first results in a way that a single top-k retrieval cannot.

Cursor subsequently hired engineers behind that decision, and Windsurf, Cline, Devin, and Sourcegraph Amp moved in the same direction.

Who is right

Neither, cleanly, and the honest framing is a cost curve rather than a winner.

Grep-based agentic search burns tokens. Every search is a tool call, every result goes into context, most results are noise, and the model pays to read them all. Critics point out that this is genuinely expensive at scale, which is true and is the main argument for keeping an index around.

Embeddings burn correctness. They are fuzzy, they go stale, and they give you a plausible neighborhood rather than the actual answer.

What is converging in practice is hybrid: exact tools when you know the string, semantic retrieval to answer "where is the thing that does X" when you do not know what it is called, and a reranker to clean up whatever the first pass produced. Cursor's own semantic search work reports meaningful accuracy gains from adding semantic retrieval back alongside the exact tools, which is the same conclusion from the other direction.

The deeper lesson is that this was never a retrieval-algorithm question. It was a question about whether the model should be handed context or should go get it. Everyone underestimated how much better "go get it" works once the model is good enough to plan a search.

Stage 6: MCP, the part that is boringly well designed

Your tools are now interesting. Read a file, run a command, query Postgres, hit Sentry, open a browser. Every IDE building every integration itself is the N times M problem again.

Model Context Protocol is the LSP-shaped answer to it, and the resemblance is not accidental. Same problem, same shape of solution.

It runs on JSON-RPC 2.0. Every message has jsonrpc: "2.0", a method, params, and an id for requests. Stateful session, capability negotiation at connect time, so subsequent calls stay small and fast, and the server can push notifications rather than being polled.

Three primitives:

  • Tools. Executable actions. Client calls tools/list, gets back an array of definitions with name, description, and JSON Schema input. Calls tools/call to invoke.
  • Resources. Read-only data, addressed by URI. Files, records, documents.
  • Prompts. Reusable templates the server offers the client, typically surfaced as slash commands.

The spec has moved fast. The 2025-06-18 revision added outputSchema, so a server can declare the shape of what it returns, plus structuredContent, so a result carries an actual JSON object alongside the human-readable text instead of forcing the model to parse prose. It also added annotations, behavioral hints such as whether a tool is read-only or destructive, which is how a client decides what to auto-approve. Later revisions continued along the 2025-11-25 line.

Three things worth internalizing:

Descriptions are prompt text. A tool's description goes into the model's context verbatim. It is not documentation for humans. It is the instruction that determines whether the tool gets called correctly, or at all. Badly written descriptions are the single most common cause of "the model ignores my tool."

The schema is the guardrail. Because arguments are validated against JSON Schema before execution, a malformed call fails at the protocol boundary and can be handed back to the model to retry, rather than reaching your code.

Every server is trust. An MCP server's tool descriptions and its returned content both enter the model's context. A malicious or compromised server can put instructions there. This is the injection surface, and it is why the security posture in Stage 8 exists.

Stage 7: Tool calling, and the thing everyone gets wrong

Here is the sentence that clears up most confusion about agents:

That is the entire agent loop. It is a while loop around an HTTP request:

the entire agent loop
loop:
  response = model(context)
  if response has no tool calls:
      break
  for each call:
      result = execute(call)          # your code, not the model's
      context.append(call, result)

Everything that feels agentic is emergent from that loop. There is no planner, no executive, no persistent process on the model side. Each turn is a fresh, stateless forward pass over a transcript that keeps growing.

Modern providers train models specifically for this, which is why reliability improved so much: constrained decoding can guarantee syntactically valid tool-call output, parallel tool calls let a model request four file reads in one turn instead of four round trips, and the model is trained to look at a result and decide whether to continue or stop.

Two failure modes are structural rather than incidental:

The loop does not terminate on its own. Nothing guarantees the model stops. Harnesses impose turn limits, token budgets, and cost ceilings. Every one of those is an external safety rail.

Context grows monotonically. Every call and every result stays in the transcript. See Stage 12.

Stage 8: The model, where your latency and your money go

The request goes out. What happens on the other side splits cleanly into two phases with completely different performance characteristics, and understanding that split explains nearly every timing behavior you have observed.

Prefill

The entire prompt is processed in one forward pass. Every token attends to itself and everything before it. This is compute bound and highly parallel, because all input tokens are known up front. It produces the KV cache: the key and value tensors for every token at every layer, which is what lets subsequent generation avoid recomputing history.

Prefill is why time-to-first-token scales with prompt size. A 100K-token context has a real, visible cost before a single character comes back.

Decode

Then tokens are generated one at a time. Each new token attends to the whole cache, then appends its own entry to it. This is memory bandwidth bound, not compute bound. The GPU is largely waiting on memory, moving the model weights and the cache around for one token's worth of arithmetic.

That asymmetry drives essentially all inference systems engineering:

  • Continuous batching. Static batching wastes the GPU badly, because requests in a batch finish at wildly different lengths and the whole batch waits for the slowest. Continuous batching works per decoding step: finished requests are evicted and queued requests injected after any step. Utilization goes way up.
  • Paged attention. Cache memory is managed in fixed-size blocks like OS virtual memory pages instead of one contiguous allocation per request, which kills fragmentation and lets identical prefixes share physical blocks.
  • Prompt caching. Falls straight out of the above. If a prefix is already computed, skip prefilling it. This is why Stage 4 obsesses over prefix stability.

Sampling

Logits come out, a distribution over the whole vocabulary at every step. Temperature flattens or sharpens it, top-p truncates the tail. For code editing, near-zero temperature is normal, because you want the most likely token, not an interesting one.

Speculative decoding, and Cursor's clever version

Decode is bandwidth bound and processes one token at a time, so it wastes most of the hardware's arithmetic capacity. Speculative decoding fixes this by guessing several tokens cheaply, then verifying all of them in a single expensive forward pass. Verification is parallel, so if the guesses were right you got several tokens for the price of one, and if they were wrong you fall back with no correctness loss. The output is identical to normal greedy decoding. This is a pure latency optimization, not a quality tradeoff.

Cursor's fast apply takes the sharpest possible version of this insight. Applying an edit to a file is a task where most of the output is already known: the unchanged parts of the file. So instead of a draft model, they speculate with a deterministic algorithm that just proposes the existing file contents as the next tokens, and validates the speculation against greedy generation, keeping the longest matching prefix. Where the model agrees with the original file, which is most of it, you fly. Where it diverges, that is the actual edit, and you pay normal cost only there.

Which brings us to the most important architectural fact in this whole article.

Stage 9: Diff generation, or: line numbers are a lie

A big model produced an intent: change this function to do that. Now something has to turn that into bytes on disk. This is a distinct problem, and it is where AI IDEs quietly diverge the most.

Why the obvious approach fails

The obvious approach is to have the model emit a unified diff with line numbers and hunk headers. It fails, and the reasons are worth spelling out because they are all consequences of things we established earlier.

  • Numbers tokenize unpredictably. Stage 1. The model is doing arithmetic on values that are not cleanly represented in its input space.
  • Position tracking degrades over long files. The model loses count.
  • Errors cascade. A unified diff is self-consistent: line numbers, hunk offsets, context lines. One wrong count and the whole patch is invalid.
  • Exact-match requirements are brittle. Cursor's write-up on this reports that deterministic matching against diffs fails at least 40% of the time.

Read those together and the conclusion is uncomfortable: the way you ask a model to express an edit changes its measured capability more than most model upgrades do.

The three families

Whole-file rewrite. Model emits the entire file with changes applied. Maximally robust: no matching, no offsets, nothing to desync. Costs output tokens proportional to file size, which was prohibitive until fast apply made those tokens nearly free. Aider's own investigation of Cursor's approach noted that full rewrites outperform diff-style edits for files under roughly 400 lines. This is why fast apply exists: it makes the dumbest, most reliable format economically viable.

SEARCH/REPLACE blocks. Model emits the exact old text and the exact new text. No line numbers, so no counting, but it requires the search text to match byte for byte, and models are imperfect at reproducing text they were shown. Cheap, and the standard for large files.

Two-model split. Big expensive model decides what to change and expresses it loosely. Small fast model, trained only on apply, turns that into the exact file. This is now the mainstream architecture, and it is a straightforward specialization argument: you are paying frontier-model prices for mechanical text transformation otherwise.

The other diff, the one you actually see

Once the new file content exists, the IDE computes a diff between old and new to render red and green in your editor. This is a completely separate operation from how the edit was expressed, and it uses classical algorithms.

Git ships four: myers (default, Eugene Myers, 1986), minimal, patience, and histogram. The split is philosophical. Myers and minimal optimize for the smallest edit script. Patience and histogram optimize for the most readable one, by anchoring on rare lines rather than common ones so that a brace or a blank line does not become a spurious alignment point. Histogram is Bram Cohen's refinement of patience with better handling of low-occurrence common elements, and it is both fast and generally the best-reading of the four. Empirical work comparing them finds histogram superior at describing code changes.

The distinction matters because the diff you see is a presentation artifact, computed after the fact. It is not what the model produced and not what was applied. A beautiful diff view can sit on top of a whole-file rewrite that discarded your unrelated edit three functions down.

Stage 10: Execution and the sandbox

Some tool calls are reads. Others run commands. That is where the security model has to be real.

Modern coding agents implement this in two independent layers, and the distinction is worth stating precisely.

Permissions are policy. Rules evaluated in the harness before a tool runs: allow, ask, deny. This is what produces the confirmation prompt. It is enforced by the agent's own code.

Sandboxing is enforcement. The kernel refuses. Nothing in the model's output, and no bug in the permission checker, changes what a syscall is allowed to do.

Concretely, Claude Code's sandbox uses bubblewrap plus seccomp on Linux (with a legacy in-process Landlock fallback when bwrap is unavailable), Seatbelt via sandbox-exec with embedded SBPL policies on macOS, and restricted token plus Job Object plus WFP network filtering on Windows. The Linux launch is two-stage: bwrap establishes the filesystem and namespace view, then PR_SET_NO_NEW_PRIVS plus a seccomp filter locks down syscalls. Seccomp unconditionally denies ptrace, process_vm_readv/writev, and the io_uring_* family, and in restricted-network mode blocks every socket family except AF_UNIX.

Those specific denials tell you the threat model. ptrace and process_vm_readv are how a sandboxed process reads or hijacks a sibling process's memory. io_uring is a well-known seccomp bypass path, because it performs I/O through a submission queue rather than through the syscalls a filter is watching.

The right mental model: permissions ask, sandboxes make the question unnecessary.

Stage 11: Patch validation, and Cursor's cleverest hack

Code is written. Is it correct?

Cheap checks first: did it parse (tree-sitter, instant), did it match the format, did the file actually change.

The real check is diagnostics, and here we collide with the asynchronous nature of LSP from Stage 3. To know whether an edit type-checks, you must apply it, wait for the language server to publish, and read the result. But the language server is watching your workspace. Applying speculative AI edits to it means your editor flickers with errors from code you never wrote, your own diagnostics get clobbered, and any save-triggered tooling fires on garbage.

Cursor's answer is the shadow workspace: spawn a hidden Electron window with show: false, loading the same workspace. The AI applies its changes there. The normal language servers, tsserver, rust-analyzer, Pylance, run against that hidden copy and publish diagnostics as usual. The agent reads them, fixes what it broke, and iterates until the diagnostics are quiet. Only then does anything touch the file you can see.

It is opt-in, and the honest tradeoff is memory: you are running a second full editor instance, which is why the guidance is to enable it only with RAM to spare.

The stated ambition goes further than lints. The goal is for the AI to have the whole LSP surface, go-to-definition included, against a workspace that is not yours.

Beyond diagnostics, validation escalates: run the linter, run the type checker, run the tests. Each rung is slower and more definitive, and each result goes back into context as another tool result, which means the agent's self-correction loop and the token growth problem are the same loop.

Stage 12: Git, and the fact that undo is a distributed systems problem

An agent editing 15 files across 4 minutes needs an answer to "put it back."

Git provides the mechanics: the index, content-addressed objects, cheap branching. But the naive integration is wrong, because the agent and the user are editing the same working tree concurrently, and git's unit of isolation is the working tree.

Approaches in the wild:

  • Checkpointing. Snapshot state before each agent turn, often into a shadow git repo or a separate ref so the user's history is untouched. Restore is a checkout. This is what most in-editor "restore checkpoint" buttons are.
  • Worktrees. git worktree gives a second checked-out directory sharing one object database. The agent works there, you work here, nobody stomps anybody. This is the standard pattern for parallel agents.
  • Full isolation. Container or cloud VM per agent, merged back as a branch or PR.

Stage 13: Context rot, and the loop eating itself

Every turn appends. Tool call, tool result, model reasoning, repeat. A single coding-agent session on SWE-bench-style tasks routinely spans millions of tokens and well over a hundred turns.

Two things go wrong, and the second is the surprising one.

Cost and latency. You re-send the entire transcript every turn. Prompt caching blunts this, which is exactly why Stage 4's ordering discipline pays off here and why "don't break the cache" is an active research topic for agentic workloads.

Context rot. Performance degrades with length in ways not explained by hitting the limit. Accumulated context is not merely expensive, it is actively harmful. The failed attempt from turn 12 is still sitting there, still being attended to, still exerting influence on turn 60. The agent is reasoning over a transcript that is mostly its own past confusion.

Mitigations, roughly in order of how established they are:

  • Compaction. At a threshold, summarize the history into a shorter state and continue from the summary plus recent turns. Now the dominant approach, and now the dominant source of "the agent forgot the thing I said at the start."
  • Structured eviction. Drop stale material by structure rather than by summarization, on the theory that summarizing lossy history lossily is not obviously the best move.
  • Addressable recall. Keep the full history externally, put pointers in context, fetch on demand. Lossless in principle, and reported to beat both summarization and RAG baselines on long-context tool-heavy tasks.
  • Context folding. Push subtasks into isolated sub-contexts and fold only their results back into the parent. This is what subagents are, architecturally.

Stage 14: Back to your screen

Tokens stream back over SSE. The client parses the stream incrementally, distinguishing prose from tool calls from code blocks while it is still arriving. Code blocks get syntax highlighted with tree-sitter, which closes the loop back to Stage 2 nicely: the same parser that had to tolerate half-written files is now highlighting half-arrived ones.

Diffs render. Files write, coalesced so the buffer layer from Stage 0 does not thrash and so tree-sitter is not reparsing on every streamed chunk. The language server sees the change, runs, publishes new diagnostics. Your file is different than it was four seconds ago.

The whole thing, in one view

LayerSystemJobFails as
0Piece table / ropeHold text, edit in the middle cheaplyStale snapshots
1BPE tokenizerText to integersSilent budget burn, broken number handling
2Tree-sitterStructure, incrementally, on broken inputNothing much, it is the sturdiest piece here
3LSPTypes, references, diagnosticsAsync, slow start, per-language gaps
4Context builderChoose and order what the model seesCache misses, mid-context blindness
5RetrievalFind the relevant 3K of 400K linesStale index, or token-expensive searching
6MCPStandard tool interfaceInjection surface
7Agent loopParse, execute, append, repeatNontermination, unbounded growth
8InferencePrefill, decode, sampleTTFT scales with prompt
9Edit format + diffIntent to bytes, bytes to red and greenThe single biggest quality lever
10SandboxContain executionThe only real defense against injection
11Shadow workspaceValidate before you see itCosts a whole extra editor
12GitUndoCannot undo side effects
13CompactionSurvive long sessionsAmnesia, context rot

What I would actually take away from this

Most of the system is not the model. Buffers, parsers, protocols, indexes, sandboxes, diff algorithms. The model is one stage of fourteen, and it is the stage you have the least control over. The products that win are winning at context assembly and edit application, not at model selection, because everyone has the same models.

The format of the answer changes the quality of the answer. A 20% to 61% swing from an edit-format change is not a footnote. It means benchmark numbers are properties of harnesses, not of models, and it means the most valuable engineering in this space is often unglamorous plumbing.

Reversals are normal. RAG was obviously correct until it was not. Diffs were obviously correct until whole-file rewrite got cheap. Both reversals came from the same direction: a cost curve moved, and an approach that was too expensive to consider became the best one. Assume the current consensus has the same shelf life.

Containment beats compliance. Nobody has solved prompt injection at the language level, and nobody is close, because the model has exactly one input channel and everything shares it. The working answer is a kernel that says no. That is why the seccomp filter list is more load-bearing than the system prompt's safety section.

Four seconds. Fourteen systems. The interesting thing is not that it works, it is that almost none of it is new. Myers is from 1986. LSP is from 2016. Piece tables predate most of the people using them. The genuinely novel layer is thin, and it is sitting on decades of unfashionable infrastructure that turned out to be exactly what it needed.

Sources

esc