# Coding Agents Can Draft the Patch. Who Pays to Prove It?

**Summary:** Coding agents make implementation cheaper, but teams still pay to verify and own every accepted change. I would delegate by the cost of catching and reversing a mistake.

- Canonical: https://markhuang.ai/news/coding-agents-patch-proof-cost
- Language: en
- Author: [Mark Huang](https://markhuang.ai/about)
- Published: 2026-09-28
- Section: News
- Tags: Coding Agents, Software Engineering, Code Review, AI Reliability, Developer Productivity
- Source: [Alex Ewerlöf Notes](https://blog.alexewerlof.com/p/coding-is-not-solved)
- License: https://creativecommons.org/licenses/by-nc/4.0/

---

![Mark inspects a component from a fast automated production line backed by a much larger testing lab](https://cdn.markhuang.ai/news/coding-agents-patch-proof-cost/hero.webp)

*The patch can arrive in seconds. The inspection line behind it still belongs to someone.*

Alex Ewerlöf's September 26 essay ["Coding is NOT solved"](https://blog.alexewerlof.com/p/coding-is-not-solved) pushes back on the idea that better coding agents have finished software engineering. His central distinction is useful: generating code has become much cheaper, while maintenance, reliability, security, and accountability still account for much of the cost.

My read is less absolute than the headline. Coding agents have removed a remarkable amount of implementation friction, and that matters. But a fast patch is not the same product as an accepted change. I would decide what to delegate by asking how expensive a wrong answer is to detect, not how impressive the first answer looks.

## The source gets the cost shift right

Ewerlöf argues that personal software and proofs of concept can tolerate risks that healthcare, finance, aviation, and other consequential systems cannot. I would not draw the border by industry alone. A small internal script can delete valuable data, while a carefully isolated tool inside a regulated company may be easy to verify. The better boundary is the acceptance test.

If a change has a crisp specification, strong tests, a reversible rollout, and a reviewer who understands the surrounding system, cheap generation is a genuine gain. If the requirement is still moving or the failure appears only under load, the agent has accelerated the part that was already easiest to see.

The phrase "coding is solved" causes trouble because it collapses typing, design, integration, operation, and ownership into one activity. The source sometimes answers that oversimplification with an equally broad claim about what models cannot do. I do not need either extreme. I need to know where the evidence stops.

## A successful session is not an owned system

[Anthropic's study of roughly 400,000 Claude Code sessions](https://www.anthropic.com/research/claude-code-expertise) offers strong evidence that agents already do substantial work. In a typical session, the human made about 70% of planning decisions while Claude made about 80% of execution decisions. Greater domain expertise was also associated with higher success.

That is far beyond fancy autocomplete. It also leaves Ewerlöf's concern intact. The study says it cannot observe whether generated code was later used, discarded, or turned into something economically valuable. Its success measures come from session transcripts and signals such as passing tests or committed work. Those are useful checks, but production ownership begins after them.

I read the planning split as a job description. The agent can carry more of the implementation while the person keeps the problem definition, tradeoffs, and acceptance criteria. Expertise did not vanish in these sessions. It changed what the user could safely hand over.

> **Info:**
>
> I will delegate aggressively when failure is cheap to detect and cheap to reverse. When either cost rises, I want tighter scope, stronger evidence, and a named human owner.

## Longer work exposes the unpaid bill

Short issue resolution can hide what happens to a codebase after the first green test. [SlopCodeBench](https://arxiv.org/abs/2603.24755) made agents extend their own solutions across 20 problems and 93 checkpoints. No evaluated agent completed a problem end to end, and the highest checkpoint solve rate was 17.2%. The authors also measured rising structural erosion in 80% of trajectories and rising verbosity in 89.8%.

One benchmark does not settle the future of coding. It does identify the failure I care about: each locally reasonable patch can make the next change harder. Review cost compounds, especially when output arrives faster than maintainers can build a mental model of it.

Even the measuring stick needs review. In February 2026, [OpenAI stopped reporting SWE-bench Verified](https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/) after finding contamination and flawed tests. In the subset it audited, at least 59.4% of frequently failed problems had tests that rejected functionally correct submissions. A benchmark can claim failure when the patch is right; a production suite can miss failure when the patch is wrong. "The tests passed" is evidence, not absolution.

## The strongest objection is that the tools moved

The source is right to reject a victory lap, but it risks freezing the capability picture. METR's early-2025 randomized trial found 16 experienced open-source developers took 19% longer with AI tools. Its [late-2025 follow-up](https://metr.org/blog/2026-02-24-uplift-update/) found signs of speedup, then declined to give a firm estimate because developers who did not want to work without AI were selecting themselves out of the experiment. The tools improved, but that experiment cannot tell us by how much.

The [Hacker News discussion of Ewerlöf's essay](https://news.ycombinator.com/item?id=49877988) exposes the disagreement. Some commenters separate coding from software engineering. Others argue that better harnesses already automate a large share of the work, or that human inconsistency makes the accountability comparison too neat. I find the first objection persuasive and the second incomplete. A human does not need to be deterministic to remain accountable for the system and the release decision.

Agents will keep moving the boundary. Better context handling, planning, and tools can change results without changing the underlying model, as I discussed in [my analysis of coding-agent harness design](https://markhuang.ai/news/stronger-coding-agents-fewer-tools). That makes permanent declarations risky in both directions.

## I would budget for proof before tokens

For each delegated task, I would write the acceptance path first. What evidence can reject a plausible but wrong patch? Who understands the code well enough to notice a missing requirement? Can the change be rolled back without corrupting state? How much future work will depend on the design it introduces?

Then I would measure cost per accepted change, including review, retries, regressions, and later repair. That metric gives coding agents full credit when they save time. It also catches the work that a cheap token bill quietly moves onto the reviewer.

Coding is not solved, but that phrase is too blunt to guide a team. Implementation is becoming cheaper and more autonomous. Teams still have to prove the change, decide whether it belongs, and own what happens after release. I would adopt the workflow only when it reduces that total bill instead of quietly moving work onto the reviewer.
