Why Agent-Written Code Needs Product Verification
Whether your tests catch real defects is not a property of the test level. It is a property of where the expectations come from and who can change them. The expectations have to come from outside the implementation loop — and those expectations need to be protected.
In my last post, I argued that when the same agent writes the implementation and its tests, both can encode the same misunderstanding. A green build does not give you independent verification. That leaves a practical question: where should we invest instead?
I think the starting point is what the product must do for users, with expectations defined outside the implementation loop. Agents change the economics of maintaining that verification, and that changes where we can put our effort.
The pyramid's economics have moved
The testing pyramid was not a law of software quality. When Mike Cohn described it in 2009, the shape followed the cost curve: unit tests were cheap to write, fast to run, and mostly reliable; end-to-end tests were expensive, slow, and brittle. You built a wide base of cheap tests and a thin layer of expensive ones. That cost optimisation made sense for human-written code.
Three factors matter here.
A green base does not establish that the product works. I've talked a bunch of times now about how at Swamp green unit tests doesn't necessarily equal the right user outcome.
The top got cheaper to maintain. Teams avoided heavy investment in E2E and integration tests because of maintenance, brittle selectors, flaky environments, tests that broke every time the UI changed. Agents handle that maintenance naturally. Updating fixtures, adjusting selectors, adapting to interface changes are routine work for an agent. E2E tests still take longer to run and require more infrastructure, but at Swamp, the ongoing cost of keeping them working is not what it was.
Our repair cycle shortened. The pyramid's economic logic assumed catching bugs late costs vastly more than catching them early. That was true when a production bug meant days of developer time and a release cycle. In our workflow, discovery-to-fix often takes minutes. An agent can diagnose and patch in a single loop and the cost gap between early and late detection shrinks. Not to zero: a defect that reaches users can still cause data corruption, broken trust, or downstream damage that no patch undoes. But faster remediation reduces one part of the cost of detecting defects at the release boundary.
You can protect test suites from agent modification, preserve TDD discipline, and enforce coverage gates. But protecting a test does not establish that it asks the right question. The investment needs to start with independent expectations about what the product must do.
Inverting the pyramid is not enough
A lot of the conversation [on X.com] right now is leading to inverting the pyramid — more E2E, fewer unit tests. That is a reasonable adjustment, but it is still the same framework: code verification at different levels of granularity. Unit tests verify functions, integration tests verify modules, E2E tests verify flows. Changing the level of a test does not necessarily change where its expectations come from.
When agents write code, the tests tend to confirm the agent's interpretation of the requirement rather than independently verifying it. The unit tests confirm that interpretation. Even E2E tests, if they are derived from the same understanding, can confirm the wrong thing convincingly.
The common defence of unit tests is that they catch edge cases and logic errors that higher-level tests miss. But when the same agent writes the logic and the edge case tests, it only tests the edge cases it already thought of. The same misunderstanding can leave an edge case absent from both the implementation and its tests. That is the independence problem, and it does not go away by moving up or down the pyramid. The edge cases that actually bite you are the ones in the user's world that were never part of the agent's interpretation.
Independent expectations can drive tests at any level. But some product failures — broken workflows, missing capabilities, misunderstood requirements — only become visible when components work together. Verification of what users actually need has to be an explicit release condition, not something you hope falls out of a green unit suite.
Product verification, not code verification
The failures that matter at our release boundary are failures to deliver what users need, even when the code runs and its unit tests pass. The feature does not do what the user actually needs, a workflow that should be possible is not, an edge case in the user's world was never part of the agent's interpretation.
Code verification asks whether the implementation matches the developer's intent. Product verification asks whether the product delivers the right capabilities to the right people. Those those people might be end users in a browser, a terminal or another team consuming your API. The verification might be UAT, contract tests, or automated checks at the API layer. The form matters less than the principle. Those are different questions, and the second one is the one the agent cannot satisfy by writing more tests against its own code.
But product-level tests are not automatically independent either. If the same agent that wrote the feature also writes the acceptance criteria, you have the same problem one level up. Independence is not a property of the test level. It is a property of where the expectations come from and who can change them. The expectations have to come from outside the implementation loop: a product definition, from user research, from domain expertise that the coding agent did not produce. And those expectations need to be protected: an agent that can edit a failing product test to make it pass has defeated the purpose.
This is not a new idea. James Whittaker at Google described an approach called ACC — Attributes, Components, Capabilities — that tested at the intersection of product qualities, user-facing components, and the capabilities each must deliver. It was product-level thinking about verification rather than code-level thinking. The specific framework matters less than the shift: define what the product must do for users, verify that it does it, and let the implementation be the agent's problem. I first heard of ACC testing when Gojko Adzic blogged about it in 2010.
That shift could be more practical now than it was then. Agents can maintain capability mappings and test mechanics (selectors, fixtures, and setup) as the product changes. Those changes still need checking to ensure they preserve the tests' expectations. Changes to assertions, expected outcomes, or covered capabilities need review outside the implementation loop. And the constraints come from the product definition, not the code and the agent did not define the product attributes and cannot rewrite what "fast" or "reliable" means to get to green.
Speed is the payoff
Faster remediation and cheaper maintenance change the tradeoffs. They do not erase them: execution cost and diagnostic precision still matter. But they make broader product verification practical in a way it was not before. When a human found a bug at the E2E level, tracing it through the codebase and getting a fix through a release cycle was enormous. That cost was part of the case for the pyramid. At agent speed, the agent can rewrite entire modules to satisfy a product-level constraint. In our workflow, these fixes often take minutes.
The fear is that you will ship broken code faster. We have had our UAT runs catch real defects at our release boundary, despite thousands of green unit tests. The unit suite did not stop them. The boundary checks did. The question is not whether defects make it past code-level tests. It is whether your verification catches them before they reach users, and how fast you can act when it does. There's argument that the unit tests are correct - that they satisfy the codebase and that it's logic is correct. But to be honest, it doesn't matter if the product isn't right...
The investment that pays off is making your product verification good, independent expectations, rooted in what users need, protected from the implementation loop. When it catches something, the agent fixes it. When it misses something, you add a capability test so it catches it next time. The verification suite grows from real product failures, not from an agent's interpretation of what might go wrong.
Even where the economics look different to ours, the independence problem remains. Cost determines how much verification you can run; it does not determine whether the expectations are trustworthy.
Where the pyramid still holds
I am not saying the testing pyramid was wrong. For human-written code at human cycle times, it was a rational allocation of effort.
There are places it still applies. Any code with externally defined correctness criteria: a cryptographic function, a pricing engine, a regulatory validation rule, a permissions model, all still benefit from unit-level verification because the expected behaviour is precise and independent of any agent's interpretation. The useful distinction is whether a unit test checks an externally defined rule or merely reproduces the implementation's assumptions.
Unit tests still have a place. What no longer holds is treating their volume or coverage as a substitute for independent verification of the product especially when the generation of code has been deferred to an agent.