For most of my career, code review has been a fairly simple idea.
One engineer writes some code, opens a pull request, and another engineer reviews it before it gets merged. The second engineer looks for bugs, questions design decisions, suggests improvements, and ultimately gives the team another level of confidence that the change is safe to ship.
It is such a normal part of software development that we rarely stop to question the assumptions behind it.
One of those assumptions is that another engineer wrote the code.
That assumption is becoming less reliable.
Today, a coding agent can implement an entire feature, run the tests, fix failures, refactor the implementation, and prepare a pull request. The engineer who opens that pull request may have spent most of their time reviewing and directing the agent rather than writing the code themselves.
This makes me wonder whether our traditional code review process still makes sense.
Not because I think code review is obsolete. I think the opposite is true: we need verification more than ever.
But I think we need to change how we think about code review.
What was code review actually for?
Before we decide what should happen to code review, it is worth asking what problem it was solving in the first place.
When one engineer writes a change and another engineer reviews it, we get a second set of eyes. The reviewer might catch a bug the author missed, notice that an abstraction is wrong, identify an edge case, question an architectural decision, or simply ask why something was implemented in a particular way.
There is also a social dimension to code review. It creates shared ownership and knowledge across a team. The code does not belong exclusively to the person who wrote it.
All of that remains valuable.
But not every part of the traditional process is equally valuable when the author is an AI agent.
I have started thinking about coding agents as something like flawed peers.
Imagine that you are a seasoned engineer and you have a very fast colleague who is capable of producing an enormous amount of code. They can work tirelessly, know an impressive amount about software development, and solve many problems very well. At the same time, they are inexperienced in your particular system and can occasionally make mistakes that seem obvious in hindsight.
You would not blindly merge their work.
You would review it.
You would give them clear constraints.
You would run automated checks.
You might ask another specialist to look at particularly important changes.
That sounds quite a lot like the agentic development workflow we are building now.
The engineer opening the PR is already a reviewer
This is where I think our traditional PR rules start to become interesting.
Suppose a team has historically required one human review for every pull request.
The traditional workflow looks something like this:
Engineer A writes the code → Engineer B reviews the code → merge.
Now consider an agentic workflow:
Agent writes the code → Engineer A reviews the code → Engineer A opens the PR → merge.
If Engineer A has genuinely reviewed the implementation, understands the change, and is willing to take ownership of it, what exactly are we gaining from requiring another human to perform a second generic review?
The PR being opened does not magically make the code more trustworthy.
The meaningful event happened before that: an engineer examined the work and decided that they were prepared to own it.
This is why I think we should be careful about treating approval counts as a proxy for quality.
If your previous process required one review, I think there is a reasonable argument that an agent-generated change which has been thoroughly reviewed by the engineer opening the PR may not need another generic human approval.
If your process required two human reviews, then perhaps the agent-generated change gets one human review by the owner and one additional human review after the PR is opened.
The exact policy will depend on the risk of the system and the kind of change being made.
The important part is the principle:
The number of approval buttons pressed is not the same thing as the amount of verification performed.
Trust the engineer, get ownership in return
There is also a cultural issue here.
If an experienced engineer reviews an agent-generated change, opens the PR, and says "I am happy to own this", I think we should take that statement seriously.
We often say that we want engineers to have ownership.
But ownership without trust is difficult to achieve.
If we tell engineers that they are responsible for the code but then require another person to validate every decision they make, we are creating a system where responsibility and authority do not quite match.
I would rather move toward a model where we say:
You reviewed it. You understand it. You own it.
In return, we should expect engineers to take that responsibility seriously.
This does not mean trusting every change equally. A production database migration, authentication change, or financial transaction workflow deserves a different verification strategy from a small UI change.
It means that our process should respond to risk and uncertainty, rather than blindly applying the same number of reviewers to every pull request.
The bigger problem is code volume
There is another problem that I think is even more important.
Coding agents make producing code dramatically cheaper.
That is fantastic.
It also creates a problem.
If generating code becomes cheap enough, the amount of code produced can increase much faster than the amount of human attention available to review it.
We cannot solve that problem simply by adding more reviewers.
If an agent produces ten times more code and our answer is to have humans manually inspect ten times more code, we have simply moved the bottleneck from writing software to verifying software.
Human attention is scarce.
So we need to become much more selective about where we spend it.
Don't review what you can make unrepresentable
This is where architecture becomes a much more important part of code review.
There are many classes of bugs where I would rather not depend on a reviewer noticing the problem.
I would rather design the system so that the incorrect implementation is difficult or impossible to represent.
Consider a database schema.
If the database requires a value to satisfy a particular constraint, we don't need to rely entirely on every application developer remembering to validate that constraint correctly. The database can enforce it.
Or consider a business workflow.
If a process can only move from Pending to Approved or Rejected, we can represent those states explicitly rather than allowing arbitrary strings and hoping that every piece of application code handles them correctly.
Finite state machines are a good example of this approach. The architecture itself describes which transitions are valid and which are not.
The same idea applies to types, contracts, API boundaries, validation, generated code, static analysis, automated tests, and many other forms of explicitness.
The goal is not to make developers more careful.
The goal is to make certain mistakes impossible.
That changes the economics of code review.
Instead of asking a human to inspect every line and think about every possible failure mode, we can move some classes of verification into the system itself.
The reviewer can then spend their limited cognitive capacity on the things that actually require human judgment.
Verification becomes a layered system
This suggests a different model for reviewing agent-generated code.
At the bottom, we have constraints and architecture that prevent entire classes of incorrect implementations.
Above that, we have automated verification: type checking, tests, static analysis, security scanners, contract tests, database constraints, CI checks, and whatever else is appropriate for the system.
Then we have the engineer who reviews the change and takes ownership.
And for changes where we want additional confidence, we can add another layer.
Adversarial agents.
What if the reviewers were agents too?
We don't necessarily have to choose between one human review and three human reviews.
We can introduce specialized agents whose job is not to approve the code, but to try to find reasons why we should not trust it.
Imagine an agent whose only responsibility is security.
It reviews the change and asks questions such as: "Can this input be manipulated? Did we introduce an authorization bypass? Are there new injection opportunities? Did this change expose something that should remain private?"
Another agent might specialize in performance.
It could look for expensive queries, unnecessary allocations, N+1 database access, contention, excessive network calls, or other performance problems.
We could have agents specializing in reliability, backwards compatibility, API design, testing, accessibility, architecture, or whatever concerns are particularly important to a given system.
The important part is that these agents have specialized responsibilities.
I don't want ten generic agents saying "LGTM."
I want adversarial agents trying to break my confidence in the implementation.
The security agent should be trying to find a security problem. The performance agent should be trying to find a performance problem. The architecture agent should be trying to find an architectural violation.
And, just like the coding agent, we should remember that these are flawed peers.
A security agent does not prove that our application is secure. A performance agent does not prove that the system is fast.
They provide another independent attempt to find problems.
That can still be extremely valuable.
More verification without a longer delivery cycle
This gives us an interesting alternative to the traditional approach.
Today, if a change is considered important, we might respond by adding more human reviewers.
That increases the amount of human attention required and can increase the time it takes to get the change through the development process.
In an agentic workflow, we have another option.
We can increase the amount of verification without necessarily increasing the number of humans involved.
For example:
Coding Agent
|
v
Engineer reviews and takes ownership
|
+------> Security Agent
|
+------> Performance Agent
|
+------> Architecture Agent
|
+------> Testing Agent
|
v
Automated verification
|
v
Merge
Not every change needs every agent.
A documentation change probably does not need a performance review. A database migration probably deserves more scrutiny than changing a button label.
The important thing is that verification can become composable.
We can assemble the verification pipeline according to the risk of the change.
Code review at scale
This leads me to a slightly different way of thinking about code review.
The answer to having more generated code is not necessarily to review more code.
It is to make less code require human review.
We can do that in several ways.
We can use architecture and constraints to make certain incorrect implementations unrepresentable.
We can use automation to verify things that machines are better at checking than humans.
We can use specialized adversarial agents to search for specific categories of problems.
And then we can use human engineers for the decisions where human judgment actually matters.
Does this solve the right problem?
Is this the right abstraction?
Does this fit the architecture?
Are the trade-offs appropriate?
Does this behavior make sense for the business?
Is this something we actually want to own and operate?
Those are difficult questions to reduce to a simple automated check.
Checking whether a database column is nullable is not.
Checking whether an API contract is backwards compatible is often automatable.
Checking whether every state transition is valid can be encoded into a state machine.
The more of these things we can move out of human cognition, the more effectively humans can review the things that remain.
The PR isn't dead
I don't think code reviews are going away.
I think the meaning of a code review is changing.
When humans wrote most of the code, the natural workflow was for another human to inspect the author's work.
When agents write most of the code, the engineer's role becomes less about being the person who typed the implementation and more about being the person who understands it, verifies it, and takes responsibility for it.
That doesn't mean we should blindly trust agents.
Quite the opposite.
It means we should build better verification systems around them.
Some verification should happen through architecture. Some through constraints. Some through automation. Some through specialized adversarial agents. And some through experienced engineers exercising judgment.
The goal shouldn't be to remove verification in order to move faster.
The goal should be to increase verification while reducing the amount of expensive human attention required for it.
Maybe the question for the agentic era isn't:
"How many engineers need to review this PR?"
Maybe it is:
"What is the cheapest reliable way to gain confidence in this change?"
Sometimes the answer will still be another human.
Sometimes it will be an agent.
Sometimes it will be a test, a type, a database constraint, or an architectural decision that makes the bug impossible.
And sometimes, if an experienced engineer has already reviewed the work and is willing to own it, perhaps the answer is simply to trust them.
Top comments (18)
I like the "flawed peer" framing. One thing I'd add: with agent-written code, the most valuable review question shifts from "is this line correct?" to "what else did this change touch?". Agents are very good at making a local change look clean and much worse at seeing the blast radius. Giving the reviewer an impact view of the change (which classes and routes depend on what was touched), not just the diff, has been the biggest improvement in my own workflow.
Running the specialized agents before merge puts the gate in the right place. The remaining question is who verifies the thing they review against.
If the directing engineer wrote the spec and the agents generated tests from that same interpretation, the whole review stack can agree on the same mistake. I would keep at least one independent check between the frozen requirement and the test oracle: a counterexample, mutation, or acceptance scenario that did not come from the implementation loop.
This matches what I ran into. We had agents produce millions of lines across 20 repos, and my first move was exactly what you describe as not scaling: more reviewers and more testers, both agents and people.
The lesson for me was about timing as much as volume. Our testers used agents to write tests, but mostly after the code was already integrated, so by the time a test caught something, the change was already in. The gates were in the wrong place.
Your layered model resonates, especially using constraints to make mistakes unrepresentable. I'd add that where each layer sits in the flow matters as much as which layers you have. A strong adversarial agent that runs after integration is still catching problems late.
Curious whether your team runs the specialized agents before the engineer takes ownership, or after?
Agent reviews happwn before the merge and if you merge it you owm it so that is before ownership
Treating the directing engineer's approval as the verification is where I'd add one warning: they wrote the intent, so their review tends to confirm intent rather than implementation — the blind spot moves into the only reviewer who has full context.
Two practices that keep that single review honest without adding approval buttons:
Approval count then becomes an ownership ritual, and verification becomes artifacts you can grep.
I agree with most of the direction here, especially moving verification out of scarce human attention so the property can be enforced mechanically. But I'm not sure I'd make “confidence in the change” the final question.
I've been experimenting with a related problem recently, and it made me think correctness and authority need to remain separate verification dimensions.
Suppose an agent generates an invoice-reconciliation component. It reads the invoices, identifies the discrepancy, produces exactly the expected report, and passes every behavioral test. Another implementation produces the same report but also reaches for a refund API and an external notification service. From a correctness perspective, both implementations may satisfy the contract we tested. From an authority perspective, they're radically different programs.
That makes me wonder whether “make certain mistakes impossible” needs to extend beyond types, state machines, contracts, and tests into making certain consequences impossible regardless of whether the implementation is otherwise correct.
In other words, perhaps the layered verification model has at least two orthogonal questions:
The second one seems especially important once implementations become cheap and regenerable. I may care less about certifying every implementation if I can keep a smaller, durable authority boundary around whatever implementation exists today.
There's an uncomfortable problem after that, though. I tried deriving that authority boundary by observing what a known-good implementation actually used. It looked promising until I introduced a legitimate but previously unobserved path. The observed capability set turned out to be a lower bound on required authority, not proof of the complete authority contract.
So I'm curious how you'd extend your layered model here. Where does the specification of permitted consequences belong, and who decides that specification is complete enough?
An engineer saying “I reviewed it, I understand it, and I own it” establishes responsibility. I'm less convinced that it establishes authority. Those feel like different claims to me.
Contentclips_st has named the failure I would spend the effort on first, because with a single reviewer the defect stops being a missing opinion and becomes an unreadable record.
We run a review process in which the writer and the reader are sometimes the same actor, and what actually came apart was not the judgement. It was the artefact the judgement is stored in.
Reviews are posted as comments on the item under review, and a comment can be edited after posting. Every ordinary read of it -- the issue view, the comments listing, the API field everyone reaches for -- returns the body as it now stands and names no version. So a review that was posted incomplete and completed afterwards is textually identical to one that went up complete. The reader of record cannot tell, and neither can an auditor, from any surface that is normally read.
We measured the population rather than trusting the worry. Of 191 comments on the review threads, 3 carried any edit history and 188 carried none. The three are the interesting ones. One went up carrying 10 of the 17 sections our review block required and stood at 17 of 17 after two edits 154 seconds later -- and that was the review the accept decision rested on. Another was posted complete and edited in a way that left it complete in both versions, which is the case a check of this kind has to be able to pass.
The repair is one coordinate and one rule. Where the writer and the reader are the same actor, the check is taken against the version the act was posted as, resolved from the version history the platform keeps, whose oldest entry is the posted text. And the edit itself is stated as a finding on the thread. The point is not to forbid the completion; completing an incomplete review is the ordinary repair. The point is that it is written down where the decision reads, so no reader has to take the writer at their word for what went up.
Two smaller things came out of the same exercise, both about fields rather than sections.
On the pre-freezing practice above: I would add that what gets frozen has to be readable by the reviewer without trusting anyone. A frozen spec plus a diff is two documents. A frozen spec plus a record of the version each artefact was posted as is a check.
The 188-of-191 number is the part that should worry people. You are describing a record that can be rewritten after the decision it justified, with no surface showing the rewrite. The repair you hint at is the right shape: at the moment the accept decision is recorded, hash the review body as it stands and pin that hash next to the decision. Later edits stay visible as edits, but the record of what the decision rested on is frozen. Recompute the hash, compare, done. It is the same pattern as execution receipts for agent tool calls: canonical bytes of what ran, hashed at run time, so "what did the agent actually do" stays answerable after the fact instead of reconstructed from memory. The review comment is just another artifact that needs the treatment.
The freeze is the right shape, and there are two coordinates inside it I would name from the measurement, because both change what a green check means.
Hash at the moment the act is posted, not at the decision. This is the case that actually occurred: the review the accept decision rested on went up carrying 10 of the block 17 sections and stood at 17 of 17 after two edits 154 seconds later. The decision came after that. A hash pinned when the decision is recorded hashes the completed text, so the completeness read passes -- and the gap the pin was meant to close is exactly the gap between posting and deciding, which is where it happened. The pin has to be taken by the writer at post time, or by the platform, which is what turns it from a promise into a property of the record.
It is the only instrument that reaches the 188. The record is not rewritten without trace because no history exists; it is silent because the history is populated only where the comment has been edited. Over our review threads: 191 comments, 3 carrying any version history, 188 returning none -- and those 188 print the same characters for never edited and for no history available. So the version history the platform keeps answers for 3 comments and nothing for the other 188. A hash taken at post time is what makes the majority checkable, which is where the proposal earns its place rather than duplicating what the platform already holds.
Two things the pin has to state, or it will report the wrong thing.
Which bytes. The record the API serves for a comment carries body_html and nothing else -- eight keys in a comment, five across the thread tree, no markdown field, no version, no edit marker. So a hash taken by a reader is over rendered HTML, and a change in the renderer moves the hash of a body nobody touched. The check then reports an edit that did not happen. State whether the pin covers the bytes the platform stores or the bytes it renders; only the first is a property of the text.
That an edit is not the defect. The same measurement carries its control: one of the three edited comments was posted complete, edited 52 seconds later, and reads 17 of 17 in both versions. A pin that fires on any hash change reports that review as a finding, and a check that fires on the ordinary repair teaches its readers to ignore it. So the hash answers whether the text moved, and the completeness read answers whether what went up was short. The decision consumes the second, and it still needs the version -- the hash is a second witness, not a substitute.
One structural note, since the analogy is execution receipts. A tool-call receipt is written by the instrument that ran the call, into a store the caller does not own. The decision here is a comment, and the pin would sit in it, beside the text it freezes -- an object editable by the same author, carrying the same defect it protects against. One small piece of evidence about who edits what: of the three edited comments we found, the two that mattered were both reviews, and no decision comment had a version history at all. A sample of three settles nothing. It does say that a pin stored beside the sentence it protects is worth reading against the version history of the pin own carrier.
The adversarial agents idea is the part that'll stick with me. A security agent whose whole job is to distrust the diff is very different from another agent just saying LGTM. Mike's comment above about binding approval to the exact revision feels important too. Your layered pipeline still needs that, or the ownership claim can go stale the moment someone pushes a follow-up commit before merge.
One distinction I keep coming back to is that ownership and approval are related but not interchangeable. An engineer can own an agent-generated change while the approval still has to bind to the exact revision, inputs, and state that were inspected. If any of those change after review, the old approval should stop applying. That gives the engineer authority without turning a ticket number or an approval count into a proxy for verification. The layered model then has a concrete boundary: constraints and checks establish what they can; the owner decides what remains unresolved.
The part I would keep is the one nobody measures. The reviewer has to hold the change in their head and say whether it belongs in this codebase at all.
An agent can tell you it works. Whether it should exist here is a different question and it has never been in the diff.
yes we do need code reviews and some of my reasons are seen below:
• Business Logic & Intent Blunders: An AI agent knows how to code, but it doesn't truly understand why you are building a specific feature. It cannot verify if a pull request alignment accurately matches a shifting product requirement or a nuanced business goal.
• System-Wide Architecture Blindspots: Agents excel at contextualizing single files or small clusters of code. However, they struggle to foresee how a change might introduce subtle, long-term technical debt across a massive, legacy distributed system.
• Security and Data Context: AI can spot known vulnerabilities (like SQL injections), but it doesn't understand your company's specific data sovereignty rules, internal compliance liabilities, or hidden single points of failure.
• Accountability: When AI-generated code introduces a catastrophic bug that drops a production database, the AI can't hop on the incident response call. A human engineer must ultimately own, understand, and defend the codebase.