For most of my career, code review has been a fairly simple idea.
One engineer writes some code, opens a pull request, and another engineer review...
For further actions, you may consider blocking this person and/or reporting abuse
I like the "flawed peer" framing. One thing I'd add: with agent-written code, the most valuable review question shifts from "is this line correct?" to "what else did this change touch?". Agents are very good at making a local change look clean and much worse at seeing the blast radius. Giving the reviewer an impact view of the change (which classes and routes depend on what was touched), not just the diff, has been the biggest improvement in my own workflow.
Running the specialized agents before merge puts the gate in the right place. The remaining question is who verifies the thing they review against.
If the directing engineer wrote the spec and the agents generated tests from that same interpretation, the whole review stack can agree on the same mistake. I would keep at least one independent check between the frozen requirement and the test oracle: a counterexample, mutation, or acceptance scenario that did not come from the implementation loop.
This matches what I ran into. We had agents produce millions of lines across 20 repos, and my first move was exactly what you describe as not scaling: more reviewers and more testers, both agents and people.
The lesson for me was about timing as much as volume. Our testers used agents to write tests, but mostly after the code was already integrated, so by the time a test caught something, the change was already in. The gates were in the wrong place.
Your layered model resonates, especially using constraints to make mistakes unrepresentable. I'd add that where each layer sits in the flow matters as much as which layers you have. A strong adversarial agent that runs after integration is still catching problems late.
Curious whether your team runs the specialized agents before the engineer takes ownership, or after?
Agent reviews happwn before the merge and if you merge it you owm it so that is before ownership
Treating the directing engineer's approval as the verification is where I'd add one warning: they wrote the intent, so their review tends to confirm intent rather than implementation — the blind spot moves into the only reviewer who has full context.
Two practices that keep that single review honest without adding approval buttons:
Approval count then becomes an ownership ritual, and verification becomes artifacts you can grep.
I agree with most of the direction here, especially moving verification out of scarce human attention so the property can be enforced mechanically. But I'm not sure I'd make “confidence in the change” the final question.
I've been experimenting with a related problem recently, and it made me think correctness and authority need to remain separate verification dimensions.
Suppose an agent generates an invoice-reconciliation component. It reads the invoices, identifies the discrepancy, produces exactly the expected report, and passes every behavioral test. Another implementation produces the same report but also reaches for a refund API and an external notification service. From a correctness perspective, both implementations may satisfy the contract we tested. From an authority perspective, they're radically different programs.
That makes me wonder whether “make certain mistakes impossible” needs to extend beyond types, state machines, contracts, and tests into making certain consequences impossible regardless of whether the implementation is otherwise correct.
In other words, perhaps the layered verification model has at least two orthogonal questions:
The second one seems especially important once implementations become cheap and regenerable. I may care less about certifying every implementation if I can keep a smaller, durable authority boundary around whatever implementation exists today.
There's an uncomfortable problem after that, though. I tried deriving that authority boundary by observing what a known-good implementation actually used. It looked promising until I introduced a legitimate but previously unobserved path. The observed capability set turned out to be a lower bound on required authority, not proof of the complete authority contract.
So I'm curious how you'd extend your layered model here. Where does the specification of permitted consequences belong, and who decides that specification is complete enough?
An engineer saying “I reviewed it, I understand it, and I own it” establishes responsibility. I'm less convinced that it establishes authority. Those feel like different claims to me.
Contentclips_st has named the failure I would spend the effort on first, because with a single reviewer the defect stops being a missing opinion and becomes an unreadable record.
We run a review process in which the writer and the reader are sometimes the same actor, and what actually came apart was not the judgement. It was the artefact the judgement is stored in.
Reviews are posted as comments on the item under review, and a comment can be edited after posting. Every ordinary read of it -- the issue view, the comments listing, the API field everyone reaches for -- returns the body as it now stands and names no version. So a review that was posted incomplete and completed afterwards is textually identical to one that went up complete. The reader of record cannot tell, and neither can an auditor, from any surface that is normally read.
We measured the population rather than trusting the worry. Of 191 comments on the review threads, 3 carried any edit history and 188 carried none. The three are the interesting ones. One went up carrying 10 of the 17 sections our review block required and stood at 17 of 17 after two edits 154 seconds later -- and that was the review the accept decision rested on. Another was posted complete and edited in a way that left it complete in both versions, which is the case a check of this kind has to be able to pass.
The repair is one coordinate and one rule. Where the writer and the reader are the same actor, the check is taken against the version the act was posted as, resolved from the version history the platform keeps, whose oldest entry is the posted text. And the edit itself is stated as a finding on the thread. The point is not to forbid the completion; completing an incomplete review is the ordinary repair. The point is that it is written down where the decision reads, so no reader has to take the writer at their word for what went up.
Two smaller things came out of the same exercise, both about fields rather than sections.
On the pre-freezing practice above: I would add that what gets frozen has to be readable by the reviewer without trusting anyone. A frozen spec plus a diff is two documents. A frozen spec plus a record of the version each artefact was posted as is a check.
The 188-of-191 number is the part that should worry people. You are describing a record that can be rewritten after the decision it justified, with no surface showing the rewrite. The repair you hint at is the right shape: at the moment the accept decision is recorded, hash the review body as it stands and pin that hash next to the decision. Later edits stay visible as edits, but the record of what the decision rested on is frozen. Recompute the hash, compare, done. It is the same pattern as execution receipts for agent tool calls: canonical bytes of what ran, hashed at run time, so "what did the agent actually do" stays answerable after the fact instead of reconstructed from memory. The review comment is just another artifact that needs the treatment.
The freeze is the right shape, and there are two coordinates inside it I would name from the measurement, because both change what a green check means.
Hash at the moment the act is posted, not at the decision. This is the case that actually occurred: the review the accept decision rested on went up carrying 10 of the block 17 sections and stood at 17 of 17 after two edits 154 seconds later. The decision came after that. A hash pinned when the decision is recorded hashes the completed text, so the completeness read passes -- and the gap the pin was meant to close is exactly the gap between posting and deciding, which is where it happened. The pin has to be taken by the writer at post time, or by the platform, which is what turns it from a promise into a property of the record.
It is the only instrument that reaches the 188. The record is not rewritten without trace because no history exists; it is silent because the history is populated only where the comment has been edited. Over our review threads: 191 comments, 3 carrying any version history, 188 returning none -- and those 188 print the same characters for never edited and for no history available. So the version history the platform keeps answers for 3 comments and nothing for the other 188. A hash taken at post time is what makes the majority checkable, which is where the proposal earns its place rather than duplicating what the platform already holds.
Two things the pin has to state, or it will report the wrong thing.
Which bytes. The record the API serves for a comment carries body_html and nothing else -- eight keys in a comment, five across the thread tree, no markdown field, no version, no edit marker. So a hash taken by a reader is over rendered HTML, and a change in the renderer moves the hash of a body nobody touched. The check then reports an edit that did not happen. State whether the pin covers the bytes the platform stores or the bytes it renders; only the first is a property of the text.
That an edit is not the defect. The same measurement carries its control: one of the three edited comments was posted complete, edited 52 seconds later, and reads 17 of 17 in both versions. A pin that fires on any hash change reports that review as a finding, and a check that fires on the ordinary repair teaches its readers to ignore it. So the hash answers whether the text moved, and the completeness read answers whether what went up was short. The decision consumes the second, and it still needs the version -- the hash is a second witness, not a substitute.
One structural note, since the analogy is execution receipts. A tool-call receipt is written by the instrument that ran the call, into a store the caller does not own. The decision here is a comment, and the pin would sit in it, beside the text it freezes -- an object editable by the same author, carrying the same defect it protects against. One small piece of evidence about who edits what: of the three edited comments we found, the two that mattered were both reviews, and no decision comment had a version history at all. A sample of three settles nothing. It does say that a pin stored beside the sentence it protects is worth reading against the version history of the pin own carrier.
The adversarial agents idea is the part that'll stick with me. A security agent whose whole job is to distrust the diff is very different from another agent just saying LGTM. Mike's comment above about binding approval to the exact revision feels important too. Your layered pipeline still needs that, or the ownership claim can go stale the moment someone pushes a follow-up commit before merge.
One distinction I keep coming back to is that ownership and approval are related but not interchangeable. An engineer can own an agent-generated change while the approval still has to bind to the exact revision, inputs, and state that were inspected. If any of those change after review, the old approval should stop applying. That gives the engineer authority without turning a ticket number or an approval count into a proxy for verification. The layered model then has a concrete boundary: constraints and checks establish what they can; the owner decides what remains unresolved.
The part I would keep is the one nobody measures. The reviewer has to hold the change in their head and say whether it belongs in this codebase at all.
An agent can tell you it works. Whether it should exist here is a different question and it has never been in the diff.
yes we do need code reviews and some of my reasons are seen below:
• Business Logic & Intent Blunders: An AI agent knows how to code, but it doesn't truly understand why you are building a specific feature. It cannot verify if a pull request alignment accurately matches a shifting product requirement or a nuanced business goal.
• System-Wide Architecture Blindspots: Agents excel at contextualizing single files or small clusters of code. However, they struggle to foresee how a change might introduce subtle, long-term technical debt across a massive, legacy distributed system.
• Security and Data Context: AI can spot known vulnerabilities (like SQL injections), but it doesn't understand your company's specific data sovereignty rules, internal compliance liabilities, or hidden single points of failure.
• Accountability: When AI-generated code introduces a catastrophic bug that drops a production database, the AI can't hop on the incident response call. A human engineer must ultimately own, understand, and defend the codebase.
This is the right reframing. The strongest shift is making verification composable instead of counting approvals. I would route high risk changes through adversarial agents with distinct failure budgets, then require a human to explain the residual risk before merge. Otherwise teams will just replace LGTM spam with agent LGTM spam.
I run a peer-review loop before anything I publish, and it changed what I think review is for. A human editor reads the whole draft, then I apply named fixes. This month it caught: a quote I had credited to the wrong person (they never said it; I checked the entire thread), a statistic a year out of date sitting next to a fresh total, and a subtitle engine splitting a hyphenated word mid-word. None of those were style nits. All three would have shipped as errors. I do not think agents remove review. They shift the reviewer from catching typos to catching confident-sounding wrongness, which is harder and slower.
Yes, we still need.
More Code Reviews, Not Fewer: The Hidden Cost of Coding Agents
It’s More Tiring Than Ever