DEV Community

Cover image for The best engineer on my team ships the least code.
Info Inlet
Info Inlet

Posted on

The best engineer on my team ships the least code.

The best engineer on my team almost got a bad review last quarter for shipping too little code.

I'm not being cute. I pulled the numbers before his review because I knew the numbers were going to be used against him. Fewest commits on the team. Smallest diffs. Whole sprints where his name was on almost nothing in the changelog. If you ranked the team by anything a dashboard can count, he came out near the bottom, and it wasn't close.

He's also the single reason we didn't lose a paying customer's money in March. One comment on a pull request. No code. Just "wait — what does this do if the webhook retries?" And the thing that was about to ship, clean and green and approved by two other people, would have quietly eaten payments on a bad night.

So I had a person whose output said coast and whose actual value said irreplaceable, and a review process that could only see the first one. That gap is the whole thing I want to talk about, because it isn't a quirk of my team. We built it on purpose, over twenty years, and then the ground moved under it and we didn't update the instrument.

For twenty years, output was a fair proxy. That's the part nobody questions.

Here's why "measure the code" was never stupid. For most of the history of this job, writing the code was the hard part. It was slow, it was scarce, and it was bottlenecked on the engineer. If you shipped a lot of good code, it genuinely meant something — you had the skill, you understood the system, you did the work nobody else could do as fast. Output was expensive, so output was a decent stand-in for value. Not perfect, but honest enough to run a company on.

That's the assumption buried under every perf review, every promo packet, every "impact = things shipped" rubric in the industry. Output is scarce, so output is signal. It was true for so long that we stopped noticing it was an assumption at all. It just felt like what engineering is.

Now say the quiet part. The thing that assumption rested on — that producing the code is the scarce, hard, bottleneck part — is the exact thing that stopped being true.

The instrument is measuring the one thing that stopped being scarce

Code is cheap now. Not free, not always correct, but cheap — a competent diff for a well-scoped task is minutes of generation, not days of typing. The bottleneck moved. It is no longer "can we produce the code." It's "can we tell which of the confident, clean, plausible diffs in front of us is quietly wrong."

And look what that does to the dashboard. We are still ranking engineers by lines, commits, PRs, story points closed — by output — at the exact moment output became the cheapest input in the building. We kept the twenty-year-old proxy and threw away the twenty-year-old reason it worked. It's like grading chefs by how fast they can boil water, the year the instant kettle shipped. The number still moves. It just stopped meaning anything.

That's the trap in one line:

Output used to be scarce, so we measured it. It stopped being scarce, and we never stopped measuring it.

The engineer who leans on the machine and ships a mountain looks like your top performer. The engineer who reads every one of those mountains and catches the one that loses money looks like he's barely working. The instrument has the ranking exactly backwards, and it's confident about it.

What actually became scarce

When code was the bottleneck, judgment was bundled into the act of writing it — you couldn't produce the thing without making a thousand small decisions along the way, so the two came stapled together and you never had to measure them separately. Abundance ripped the staple out. Now you can have all the output you want without anyone exercising judgment over it, which means judgment is suddenly its own scarce input — standing alone, un-bundled, for the first time.

And the scarcest form of judgment is the one with the lowest status: the "no." The ability to look at a clean, confident, approved-by-two-people diff and say that breaks under a retry, don't ship it. That skill doesn't produce anything. It prevents. And prevention is invisible to every tool we use to decide who's valuable.

My quiet engineer is the best person on the team because he has the most of the thing that got scarce and the least of the thing that got cheap. The review process reads that as underperformance, because the review process was built in the world where those two were the same thing.

The bug that proves it — and the review that caught it

Let me make it concrete, because I lived the before-version of this once and it's why I now watch for the after.

Years ago, on a different team, I shipped a write path that acknowledged a request before it had actually persisted the row. The code was clean. Idiomatic, typed, tested, reviewed, green. By every measure a dashboard could take, it was a productive, high-quality piece of work, and it made my output numbers look great that sprint.

Then one ordinary day a retry hit at exactly the wrong moment. The "got it" went out, the save never landed, and a paying customer got locked out of their own account with nothing in the logs to say they'd ever been there. Ack before persist. I can still feel the phone call.

Now back to March. Same bug, different seat. The diff in front of us acked the payment webhook before it wrote the row — clean, green, already approved. Two engineers who had shipped a lot that sprint looked right at it and moved on, because it reads fine; you only catch it if something in you is actively asking "what happens on the retry." My quiet engineer asked. One sentence on the PR. We reordered two lines. Nothing shipped under his name that day — his output for that interaction was negative, he subtracted a line.

That one sentence was worth more than everything the rest of us shipped that sprint combined. And there is no field on the review form where it goes.

The cruel math: prevention has no throughput and no attribution

Here's why this doesn't self-correct, and why it's about to get much worse.

A disaster that ships is a visible, countable event. It has a postmortem, a Jira ticket, a name attached, a number in the incident log. A disaster that doesn't ship is a non-event. Nothing happens. There is no artifact, no ticket, no graph that goes up. You cannot put a prevented outage on a dashboard, because the whole nature of a prevented outage is that there's nothing there to point at.

So the person whose job is the "no" lives inside a brutal asymmetry:

  • When they're right and block the bad diff, nothing happens — so they get no credit. The absence of a disaster looks exactly like a quiet week.
  • When they're wrong and block something that was fine, they're the annoying bottleneck who slowed the team down — visible, countable, held against them.

Right is invisible. Wrong is expensive. That payoff structure doesn't just fail to reward the skill — it actively teaches every rational person to stop using it, because silence is the only move with no downside. The instrument doesn't just miss the best engineer. It quietly trains him to become an average one.

And this is the part that should scare anyone running a team right now: when the layoff list gets made off a productivity dashboard, the person with the fewest commits and the most prevented-but-invisible disasters is sitting at the top of the cut list. We are about to fire our smoke detectors for never having started a fire.

To be clear — this is not "shipping is bad"

I want to be careful here, because the lazy version of this take is "real engineers don't write code, they think deep thoughts," and that's garbage.

Shipping matters. I ship all day. The machine writes most of my code now and I'd never go back — the typing was never the hard part. A team where nobody produces anything isn't full of brilliant skeptics; it's dead. Output is table stakes. You still have to deliver, and the "no" with nothing ever built behind it is just a person slowing down a room.

The claim is narrower and, I think, harder to argue with: output stopped being the scarce thing, so it stopped being the thing worth measuring people by. When the hard part was production, ranking by production was fair. Now the hard part is judgment, and we're still ranking by production — rewarding the abundant input and starving the scarce one. The fix isn't to stop shipping. It's to stop mistaking the thing the machine now does for the thing you're actually paying a human for.

What I do now, since the dashboard won't do it for me

None of this is abstract for me anymore — it changed how I run reviews:

  • I count the "no"s. When someone blocks a diff and is right, that goes in the review as a win, in writing, the same way a shipped feature would. If I don't record it, the dashboard's silence wins by default and the save never happened as far as the company is concerned.
  • I make prevention leave a receipt. A blocked bad diff gets a one-line note on the PR: here's the input that would've lost money, here's what we changed. That artifact is the only way a prevented disaster becomes something a manager can point at later. The skeptic who hands you a repro is cheap to keep. The skeptic who blocks on vibes is the first one cut — not because they're wrong, but because there's nothing to show they were right.
  • I stopped reading a small diff as a small contribution. The most valuable thing a senior did this week might be the three lines they stopped from shipping. I go looking for that on purpose now, because the instrument never will.

Why this is the exact reason I build the way I do

One level up, it's the same problem, and it's why I build an agent platform the way I do.

The whole industry right now wants to cheer for the thing that produces — look how much it ships, look how clean, look how fast. But an agent that generates is just doing the skill that stopped being scarce, at a scale no human can match, and then — worse — it signs off on its own work in the same confident voice whether the diff is a masterpiece or a disaster. It's my high-output engineer with none of my quiet one's instinct: infinite throughput, zero "no."

So I never let the thing that writes the code be the thing that blesses it. There's an author that produces the diff — cheap, fast, endless, the abundant thing. There's a separate skeptic whose entire job is to try to break that diff rather than admire it — the "no" given its own seat, institutionalized, so it doesn't depend on one underappreciated human happening to be in the room. And there's a human on the merge button who owns the call and sees the blast radius neither agent can. Author, skeptic, human. The author has infinite throughput and I measure it by almost nothing. The skeptic has zero throughput and it's the whole product. That's the shape of xenition, and it's the same lesson my review numbers taught me: the moment production got cheap, the thing worth paying for stopped being production.

My quiet engineer figured out, before I did, that the valuable move in 2026 is the one the dashboard can't see. The least code on the team, and the reason the team still has that customer. I almost let a number tell me he was coasting.

Top comments (0)