I ran a stupid little experiment a few months ago. I built the same app twice.
The first time, the slow way — by hand, the way I used to build everything. Typing most of it, reaching for the AI only when I got stuck. It took most of a week of evenings.
Then, a month later, half out of boredom, I threw the whole thing away and rebuilt it from scratch. Same features, same spec. Except this time I let the AI do the writing. I described what I wanted, it produced, I steered. It took an afternoon.
The second version is better. Cleaner code, fewer rough edges, more consistent naming. It even passed a couple of edge-case tests the first one didn't. By every number you'd put on a dashboard, the AI version won, and it wasn't close.
And I trust it less. Not a little less — a lot less.
That gap is the whole thing I've been chewing on since, because it isn't supposed to exist. Better code is supposed to earn more trust, not less. Mine did the exact opposite.
Speed and trust weren't uncorrelated. They were inverted.
I went in assuming trust would track quality. Write better code, trust it more — obviously. The AI version is higher quality. So why do I open the hand-built one with ease, and open the AI one with a small flinch, every time?
The faster I built, the less I knew where it would break. Those two moved together, in the wrong direction, and once I saw it I couldn't unsee it. Which means trust was never a property of the code at all. It was a property of something else — something the fast version skipped without telling me.
What "trust" actually turned out to be
When I built version one by hand, I ended up with something that never made it into the repo: a map.
I knew which parts were load-bearing and which were duct tape. I knew the one function that quietly assumed the list was already sorted. I knew the exact spot where, if the same request came in twice, something ugly happened. Nobody wrote any of that down. It lived in my head, and I'd built it without noticing — one decision at a time, because every one of those lines had made me pause. Ack here, or after the save? Trust this input, or validate it? What happens if this gets called twice? Each pause left a little pin in a mental map of where the bodies were buried.
That map is what trust is. Not "the code is correct" — I can't actually prove that about either version. Trust is "I know where this breaks." I had that map for version one. I did not have it for version two, and no amount of the AI's cleaner code was going to give it to me.
The typing was never the point. It was where the map came from.
Here's the part that actually unsettled me.
For years I thought the value of writing code by hand was the code. The craft. The output. It wasn't. The AI writes better code than my hands do — version two settled that argument. The typing was a side effect. What the typing secretly produced was the map. Every keystroke was a small decision, and the decisions were the thing I was really accumulating.
I thought I was manufacturing code. I was manufacturing the knowledge of where it breaks, and the code was the byproduct. AI handed me the byproduct for free and quietly kept the thing I actually needed.
The one line that proves it
Let me make it concrete, because I can point at the exact line.
In both versions there's a write path that acknowledges the request before it persists the row. In version one, I remember writing that line. I remember pausing on it. I remember moving the ack down below the save, because some part of me went cold when I looked at the order.
That cold feeling has a source. Years ago I shipped exactly that bug — a write path that told the client "got it" before it had saved anything. The code was clean, idiomatic, tested, green. Then one ordinary day a retry hit at the wrong moment: the "got it" went out, the save never landed, and a paying customer got locked out of their own account with nothing in the logs to say they'd ever been there. Ack before persist. I can still feel the phone call.
So when I typed that line by hand in version one, the scar fired, and I fixed the bug before it existed.
In version two, the AI wrote the same line — ack before persist — confidently, cleanly, and I scrolled right past it. Not because I'd gotten dumber in a month. Because I never typed it. The decision never passed through me, so the scar never got its chance to fire.
That's the moment it clicked. The typing wasn't only building the map. It was the checkpoint — the turnstile every line used to pass through, where each of my scars got a vote on the way in. Automate the typing and you don't just lose the map. You remove the one place my judgment used to get consulted automatically, for free, on every single line.
To be clear — this is not "write it by hand"
I'm not telling you to go back to typing everything. I shipped version two, not version one. The AI code is better and I'd never hand-roll that app again. The typing was never worth defending for its own sake.
That's actually the trap: thinking the fix is to slow down and type more, like the keystrokes were sacred. They weren't. The problem isn't that the AI writes the code. The problem is that we got the code and assumed the map came stapled to it — the way it always had, for twenty years, because for twenty years you couldn't get the code without doing the work that produced the map.
The map used to be free. It came bundled with the keystrokes whether you wanted it or not. Now it's unbundled, and most people are shipping the code without ever paying for the map — and calling the speed a pure win.
So I buy the map separately now, on purpose
The job changed. I can't get the map from the typing anymore, so I buy it deliberately, after the AI writes:
I read the diff like a stranger wrote it. A stranger who's trying to make me look good in a demo and does not care what happens at 2am. Because that is exactly what wrote it. Admiration is the enemy here — the cleaner it looks, the harder I read.
For every file, I make myself finish "this breaks when ___." Out loud, specifically. If I can't finish the sentence, that's not a pass — it's a hole in the map, and the hole is the part I don't understand yet. "Looks good" is not a map. It's the absence of one.
I ship it small. Behind a flag, to a canary, to one internal user. So the parts of the map I missed get filled in where filling them in is cheap — instead of at 2am, on a real customer, the expensive way.
None of that is typing. It's the thing the typing used to do for me, done on purpose instead of as an accident. Slower than just accepting the diff. Still a fraction of the week I spent on version one. And it gets me the one thing the AI genuinely could not hand me: knowing where the thing breaks.
Why this is the exact reason I build the way I do
One level up, it's the same problem.
An agent that writes code is fast, and clean, and produces a beautiful artifact with no map attached. And worse than mapless — it signs off on its own work. The thing that wrote the diff has no scar to fire and nothing to be suspicious of. It's version two with the confidence of version one and none of the earned caution. The most dangerous possible combination: fluent, fast, and completely unable to distrust itself.
So I never let the thing that writes the code be the thing that blesses it. There's an author that produces the diff — fast, clean, mapless. There's a separate skeptic whose entire job is to build the map the author skips: read the diff as a stranger, finish "this breaks when ___," try to make it fall over. And there's a human on the merge button, who owns the call and keeps collecting the real scars. Author, skeptic, human. The skeptic is the typing turned into a seat — the checkpoint every line used to pass through, rebuilt as its own job so it doesn't vanish the moment the author got fast. That's the whole shape of xenition.
I'll keep letting it write the code. The typing was never the point — version two proved that for good. But I stopped pretending the map comes free with the keystrokes, because it doesn't anymore. The code got faster. The trust didn't come with it. So now I pay for the trust on purpose, as its own line item, because it stopped being a side effect of the work.
I built the same app twice. The fast version is the one I ship. The slow version is the one that taught me what the fast one forgot to include.
Top comments (6)
The detail I keep coming back to is that version two passed edge-case tests version one didn't and still shipped ack-before-persist. That isn't a contradiction, it's what that bug looks like to a test suite: the failure needs a retry or a crash to land between the ack and the save, and a test that calls the handler and checks the response never opens that gap. So the map you lost isn't only where things break, it's which breakages your tests are structurally blind to. A cheap partial substitute might be walking each write path line by line asking "what has the client been told here, and what has actually been saved?", since ordering bugs are the ones scrolling hides best.
Yeah, you put your finger on the exact thing that bugged me about it. The green test suite wasn't lying, it just wasn't being asked the right question. It checked "did the handler return the right thing" and never "did the right thing return before the save landed." The gap only exists in wall-clock time, and a unit test collapses wall-clock time to zero — so the bug is invisible by construction, not by accident.
And "which breakages your tests are structurally blind to" is better than how I put it. I was treating the map as one thing — where it breaks — when it's really two: where it breaks, and where your safety net has holes you can't see from inside the net.
Your line-by-line fix is the move. Walk each write path asking "what has the client been told at this point, and what's actually durable right now?" Ordering bugs hate that because it forces you to read in execution order instead of top-to-bottom — and scrolling only ever gives you top-to-bottom. That mismatch is exactly where the eye slides past.
Might steal "structurally blind." Good comment.
The discussion around tests being "structurally blind" hits the nail on the head, but there's a second trap people probably fall into: "we'll just catch this with E2E or negative test cases."
In reality, they almost never do.
Negative tests look for malformed payloads or missing auth, whereas
ack-before-persistships 100% valid data. And in a healthy CI pipeline, your E2E test sends the payload, gets the premature 200 OK, queries the DB a few milliseconds later—and by then, the row has already quietly landed. 100% green, 100% false confidence.The only test that actually bought the "failure map" back for me was temporal fault injection: intentionally holding the connection or killing the worker in that tiny micro-gap between the response and the durable write, and asserting that the UI doesn't prematurely celebrate. Unless you deliberately pry that 2ms gap open in tests, the AI's blind spot simply becomes your 2am fire.
Interesting framing: trust as a map of where things break. I ran a similar experiment from the other side: built a full app (backend, web, PWA, Chrome extension) without ever reading the code, and had to get the "map" some other way.
For me it came from design instead of typing: written module boundaries (a short note per module on what it owns and must not touch), standards in files, and constantly using the product like a user. It is not proof of correctness, and bugs I found along the way went into a known-issues file rather than being fixed on the spot. But it did make me trust the result more than I expected.
Wrote it up here as one more data point: dev.to/nomad4tech/i-built-a-time-t...
This is the version of the argument I was hoping someone would make, because it's the part I left implicit. If the map is the real deliverable, then typing was only ever one way to produce it — and you found another. Written module boundaries, "what it owns and must not touch," standards in files: that's the map, just authored up front instead of accumulated line by line. You paid for it deliberately, which is the whole point I was fumbling toward.
The known-issues file is the detail I like most. You resisted the reflex to fix on the spot, and that's exactly the discipline — a fix closes the ticket but doesn't necessarily enter the map, whereas writing it down is the map getting drawn. Most people would've patched and forgotten, and the knowledge would've evaporated the same way version two's did for me.
The one place I'd poke: boundaries and standards tell you where things are supposed to break — the planned seams. My scar stuff is about the breaks nobody designed, the ack-before-persist that looks fine against every rule you wrote down. "Constantly using the product like a user" is probably what catches those for you, the way the 2am phone call caught it for me. Curious whether the bugs in your known-issues file came more from the design notes or from just living in the thing.
Reading your write-up now. Good to see it from the other side.
The ack-before-persist example nails it, clean idiomatic code with green tests that only blows up on a mis-timed retry. Your framing that trust is the mental map of where it breaks, not a property of the code, is what I keep failing to explain about AI-generated code: when you didn't type it, you never built that map. That reconstruction is exactly what tracing and evals end up doing after the fact.