I have a bookmark folder called prompting. Forty-one tabs in it.
"The 12 prompts that 10x your engineering." "Context engineering for agents, explained." "The system prompt that changed how I ship." A course I finished at 1am, certain I'd finally cracked the thing — that if I just learned to ask well enough, the output would come out right.
I built that folder the same way I once built a GitHub full of dead repos: convinced the scarce skill was the one everybody was selling me.
It wasn't. And I want to say the thing nobody says at the top of the "master prompt engineering, learn context engineering, become an AI-native engineer" sermon:
Getting better at prompting is getting better at the wrong skill. Not a useless one — the wrong one. And the reason is worse than "the models will get better and the prompt won't matter." That's true, but it's the boring half. The sharp half is that even a perfect prompt can't touch the thing that was actually going to hurt you.
Stay with me, because this isn't a doomer post. There's exactly one skill in this whole stack that's still worth every hour — and the logic is what tells you which one.
The three things we tell ourselves about why prompting is THE skill
Nobody grinds prompt guides for no reason. You do it because you believe one of three things:
- It'll make the output better. ("Better prompt, better code. Obviously.")
- It'll make me employable. ("'AI-native,' 'prompt engineer' — this is the moat now.")
- It's the new literacy. ("This is the fundamental skill of the era. Learn it or fall behind.")
All three are real motivations. All three are, as usually practised, pointed at the wrong target. Let me take them one at a time, because the miss is instructive — it points straight at the one skill that survives.
Belief #1, because it's the load-bearing one: a better prompt can't catch the bug that hurts you
Let's do the one everybody believes hardest.
Yes, a better prompt gets you better-looking output. Cleaner structure, fewer obvious mistakes, code that reads like a senior wrote it. That part is real. Here's the part the courses skip:
A better prompt raises how convincing the output is. It does nothing to whether it's correct. And those two dials were never connected.
Think about the failure mode that has actually cost you. It was never "the AI misunderstood me and produced obvious garbage" — you catch that in two seconds. The one that hurts is the opposite: the output that matched your request perfectly, read beautifully, passed the tests you thought to write, and was still wrong in a way you only discovered in production. Plausible. Confident. Broken.
Now watch what a better prompt does to that failure. It makes the output more polished, more authoritative, more obviously-fine at a glance. Which means the wrong ones get harder to catch, not easier. You didn't improve your defense. You upgraded the disguise on the thing you were supposed to be defending against.
That's the whole trap in one line:
A better prompt just gets you to a more convincing wrong answer, faster.
You cannot prompt your way out of this, by construction. The correctness of the answer isn't decided at ask time. It's decided at review time, on your side of the desk, by whether you can look at a clean, confident diff and say no. Prompting optimizes the question. The bug lives in the answer.
Belief #2: prompting is recall wearing a new hat
"But it's a moat. 'AI-native' is what gets hired now."
Here's what nobody wants to hear about their new favorite skill: prompting is recall. The same recall AI just made worthless, re-manufactured one layer up and sold back to you as expertise.
For twenty years, "knowing how to code" was quietly two different things bundled together: recall — the syntax, the API surface, the flag order, the incantation — and judgment — knowing what to build, what to distrust, what breaks at 2am when a real person does something strange. AI ate recall first, and good riddance; it was never the valuable part. Testing for it in 2026 is testing for penmanship.
So what did we do? We rebuilt recall at a higher altitude. "The exact phrasing that makes the model behave." "The context pattern that gets the right output." "The system-prompt incantation." It's the same muscle — memorize the magic words, produce the output — just moved up the stack. We fled the thing AI made free and ran straight into a fresh version of it.
And this version has a shrinking half-life. Every model release infers your intent better than the last. The gap between a novice prompt and an expert prompt narrows with each update, because closing that gap is literally what the labs are optimizing. You are grinding a skill the vendor is actively deleting. You cannot build a moat on the exact thing your supplier ships as a feature next quarter.
Belief #3: the fundamental was never the asking
"Fine — but it's the new literacy. The base skill everything else sits on."
The fundamental skill of working with a system that produces confident, plausible, occasionally-wrong output was never how you ask it. It's what you do with the answer.
Prompting is a UI over the model — and UIs get better on their own, without you. Distrust is a stance toward the output — and it's the one thing that doesn't improve while you sleep. One is a feature. The other is a muscle only you can carry into the room.
There's a completely different skill hiding under "get good at AI," and it only shows up under one specific condition — when the output is wrong and everything about it looks right:
A better prompt produces more convincing output. Only distrust decides whether it's right. The asking is free now — the refusing is the entire job.
You don't earn that by writing better prompts. You earn it by shipping a confident answer to a real person, watching it break, feeling the consequence — and carrying the scar into every diff after. That's the forge. Everything before it is a tutorial with better phrasing.
The moment a perfect prompt handed me a perfect-looking disaster
Let me make this concrete, because I earned it the expensive way.
I once shipped a write path that acknowledged the request before it had actually persisted the row. The prompt that produced it was clean. The code that came back was clean — it read like something a careful engineer wrote. In the demo, in the tests, on my machine: flawless. Exactly what I asked for.
Then one day a retry hit at the wrong moment. The ack went out, the save didn't land, and a paying customer got locked out of their own account with no record they'd ever done the thing. Ack-before-persist. I can still feel it.
No prompt was going to save me there. A better prompt would have made it worse — cleaner code, more convincing, even harder to doubt. The thing that would have caught it wasn't a sharper question. It was the trained reflex to look at a diff that acknowledges before it persists and think that breaks under a retry — before a real human's bad night taught it to me.
That reflex is distrust. It's the only skill in this whole stack that AI can't hand you, can't improve for you, and can't make obsolete. It's the one line on my résumé that would've been worth an entire interview — and no course sells it, because you can't.
Be clear about what I'm actually saying — because it's not "stop prompting"
I am not the guy telling you prompting is beneath you. I prompt all day. AI writes most of my code and I'd never go back — the typing was never the hard part, and neither is the asking. Prompt well. A sloppy prompt wastes everyone's time, yours included.
What I'm saying is narrower and, I think, freeing:
Prompting is table stakes, not the edge. It's the cost of entry now, like knowing how to use a keyboard — necessary, and worth exactly zero as a differentiator, because everyone has it and the tool keeps closing the gap. Grinding it harder is polishing a skill whose ceiling the vendor lowers every month. The edge — the only part that compounds — is what you do after the model answers:
- After the AI answers, before you accept it, try to break it. Ask what it does under a retry, at 2am, when the input is hostile, when the customer does the dumb thing. Attack the answer instead of admiring it.
- Make something other than the author be the skeptic. The thing that produced the diff is the worst possible judge of the diff — it's proud of it. The doubt has to come from a different seat.
- Ship it to someone who can hurt you. That's the only way the wrong ones actually cost you, and cost is the only thing that grows the muscle. A distrust you never had to use isn't a skill. It's a slogan.
Do that, and you're building the one thing that survives every model release. Skip it, and you're the guy with a beautiful prompt library who can't tell when the beautiful answer is going to lock a customer out at 2am.
Why this is the exact reason I build the way I do
Here's the part that goes one level up, because it's the same logic.
I build an agent platform, and the temptation in this whole industry right now is to worship the author — the thing that generates. Look how well it responds to a good prompt! Look how clean the output is! But an author that produces gorgeous, confident output is producing gorgeous, confident output whether or not it's right — and a better prompt only turns the polish up, never the truth. The convincing-ness and the correctness are two different dials, and the whole market is cranking the one that doesn't matter.
So I never let the thing that writes the code be the thing that blesses it. There's an author that produces the diff — prompt it as well as you like. There's a separate skeptic whose only job is to try to break the diff rather than admire it — institutionalized distrust, the "no" made into its own seat. And there's a human on the merge button who can see the blast radius the model can't. That separation — author, skeptic, human — is the whole shape of xenition, and it's the same lesson the bookmark folder taught me: asking is free and getting freer; deciding whether the answer survives contact with reality is the entire job. The prompt is cheap. The "no" is the product.
Your folder of prompt guides isn't a sign you're serious. It's a sign you were optimizing the free thing. Stop grinding the question. Learn to doubt the answer.
Honest question for the comments: what's the most convincing, cleanest, best-prompted AI answer you ever rejected — and how did you know to say no? I want to hear about the one your distrust caught, not the one your prompt produced. 👇
(If this made you close a few prompt-guide tabs with a little less guilt, a ❤️ and a 🔖 help it reach the next person grinding the wrong skill on a Saturday.)
Top comments (28)
I think there is an important distinction between AI literacy and prompt engineering. Someone can write a very sophisticated prompt and still struggle to evaluate whether the result is actually useful, correct, or appropriate for the context.
The more valuable skill is understanding the domain, defining the constraints, providing the right context, and knowing how to verify the output. Prompting is a tool. Judgment is the skill behind the tool.
"Judgment is the skill behind the tool" — that's the sentence the whole post was trying to earn, and you got there in one line. Steal it back from me anytime.
The distinction you're drawing is the one I think most people miss because both things look like "being good at AI." A sophisticated prompt and sound judgment produce similar-looking artifacts on a good day — clean output, confident delivery. They only diverge on the bad day: the moment the result is plausible and wrong. That's the exact instant a great prompt does nothing for you and domain judgment does everything. The prompt got you a beautiful answer; only the judgment can tell you it's going to break under a retry at 2am.
And I'd add one teeth to your list — "knowing how to verify the output" is doing a lot of quiet work there. Verification isn't just ability, it's willingness to say no to something clean. Plenty of people can spot the flaw and still ship it because the diff looked senior and the deadline was real. The judgment includes the spine. That part's even harder to teach than the domain knowledge, which is maybe why nobody sells it.
You basically wrote the thesis in a shorter, better form than I did. 🙏
Exactly. I think that’s the part that gets underestimated.
AI can make a wrong answer look incredibly convincing, so verification isn’t just about finding errors. It’s also being willing to stop and say, “No, this looks good, but it doesn’t hold up.”
That’s where judgment becomes much more than a prompt skill. The prompt helps you produce the output. Judgment decides whether that output deserves to move forward. 🎯
"Judgment decides whether that output deserves to move forward" — that's the cleanest one-line spec for the whole thing I've seen. Producing the output and authorizing it are two different acts, and the entire industry keeps collapsing them into one because the prompt makes them feel like one continuous motion: ask, receive, ship. You've split them back apart, and the split is where all the value hides.
The one thing I'd add, because you're circling it: the convincing-ness actively fights the verification. It's not neutral. The more polished the output, the more your own instinct to wave it through — so the exact quality that makes an answer look ready is the quality that makes saying "no, it doesn't hold up" hardest. Verification isn't just work; it's work performed against your own pull toward the clean-looking thing. That's why it's rare. Finding the flaw takes skill. Overruling a beautiful diff that everyone's ready to merge takes something closer to nerve.
Which is exactly your point — judgment is more than a prompt skill because it operates in the one moment the prompt can't reach: after the answer looks done, when the only thing standing between "convincing" and "shipped" is a person willing to be the friction. The prompt gets you to the door. Judgment is what decides whether the door opens.
Genuinely one of the sharpest exchanges I've had in a comment section — you kept handing me tighter versions of my own argument. 🙏🎯
Another angle I find interesting is that AI itself shouldn't become the system. AI can generate code, suggest solutions, analyze data, or produce content. But a real workflow still needs decisions about where AI is used, what happens before it, how the output is tested, and what happens when the output is wrong.
So I would think of it as:
System → AI → Output → Verification → Decision
rather than:
Prompt → AI → Done
The interesting engineering challenge is designing the whole system around AI, not just learning how to talk to the model.
This is exactly the level-up, and I think you've named the thing the "learn to prompt" framing quietly hides: the model is a component, not the architecture. The second you write Prompt → AI → Done, you've smuggled in the assumption that the output is the deliverable. It isn't. The deliverable is a decision you're willing to stand behind — and the AI touches only one box in that chain.
What I'd push on in your diagram: the interesting failures don't live in the AI box, they live in the arrows. AI → Output is easy. It's Output → Verification where everything is actually won or lost — because verification isn't a step you add, it's a seat you have to staff with something other than the author. The thing that generated the output is structurally the worst judge of it; it's proud of it. So the arrow only works if the skeptic is a different seat than the generator. That's the whole reason I build with author → skeptic → human instead of one confident box.
And your last line is the part I'd underline twice: "designing the whole system around AI, not just learning how to talk to the model." That's the migration nobody's selling a course on — because you can't grind it in a weekend. Prompting is table stakes inside the AI box. The engineering is everything wrapped around it. You get it.
Exactly. That “author → skeptic → human” distinction is important.
I’d even say verification is part of the architecture, not just a final checkpoint. If the same system generates, evaluates, and approves its own output, we’ve basically built a loop with no real external constraint. AI can accelerate the work, but the system still needs a mechanism that can challenge the output before a human has to stand behind the decision.
That’s where I think the real engineering starts. 🎯
"A loop with no real external constraint" — that's the failure mode stated with more precision than I managed in the whole post. A system that generates, evaluates, and approves its own output isn't verified; it's self-certifying, and self-certification is just confidence with extra steps. The evaluation feels like a check but it inherits the exact blind spots of the thing being checked, because it is the thing being checked wearing a second hat. You haven't added a constraint. You've added a mirror and called it a jury.
And your reframe — verification is part of the architecture, not a final checkpoint — is the load-bearing correction. The moment you treat it as a step at the end, you've already lost, because a step can be skipped, softened, or rushed under a deadline, and it sits downstream of all the momentum ("it's basically done, just needs a review"). Baked into the architecture, the skeptic isn't a gate the work flows toward — it's a structural pressure the work is subjected to by construction, with no path to merge that routes around it. One is a speed bump you can floor over. The other is load-bearing wall.
The word doing the real work in your comment is external. That's the whole game. The constraint has to come from outside the generator's own frame — a different seat, a different objective, ideally something that can't be talked into it at all (a deterministic check for the invariants you can name, a skeptic prompted to refute for the fuzzy blast-radius you can't). The instant the challenge originates inside the same system that produced the output, it's not a constraint, it's a preference. External-ness is the property; everything else is implementation.
"That's where the real engineering starts" — exactly. Prompting is the easy box. The architecture around the box — where the external constraint lives, how it's structured so it can't be bypassed, who ultimately stands behind the decision — that's the part that's actually hard, actually compounds, and actually can't be bought as a weekend course. You've basically written the sequel to the post in three replies. 🎯
Agree with the review-time point — and one failure mode from practice that sharpens it: if the model can see the tests before writing the code, it optimizes the oracle, not the behavior. I now write the tests against the requirement, keep them out of the model's context entirely, and run them after — otherwise you get exactly the "passes every test you thought to write, still wrong" class, because the prompt said the same thing the tests say.
Second: for anything that mutates state (file writes, migrations, deletes), I stopped using a second model call as the reviewer. Two confident reviewers just give you two confident wrong answers that agree with each other. A dumb deterministic post-check script — did the file end up valid, does the migration roll back, is the row count sane — is harder to impress but it actually refuses.
The judgment residue is real though: someone still has to hold the requirement in their head and say no to a clean diff. That's the 2am part, and no prompt folder covers it.
Both of these are load-bearing and I want to sit on each for a second because you've operationalized the thing the post left as a slogan.
The oracle point is the sharpest version of "passes every test you thought to write." If the model can see the tests, correctness collapses into test-satisfaction — and those were never the same thing, you just stopped being able to tell them apart. The prompt and the tests saying the same words is the tell: you've built a closed loop that grades its own homework. Keeping the tests out of context and running them after is the difference between an oracle and a mirror. I'm stealing "it optimizes the oracle, not the behavior."
The second point is where I'll push hardest in agreement: a second model call as reviewer is the trap I most want people to see. It feels like distrust — separate seat, second opinion — but it's the same failure mode wearing a lanyard. Two generators trained to be plausible will agree plausibly, and now you've got consensus as a disguise instead of a check. Your framing is exactly right: the dumb deterministic script is "harder to impress but it actually refuses." That's the whole property. Refusal has to come from something that can't be talked into it — and a model, by construction, can always be talked into it. The post-check doesn't have a good day and a bad day. It just has a correct answer.
Where I'd nuance it: the deterministic check and the skeptic-model aren't rivals, they're different ranges. The script nails the invariants you can name — row counts, rollback, valid-on-disk. The model-skeptic is for the fuzzy blast-radius stuff you can't reduce to an assertion ("does this ack before it persists?"). So I'd run the script as the hard gate and the skeptic as a hint generator that never gets to bless anything. Neither one is the human.
Which is your last paragraph, and it's the one I have no cleaner version of: the residue is a person holding the requirement in their head and saying no to a clean diff. Everything above just narrows how often they have to — it never removes the seat. That's the 2am part, and you're right, no folder covers it. Great comment.
Exactly - "hint generator that never gets to bless anything" is the phrase I wish I had coined. One refinement from practice: the skeptic model only earns its keep if its output is consumed as information by a human or by the script gate, never as authority. The failure mode I have seen is gradations of blessing: the script says no, the skeptic says probably fine, and someone merges anyway. So make the skeptic output a separate artifact with zero merge privileges - if reading it changes a decision, that decision was human. Which is exactly your seat that never goes away.
"Gradations of blessing" is the failure mode I've watched sink more review setups than any outright bug — and you've named the mechanism exactly. The moment the skeptic can say "probably fine," you've built a second author: something that emits a soft verdict. Now the human isn't deciding, they're adjudicating between two opinions and picking the one that lets them go home. The script's hard "no" gets laundered into a maybe by a more agreeable voice, and the merge happens on the maybe. You didn't add a check — you added a lawyer for the defense.
Your fix is the right shape, and I'd make the rule as sharp as you did: the skeptic may describe, never conclude. Its output is "here's what breaks under a retry / here's the invariant this touches" — evidence with no verdict field at all. The instant it can emit "probably fine," you've handed it a slice of the merge decision, and authority abhors a vacuum: whatever can express confidence eventually gets treated as if it earned it. So you don't soften its authority — you delete the surface it could express authority on. No verdict to defer to means nothing to defer to.
And your last line is the whole thing: "if reading it changes a decision, that decision was human." Cleanest test I've heard for whether a check is doing its job or quietly usurping it. The skeptic's entire value is that it moves information into the human's head — surfaces the retry case they'd have missed — then gets out of the way before the verdict. The second its output can be the verdict, it stops informing the seat and starts replacing it, and you're back to an author blessing its own genre of work.
Information in, authority never. The gate refuses; the skeptic testifies; the human decides. Three seats — and only one of them is allowed to say yes.
This thread's been the best part of writing the piece. 🤝
This is a compelling reframing of AI fluency: the real differentiator is increasingly judgment, not the ability to formulate increasingly elaborate prompts. I especially liked the distinction between convincing output and correct output, because polished AI-generated work can actually make errors harder to notice. The author/skeptic separation is also a powerful architectural idea—especially when the same system that creates an answer should not be trusted to be its final judge. “Asking is free; deciding whether the answer survives contact with reality is the entire job” captures the central argument beautifully.
Thank you, that's a really generous read, and honestly you've summed it up tighter than parts of the post did.
One thing I'd add on the author/skeptic split, since you zeroed in on it: the trap is assuming a second seat is enough. If the skeptic is just the same model with a "be critical" prompt, it shares the author's blind spot and cheerfully re-derives the same mistake with a frown on. The separation only pays off when the doubt comes from somewhere with a genuinely different prior. A different model, an adversarial frame, a test that actually runs, or a human who's been burned before. Independence is the real ingredient. A second opinion from the same mind is just the first opinion again.
Prompting can improve the output, but it doesn’t replace understanding the problem, validating the result, or knowing what “good” actually looks like. The real advantage comes from combining AI with strong fundamentals and judgment.
Exactly. Though I'd push on one word in there: validating. The trap is that validation feels done the moment the output looks right, and a good prompt makes it look right faster. The answers that actually cost me weren't the ones I failed to validate. They were the ones that sailed through validation because they were clean, confident, and matched what I asked for.
So "knowing what good looks like" is necessary but not quite enough on its own, because a convincing wrong answer also looks like good. The part that saved me is narrower: treating the clean answer as guilty until I've tried to break it, and making sure that doubt comes from a different seat than the thing that wrote it. Fundamentals tell you what good looks like. Distrust is what you do when good-looking turns out not to be the same as good.
I agree with the distinction between prompting and verification. A good prompt can improve the quality of the output, but it doesn't guarantee correctness. The ability to question, test, and challenge AI-generated code is becoming just as important as knowing how to generate it.
Exactly — and I'd push your own point one degree further, because you've landed on the thing most people stop just short of. You said questioning and testing AI code is becoming as important as generating it. I'd argue it quietly became more important, and here's the tell: generating gets cheaper every model release, and verifying doesn't. The vendor ships you a better generator next quarter, for free, while you sleep. Nobody ships you better judgment. So the two skills aren't just both important — they're moving in opposite directions, and the gap widens with every update.
The one nuance I'd add to "test": the tests you think to write come from the same mental model that accepted the code, so they tend to confirm it rather than attack it. The verification that actually catches things is the adversarial kind — "what breaks this under a retry, at 2am, when the input is hostile" — coming from a seat that's trying to prove the answer wrong, not right. That's the muscle, and you clearly already have it. Thanks for reading. 🙏
Out of curiosity Great post by the way. Did you prompt Claude to help draft the skeptic architecture, or did you write all 2,000 words of anti-prompting rhetoric by hand 🤔
Better prompting doesn't replace verification; it reduces the verification chaff. The two aren't mutually exclusive, so I wouldn't dismiss either as a skill worth mastering.
Ha — fair jab, and I'll take it straight: I draft with AI all day, this post included. That's kind of the point of the post, right? The asking is free. If I'd hand-typed 2,000 words to prove typing matters I'd be arguing against my own thesis. What you're reading is the after — the part where I read it back and cut the paragraphs that were confident and wrong. The distrust is the labor; the draft was the cheap part. You caught me doing exactly the thing I'm describing. 🙂
Now the real point, because it's a good one and I don't want to strawman it: you're right that they're not mutually exclusive, and I never said stop prompting — prompting is table stakes, worth doing well. But I'd push on "reduces the verification chaff," because I think it does something sneakier than that.
Better prompting reduces the volume of things to check, yes. But it doesn't reduce them uniformly — it clears out the obvious wrong answers, the ones you'd have caught in two seconds anyway. What's left after a great prompt is a higher-concentration pile: fewer outputs, but the survivors are more polished, more authoritative, more obviously-fine at a glance. So the chaff you removed was the cheap-to-catch chaff. The wheat that remains is harder to tell from the poisoned wheat. Your hit rate on the dangerous class can actually go down even as your total error count goes down, because the dangerous class is the one better prompting disguises best.
So I'd reframe rather than dismiss: prompt well, absolutely — just don't let the smaller pile fool you into checking it less, because the smaller pile is more dangerous per item. Both skills, sure. But they scale in opposite directions, and only one of them gets easier when the model improves. Guess which one I'd spend my next hour on.
Good comment — you made me sharpen it. 🙏
Curious what practicing review actually looks like for you. Re-reading old diffs that shipped bugs?
Good question, because "get better at distrust" is uselessly abstract if I don't say what the reps actually are. Re-reading old diffs is part of it, but it's the weakest part — a bug you already know the ending of doesn't scare you anymore. The scar came from the consequence, and re-reading doesn't reproduce the consequence.
What actually trains it, for me:
Pre-mortem the diff before I accept it. Before merging anything the model wrote, I make myself finish the sentence "this breaks when ___." Out loud, specifically. Not "it might have edge cases" — "the ack fires before the row persists, so a retry at this line locks the user out." If I can't name a concrete failure, I don't trust the diff yet, I just haven't found the failure. Most of the reps are here, and they're free.
Keep a personal list of failure shapes, not bugs. Not "that time the migration broke" — the transferable shape: ack-before-persist, read-after-write across a replica, retry that isn't idempotent, the empty-array that should've been an error. Bugs don't repeat; shapes do. That list is the actual asset. Every real incident adds one line to it, and the line pays out on every future diff.
Read incident write-ups like the one I linked — other people's postmortems are borrowed scars. Cheaper than earning your own, though weaker per unit. You don't get the 2am adrenaline, but you get the shape.
The one that works best and costs the most: ship something that can actually hurt you, to a real person, and watch it. Not a toy. The muscle only grows where the consequence is real — a distrust you never had to use is a slogan, not a skill.
So: less "re-read old diffs," more "practice naming the failure before it happens, and keep a library of the shapes that have burned people." The old diffs are how you seed the library. The pre-mortem is how you use it.
What's your version — do you have a place you keep the shapes, or is it still in your head?
The line I keep coming back to is that a better prompt raises how convincing the output is, not how correct it is. That split is exactly where we got burned: the answer reads like a senior wrote it, and the tool call behind it was wrong.
The skill that survived for me is verification, not formulation. Assert on the tool path and the side effects instead of the prose, and run the same scenario a few times, because a single run of an agent test moves enough to make a 3-point score difference meaningless.
Prompting formats the request. Whether the result is true is an evaluation problem, and nobody is selling a course on that.
Yeah — you named the exact seam. "Verification, not formulation" is the whole thing in three words, and I want to push on the part you added at the end because it's the part almost nobody operationalizes: assert on the tool path and the side effects, not the prose.
That's the tell. The prose is what the model is optimized to make convincing, so grading the prose is grading the disguise. The tool call, the write that did or didn't land, the retry that fired at the wrong moment — that's where the truth lives, and it's the one layer a better prompt can't dress up. My ack-before-persist bug read flawless as prose. It only ever confessed at the side-effect layer.
And your point about running the scenario a few times is the one that separates people who actually do eval from people who say "eval": a single green run isn't a pass, it's one sample from a distribution. If a 3-point score can swing on rerun, then a single run was never evidence — it was a vibe with a number attached. Distrust that survives is statistical, not anecdotal.
The reason nobody's selling the course is the same reason it's the moat: formulation has a clean ceiling the vendor keeps raising for you, and evaluation has no ceiling and no shortcut — you earn it by shipping something that breaks a real person and carrying the scar into the next diff. One's a feature. The other's a scar. Only one of those goes on a résumé and means anything.
Verification is just distrust with assertions bolted on. Same seat, same "no" — you've just made it reproducible. That's the upgrade.
Great essay thank you!
Thank you, Mia — that genuinely means a lot. 🙏
If you take one thing from it into your week, let it be the smallest habit: after the next clean-looking AI answer, before you accept it, ask "what breaks this under a retry, or when the user does the dumb thing?" That one question is the whole essay in practice. Glad it landed with you. 🔖