I have a bookmark folder called prompting. Forty-one tabs in it.
"The 12 prompts that 10x your engineering." "Context engineering for agents, expl...
For further actions, you may consider blocking this person and/or reporting abuse
I think there is an important distinction between AI literacy and prompt engineering. Someone can write a very sophisticated prompt and still struggle to evaluate whether the result is actually useful, correct, or appropriate for the context.
The more valuable skill is understanding the domain, defining the constraints, providing the right context, and knowing how to verify the output. Prompting is a tool. Judgment is the skill behind the tool.
"Judgment is the skill behind the tool" β that's the sentence the whole post was trying to earn, and you got there in one line. Steal it back from me anytime.
The distinction you're drawing is the one I think most people miss because both things look like "being good at AI." A sophisticated prompt and sound judgment produce similar-looking artifacts on a good day β clean output, confident delivery. They only diverge on the bad day: the moment the result is plausible and wrong. That's the exact instant a great prompt does nothing for you and domain judgment does everything. The prompt got you a beautiful answer; only the judgment can tell you it's going to break under a retry at 2am.
And I'd add one teeth to your list β "knowing how to verify the output" is doing a lot of quiet work there. Verification isn't just ability, it's willingness to say no to something clean. Plenty of people can spot the flaw and still ship it because the diff looked senior and the deadline was real. The judgment includes the spine. That part's even harder to teach than the domain knowledge, which is maybe why nobody sells it.
You basically wrote the thesis in a shorter, better form than I did. π
Exactly. I think thatβs the part that gets underestimated.
AI can make a wrong answer look incredibly convincing, so verification isnβt just about finding errors. Itβs also being willing to stop and say, βNo, this looks good, but it doesnβt hold up.β
Thatβs where judgment becomes much more than a prompt skill. The prompt helps you produce the output. Judgment decides whether that output deserves to move forward. π―
"Judgment decides whether that output deserves to move forward" β that's the cleanest one-line spec for the whole thing I've seen. Producing the output and authorizing it are two different acts, and the entire industry keeps collapsing them into one because the prompt makes them feel like one continuous motion: ask, receive, ship. You've split them back apart, and the split is where all the value hides.
The one thing I'd add, because you're circling it: the convincing-ness actively fights the verification. It's not neutral. The more polished the output, the more your own instinct to wave it through β so the exact quality that makes an answer look ready is the quality that makes saying "no, it doesn't hold up" hardest. Verification isn't just work; it's work performed against your own pull toward the clean-looking thing. That's why it's rare. Finding the flaw takes skill. Overruling a beautiful diff that everyone's ready to merge takes something closer to nerve.
Which is exactly your point β judgment is more than a prompt skill because it operates in the one moment the prompt can't reach: after the answer looks done, when the only thing standing between "convincing" and "shipped" is a person willing to be the friction. The prompt gets you to the door. Judgment is what decides whether the door opens.
Genuinely one of the sharpest exchanges I've had in a comment section β you kept handing me tighter versions of my own argument. ππ―
Another angle I find interesting is that AI itself shouldn't become the system. AI can generate code, suggest solutions, analyze data, or produce content. But a real workflow still needs decisions about where AI is used, what happens before it, how the output is tested, and what happens when the output is wrong.
So I would think of it as:
System β AI β Output β Verification β Decision
rather than:
Prompt β AI β Done
The interesting engineering challenge is designing the whole system around AI, not just learning how to talk to the model.
This is exactly the level-up, and I think you've named the thing the "learn to prompt" framing quietly hides: the model is a component, not the architecture. The second you write Prompt β AI β Done, you've smuggled in the assumption that the output is the deliverable. It isn't. The deliverable is a decision you're willing to stand behind β and the AI touches only one box in that chain.
What I'd push on in your diagram: the interesting failures don't live in the AI box, they live in the arrows. AI β Output is easy. It's Output β Verification where everything is actually won or lost β because verification isn't a step you add, it's a seat you have to staff with something other than the author. The thing that generated the output is structurally the worst judge of it; it's proud of it. So the arrow only works if the skeptic is a different seat than the generator. That's the whole reason I build with author β skeptic β human instead of one confident box.
And your last line is the part I'd underline twice: "designing the whole system around AI, not just learning how to talk to the model." That's the migration nobody's selling a course on β because you can't grind it in a weekend. Prompting is table stakes inside the AI box. The engineering is everything wrapped around it. You get it.
Exactly. That βauthor β skeptic β humanβ distinction is important.
Iβd even say verification is part of the architecture, not just a final checkpoint. If the same system generates, evaluates, and approves its own output, weβve basically built a loop with no real external constraint. AI can accelerate the work, but the system still needs a mechanism that can challenge the output before a human has to stand behind the decision.
Thatβs where I think the real engineering starts. π―
"A loop with no real external constraint" β that's the failure mode stated with more precision than I managed in the whole post. A system that generates, evaluates, and approves its own output isn't verified; it's self-certifying, and self-certification is just confidence with extra steps. The evaluation feels like a check but it inherits the exact blind spots of the thing being checked, because it is the thing being checked wearing a second hat. You haven't added a constraint. You've added a mirror and called it a jury.
And your reframe β verification is part of the architecture, not a final checkpoint β is the load-bearing correction. The moment you treat it as a step at the end, you've already lost, because a step can be skipped, softened, or rushed under a deadline, and it sits downstream of all the momentum ("it's basically done, just needs a review"). Baked into the architecture, the skeptic isn't a gate the work flows toward β it's a structural pressure the work is subjected to by construction, with no path to merge that routes around it. One is a speed bump you can floor over. The other is load-bearing wall.
The word doing the real work in your comment is external. That's the whole game. The constraint has to come from outside the generator's own frame β a different seat, a different objective, ideally something that can't be talked into it at all (a deterministic check for the invariants you can name, a skeptic prompted to refute for the fuzzy blast-radius you can't). The instant the challenge originates inside the same system that produced the output, it's not a constraint, it's a preference. External-ness is the property; everything else is implementation.
"That's where the real engineering starts" β exactly. Prompting is the easy box. The architecture around the box β where the external constraint lives, how it's structured so it can't be bypassed, who ultimately stands behind the decision β that's the part that's actually hard, actually compounds, and actually can't be bought as a weekend course. You've basically written the sequel to the post in three replies. π―
Agree with the review-time point β and one failure mode from practice that sharpens it: if the model can see the tests before writing the code, it optimizes the oracle, not the behavior. I now write the tests against the requirement, keep them out of the model's context entirely, and run them after β otherwise you get exactly the "passes every test you thought to write, still wrong" class, because the prompt said the same thing the tests say.
Second: for anything that mutates state (file writes, migrations, deletes), I stopped using a second model call as the reviewer. Two confident reviewers just give you two confident wrong answers that agree with each other. A dumb deterministic post-check script β did the file end up valid, does the migration roll back, is the row count sane β is harder to impress but it actually refuses.
The judgment residue is real though: someone still has to hold the requirement in their head and say no to a clean diff. That's the 2am part, and no prompt folder covers it.
Both of these are load-bearing and I want to sit on each for a second because you've operationalized the thing the post left as a slogan.
The oracle point is the sharpest version of "passes every test you thought to write." If the model can see the tests, correctness collapses into test-satisfaction β and those were never the same thing, you just stopped being able to tell them apart. The prompt and the tests saying the same words is the tell: you've built a closed loop that grades its own homework. Keeping the tests out of context and running them after is the difference between an oracle and a mirror. I'm stealing "it optimizes the oracle, not the behavior."
The second point is where I'll push hardest in agreement: a second model call as reviewer is the trap I most want people to see. It feels like distrust β separate seat, second opinion β but it's the same failure mode wearing a lanyard. Two generators trained to be plausible will agree plausibly, and now you've got consensus as a disguise instead of a check. Your framing is exactly right: the dumb deterministic script is "harder to impress but it actually refuses." That's the whole property. Refusal has to come from something that can't be talked into it β and a model, by construction, can always be talked into it. The post-check doesn't have a good day and a bad day. It just has a correct answer.
Where I'd nuance it: the deterministic check and the skeptic-model aren't rivals, they're different ranges. The script nails the invariants you can name β row counts, rollback, valid-on-disk. The model-skeptic is for the fuzzy blast-radius stuff you can't reduce to an assertion ("does this ack before it persists?"). So I'd run the script as the hard gate and the skeptic as a hint generator that never gets to bless anything. Neither one is the human.
Which is your last paragraph, and it's the one I have no cleaner version of: the residue is a person holding the requirement in their head and saying no to a clean diff. Everything above just narrows how often they have to β it never removes the seat. That's the 2am part, and you're right, no folder covers it. Great comment.
Exactly - "hint generator that never gets to bless anything" is the phrase I wish I had coined. One refinement from practice: the skeptic model only earns its keep if its output is consumed as information by a human or by the script gate, never as authority. The failure mode I have seen is gradations of blessing: the script says no, the skeptic says probably fine, and someone merges anyway. So make the skeptic output a separate artifact with zero merge privileges - if reading it changes a decision, that decision was human. Which is exactly your seat that never goes away.
"Gradations of blessing" is the failure mode I've watched sink more review setups than any outright bug β and you've named the mechanism exactly. The moment the skeptic can say "probably fine," you've built a second author: something that emits a soft verdict. Now the human isn't deciding, they're adjudicating between two opinions and picking the one that lets them go home. The script's hard "no" gets laundered into a maybe by a more agreeable voice, and the merge happens on the maybe. You didn't add a check β you added a lawyer for the defense.
Your fix is the right shape, and I'd make the rule as sharp as you did: the skeptic may describe, never conclude. Its output is "here's what breaks under a retry / here's the invariant this touches" β evidence with no verdict field at all. The instant it can emit "probably fine," you've handed it a slice of the merge decision, and authority abhors a vacuum: whatever can express confidence eventually gets treated as if it earned it. So you don't soften its authority β you delete the surface it could express authority on. No verdict to defer to means nothing to defer to.
And your last line is the whole thing: "if reading it changes a decision, that decision was human." Cleanest test I've heard for whether a check is doing its job or quietly usurping it. The skeptic's entire value is that it moves information into the human's head β surfaces the retry case they'd have missed β then gets out of the way before the verdict. The second its output can be the verdict, it stops informing the seat and starts replacing it, and you're back to an author blessing its own genre of work.
Information in, authority never. The gate refuses; the skeptic testifies; the human decides. Three seats β and only one of them is allowed to say yes.
This thread's been the best part of writing the piece. π€
This is a compelling reframing of AI fluency: the real differentiator is increasingly judgment, not the ability to formulate increasingly elaborate prompts. I especially liked the distinction between convincing output and correct output, because polished AI-generated work can actually make errors harder to notice. The author/skeptic separation is also a powerful architectural ideaβespecially when the same system that creates an answer should not be trusted to be its final judge. βAsking is free; deciding whether the answer survives contact with reality is the entire jobβ captures the central argument beautifully.
Thank you, that's a really generous read, and honestly you've summed it up tighter than parts of the post did.
One thing I'd add on the author/skeptic split, since you zeroed in on it: the trap is assuming a second seat is enough. If the skeptic is just the same model with a "be critical" prompt, it shares the author's blind spot and cheerfully re-derives the same mistake with a frown on. The separation only pays off when the doubt comes from somewhere with a genuinely different prior. A different model, an adversarial frame, a test that actually runs, or a human who's been burned before. Independence is the real ingredient. A second opinion from the same mind is just the first opinion again.
Prompting can improve the output, but it doesnβt replace understanding the problem, validating the result, or knowing what βgoodβ actually looks like. The real advantage comes from combining AI with strong fundamentals and judgment.
Exactly. Though I'd push on one word in there: validating. The trap is that validation feels done the moment the output looks right, and a good prompt makes it look right faster. The answers that actually cost me weren't the ones I failed to validate. They were the ones that sailed through validation because they were clean, confident, and matched what I asked for.
So "knowing what good looks like" is necessary but not quite enough on its own, because a convincing wrong answer also looks like good. The part that saved me is narrower: treating the clean answer as guilty until I've tried to break it, and making sure that doubt comes from a different seat than the thing that wrote it. Fundamentals tell you what good looks like. Distrust is what you do when good-looking turns out not to be the same as good.
I agree with the distinction between prompting and verification. A good prompt can improve the quality of the output, but it doesn't guarantee correctness. The ability to question, test, and challenge AI-generated code is becoming just as important as knowing how to generate it.
Exactly β and I'd push your own point one degree further, because you've landed on the thing most people stop just short of. You said questioning and testing AI code is becoming as important as generating it. I'd argue it quietly became more important, and here's the tell: generating gets cheaper every model release, and verifying doesn't. The vendor ships you a better generator next quarter, for free, while you sleep. Nobody ships you better judgment. So the two skills aren't just both important β they're moving in opposite directions, and the gap widens with every update.
The one nuance I'd add to "test": the tests you think to write come from the same mental model that accepted the code, so they tend to confirm it rather than attack it. The verification that actually catches things is the adversarial kind β "what breaks this under a retry, at 2am, when the input is hostile" β coming from a seat that's trying to prove the answer wrong, not right. That's the muscle, and you clearly already have it. Thanks for reading. π
Out of curiosity Great post by the way. Did you prompt Claude to help draft the skeptic architecture, or did you write all 2,000 words of anti-prompting rhetoric by hand π€
Better prompting doesn't replace verification; it reduces the verification chaff. The two aren't mutually exclusive, so I wouldn't dismiss either as a skill worth mastering.
Ha β fair jab, and I'll take it straight: I draft with AI all day, this post included. That's kind of the point of the post, right? The asking is free. If I'd hand-typed 2,000 words to prove typing matters I'd be arguing against my own thesis. What you're reading is the after β the part where I read it back and cut the paragraphs that were confident and wrong. The distrust is the labor; the draft was the cheap part. You caught me doing exactly the thing I'm describing. π
Now the real point, because it's a good one and I don't want to strawman it: you're right that they're not mutually exclusive, and I never said stop prompting β prompting is table stakes, worth doing well. But I'd push on "reduces the verification chaff," because I think it does something sneakier than that.
Better prompting reduces the volume of things to check, yes. But it doesn't reduce them uniformly β it clears out the obvious wrong answers, the ones you'd have caught in two seconds anyway. What's left after a great prompt is a higher-concentration pile: fewer outputs, but the survivors are more polished, more authoritative, more obviously-fine at a glance. So the chaff you removed was the cheap-to-catch chaff. The wheat that remains is harder to tell from the poisoned wheat. Your hit rate on the dangerous class can actually go down even as your total error count goes down, because the dangerous class is the one better prompting disguises best.
So I'd reframe rather than dismiss: prompt well, absolutely β just don't let the smaller pile fool you into checking it less, because the smaller pile is more dangerous per item. Both skills, sure. But they scale in opposite directions, and only one of them gets easier when the model improves. Guess which one I'd spend my next hour on.
Good comment β you made me sharpen it. π
Curious what practicing review actually looks like for you. Re-reading old diffs that shipped bugs?
Good question, because "get better at distrust" is uselessly abstract if I don't say what the reps actually are. Re-reading old diffs is part of it, but it's the weakest part β a bug you already know the ending of doesn't scare you anymore. The scar came from the consequence, and re-reading doesn't reproduce the consequence.
What actually trains it, for me:
Pre-mortem the diff before I accept it. Before merging anything the model wrote, I make myself finish the sentence "this breaks when ___." Out loud, specifically. Not "it might have edge cases" β "the ack fires before the row persists, so a retry at this line locks the user out." If I can't name a concrete failure, I don't trust the diff yet, I just haven't found the failure. Most of the reps are here, and they're free.
Keep a personal list of failure shapes, not bugs. Not "that time the migration broke" β the transferable shape: ack-before-persist, read-after-write across a replica, retry that isn't idempotent, the empty-array that should've been an error. Bugs don't repeat; shapes do. That list is the actual asset. Every real incident adds one line to it, and the line pays out on every future diff.
Read incident write-ups like the one I linked β other people's postmortems are borrowed scars. Cheaper than earning your own, though weaker per unit. You don't get the 2am adrenaline, but you get the shape.
The one that works best and costs the most: ship something that can actually hurt you, to a real person, and watch it. Not a toy. The muscle only grows where the consequence is real β a distrust you never had to use is a slogan, not a skill.
So: less "re-read old diffs," more "practice naming the failure before it happens, and keep a library of the shapes that have burned people." The old diffs are how you seed the library. The pre-mortem is how you use it.
What's your version β do you have a place you keep the shapes, or is it still in your head?
The line I keep coming back to is that a better prompt raises how convincing the output is, not how correct it is. That split is exactly where we got burned: the answer reads like a senior wrote it, and the tool call behind it was wrong.
The skill that survived for me is verification, not formulation. Assert on the tool path and the side effects instead of the prose, and run the same scenario a few times, because a single run of an agent test moves enough to make a 3-point score difference meaningless.
Prompting formats the request. Whether the result is true is an evaluation problem, and nobody is selling a course on that.
Yeah β you named the exact seam. "Verification, not formulation" is the whole thing in three words, and I want to push on the part you added at the end because it's the part almost nobody operationalizes: assert on the tool path and the side effects, not the prose.
That's the tell. The prose is what the model is optimized to make convincing, so grading the prose is grading the disguise. The tool call, the write that did or didn't land, the retry that fired at the wrong moment β that's where the truth lives, and it's the one layer a better prompt can't dress up. My ack-before-persist bug read flawless as prose. It only ever confessed at the side-effect layer.
And your point about running the scenario a few times is the one that separates people who actually do eval from people who say "eval": a single green run isn't a pass, it's one sample from a distribution. If a 3-point score can swing on rerun, then a single run was never evidence β it was a vibe with a number attached. Distrust that survives is statistical, not anecdotal.
The reason nobody's selling the course is the same reason it's the moat: formulation has a clean ceiling the vendor keeps raising for you, and evaluation has no ceiling and no shortcut β you earn it by shipping something that breaks a real person and carrying the scar into the next diff. One's a feature. The other's a scar. Only one of those goes on a rΓ©sumΓ© and means anything.
Verification is just distrust with assertions bolted on. Same seat, same "no" β you've just made it reproducible. That's the upgrade.
Great essay thank you!
Thank you, Mia β that genuinely means a lot. π
If you take one thing from it into your week, let it be the smallest habit: after the next clean-looking AI answer, before you accept it, ask "what breaks this under a retry, or when the user does the dumb thing?" That one question is the whole essay in practice. Glad it landed with you. π
Exactly! I think thatβs the part that gets underestimated.
Underestimated is exactly the word β and I think I know why: it's invisible on every metric that looks good. Generating shows up as velocity, tickets closed, features shipped β all the numbers that point up and to the right. Verifying only shows up as absence β the outage that didn't happen, the customer who never got locked out. You can't screenshot a disaster you prevented. So the skill that matters most is the one that leaves no evidence when it's working, which is exactly how it ends up underestimated. Appreciate you saying it out loud. π