The last ten bugs I fixed, I fixed the same way. Stack trace in. Fix out. Apply, run, green, move on. Each one took minutes. It felt like getting b...
For further actions, you may consider blocking this person and/or reporting abuse
My system has always been prompt with an assumption. You usually have an idea of what went wrong, so test that theory, eg. 'OOM - look for memory leaks in the last API call you made', it understands OOM resolved is the goal, your first instinct is it's entry point. Doesnt mean you're still practicing looking at an async loop and properly discerning what should be threadpooled, but you know where the bug is and you know AI can fix it with that guidance.
A backup method that works surprisingly well, is ask AI to write up a mermaid diagram of the architecture for the feature bugging. Then ask it 'in this architecture, we're experiencing this bug... Inspect the hot sites where this is most likely caused'. LLMs are literally glorified pattern matchers, so they are shockingly well suited at digging with instruction.
Both of these are great, and what I like most is that you've already spotted the seam yourself in the first one — so let me build on it rather than just nod.
On "prompt with an assumption": this is "guess before you paste" done in the wild, and done well — you're not asking the AI what broke, you're handing it a suspect and a location ("OOM → memory leak in the last API call") and letting it do the fast recall work of confirming and patching. That's exactly the right division of labor: you form the hypothesis, it executes the search. But — and you said this yourself, which is why I trust the rest of your comment — "doesn't mean you're still practicing looking at an async loop and discerning what should be threadpooled." That caveat is the whole post in one clause. You kept the where-generating muscle (you knew it was the entry point). What quietly stops getting reps is the deeper why — the "should this even be async here" judgment that you only build by sitting in it. So your system keeps the top of the loop alive and lets the bottom soften. Worth knowing which half you're preserving, and you clearly already do.
On the mermaid-diagram backup: this one's genuinely clever and I've had it work too, but notice why it works, because it proves the thesis instead of dodging it. "Glorified pattern matcher" is right — and a pattern matcher pointed at "hot sites where this bug is most likely caused" is doing recall over the architecture, ranking candidate locations by prior. That's real value. But look at what you still supplied: the bug description, the symptom, and the judgment to say "inspect the hot sites" instead of "fix it." You're still forming the hypothesis-space; the AI is just ranking suspects inside a space you defined. The moment the true cause is not a hot site — the low-prior, three-layers-away race that nobody's pattern-matcher flags because it's rare by construction — the diagram trick goes quiet, and you're back to the hand loop. Which is the exact 5% the post is about. Your technique is a fantastic force-multiplier on top of a hypothesis; it just can't be the thing that generates one.
Net: you've basically built the healthy version of the workflow — human forms the suspect, AI does the fast search — and you're honest about the one muscle it doesn't train. That's more than most people in this thread. Steal-worthy, the mermaid one especially. 🙏
and you get a nice little diagram from it, so you can refresh yourself on the architecture at a glance. Sometimes, the bug is because a node in the diagram isnt fully fleshed out yet. So sometimes it's more worth continuing than fixing.
That's a genuinely nice second-order payoff, and the last bit is sneakily deep: "sometimes the bug is because a node isn't fully fleshed out yet." That's the difference between a defect and a gap — the code doesn't do the thing because the thing was never actually built, not because it was built wrong. And those two want opposite responses. A defect you patch; a gap you keep building. Reaching for a fix on a gap is how you get a pile of special-case patches spackled over a hole that just needed the node finished — the classic "why is this function 400 lines of edge cases" that's really an unbuilt abstraction wearing bug-fixes as a disguise.
And notice: deciding "this is unfinished, not broken" is itself a judgment call the AI won't make for you. Point a pattern-matcher at a gap and it'll cheerfully generate a plausible patch — it has no concept of "this was supposed to be more than it is," because completeness lives in your intent, not in the code. So the diagram's real gift isn't the hot-site ranking; it's that seeing the architecture at a glance lets you make the continue-vs-fix call that no amount of prompting surfaces on its own.
You've turned the debugging tool into an architecture-review tool, which is the healthiest possible drift. The bug becomes a prompt to look at the whole shape, not just the broken line. That's the muscle staying alive right there. 🙏
Nine years in and the bug with no stack trace is still the one that finds out whether I understand the system or only its error messages.
Knowing where to look first is the whole skill, and nothing trains it except losing it for a while and noticing.
▎ Nine years in and the bug with no stack trace is still the one that finds out whether I understand the system or only its error messages.
That line — "understand the system or only its error messages" — is the cleaner version of the whole post. Thank you for it.
And you've named the thing I only circled: the skill can't be taught directly, only recovered. You lose it, you notice the hole, and the noticing is the lesson. Which is a little terrifying, because the noticing only happens when a real one lands — there's no drill for it, no sandbox. The empty log file is the exam and the only place the material was ever covered.
The "guess before you paste" habit is my hack around that: a way to keep loading the muscle on the cheap bugs so the expensive one isn't the first rep in a year. But you're right that it's a substitute for the real trainer, which is loss.
Nine years and it still finds you out — that's the honest part most people won't say out loud. It never stops being the bug that checks whether you actually know the system. Appreciate you dropping this. 🙏
Honestly, mine is still ongoing 😂
It stems from vibecoding a WSL2 system and trying to integrate a local AI through Ollama. The problem isn't really that it doesn't work, it's that it doesn't consistently do what I tell it to do.
I'll give it specific code instructions, it'll follow some of them, break the output, ignore another instruction, or just give me something that makes no sense. Then Claude starts guessing at what's wrong, and I have to keep telling it, “stop guessing, let's actually figure out what's happening.”
I've made some progress. I've even had to use Python as a sort of control layer to take some of the workload off the local model or prompt it into doing specific things. But even then, it doesn't always execute exactly what I intended.
So I ended up changing the scope. For now, I'm mostly using the local AI for simple file editing and straightforward questions rather than letting it run the whole workflow.
I haven't had the breakthrough yet, but I'm still working on it.
And honestly, this is probably the most interesting debugging I've done recently because there isn't really a clean error message telling me where the problem is. I actually have to figure out which part of the system is failing instead of just asking AI to fix whatever looks wrong.
This is the exact bug the post is about, and you even said the sentence yourself: "stop guessing, let's actually figure out what's happening." That instruction — the one you keep having to give Claude — is the hypothesis-forming muscle running out loud. You're doing the rep. That's why this feels like the most interesting debugging you've done in a while: nobody handed you the suspect.
And notice what kind of bug it is. It's not "it's broken," it's "it's inconsistent" — follows some instructions, drops others, sometimes produces nonsense. That's the ack-before-persist shape all over again: the symptom is miles from the cause, and there's no red line to paste because nothing technically crashed. A non-deterministic system doing a plausible-but-wrong thing is the purest version of the hard 5%.
One thought from the trenches, since you asked for the mechanism and not the fix: the Python control layer you added is the right instinct, and it maps directly to the "author vs. skeptic" split from the post. Right now the local model is both writing and deciding it's done — same failure as letting the fixer certify its own fix. The more you can make Python the skeptic — assert the output actually matches the instruction, reject and re-prompt on mismatch instead of trusting the first pass — the less you're relying on the model to be consistent, and the more you're forcing consistency around it. Narrowing scope to file edits and simple Q&A is the same move: you shrank the surface until the failure was observable. That's not giving up, that's isolating the variable.
No clean error message, figuring out which part is failing instead of asking AI to patch whatever looks wrong — that's the whole craft. You're not stuck. You're training. Keep the log of this one. 🔧
The “guess before you paste” rule is probably the most interesting part of this.
I think there’s another thing that gets lost when AI handles the easy debugging loop: not just the hypothesis itself, but the trail of hypotheses that were rejected along the way.
The final fix tells you what worked. But when a similar bug appears months later, knowing what was considered, what was ruled out, and why can be just as valuable as the fix itself.
Maybe preserving that reasoning trail is part of keeping the debugging skill alive, rather than just preserving the solution.
This is a genuinely new angle and I think you've found the load-bearing thing I left implicit. The fix is the destination. The rejected hypotheses are the map — and the map is what you actually reuse.
Here's why it bites harder than it first looks: when you rule something out, you're not just crossing off a suspect, you're encoding a fact about how the system actually behaves under stress — "it wasn't the cache, because the value was already correct on the second read, which means the write did land." That negative result is a load-bearing piece of system knowledge, and it's invisible in the diff. The final patch says await moved two lines up. It says nothing about the four theories you killed to know that was the fix — and those four are exactly what you need when the next weird one shows up in the same subsystem.
And this is precisely what the paste-the-symptom loop destroys. When AI hands you the answer, you don't just skip forming the winning hypothesis — you skip the whole search tree. You never generated the four wrong ones, so you never learned the four facts that ruling them out would have taught you. The reasoning trail doesn't get lost; it never gets built. You get the fix with none of the map.
So I'd fold your point straight into the "keep a log of the far ones" habit and sharpen it: don't log what fixed it. Log what you suspected and why you were wrong. "Thought it was the cache — wasn't, because X. Thought it was a race on the read — wasn't, because Y." That trail is the thing that compounds; the fix is disposable the moment it's merged. You've basically named the difference between preserving a solution and preserving the ability to solve — and only one of those survives a similar bug six months out.
Might steal "preserving the reasoning trail, not just the solution" for the next piece, with credit. Great comment.
Answering your honest question: a prod-only duplicated-order bug. The checkout occasionally wrote two rows for one order — only under load, only in prod. Every log line showed a single request id and reported success. No exception, nothing to paste; the AI suggestions were generic "add idempotency" advice.
What cracked it was running the falsify-narrow loop by hand, exactly the rep you describe. Cheapest falsifiable theory first: "if a retry re-enters the handler, what state would prove it did?" That led to the actual cause: the client timeout budget (3s) was shorter than the worst-case commit path (p99 4.2s), so the client retried while request one was mid-commit, and the dedupe check ran before the first write was visible — both attempts saw "no order yet" and proceeded. Symptom three layers from cause; no stack trace would ever have pointed there.
The transferable bit: writing the suspect down before touching anything turns "what could it be" into a cheap binary test — "would theory X be disproven by observation Y?" — which is what keeps the loop moving when there's no error to steer by.
And it sharpens your dashboards point: the metric that would have caught mine was client timeout vs commit p99 — a ratio, not a rate. No dashboard had it, so everything stayed green while the bug shipped.
This is the one. The pinned answer to the honest question, and it does the thing the whole post was betting someone would: it's a bug three layers from its symptom, with nothing to paste, cracked by the exact loop AI can't run for you. Timeout budget shorter than the p99 commit path, so the client's own retry races the dedupe check before the first write is visible — both attempts see "no order yet" and proceed. That's not a bug you read. It's a bug you hypothesize, and the generic "add idempotency" the AI offered is the tell: it pattern-matched the shape ("duplicate writes → idempotency") without ever forming the mechanism, so it prescribed the treatment for the wrong disease. Idempotency is downstream of why two requests existed, and only the loop gets you upstream of that.
The transferable bit you pulled out is the part I'd underline for anyone reading: writing the suspect down before you touch anything turns "what could it be" into a binary test — "would theory X be disproven by observation Y?" That reframe is the whole engine. An un-written hypothesis is a vibe you wander around inside; a written one has an edge you can push against, and the edge is what keeps the loop moving when there's no exception to steer by. Vague dread → falsifiable claim is the single conversion that separates debugging from staring. You said it cleaner than I did.
And your closing point is the sharpest thing in this thread, so I'm going to sit on it: the metric that would've caught yours was a ratio, not a rate — client timeout vs commit p99. That reframes my dashboards section entirely. It's not just that the hard bugs are rare and invisible; it's that the signal for them lives in relationships between numbers nobody thought to divide, while every dashboard ships the raw counts that stay reassuringly green. The rate said "healthy." The ratio — had anyone plotted it — said "you are one latency spike from double-charging people." Rates are what you watch; ratios are where the bugs actually live, and almost nobody instruments the second kind until after the incident writes the query for them.
Genuinely one of the best debugging war stories I've gotten on any of these. Mind if this gets the pin?