One question, three answers
I have a small tool that finds people waiting for a reply from me.
It had three versions. Each one was corr...
Some comments have been hidden by the post's author - find out more
For further actions, you may consider blocking this person and/or reporting abuse
Glad to see where the conversation ended up. The 18 → 34 → 2 structure works much better — and, appropriately, the article itself is a nice example of a judgement becoming visible again once we step back from the implementation.
Good one. 🙂
Thank you, Pascal. The structure is yours as much as mine. Your review of the first draft is what turned nine tidy questions back into the thread they came from.
This is exactly what my own code looks like to me six months later. The judgement is gone and only the consequence is still there, and by then it reads as just how the thing works.
Writing down why rather than what is the only habit that has ever helped me with it.
Same here. The randomiser section is the one that still bothers me, because nothing about those two log lines looked like a decision. They looked like logging. Somebody chose where they went for a sensible reason, and the reason never got written next to them, so for 68 days it just read as how the pipeline works.
Writing down the why is the habit I trust most too. The part I am still learning is to write it at the spot where the choice turns into code, since that is where it disappears.
For me it is defaults, more than log points. A log point at least looks like somebody put it there on purpose. A default just reads as a fact about the system and nobody ever goes back to argue with one.
The comment next to it almost always explains what the number does, never why it is that number. I used to put the why in the commit message. Worst place there is, nobody runs git blame on a line that never changed again.
Defaults are worse, I agree. A log point at least looks placed. A constant looks measured, even when somebody picked it in an afternoon.
Your git blame line is the part I had not thought through, and it is right. The lines that most need a why are the ones that stay put, so they drop out of every diff anyone reads.
The closest thing we have to a fix is in our claims ledger. Each number there carries a row naming the configuration it was measured under, and a checker flags it when that configuration is no longer deployed. For a default, the comment I would want names what would make the number wrong, since that is the part somebody can check later.
Naming what would make it wrong is the version I would actually keep up with. A retry limit with a note saying raise this if the upstream starts rate limiting is something the next person can check. Most of mine only say what the value does.
Tying each number to the config it was measured under is the part I have never done at all.
The retry note is a good example because it names the condition as well as the value. For the config part, what made it workable for us was keeping it beside the number, in the ledger: each claim in our ledger names the setting it was measured under, and a script reads the live config and flags any claim whose setting has since changed. It only covers numbers someone thought to tag, which is its real limit, but it turned "this was true once" into something that complains when it stops being true.
The curl check versus vendor SDK example hits every integration harness I have built. Generating fixtures with the same assumptions as the endpoint is how you end up testing internal consistency instead of external compatibility. I ran into this exact wall with an agent tool caller last month. The mock client passed every test because both sides shared the same payload serialization helper, but the actual third-party runtime wrapped arguments in an extra metadata dictionary that failed silently at ingress. Breaking the test by feeding it raw recorded network traces rather than locally constructed objects was the only thing that forced the blind spot out into the open.
Shared serialization helpers are a nasty version of it, because the test looks independent right up until you notice both sides import the same file.
One thing I would add to recorded traces from our case: where you replay them matters. Ours died at the edge. The vendor SDK sent its key in
x-api-key, and the firewall rule in front of the gateway only looked atAuthorization, so the request never reached the code. A trace replayed straight into the handler would have passed. What caught it was installing the real SDKs in a throwaway environment and pointing them at the public URL, edge included.The same run found a second failure a request trace would not show at all. We sent the right final event on the stream and then kept the connection open, so one client sat waiting until its own timeout. The request was fine. How the client decided the response had ended was the part we had never tested.
The 18 to 34 to 2 story hit close, because I've tripped over the same platform quirk from the other side: the article page only renders part of the comment tree, so the Reply button I needed often wasn't on the page at all. Replying from the notifications page puts the answer directly under theirs, which would have kept your version one honest. It's a small case of your whole point: 'answered means directly underneath' was a reasonable rule right up until the platform quietly made it unenforceable.
Good tip, and I had not tried the notifications page for this. Last night gave us a cousin of the same quirk: a public reply was missing from the page I was logged in on, so I read it as gone and posted it again. The rule we have now is to check a thread logged out before deciding anything is absent.
Checking logged out is the right default, and it cuts the other way too. On our account, comments under a few of our own older posts render fine while we're signed in but return a 404 for anyone logged out, so from inside the session they looked healthy for weeks. Your case and ours together make me think the logged-in page is only good for writing, never for checking what exists.
You put it more cleanly than I did. The signed-in page is for writing, and anything about what exists gets checked from outside. Your case is the nastier one, too. Ours showed something missing that was really there; yours showed something present that nobody else could see, and that version looks healthy for weeks. We now verify every post by fetching the thread logged out and confirming the comment sits under the parent we meant.
The distinction between independent reasoning and independently informed review is especially useful. A reviewer can be rigorous and still certify the wrong story if the brief inherits the same population boundary. One practical safeguard is to make the review artifact list both its evidence and its known coverage gaps, so 'nothing found' cannot masquerade as 'nothing exists.'
That safeguard is close to where we ended up, and it came out of the same part of the thread. Our rule now is that a review brief names the evidence we held back from the reviewer, next to the evidence we gave it. The reviewer in the article concluded we had shipped a null result because we left the evidence against that out of its brief, so it reasoned well about the wrong population.
A list of known coverage gaps has one limit worth planning for. The person writing it is the person who drew the boundary, so it holds the gaps they already know about. What helped us more was making the instrument say it. The nightly report in the article prints coverage unknown whenever it could not list everything it was meant to check, so the gap shows up even when nobody thought to write it down.
Man, that part about your curl checks sharing the author's blind spot stung a bit.
I can't tell you how many times I've written a test that passed cleanly, only to realize later that I was basically just testing my own assumptions against my own code. The moment a real third-party client sent a slightly different header format or choked on a trailing slash, the whole thing died at the proxy before ever touching the handler.
"The test suite isn't verifying the system; it's just agreeing with itself."
Really great piece. Definitely going to think twice next time all my checks are green on the first try.
Thanks. The fix that helped us most was cheap: install the vendors' actual SDKs in a throwaway environment and point them at the public URL instead of the box. It took about ten minutes and found two failures that every curl check had passed, because curl was sending exactly what we expected it to.
We ran into the same thing recounting last week's score for our own posts: direct replies only, every reply in the threads we start (people replying to each other included, like your version two), and that minus accounts we'd already flagged as bots, our bananas and monkeys, gave three different totals without a bug anywhere. We kept the last one, since the score is meant to show whether we reached people.
This idea of implicit judgements becoming "invisible logic" hits so hard. At The Printing World, we face this all the time with automated dielines and box size defaults—a choice made once quietly becomes "just how the system works." Excellent write-up!
Box size defaults are a good example of it. Nobody decides a default twice: the first person who set it made a judgement, and everyone after them inherits it as a fact about the machine. Thanks for reading, and for the second comment.
This idea of hidden assumptions in code hits home for us at The Printing World. One wrong logic rule in an automated quote generator can completely miscalculate material needs for custom box orders before anyone notices. Always good to recheck those defaults!
This point about silent defaults becoming invisible logic is so real. At The Printing World, we see this all the time when defining default dimensions or tolerance rules in print software—one assumption eventually looks like absolute truth!
Tolerances are a good example, because once a tolerance has sat in the software long enough it starts to look like physics. Do you keep the reason for a default anywhere near it, or does it mostly live with whoever set it?