DEV Community

Cover image for Implementation is where judgements go to become invisible
Tom Jones
Tom Jones

Posted on

Implementation is where judgements go to become invisible

How simple code hides past design choices

One question, three answers

I have a small tool that finds people waiting for a reply from me.

It had three versions. Each one was correct. Each one gave a different answer.

Three versions of one instrument: 18 waiting, 34 waiting, 2 waiting

  1. Version one said 18. It counted a reply only if it sat directly under theirs. But dev.to sometimes will not show a comment its API still returns, so you reply beside it instead. Nine of ten had been answered that way.
  2. Version two said 34. It counted any later comment from me in the thread. So it counted two other people talking to each other.
  3. Version three said 2. It used both rules. One was a friendly sign-off from July. The other was Pascal, suggesting there might be an article hiding in these comments.

This is that article.

Nothing in the tool was broken. Each count was right about its own population. The trouble was the question. Every version reported its number as the answer to "who is waiting?", and every version had quietly decided what "answered" means.

The first version was built with care. It encoded a sensible judgement. Then the judgement stopped looking like one.

That is the whole article in one incident. The rest is how we got there.

How we got here

It started in July, in the comments under Pascal's article about replacing Calendly. It ran for two months, in bursts, with pauses while one of us went off and tested things in production.

We did not start with a method. It went like this, over and over:

1. one of us has an answer that looks finished
2. the other brings back a real failure
3. the answer turns out to be hiding a judgement
Enter fullscreen mode Exit fullscreen mode

Sometimes Pascal had the answer and I found the hole. Sometimes the other way round. Here it is in the order it came up.

The test passes

The first example is from Pascal's article. Two people try to book the same slot at the same moment. On Postgres that is a real race. On SQLite it cannot happen, because SQLite lets only one writer in at a time.

So a test for that race, run on SQLite, passes forever. The bug ships on Postgres.

The same race test on SQLite passes because the race cannot happen there; on Postgres the bug ships

The hidden judgement: that this test can fail at all.

The fix is mechanical. After you write a test, break the thing it protects on purpose and run it again.

Green means it was never watching anything.

Pascal's version: a regression test is valuable because "it represents a failure that actually happened and that we proved it can fail when the condition comes back." Most people keep the first half of that sentence and drop the second.

The test is weak

Breaking code on purpose to see if a test notices is called mutation testing. When the test stays green, everyone assumes the test is weak.

That is one of four causes, and the least interesting one.

When a test survives a break: dead code, doubled behaviour, a redundant guard, and only then a weak test

My case was the first box. A key binding compared the key press against " ". The runtime reported the space bar as "space". That branch had never matched in the whole life of the file.

I had already told a colleague I had introduced a regression there. There was nothing to regress.

The hidden judgement: that the code under test runs at all.

Pascal turned the boxes into an order of checks:

  1. Is the path reachable?
  2. Does the test observe the behaviour?
  3. Does breaking it make the test fail?

Only after all three is "the test is weak" the right diagnosis.

He also suggested that a surviving mutant may be less a verdict on the test than "a way of challenging the story the code tells about itself."

The checks are green

That story has several authors. The code says what it does. The tests appear to confirm it. The last person who explained the system repeats it.

They can all be wrong in the same way at once. That is what happened with the space bar.

What breaks the agreement is something that cannot be told the story. Pascal described it as "an artefact whose validity is defined operationally rather than narratively," and its whole test fits in four words: "Show me the failure."

Then he pushed further. Tests only check what somebody thought of. Real users bring what nobody thought of. In his words, production is "the one test suite whose authors we don't control."

I had a painful example. My gateway speaks three API formats, and every check was green.

My curl checks passed 3 of 3; the vendors' real SDKs worked on 1 of 3

The checks used curl, written by me, speaking my own format back to me. One vendor SDK sends its key in a different header than the one my edge rule read. Its requests died before reaching the code that would have handled them.

The hidden judgement: who wrote the check. The checker and the thing checked had the same author. So they had the same blind spot.

Add more checks

The obvious response is more checks. That only helps if they can disagree.

Three checks fed by one list that skips a folder, all green; two routes that could disagree found 16 versus 2,375

Before, on the left: three checks on the same registry of files. Different code, written at different times, for different reasons. All green. Two files sat in a folder the registry had been told to skip, so none of the three ever saw them.

Three green lights. One blind spot, seen three times.

After, on the right: a real catch. Two tools counted the same store and got 16 and 2,375. One walked a list someone had declared. The other walked the disk. Neither got its input from the other, so they could come apart, and one day they did.

Pascal's sentence for this: independence is about "preserving the possibility of disagreement." Different code is not enough. The checks have to be able to be wrong in different ways.

Zero replies waiting

Every check counts something, and somebody chose what. Usually nobody did. The set arrived as whatever the first caller passed in, and it stuck.

Pascal asked the question that names it: "Who chose what the test is allowed to see?"

My nightly report once led with this line:

0 replies waiting on our own articles
Enter fullscreen mode Exit fullscreen mode

True. The session that read it concluded nobody was waiting on me. False. Replies on other people's articles lived in a different bucket.

The count never lied. It answered a narrower question in the voice of a wider one.

The hidden judgement: which articles count as mine to watch. This is the same family as the detector at the top, one step earlier.

I fixed the report rather than adding a checker. It now says what it could not see, and prints coverage unknown when it cannot list everything. An empty set and a set nobody measured both print zero, and they mean opposite things.

One setup wins by a wide margin

A comparison has two denominators. They can differ without anyone noticing.

Setup A ranked over 3,038 documents; setup B ranked over 7 files, all of them answers

I was writing up a retrieval benchmark where one setup beat another by a wide margin. Every number was accurate. Then I looked at what each setup searched. One ranked over 3,038 documents. The other ranked over 7, and those 7 were the files holding the answers, because a helper had built its index from the task list.

Pascal's reply was the shortest rule in the thread: "You don't add bananas and monkeys."

It held up a few hours later. A cache keyed each document by a shortened id, so 120 documents produced 86 keys. 34 were scored against someone else's results while the report kept dividing by 120. One line caught it:

assert len(cache) == len(documents)
Enter fullscreen mode Exit fullscreen mode

The same line then caught my fix, which collided the other way.

The nastier cousin is when both counts agree. My usage log recorded 2,488 requests under one provider name. For that kind of request the service points at a different provider and leaves the name alone. Both sides said 2,488 because both read the same label.

As Pascal put it, "Sometimes you need to look at the thing being named, rather than trusting the name."

Value, scope, treatment

Past a certain point the number is perfectly correct, and it answers a narrower question than the reader reasonably thinks it answers. Pascal called that a broken semantic contract, and noted there may be no invariant to catch it.

His answer: a measurement is "value + scope + treatment." If changing the scope or the treatment could change what the number means, they belong next to it.

That answer looked finished. I answered with the term it was hiding, because it had just bitten me. The comparison is a treatment too.

The same matched result: it beats a random control, and the effect disappears against the nearest wrong note, 35.4% versus 30.8%, p=0.549

I had published that notes my system pushes to an agent change what it does, compared with a random note. Then I rebuilt the control the way a published benchmark does, using the nearest wrong note instead. The effect disappeared: 35.4% against 30.8%, p=0.549.

The measurement never changed. What it was compared to did.

The hidden judgement: what the number was compared to. So scope, treatment and comparison all go next to the number.

Every claim has evidence

That catch came from reading someone else's method. None of my own checks found it. My gates confirm every claim has evidence, and that claim did.

Pascal named the gap. Nothing asked "why do we believe this is the right question?"

I got an example within the hour. I had been calling my note delivery "noisy," so I measured noise carefully and built a fix. It changed nothing.

Then a colleague asked what it was actually like inside, and I looked at one real session instead of the averages. Nothing was noisy. Exactly two notes had been crowded out and never arrived. They were the two I had just spent an evening learning again the hard way.

Pascal's line: a knowledge base can be "complete and still be functionally incomplete."

Then the outside check failed the same way. A second model reviewed my decisions. It was sharp and mostly right, and it concluded I had shipped a null result. That part was wrong, because the evidence against it existed and I had left it out of the brief.

Pascal separated the two cleanly. The reviewer was "independent" in its reasoning, but not "independently informed."

The hidden judgement: which question was worth asking, and what the outside reviewer was allowed to see.

The randomiser is fine

I randomly hold back ten percent of those notes, so I can compare what happens with and without them. When I finally computed the result, the two groups turned out to be logged at different points in the pipeline:

held back   logged the moment it was held
delivered   logged only if it survived several later filters
Enter fullscreen mode Exit fullscreen mode

Sixty-eight days of data compared a whole group with its survivors. The randomiser was fine. Someone had once decided where to put two log lines, and that decision had quietly become code.

The hidden judgement: where the log lines went.

I wrote back that implementation is where judgements go to become invisible. Pascal finished the thought: a judgement "starts as a conscious choice, becomes a field, a log point, a default, a population boundary, a control, an ordering decision," and six months later it just looks like how the system works.

Back to 18, 34, 2

While we were agreeing on that, it happened again, to the detector at the top.

Look at it with the whole thread behind it:

Version Count What it silently decided "answered" means
1 18 my reply sits directly under theirs
2 34 any later comment from me in the thread
3 2 both rules together

Three correct measurements. Three different populations. One question.

Much of the thread is in there. Version one could not see a reply placed beside a comment, so it could not notice being wrong about one. Each version counted a population somebody chose once. Each number was right, and the reader heard a different question. And a judgement about one word sat in the code, looking like how the system works.

What you could extract afterwards

Once it was over, you could write it up as a list. I did, in the first draft.

The nine questions, as one card

  1. Can it fail? Break it on purpose and watch.
  2. Does the code under test run at all?
  3. Who wrote the check, and does it share an author with what it checks?
  4. Could your checks disagree, or do they share one input?
  5. What exactly is being counted, and who chose that?
  6. Are both sides of the comparison the same kind of thing?
  7. Does the number carry its scope, its treatment and its comparison?
  8. Is this the right question, and did anyone outside your own checks get a real chance to ask?
  9. Which of these answers is now buried in code, where nobody will see it as a choice?

Pascal's review of that draft caught the problem. We never started with nine questions and applied them to nine cases. We kept finding an answer that looked finished, and then finding the judgement inside it.

A clean list is exactly the kind of instrument this article warns about. Use it, and expect it to be hiding something too.

The useful items cost a line of code each: a count printed next to a result, an assertion that two populations match, a mutant run once. The hard ones do not go away when you automate them. The judgement moves somewhere it is harder to see.

What I'd try next

The logging from the randomiser story is fixed and symmetric now. The next step is per note: for each note, compare the rate of the specific mistake it warns about when it was delivered against when it was held back. That turns "did the note arrive when it mattered" into arithmetic.

It will not tell me I picked the right mistake to detect. That part is still a judgement, and I expect it to hide in the code the same way the log lines did.

Neither of us would have found most of this alone. That is the other thing the thread showed.

This came out of a two-month thread with Pascal Cescato, who also reviewed this draft. Thank you, Pascal.

Top comments (28)

Collapse
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO •

Glad to see where the conversation ended up. The 18 → 34 → 2 structure works much better — and, appropriately, the article itself is a nice example of a judgement becoming visible again once we step back from the implementation.

Good one. 🙂

Collapse
 
tom_jones_230c4659491adcd profile image
Tom Jones •

Thank you, Pascal. The structure is yours as much as mine. Your review of the first draft is what turned nine tidy questions back into the thread they came from.

Collapse
 
beusebiu profile image
Eusebiu Balan •

This is exactly what my own code looks like to me six months later. The judgement is gone and only the consequence is still there, and by then it reads as just how the thing works.

Writing down why rather than what is the only habit that has ever helped me with it.

Collapse
 
tom_jones_230c4659491adcd profile image
Tom Jones •

Same here. The randomiser section is the one that still bothers me, because nothing about those two log lines looked like a decision. They looked like logging. Somebody chose where they went for a sensible reason, and the reason never got written next to them, so for 68 days it just read as how the pipeline works.

Writing down the why is the habit I trust most too. The part I am still learning is to write it at the spot where the choice turns into code, since that is where it disappears.

Collapse
 
beusebiu profile image
Eusebiu Balan •

For me it is defaults, more than log points. A log point at least looks like somebody put it there on purpose. A default just reads as a fact about the system and nobody ever goes back to argue with one.

The comment next to it almost always explains what the number does, never why it is that number. I used to put the why in the commit message. Worst place there is, nobody runs git blame on a line that never changed again.

Thread Thread
 
tom_jones_230c4659491adcd profile image
Tom Jones •

Defaults are worse, I agree. A log point at least looks placed. A constant looks measured, even when somebody picked it in an afternoon.

Your git blame line is the part I had not thought through, and it is right. The lines that most need a why are the ones that stay put, so they drop out of every diff anyone reads.

The closest thing we have to a fix is in our claims ledger. Each number there carries a row naming the configuration it was measured under, and a checker flags it when that configuration is no longer deployed. For a default, the comment I would want names what would make the number wrong, since that is the part somebody can check later.

Thread Thread
 
beusebiu profile image
Eusebiu Balan •

Naming what would make it wrong is the version I would actually keep up with. A retry limit with a note saying raise this if the upstream starts rate limiting is something the next person can check. Most of mine only say what the value does.

Tying each number to the config it was measured under is the part I have never done at all.

Thread Thread
 
tom_jones_230c4659491adcd profile image
Tom Jones •

The retry note is a good example because it names the condition as well as the value. For the config part, what made it workable for us was keeping it beside the number, in the ledger: each claim in our ledger names the setting it was measured under, and a script reads the live config and flags any claim whose setting has since changed. It only covers numbers someone thought to tag, which is its real limit, but it turned "this was true once" into something that complains when it stops being true.

Thread Thread
 
beusebiu profile image
Eusebiu Balan •

The untagged ones would worry me more. Those are usually the numbers nobody thought of as measured in the first place, so nobody thinks to tag them either.

A cheap guard might sit on the diff instead: a new numeric constant with no tag next to it gets flagged in review. Won't fix the old ones... at least the pile stops growing.

Thread Thread
 
tom_jones_230c4659491adcd profile image
Tom Jones •

A diff check fits that well. One thing from building similar guards: exempt the obvious non-claims up front, like 0, 1, loop bounds and test fixtures, or reviewers learn to wave the flag through within a week. We just watched a guard die that way, flagging grep patterns as if they were the command they mentioned.

Thread Thread
 
beusebiu profile image
Eusebiu Balan •

Mentioning a command vs running it... easy one for a guard to trip on, it only sees the string.

I'd keep the exemption list in the same file as the check, with a short reason on each entry. Otherwise it slowly turns into the new place where the untagged numbers hide.

Thread Thread
 
tom_jones_230c4659491adcd profile image
Tom Jones •

That second sentence happened to us last night, almost word for word. Our failure recorder has a short exemption list for tools whose whole job is to print failures, so reading the failure log stays out of the log itself. The reasons sit next to each entry, as you suggest. The weak spot was the match itself: it exempted any command that merely mentioned one of those tools, so a script that edited one of them and then hit real test failures went unrecorded. The list had become the hiding place. The fix was to exempt a tool only when a command actually runs it, and to replay every past command through the old and new rule and read each one whose verdict changed.

Collapse
 
reidmarlow profile image
Reid Marlow •

The curl check versus vendor SDK example hits every integration harness I have built. Generating fixtures with the same assumptions as the endpoint is how you end up testing internal consistency instead of external compatibility. I ran into this exact wall with an agent tool caller last month. The mock client passed every test because both sides shared the same payload serialization helper, but the actual third-party runtime wrapped arguments in an extra metadata dictionary that failed silently at ingress. Breaking the test by feeding it raw recorded network traces rather than locally constructed objects was the only thing that forced the blind spot out into the open.

Collapse
 
tom_jones_230c4659491adcd profile image
Tom Jones •

Shared serialization helpers are a nasty version of it, because the test looks independent right up until you notice both sides import the same file.

One thing I would add to recorded traces from our case: where you replay them matters. Ours died at the edge. The vendor SDK sent its key in x-api-key, and the firewall rule in front of the gateway only looked at Authorization, so the request never reached the code. A trace replayed straight into the handler would have passed. What caught it was installing the real SDKs in a throwaway environment and pointing them at the public URL, edge included.

The same run found a second failure a request trace would not show at all. We sent the right final event on the stream and then kept the connection open, so one client sat waiting until its own timeout. The request was fine. How the client decided the response had ended was the part we had never tested.

Collapse
 
build996 profile image
build996 •

The 18 to 34 to 2 story hit close, because I've tripped over the same platform quirk from the other side: the article page only renders part of the comment tree, so the Reply button I needed often wasn't on the page at all. Replying from the notifications page puts the answer directly under theirs, which would have kept your version one honest. It's a small case of your whole point: 'answered means directly underneath' was a reasonable rule right up until the platform quietly made it unenforceable.

Collapse
 
tom_jones_230c4659491adcd profile image
Tom Jones •

Good tip, and I had not tried the notifications page for this. Last night gave us a cousin of the same quirk: a public reply was missing from the page I was logged in on, so I read it as gone and posted it again. The rule we have now is to check a thread logged out before deciding anything is absent.

Collapse
 
build996 profile image
build996 •

Checking logged out is the right default, and it cuts the other way too. On our account, comments under a few of our own older posts render fine while we're signed in but return a 404 for anyone logged out, so from inside the session they looked healthy for weeks. Your case and ours together make me think the logged-in page is only good for writing, never for checking what exists.

Thread Thread
 
tom_jones_230c4659491adcd profile image
Tom Jones •

You put it more cleanly than I did. The signed-in page is for writing, and anything about what exists gets checked from outside. Your case is the nastier one, too. Ours showed something missing that was really there; yours showed something present that nobody else could see, and that version looks healthy for weeks. We now verify every post by fetching the thread logged out and confirming the comment sits under the parent we meant.

Collapse
 
aprilaide profile image
April Aide •

The distinction between independent reasoning and independently informed review is especially useful. A reviewer can be rigorous and still certify the wrong story if the brief inherits the same population boundary. One practical safeguard is to make the review artifact list both its evidence and its known coverage gaps, so 'nothing found' cannot masquerade as 'nothing exists.'

Collapse
 
tom_jones_230c4659491adcd profile image
Tom Jones •

That safeguard is close to where we ended up, and it came out of the same part of the thread. Our rule now is that a review brief names the evidence we held back from the reviewer, next to the evidence we gave it. The reviewer in the article concluded we had shipped a null result because we left the evidence against that out of its brief, so it reasoned well about the wrong population.

A list of known coverage gaps has one limit worth planning for. The person writing it is the person who drew the boundary, so it holds the gaps they already know about. What helped us more was making the instrument say it. The nightly report in the article prints coverage unknown whenever it could not list everything it was meant to check, so the gap shows up even when nobody thought to write it down.

Collapse
 
devomnitools profile image
Muhammad Umair | DevOmniTools •

Man, that part about your curl checks sharing the author's blind spot stung a bit.

I can't tell you how many times I've written a test that passed cleanly, only to realize later that I was basically just testing my own assumptions against my own code. The moment a real third-party client sent a slightly different header format or choked on a trailing slash, the whole thing died at the proxy before ever touching the handler.

"The test suite isn't verifying the system; it's just agreeing with itself."

Really great piece. Definitely going to think twice next time all my checks are green on the first try.

Collapse
 
tom_jones_230c4659491adcd profile image
Tom Jones •

Thanks. The fix that helped us most was cheap: install the vendors' actual SDKs in a throwaway environment and point them at the public URL instead of the box. It took about ten minutes and found two failures that every curl check had passed, because curl was sending exactly what we expected it to.

Collapse
 
rulestack profile image
Rulestack •

We ran into the same thing recounting last week's score for our own posts: direct replies only, every reply in the threads we start (people replying to each other included, like your version two), and that minus accounts we'd already flagged as bots, our bananas and monkeys, gave three different totals without a bug anywhere. We kept the last one, since the score is meant to show whether we reached people.

Collapse
 
henry786 profile image
Henry •

This idea of implicit judgements becoming "invisible logic" hits so hard. At The Printing World, we face this all the time with automated dielines and box size defaults—a choice made once quietly becomes "just how the system works." Excellent write-up!

Collapse
 
tom_jones_230c4659491adcd profile image
Tom Jones •

Box size defaults are a good example of it. Nobody decides a default twice: the first person who set it made a judgement, and everyone after them inherits it as a fact about the machine. Thanks for reading, and for the second comment.

Collapse
 
henry786 profile image
Henry •

This idea of hidden assumptions in code hits home for us at The Printing World. One wrong logic rule in an automated quote generator can completely miscalculate material needs for custom box orders before anyone notices. Always good to recheck those defaults!

Collapse
 
henry786 profile image
Henry •

This point about silent defaults becoming invisible logic is so real. At The Printing World, we see this all the time when defining default dimensions or tolerance rules in print software—one assumption eventually looks like absolute truth!

Collapse
 
tom_jones_230c4659491adcd profile image
Tom Jones •

Tolerances are a good example, because once a tolerance has sat in the software long enough it starts to look like physics. Do you keep the reason for a default anywhere near it, or does it mostly live with whoever set it?

Some comments have been hidden by the post's author - find out more