Some months ago I wrote here about one bad instance hiding inside a fleet average. This is the opposite failure, and it happened to the same checkout service.
In June the partner API that calculates sales tax began failing about one call in eight. We heard from support two hours in, and then found an alert rule written for exactly that failure, which had watched it happen and said nothing.
The rule fired when any checkout pod logged more than forty upstream errors a minute. It dated from three years earlier, when checkout ran on four large pods and forty errors a minute on one of them meant something was plainly wrong. It had fired correctly before, so nobody had looked at it since.
In the spring a cost exercise moved checkout to smaller pods on cheaper nodes, and we let the autoscaler choose how many. On a normal afternoon that is now between thirty and thirty eight. During the incident the service produced about six hundred upstream errors a minute, spread across thirty four pods, which is roughly eighteen each. Every pod stayed comfortably under forty. The rule checked thirty four series every minute and every one of them passed.
A dashboard panel showing the worst pod's error count had drifted the same way: it now showed one pod's shrinking share of the problem.
Nothing in the rule mentioned a replica count, but its threshold assumed one. It was a service level limit divided by four, written down as a constant, and the fleet had grown underneath it while the constant stayed put.
We went through every alert rule we own and split them into two kinds. Rules that ask whether the service is failing aggregate across all instances first and compare a ratio: upstream errors over upstream calls for the whole service, paging above two percent for five minutes. Rules that ask whether one instance is misbehaving compare it with its siblings, which works the same at four pods or forty. Any rule still holding an absolute count per instance must now state, in its description, the instance count it was sized for, and a monthly job flags every rule where that figure and the current fleet differ by more than half.
An autoscaler changes the divisor in every per instance threshold you have, and it does not tell the people who chose the numerator.
– Sergey Shinder
Top comments (0)