My health check passed for weeks while reading nothing
One of my monitoring checks reported "all clear" on every run for weeks. It was reading an empty list and calling that good news. Here is the whole bug: // The alarm: is any real customer on our email provider's supp
One of my monitoring checks reported "all clear" on every run for weeks. It was reading an empty list and calling that good news.
Here is the whole bug:
// The alarm: is any real customer on our email provider's suppression list?
const res = await fetch(
"https://api.provider.com/v3/suppressions?limit=500",
{ headers: { "api-key": key } },
);
const body = await res.json();
const blocked = body.contacts ?? []; // <-- the bug lives here
The provider caps limit at 100. Asking for 500 returns HTTP 400 with {"code":"out_of_range"}. There is no contacts key on an error body, so ?? [] turned a failed request into an empty list, and the step reported:
ok | suppression list holds no real customer — 0 blocked
Zero blocked. Beautifully green. It had never successfully read the list in its life.
Why this class of bug survives
The check would have gone green in exactly the same way if every single customer had been suppressed. Its output was completely disconnected from the thing it was watching, and it had no way to say so.
That is worse than having no check at all, because the green tick is doing active harm — it is answering a question you now believe you have covered.
I think of these as instruments that have stopped touching what they claim to measure. They are hard to spot because:
- they never fail, so they never appear in an incident
- they never flap, so they look like your most stable check
- the code reads fine —
?? []is idiomatic and looks defensive
That last one matters. Reading the diff would not have caught this. contacts ?? [] is the kind of line you skim past approvingly.
How it actually got caught
Not by review. By a second check that disagreed.
I had just added a different step that picks a genuinely suppressed address from the live list and asserts that our login form refuses it with a reason. It ran in the same fifteen-minute cycle and found an address. The older step, in the same run, said zero.
Two checks over the same data, one saying "here is one" and the other saying "there are none". That contradiction is what made it visible.
The two fixes
A failed read is not an empty result.
if (res.status !== 200) {
return {
ok: false,
detail: `provider would not return the list (HTTP ${res.status})`,
};
}
An alarm that cannot read its input has to fail, loudly. "I don't know" and "nothing is wrong" are completely different answers and only one of them is safe to report as green.
Respect the API's real limits. limit=100 and paginate. Worth actually checking the cap rather than picking a number that feels large enough — the failure mode of guessing high was silent.
Two things I would do differently
Test the check against a value you know is bad. Every one of these I have found had never been run against a true positive. If your alarm has never once fired on purpose, you do not know it can. Point it at a known-broken input and watch it go red before you trust it green.
I made this exact mistake twice in one day. The replacement check I wrote first looked up a single address via GET /suppressions/{email} and treated 200 as suppressed. That path returns 404 for suppressed and healthy addresses alike, because it only exists for DELETE. It would have reported "fine" for every customer, forever, without ever failing. I only found out because I tested it against an address I knew was on the list.
Prefer checks that can disagree with each other. A single check has no way to know it has gone deaf. Two independent checks over the same fact do — the contradiction is the signal, and it costs nothing to look for.
Where to look tonight
Go through your monitoring and find every place where a failure could be read as an empty or zero result:
-
?? []or|| []on a parsed API response -
data?.rows?.length ?? 0where 0 means healthy - a
try/catchthat returns a default on parse failure - any count-based alert where "we got nothing back" and "there is nothing to report" produce the same number
Each one is a check that can quietly stop checking, and stay green while it does.
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.