Everything was green. Except the watchman himself.
The safety net that should report failing crons had itself been broken for two weeks: 258 errors, zero alerts. And an UNKNOWN in my dashboard turned out to be an assumption nobody had tested. On meta-monitoring, dead man's switches and why a green dashboard is a claim.

In mid-August a routine run of one of my agents did something no dashboard had done that month: it brought bad news. The cron guardian, the safety net that is supposed to switch off failing scheduled jobs in my stack, had itself been broken for more than two weeks. 258 consecutive errors, neatly logged, read by nobody. No alert came, because alerting was the job of precisely the component that was broken.
Earlier I wrote about how my AI agent invented successes. This is the next layer of the same problem: the check on those claims can fail silently too. That same week another measurement turned out to have been lying for a month, with a tidy excuse attached.
Five weeks of silence
The reconstruction afterwards was sobering. On 11 July the last backup commit went to my vault repo. On 29 July the cron guardian did its last successful check and started failing itself. After that the errors piled up. Day after day. Until another agent happened to stumble on it on 14 August during a usage report. Added up: 34 days without a backup, the last two-plus weeks of which also without a working safety net.
The uncomfortable part is not that something broke. Things break. The uncomfortable part is that the discovery depended on chance, while the evidence had been sitting there the whole time.
The damage in numbers
The safety net and the measurement failed at the same time; that is the only reason it stayed invisible for so long.
UNKNOWN is not a data point, it is an assumption
The second find was hidden deeper. In my infrastructure overview the backup freshness had been on UNKNOWN for 33 days, with a tidy reason: "no GitHub access". That sounded like a technical limitation, so nobody looked at it anymore. You get used to an UNKNOWN surprisingly fast.
When I checked, the token that check needed had simply been there all along. The check had just never been tested with that token. One HTTP request later I had a status 200 and a view of the real problem: 34 days without a backup commit.
Since then I read every UNKNOWN in a dashboard differently. It is rarely a missing data point. It is usually an assumption nobody ever tested, with a plausible excuse on top.
What it cost to fix
Measurement gaps are rarely expensive to close; they are expensive to discover.
Who watches the watchman: meta-monitoring in practice
The Romans already had a phrase for it: quis custodiet ipsos custodes. In SRE language the answer is heartbeat monitoring, also known as a dead man's switch. Your safety net must emit a heartbeat of its own. If it stops, an alarm should go off somewhere else. For cron job monitoring this is the standard pattern, and Google's SRE book says essentially the same about every monitoring layer: if nobody notices the monitoring going down, you don't have monitoring, you have decoration.
My own implementation is deliberately small. Every guardian, every health check and every scheduled AI agent now answers one extra question: when did you last run yourself? Those timestamps land in one table, and one weekly run looks only at that. No fourth layer on top. At that point a human takes over, and that is how it should be.
In fairness: that weekly run is two days old and has not proven itself yet. Its silence should now stand out within a week instead of within a month. The first weeks will show whether that holds.
What this means for your marketplace operation
Translate this to e-commerce and it becomes concrete immediately. You have a monitor on your order flow, an alert on feed errors, maybe a tool watching the Buy Box. But when did you last check that those watchers are still running themselves? A feed monitor that crashed in July produces exactly the same thing in August as a healthy monitor with no problems: silence.
That is why, next to my own order monitoring on Channable, there is a second check that only confirms the first one is still alive. And every UNKNOWN or "could not measure" in a weekly report I treat as an action item, not a footnote.
The nuance: not every monitor deserves a monitor
You can overdo this too. Monitor the monitors, give that monitor another monitor, and you build a tower nobody maintains that becomes the outage itself. The boundary that works for me: two automated layers, and above that a human with a fixed weekly look at one overview. More layers mostly add a feeling of safety, not safety. And yes, my new meta-check can break too. The difference is that its silence is now a visible deviation in an overview I look at weekly, not silence among other silences.
A green dashboard is a claim, not a fact. Treat it as one.
Sources and further viewing
Want to spar about your marketplace strategy?
No hype. A sober look at where your growth is and where margin leaks away.
Get in touch