Back to Selected Work
Autonomous Systems

The Great Test, Part 10: The Day the Fuse Box Melted

Category
Autonomous Systems
Client
crimesandmyths.com (internal experiment)
Role
Incident record for the breaker churn regression, the lifetime fuse audit, and the two-tier redesign, day seventeen.
Timeline
Oct 2026
The Great Test, Part 10: The Day the Fuse Box Melted

The Unorthodox Angle

The safest-sounding fix in the world, make the safety record a singleton, was the one that melted down: two new breaker records every fifteen seconds, forever. And the machine's proudest number, an article that refused to give up and redrafted itself 82 times, was actually its most expensive bug. The morning's lesson was that reliability theater has two costumes: fuses that trip on arithmetic, and persistence that never checks the meter.

01

The Problem

The breaker redesign began as a five-minute repair and briefly became the worst incident since the original deadlock. The fix for the stale-singleton problem, make the breaker one upserted record, deployed at 12:32 UTC and did the opposite: evaluation began inserting a new breaker record roughly every fifteen seconds, unbounded, twenty-five and climbing within minutes, while the reads meant to enforce safety drank from a firehose of contradictory state. Mid-churn the machine also tripped on a reason that could never resolve: dead-letter depth 35, the stable pre-restart backlog that had never been baselined, a fuse that would fire on every evaluation forever. The audit that followed the stop-the-bleed was the experiment's first full accounting, and it found the uncomfortable arithmetic underneath eighteen days of trips: sixteen lifetime trips, eleven of them verification failure rate, steady state around 30 percent against a 25 percent threshold on a ten-article window. Three failures in ten articles trips the fuse. The breaker could mathematically never stay closed; every batch re-tripped it, not because the machine was broken but because the fuse was designed to detect a trend using a sample that could not contain one. The audit also caught the machine's proudest number hiding its most expensive bug: an article that had redrafted itself 82 times rather than quit, and a second at 81, together a quiet credit hemorrhage wearing the costume of persistence.

02

The Approach

Stop the bleeding, then redesign. The evaluation path was rewritten to upsert one fixed breaker record and never insert, and the churn died at 12:37 UTC, five minutes after it began. Then came the redesign the incident deserved. Two tiers of fuses: a per-article problem now dead-letters that article alone, while a systemic signal throttles only the stage it implicates, never the whole machine. The verification window widened from ten to thirty articles at the same 25 percent threshold, tripping only on two consecutive checks, so the fuse measures a trend instead of a bad morning. Dead-letter counting baselined and growth-only, so old stable backlog can never trip it again. The call budget removed from global trips entirely, an article that overspends is that article's problem. Budget advisories idle the writer and gatherer instead of killing the heartbeat. Auto-clear once two consecutive checks come back healthy. Redrafts capped at five, with one redraft guaranteed before any final fail. The same day, a second seventeen-day wound closed: the IndexNow key file, which had been serving the site's HTML shell to Bing with a content-type of text/html, finally went out as plain text with the correct key body, and Bing accepted the first URL submissions in the site's history.

03

The Outcome

By nightfall the machine held its cleanest bill of health to date: churn stopped, the breaker stable, publishing resumed, and the fuse box rebuilt so that its failures are sized to their causes. The audit's lesson generalizes beyond this machine: a safety system whose trip condition is mathematically unavoidable is not a safety system, it is a scheduled outage. The redesign's bet is that the machine can now fail at the size of the problem, an article, a stage, a night, rather than the size of everything. And the IndexNow fix closed a seventeen-day arc that deserves its own line: from a single rate-limited submission on day two, through weeks of the key file answering every verification request with a web page, to Bing finally saying yes, three times, at 12:39 on the worst morning of the month. The site's first URL was accepted by an external system on the same day its own fuse box melted. The experiment has stopped being surprised by that kind of simultaneity.

04

The Metrics

Churn peak: 2 new breaker records roughly every 15 seconds from 12:32:33 UTC, 25-plus within minutes, stopped 12:37:33 UTC; entity stable at 5 records after. Lifetime audit: 16 trips, 11 of them verification_failure_rate, steady-state ~30 percent against a 25 percent threshold on a 10-article window. Runaway redrafts: Ronald Tammen at 82, Michael Rockefeller at 81, cap set at 5. IndexNow: 17 days of 100 percent failed submissions (429s and mislabeled no-responses), root cause key file served as text/html; first accepted submissions, HTTP 202, at 12:39 UTC the same day. First post-churn publish, Adam Walsh, 12:39:37 UTC. Redesign deployed and verified the same evening.

Skills

incident responsesystems auditsafety designfailure isolationcost engineering
the great testai systemsincident responsereliability engineeringexperiment