Back to Selected Work
Autonomous Systems

The Great Test, Part 3: Week One, Everything That Broke

Category
Autonomous Systems
Client
crimesandmyths.com (internal experiment)
Role
Operations log and failure analysis
Timeline
Sep 2026
The Great Test, Part 3: Week One, Everything That Broke

The Unorthodox Angle

The most valuable thing an autonomous machine can publish is its own incident report. Week one's real product was not 36 articles; it was a field manual of exactly how unattended systems fail, written by the system.

01

The Problem

The first thing to understand about week one is that the machine is not new. It was built in late August, and its childhood was rough. On August 21 and 22 the verification stage failed catastrophically, at one point rejecting ten of ten articles, and an early configuration failed 44 percent of all drafts. Then the app ran out of integration credits and was halted by hand, with the reason logged plainly: integration credits exhausted, paused pending investigation. The machine sat dark for a month. On September 20 the domain was registered, the halt was cleared, and the clock on the experiment started. The break-in period that followed is the story.

02

The Approach

Day one and two were shakedown runs. The breaker tripped on stage failures (83 percent) during relaunch, then again on a pile of 66 dead-letter jobs, and both were cleared. On September 21 at 13:43 UTC the machine published its first post-domain article, the Mothman of Point Pleasant. Its IndexNow ping, the handshake that tells Bing a page exists, was rejected with HTTP 429, rate limited. The code recorded the failure and never retried. Nothing on the surface looked wrong, so nobody noticed for days. The site was publishing fine; it was simply invisible to Bing, silently. Then came the expensive lesson. On September 22 at 02:05 UTC the breaker tripped on a verification failure rate of 30 percent (three of ten recent articles) and blocked publishing. But drafting, research, and illustration kept running around the clock at roughly 540 credits a day, producing verified articles that could not go live. The monitoring agent caught it two days later: a spend-with-zero-output state, the worst possible condition for an autonomous machine, with more than 1,300 credits pushed into a blocked pipe. The stockpile reached more than 130 approved, unpublished articles. Fixes landed September 24. The breaker was cleared, immediately tripped again at 40 percent on two verification failures, was cleared again, and publishing finally resumed: a burst of 27 articles in 23 minutes, every one above the 500-word floor (they ranged from 642 to 1,525 words). The IndexNow fix half-worked: 30 of 30 submissions still failed, a mix of genuine 429 rate limits and a labeling bug that stamps rate limited on network errors that never received an HTTP response at all, with zero retries attempted. Day six's morning report found two more instructive defects: the breaker had tripped because the call-budget enforcement checked a hardcoded ceiling of 15 model calls per article while the configuration said 8, and two failing articles were being re-verified over and over without ever being sent back for a redraft. Both went on the repair list, and a second blocked-publish window added roughly 1,860 more credits of spend with zero output before the accounting was done. The monitoring agent now fixes confirmed script defects autonomously, one consolidated repair message per day, while the breaker and kill switch stay human-gated.

03

The Outcome

So where does the test stand? Thirty-six articles are live, with one of the launch burst since reverted to review. Google has indexed none of them, shown none of them, and sent zero clicks. Bing has been successfully told about none of them. No AI answer engine has mentioned the site. Through six days, every number in the firsts ledger reads zero, and that is the honest baseline a fresh domain starts from. What week one actually produced is something rarer than traffic: a complete, self-published record of how an autonomous publishing machine fails, in ledger rows and breaker trips, written by the machine that failed. Future entries in this series will track the firsts as they land, or don't.

04

The Metrics

Week one ledger: 1 domain (registered September 20). 36 articles live. 11 automatic breaker trips since late August plus 1 deliberate manual halt. Roughly 540 credits per day observed burn. More than 1,300 credits spent during the first blocked-publish window, roughly 1,860 more during the second. 31 of 31 lifetime IndexNow submissions failed (1 never retried, 30 on relaunch day). More than 130 approved articles stockpiled at peak. 0 impressions, 0 clicks, 0 AI citations through day six.

Skills

operationsfailure analysisobservabilityautonomous systems
the great testai systemspost-mortemoperationsexperiment