Back to Selected Work
Autonomous Systems

The Great Test, Part 7: Teaching the Machine to Fail Small

Category
Autonomous Systems
Client
crimesandmyths.com (internal experiment)
Role
Architecture redesign, cost engineering, and incident record for the split-pipeline overhaul
Timeline
Sep 2026
The Great Test, Part 7: Teaching the Machine to Fail Small

The Unorthodox Angle

Every instinct in systems work says reliability comes from adding: more checks, more retries, more intelligence. This overhaul came almost entirely from subtraction. The machine was made dumber on purpose, scripts where judgment was not needed, one dossier instead of five fetches, and it was given permission to do absolutely nothing when there was nothing worth doing. The deepest fix was not teaching it to write better; it was teaching it to fail in pieces small enough to afford.

01

The Problem

By the morning of day nine, the machine had a biography of failure that read like a warning label. In a single week it tripped its own circuit breaker roughly a dozen times, most recently at 3:34 in the morning because three of its last ten drafts had failed self-inspection. The trip was fair. What followed was not: the breaker governed publishing, but nothing governed drafting, so the machine spent the rest of the night writing articles nobody was allowed to publish, at ten to twenty metered calls every thirty minutes, around the clock. This was the monolith's original sin, and it had repeated three times in nine days: a burst of blocked-publish spend that added up to more than three thousand credits, roughly a third of the machine's entire monthly budget, exchanged for nothing. Against that ledger stood twenty-four live articles and a queue of a hundred and thirty-two more, written, verified, and shelved. The week also exposed how poorly the machine handled its own good work. A published article, The Red Room Curse, was pulled back from live status overnight by a stale redraft job that rewrote its title on the way down, the third variation of the same class of bug: work that should never have touched a finished piece. A phantom nightly production batch, scheduled for two in the morning, turned out never to have run at all; every publish the machine had ever made slipped out through the side door of the maintenance heartbeat. A cache built to make the machine smarter filled for two days with beautiful, diverse sources and was consulted zero times, because the part that was supposed to read it was never wired. And when a fresh draft called The Dybbuk Box failed verification twice, the machine discarded a piece that was, by its own measurements, most of the way done. The problem was never any single bug. The problem was that one circuit carried everything, so every failure was a system failure, and the machine's response to trouble was to keep its most expensive habit running while its only valuable output sat blocked.

02

The Approach

The redesign splits the machine into two automations that no longer share a fuse. The Gatherer runs at night and does everything cheap: it curates topics, finds sources, and fetches, but instead of streaming its findings straight into a drafting session, it parks them in a dossier, one record per topic, as thin as the filesystem allows: a single large research document in plain markdown, a structured claims file in JSON, and the list of source URLs. The Writer runs separately and never touches the internet. It picks a dossier, writes the article from it alone, checks every sentence against the claims file, and publishes. Its failures are its own; the Gatherer keeps working through a bad writing night, and a bad research night never blocks a finished article from going live. The second change is a small mercy with large consequences. When a draft fails verification even after its two redrafts, the machine no longer throws it away. The article parks in a state called needs_work, carrying its draft, its pointer to the dossier, and the exact findings that sank it, and the next writing cycle resumes it before starting anything new, with a fresh pass against the same research. Dead letters are now reserved for genuine structural failures, a bad dossier, an unreachable topic, a budget breached twice. The breaker itself survives, but scoped to its one honest job: it gates the Writer, it fires on the systemic signals it was always meant to judge, and nothing else in the machine trips it. The third change is a diet. The meter had been quietly double-counting, logging identical calls twice and re-uploading the same illustration per article, so the first fix was honest accounting. After that came batching: one call now curates ten topics at once, one extraction call handles a batch of sources rather than a source at a time, and composition happens in a single pass that produces the article, its metadata, and its image briefs together. A cache that had spent days writing without ever being read is wired into the fetch path, so a source seen once never costs a call twice. The idle case, which used to be the expensive case, is now the cheapest: a writer with nothing to do logs a single line and spends nothing, where the old machine drafted into a blocked pipe all night. And a daily ledger watches the total: if the machine passes 300 credits in a day, it downshifts to cache-only work and finishes what is in flight rather than pretending nothing happened. The target, end to end, is about eight metered events per published article, against the twenty-plus the monolith spent by default.

03

The Outcome

The honest word for all this is not success but bet. The split landed on day nine, and nothing about it is proven until the machine survives a full night under the new rules: the first cache hit, the first article resumed from needs_work instead of re-researched, the first publish that costs single digits. The scoreboard the experiment actually cares about, Google's index, Bing's crawler, the answer engines, has not moved; twenty-four articles are live and none of them are indexed, and no amount of architecture fixes that. But the shape of the failure curve has changed, and that was the point. The monolith failed like a power grid: one fault blacking out the region. The split machine is meant to fail like a good kitchen: one burnt dish goes in the bin, the service continues, the recipe gets a note pinned to it. Whether the new machine earns relevance faster than the old one is unmeasurable this early. What is already measurable is that it can now have a bad night without having a bad week, and that the cost of its bad nights stopped being the cost of its good ones.

04

The Metrics

Overhaul ledger, days eight and nine: 13 breaker trips lifetime before the split, including one for drift flags and one legitimate 30 percent verification failure rate that cost the machine its publishing window. 540 to 900 credits per day of observed burn in the monolith era, with three separate spend-with-zero-output windows totaling more than 3,000 credits. 132 articles sitting in review at peak, unpublished. One cache that wrote for two days and scored zero hits. 31 dead-letter jobs at the moment of the redesign. Post-overhaul targets: 8 metered events per article, under 300 credits per day at full publishing pace, cache hits above zero, and a resumed article costing a fraction of a fresh one. First verification pending the first nights under the new architecture.

Skills

systems architecturecost engineeringautomationfailure isolationobservability
the great testai systemsarchitecturecost engineeringexperiment