Back to Selected Work
Autonomous Systems

The Great Test, Part 2: Anatomy of an Unattended Publishing Machine

Category
Autonomous Systems
Client
crimesandmyths.com (internal experiment)
Role
Systems architecture and honest automation design
Timeline
Aug 2026
The Great Test, Part 2: Anatomy of an Unattended Publishing Machine

The Unorthodox Angle

The design bet: honesty is a systems property, not a prompt. Verification limits, budget ceilings, and breaker trips live in code and database tables where a persuasive model cannot charm them away, and the machine's own spend ledger doubles as its lie detector.

01

The Problem

Everyone talks about AI agents as if the hard part is generating the text. It isn't. Language models got good, and the writing is now the easy stage. The hard part is operations: a machine that publishes on its own also fails on its own, at three in the morning, while spending real metered money. If you automate only the writing, you get a factory that can produce confident nonsense at scale and stop silently the first time something upstream changes. So the design problem for crimesandmyths.com was never how do we generate articles. It was: how do you build a publishing machine whose failures are loud, whose spending is visible, and whose honesty does not depend on the mood of a model?

02

The Approach

The machine is a ten-stage chain. Topic selection pulls from a ranked topic bank. Research is paid: real API calls to real sources. Source ingestion extracts structured claims, quotes, dates, and entities from every page fetched, and scores each source for reliability over time. Synthesis reconciles claims across sources, marking agreements, disputes, and confidence levels. Drafting writes the article. Then comes the part most pipelines skip: verification, a second model pass that checks the draft against the extracted claims in exactly five categories, named-person facts, attributions, dates, locations, and numbers, outcomes, and direct quotes. Narrative color is explicitly allowed. Only genuine contradictions with the sources fail. Failed drafts go to a redraft loop, and an article that cannot pass eventually lands in a dead-letter queue instead of on the site. Illustration adds three images per article, with the hero image getting a machine vision check for subject and historical era before it ships. A deterministic SEO normalizer then enforces titles, slugs, schema markup, and metadata in code, because SEO rules that live only in prompt text get violated the moment the model feels creative. Auto-publish follows, and an IndexNow ping tells Bing the page exists. The skeleton is about forty database tables: articles, topics, sources, claims, evidence items, verification reports, generated images, a durable job queue with leases, retries, and dead letters, a publish queue, an IndexNow submission ledger, a credit ledger, a growth scoreboard, a circuit breaker, and a configuration store. Three automations drive it: a heartbeat every thirty minutes that claims queued jobs, a production batch at 02:00 UTC, and an IndexNow verification that fires on every publish. The economics are metered to the call. Every integration call bills credits multiplied by model tier, and every automation run costs 0.1 credits, against a 10,000 credit monthly cap drawn from a 20,000 credit shared workspace pool. A budget governor paces spend across the cycle, with a halt floor at 150 credits remaining. And the honesty is architectural. A circuit breaker trips on statistical evidence, a verification failure rate above 25 percent, a stage failure rate above 30 percent, a dead-letter pile, and blocks publishing until a human decides. A kill switch can stop everything. The spend ledger is a first-class table with a standing rule for reading it: no credit rows means the meter is broken, not that nothing was spent. The machine cannot quietly fail and cannot quietly lie.

03

The Outcome

Around the machine sits a monitoring layer: a separate agent that reads the vitals every morning and is allowed to repair confirmed script defects with one consolidated fix message a day, but is explicitly forbidden from weakening verification thresholds or resetting the breaker. Safety stays human-gated; repair becomes automatic. That division, automatic repair with human-gated safety, is the operating philosophy, and week one tested it immediately.

04

The Metrics

About 40 entities in the data model. 10 pipeline stages from topic selection to IndexNow ping. 5 verification categories. 3 images per article with 1 machine vision check. 3 scheduled automations (30-minute heartbeat, 02:00 UTC production batch, publish-triggered IndexNow verify). Breaker thresholds: 25% verification failures, 30% stage failures, 5 dead-letter jobs. 0.1 credits per automation run. 10,000 credit monthly cap with a 150 credit halt floor.

Skills

systems architectureautomationobservabilitycost engineeringcontent pipeline
the great testai systemsarchitectureautomationobservability