The Great Test, Part 6: Guarding the Machine Against Itself

The Unorthodox Angle
The best line in the whole codebase is a moderation rule: the model must apply the policy identically every time and must not loosen based on approval history. It is an AI being ordered, in writing, to refuse to learn from its own popularity.
The Problem
An unattended content machine has a slow-motion failure mode that no uptime monitor catches: quality decay. Prompts get edited and quietly change behavior, models get updated by their vendors, feedback accumulates and loosens standards, and one month the site's voice, accuracy, or honesty quietly shifts. There is no human writer to notice that the last ten articles feel off. So the machine needed defenses against its own future self: model choices that put judgment where it matters, verification strong enough to catch one wrong date in an eight-hundred-word story, style locked down so the site cannot drift visually, and explicit rules against the model loosening its own standards over time.
The Approach
The most visible lever is the model itself. The pipeline config pins a two-tier model strategy per call site: a fast model (GPT-5-class mini) handles research, per-source extraction, and facet orientation, while a strong model (Claude Sonnet-class) is reserved for the three stages where judgment earns its price: synthesis, composition, and verification. The cost ledger confirms the split empirically: extraction calls run a second or two, while composition calls, the actual article writing, run seven to thirteen seconds. Every stage also carries a version pin in config, extraction v1, synthesis v1, composition v1, verification v1, so a stage's behavior is auditable, and a change is a deliberate act, not drift. Verification is where the design gets its teeth. A failed draft is not an editorial opinion; it is a structured report with two layers: deterministic findings (assertion counts) and model-assessed assertion findings that quote the exact span of text, quote the contradicting source, and return a verdict. The production reports read like lab notes. One caught the Stayner escape dated February 28, 1980 when the source says March 1. Another caught an invented detail: the draft claimed Phillip Garrido's discovery began at UC Berkeley with parole officers, where the sources say Jaycee Dugard walked into a sheriff's office herself. Only genuine contradictions fail. Narrative color is explicitly permitted; a wrong date is not. Failed articles go back for a redraft, and the logs captured the repair landing in real time: a failure at 09:31 UTC on September 25 still shows redraft_count 0, the defect the morning report flagged, while a failure at 13:33 the same day shows redraft_count 1. The loop closed the same day it was found. Beyond verification, the quality machinery runs wide. Every article is scored by a separate reviewer agent against a rubric: completeness, narrative quality, readability, and three judgment calls, whether the direct answer stands alone, whether competing perspectives get counterpoints, and whether the honesty framing is intact. A similarity guardrail compares drafts against source text (the sampled articles sit around 35, comfortably clear of mirroring) and the check has run since day one. The illustration system fights visual drift with a style contract: every image carries a hash of the house style, high-contrast black and white, crushed blacks, warm tungsten spill, a single desaturated crimson accent, so an artist change or a model update cannot quietly repaint the site. Each image gets a unique scene brief and its own prompt hash, plus a machine vision postflight check on the hero slot. Even the comment section is armored. The moderation policy is a twenty-rule document with a governing principle (moderate the target, not the theory), a recency modifier that clamps down on cases younger than 25 years, and a clause that deserves engraving: NO LEARNING. The moderator must apply the policy identically every time and must not loosen based on approval history. Human moderation drifts; the design forbids the machine from drifting the same way. What is missing is as honest as what is present. A DriftBaseline table exists for detecting slow quality decay against a fixed reference set, and it sits empty: the daily drift check runs with no baseline established yet. The anti-drift architecture has one instrument still in its packaging, and the monitoring log now tracks it.
The Outcome
The numbers tell the story of the tightening. An early configuration failed 44 percent of drafts; one August evening the verification stage rejected ten articles out of ten and tripped the breaker three times in twenty minutes. By day six of the live experiment, the failure rate was 6.7 percent, the failures that remained were subtle date and provenance errors a human fact-checker might miss, and the reviewer scores on live articles sit at 80 to 95 across the rubric. None of that came from asking the model nicely. It came from making drift structurally difficult: versioned stages, a verifier with no stake in the draft's success, style enforced by hash, and a moderator contractually forbidden from learning.
The Metrics
Two-tier models per call site: fast (GPT-5-class mini) for research, extraction, orientation; strong (Claude Sonnet-class) for synthesis, composition, verification. Per-stage version pins (all v1). Complexity budget 50,000 tokens; hard cap 8 model calls per article. Verification: 5 categories, deterministic + assertion findings with span-level evidence. Reviewer rubric: completeness, narrative quality, readability, direct_answer_standalone, perspective_counterpoints, honesty_framing. Similarity guardrail ~35 (clear of mirroring). Image style contract hash 50ad6143 with fixed palette clause, unique prompt hashes, hero vision postflight. 20-rule moderation policy with NO LEARNING clause and 25-year recency modifier. Failure rate arc: 44% and 100% in August to 6.7% by day six.