The Great Test: Our Journey to First Contact

The Unorthodox Angle
Most experiments publish their findings at the end, cleaned up. This one published the story while it was happening, failures first, and the failures turned out to be the plot: a machine built to earn the respect of other machines had to survive itself for eighteen days before anything outside would answer. The overview of that arc sits above the eleven-part series, but the honest order to read the parts in is the order they happened.
The Problem
On September 20, 2026, we registered a fresh domain and handed it to a machine. crimesandmyths.com would pick its own stories, research them, write them, fact-check itself, illustrate them, and publish, with no human writing a word and every page labeled as AI-authored. The question was blunt and the stakes were real: how long does agent-created content on a brand-new domain take to earn relevance in Google and the AI answer engines, honestly, with no ads, no paid traffic, and no tricks, in a year when the answer engines were reorganizing search so that fewer people ever arrive at a website at all? The premise looked hopeless from the outside. New domains have no reputation. Google's policies target scaled AI content. The site had a hard budget of ten thousand metered credits a month, a monitoring agent that reports every number whether it flatters the experiment or not, and a scoreboard of firsts, first crawl, first index, first impression, first click, first ranking, first AI citation, that started at zero in every column. What follows is the whole arc in one place: the eighteen days between a blank domain and the moment an AI answer engine used the site to answer someone's question. The detailed record lives in the eleven-part series this overview sits above; the failures, the redesigns, the crises and the numbers are all documented there, with evidence.
The Approach
The journey had six phases, and each one taught the machine something it could not have learned any other way. The monolith, days one to eight. The first version carried everything on one circuit, and it failed like a power grid: one fault blacking out the region. Two days in, its circuit breaker tripped on a 30 percent verification failure rate and stayed open for days while drafting kept spending. Roughly 320 credits a day burned with zero output, more than three thousand across the experiment's worst windows, while 132 approved articles sat shelved behind the dam. A failed draft named Dybbuk Box, most of the way done, was thrown away entirely. A cache filled for two days and was read zero times. The machine was expensive, wasteful, and above all fragile. The redesign, days eight and nine. The monolith was split into a Gatherer, which researches at night and parks one dossier per topic, and a Writer, which drafts from the dossier alone and publishes through a breaker scoped to itself. Failed drafts stop being discarded and start being parked, with their findings attached, to resume next cycle. A credit diet, batching, deduplication, idle discipline, cut the cost of an article from roughly twenty metered events to under three. The machine got cheaper exactly when it stopped trying to be clever. The siege, day eleven. Google Safe Browsing classified the domain as deceptive, social engineering class, on a ten-day-old site whose audit found nothing malicious at all. The causes were structural: public registration pages, an admin route answering strangers, horror imagery an automated watcher could misread. Five fixes landed and verified the same morning, a review was requested honestly, and the experiment wrote its launch rules for every future domain. The drain, days fifteen and sixteen. A broken self-healing loop, drafts re-verified without ever being redrafted, was repaired, and the stockpile behind the dam finally moved: 126 articles shipped across two days, then the machine settled into a calm cadence of about one publish per hour under a governor that watches Google's indexation rate and steps the pace to match the domain's digestion. The meltdown, day seventeen. The breaker repair meant to make safety records singletons instead inserted a new one every fifteen seconds, unbounded, a safety system melting itself down in public. The stop-the-bleed fix landed within minutes, the lifetime audit found that eleven of sixteen trips were a fuse tripping on arithmetic that could never stay closed, and the fuse box was rebuilt in two tiers: per-article problems cost one article, systemic signals throttle one stage. The same afternoon, after seventeen days of serving a web page where a key file should have been, the IndexNow integration finally handed Bing a plain text key, and Bing accepted the site's first URL submissions ever. First contact, day eighteen. The machine's first genuinely stable night, one verified article an hour, zero trips. And then the thing the whole experiment was built to catch: real traffic, and the first recorded use of the site by an AI answer engine. Within hours, the experiment had instrumented the moment, a Drive headquarters with the full narrative, one document per build and solved incident, and a KPI workbook fed by a daily sync, and rewritten its own templates for the readers it had just learned it has: descriptive heads instead of labels like Transmission, length scaled to research depth, answers first, sources cited inline, an llms.txt for the machines that read catalogs.
The Outcome
The honest summary of the journey is that the machine had to stop failing like a system before anything outside would look at it twice. Eighteen days is a short time in SEO and a long time in operations, and almost all of it was spent surviving the machine's own instincts: to keep its most expensive habit running during a blockage, to trust a fuse whose trip condition was arithmetic, to write for a category label instead of a reader. The redesign that worked was mostly subtraction. And the bet has its first data point. A site written and operated entirely by agents, clearly labeled as such, on an eighteen-day-old domain, was used by an AI answer engine to answer a question. Not indexed prominently, not ranking, not yet verifiable in search data that lags, and one observation is not a trend. But the direction of the question has changed. It is no longer whether agent-created content can earn relevance; it is how fast, and in which layer, the citation layer arriving first. The instruments are now running on both sides of that question: the machine writes to be cited, and the ledger records whether it is. The search firsts, first index, first impression, first click, first ranking, are the next milestones, and when they land they will be written up with the same candor as the fuse box melting. The series continues.
The Metrics
18 days from domain registration to first contact. 171 articles published, 10 of them pre-domain inventory. Cost per article: 1.4 to 3.0 metered credits in the current architecture, down from roughly 20 in the monolith. Verification pass rate at the end: 100 percent, redraft counts 1 to 4. 16 lifetime breaker trips before the two-tier redesign. 126 articles shipped across the two-day stockpile drain. First IndexNow submission accepted by Bing on day 17 after 17 days of failures. First real traffic and first AI answer usage on day 18, owner-verified on platform analytics; search-side firsts still pending on data that lags 2 to 3 days. Full ledger: 11-part series, plus the KPI workbook and incident docs in the experiment's Drive headquarters.