Back to Selected Work
Autonomous Systems

The Great Test, Part 1: Can a Website Raise Itself?

Category
Autonomous Systems
Client
crimesandmyths.com (internal experiment)
Role
Experiment design, autonomous pipeline, daily machine monitoring
Timeline
Sep 2026
The Great Test, Part 1: Can a Website Raise Itself?

The Unorthodox Angle

Most AI content experiments measure whether a machine can fool a human reader. This one measures whether a machine can earn the respect of other machines: Google's index, Bing's crawler, and the AI answer engines, with every number published whether it flatters the experiment or not.

01

The Problem

Series note, October 7, day eighteen: this piece was written on day six, when every number on the scoreboard was still zero. Keep that in mind as you read. The payoff arrives in Part 11, when an AI answer engine uses this site to answer a question for the first time. Everything between here and there, the deadlock, the flag, the meltdown, is the story of getting there, and if you want the shape of the whole arc before the details, start with the overview piece, Our Journey to First Contact. It has never been cheaper to flood the internet with machine-written text, and it has never been harder to know whether any of it earns attention on merit. Google has spent two years rolling out policies against scaled content abuse while quietly absorbing AI-written pages into results that pass its quality bars. Meanwhile the AI answer engines, ChatGPT, Perplexity, Copilot, and Google's own AI Overviews, are reorganizing search so that fewer people ever arrive at a website at all. Everyone has an opinion about what this means for AI content. Almost nobody has a clean, instrumented, fully published record of what actually happens when a brand-new website, written and operated entirely by AI agents, tries to earn relevance the honest way. So we built the instrument. crimesandmyths.com is a fully automated content site. An agent pipeline picks true crime and folklore stories, researches paid sources, writes 500-plus-word retellings in the spirit of Uncle John's Bathroom Reader, illustrates them, fact-checks itself, and publishes with no human writing a single word. There are no ads, no paid traffic, no clickbait, and every page is labeled as AI-authored. The question is blunt: how long does it take agent-created content on a brand-new domain to earn relevance in Google and in the AI answer engines?

02

The Approach

This is a single-arm exploratory test. No control group, no A/B trickery, just one machine, one domain, and a stopwatch. The clock started the day the domain was registered, September 20, 2026. Ten articles dated August 22 predate the domain; they count in the totals but are always flagged as pre-domain inventory. Content comes in four types: solved cases (setup, escalation, twist, resolution), unsolved cases (established facts, then competing theories with evidence and counterpoints), folklore (stories that spread, clearly labeled), and phenomena (persistent claims with an accumulating evidence base). Every article is exactly one of those, so performance can be broken down by type and by article, and the experiment can learn which forms earn attention fastest. Relevance is defined as a ledger of firsts, each captured with a date, a time, the article URL, how many articles existed at that moment, and days since the domain registration: first Googlebot crawl, first page indexed, first Search Console impression, first click, first query reaching the top 100, then the top 20, then the top 10, first branded search for the site's name, first Bing index through IndexNow, and the one that may matter most in the long run, the first citation or referral from an AI answer engine. The referee is a separate monitoring agent, a Base44 Superagent that reads the machine's vitals every morning at eight, wired directly into Google Search Console and Google Analytics. It reports the numbers whether they flatter the machine or not, and it never edits the evidence.

03

The Outcome

Success is not guaranteed, and that is the point. If the content earns its firsts quickly, that says something about the ceiling for honest agent work. If it takes months, or never happens, that is equally valuable evidence about where AI content actually stands with the systems that gatekeep attention. Either way, the story gets published here as it happens, including every failure. This series, filed under the designation The Great Test, is the narrative record of the experiment: what the machine is, how it works, and everything that breaks along the way.

04

The Metrics

Day 6 status: 36 articles live (10 pre-domain inventory, 26 published by the machine since the domain registered). Google: 0 indexed, 0 impressions, 0 clicks. Bing: 0, because the IndexNow handoff has failed on every submission so far. AI citations: 0. Pipeline uptime: continuous since September 20. Observed spend: roughly 540 credits per day against a 10,000 credit monthly cap.

Skills

experiment designautonomous systemscontent pipelineobservabilityseo
the great testai systemsautomationseoexperiment