The Great Test, Part 5: SEO That Cannot Drift

The Unorthodox Angle
The trick was never getting the model to do SEO. It was refusing to let the model do SEO. Everything Google parses is written by deterministic code, and the only thing the model contributes is the one thing it is actually good at: the story.
The Problem
Technical SEO is mostly tedious, and tedious is exactly what language models are bad at. An LLM can write a compelling meta description, but ask it to always produce one under 160 characters, always set the canonical URL, always emit valid JSON-LD with the right schema.org types, always size the social image at 1200x630, and always label its own authorship, and it will comply impressively for a while, then drift. Every AI content site eventually ships an article with a missing OG tag or a malformed schema graph. On an unattended site with no human ever glancing at the output, prompt-level SEO discipline is not a strategy; it is a countdown timer to silent breakage.
The Approach
The proof came early and cheaply. The known failure pattern, logged as a design rule during the build, is blunt: SEO rules that live only in prompt text get violated. Ask a model to always include a meta description, and eventually it will decide a particular article reads better without one. So the machine inverts the responsibility: the drafting model writes prose, and then a deterministic SEO normalizer, seo_version 1.0, pinned in the pipeline config, builds the entire technical surface in code. You can read the output on any article. Take the cow tipping piece, a representative sample straight from the database. The normalizer assembles: a canonical URL, meta title and description, OpenGraph and Twitter card images at exactly 1200x630, a table of contents with anchor slugs, and a breadcrumb trail from home to category to article. Then the structured data, a schema.org JSON-LD graph with six node types: the Article itself (with wordCount, articleSection, inLanguage, and a speakable specification pointing at the direct-answer block), the hero image, the BreadcrumbList, an ItemList of competing perspectives for the case's theories, the Organization, and the WebSite with a SearchAction. Every field, every article, identically, because code wrote it. The layer goes further than most human sites bother. Each article carries a direct_answer field, a standalone summary written for machines to quote, and the speakable CSS selector points at it so voice assistants and AI answers can find it. Every article is also served as clean markdown and raw JSON at dedicated endpoints, so the content is fetchable in the format an AI agent prefers. The sitemap regenerates on every publish, and an IndexNow ping announces each new URL to Bing. Facet tables for places, themes, and historical eras hold orientation text and indexable flags, building an internal linking skeleton that grows with the catalog. And the honesty is embedded in the metadata itself. The image nodes in the structured data carry digitalSourceType set to the IPTC code for trained algorithmic media and a creditText that reads: Illustration generated by AI. Not a photograph, not a depiction of real events. The disclosure is not a footer; it is structured data, machine-readable, in the same graph Google parses.
The Outcome
The result is structural completeness at machine scale: the 36-article catalog ships with the same technical rigor whether it is article 3 or article 3,000, and the site's AI disclosure travels inside the structured data where no redesign or template whim can quietly remove it. Two honest caveats belong in the record. First, hardcore technical SEO guarantees crawlability and machine-readability, not relevance; with zero pages indexed through day six, the experiment's early bottleneck is discovery, not markup. Second, the normalizer cannot fix what Google cannot see yet: the speakable direct answers and JSON endpoints are a bet on AI answer engines, and that bet cannot pay off until something starts citing the site.
The Metrics
Per-article technical surface, built deterministically: canonical URL, meta title and description, OG and Twitter images at 1200x630, TOC with anchors, breadcrumbs, 6-node JSON-LD graph (Article with speakable spec, ImageObject with IPTC trainedAlgorithmicMedia digitalSourceType, BreadcrumbList, competing-perspectives ItemList, Organization, WebSite with SearchAction). Direct-answer block plus dedicated markdown and JSON content endpoints. Sitemap regenerated on every publish. Facet tables for place, theme, and era. seo_version pinned at 1.0.