Consulting

Have an
interesting problem?

The best engagements aren't agency-client arrangements — they're collaborations on problems worth solving. I'm selective. The work I do commit to gets my full attention.

How I Work

Three ways to engage

01

Strategic Architecture

Systems that compound, not campaigns that expire.

02

Product & Ownership

Embedded product leadership — own the outcome.

03

Applied AI

Real workflows, real leverage. Not a deck about the future.

Pitch Me Your Problem

Tell me about the problem — not the deliverable

Selected Work

Recent engagements

The Great Test: Our Journey to First Contact
The Great Test: Our Journey to First Contact
Autonomous SystemsThe complete arc of the experiment in one narrative: from a blank domain and a fragile machine to the first use of the site by an AI answer engine, eighteen days later.

On September 20, 2026, we registered a fresh domain and handed it to a machine. crimesandmyths.com would pick its own stories, research them, write them, fact-check itself, illustrate them, and publish, with no human writing a word and every page labeled as AI-authored. The question was blunt and the stakes were real: how long does agent-created content on a brand-new domain take to earn relevance in Google and the AI answer engines, honestly, with no ads, no paid traffic, and no tricks, in a year when the answer engines were reorganizing search so that fewer people ever arrive at a website at all? The premise looked hopeless from the outside. New domains have no reputation. Google's policies target scaled AI content. The site had a hard budget of ten thousand metered credits a month, a monitoring agent that reports every number whether it flatters the experiment or not, and a scoreboard of firsts, first crawl, first index, first impression, first click, first ranking, first AI citation, that started at zero in every column. What follows is the whole arc in one place: the eighteen days between a blank domain and the moment an AI answer engine used the site to answer someone's question. The detailed record lives in the eleven-part series this overview sits above; the failures, the redesigns, the crises and the numbers are all documented there, with evidence.

Most experiments publish their findings at the end, cleaned up. This one published the story while it was happening, failures first, and the failures turned out to be the plot: a machine built to earn the respect of other machines had to survive itself for eighteen days before anything outside would answer. The overview of that arc sits above the eleven-part series, but the honest order to read the parts in is the order they happened.

The Great Test, Part 8: When Google Calls You Deceptive
The Great Test, Part 8: When Google Calls You Deceptive
Autonomous SystemsSurvival record and same-day remediation of the Google Safe Browsing social engineering flag, day eleven of the experiment.

On day eleven, ten days after the domain was registered, the experiment received its first verdict from outside, and it was brutal. Google Safe Browsing classified crimesandmyths.com as deceptive, social engineering class: a site whose pages allegedly try to trick visitors into sharing personal information or downloading software. For a ten-day-old domain whose entire thesis is earning its first impression from search engines, being flagged as a scam by the internet's immune system is close to existential. Chrome, Firefox and Safari would warn visitors away, and flagged domains are widely believed to index poorly, which meant the punishment landed on the exact quarter the experiment was built to measure. The audit that followed was thorough and came back empty. Valid TLS. No injected scripts or iframes. No forms, no password or email inputs anywhere in article content. No downloads, no comment spam, no suspicious outbound links. The site was not compromised. What it was, was structurally suspicious: a ten-day-old domain with zero reputation, public login, register and forgot-password pages, an admin route returning a live panel to anonymous visitors, and AI-generated horror imagery that happened to depict nearly a textbook version of what the classifier hunts for. Nothing malicious. Plenty misreadable.

The scariest moment of the experiment so far was not a pipeline failure. It was a verdict. The machine's own breakers, budgets and verifiers are opinions we can argue with; Safe Browsing is a wall we cannot. The instinct under fire is to hide and wait it out. The actual fix was the opposite: strip away everything a hostile automated reader could misread, then ask for the review honestly, on the record, describing the site exactly as what it is.

The Great Test, Part 9: The Night We Shipped 126 Stories
The Great Test, Part 9: The Night We Shipped 126 Stories
Autonomous SystemsStockpile drain, batch publishing, and the cadence governor that lets Google set the pace, days fifteen to eighteen.

After the split architecture landed, the machine kept its cost promises but broke a quieter one: the self-healing loop did not heal. Failed drafts were being re-verified without ever being redrafted, failures at redraft count zero, so articles that could have saved themselves never did, and every one of them kept polluting the failure window that governed publishing. The consequence was a machine that wrote cheaply, verified honestly, and then sat on its own work: more than 130 articles approved, verified, ready, and shelved while the failure window slowly decayed and the publish queue waited for permission. It was the inverse of the old disease. The monolith spent money and published nothing because it was broken. The split machine spent almost nothing and published nothing because one small loop was wired wrong, and the stockpile behind the dam kept growing.

Publishing more is usually a growth decision. Here it was an engineering one: a machine holding a library it had already paid for, behind a dam made of one broken loop. And the interesting problem was never how to publish faster. It was how to let a brand-new domain take a deep breath between batches, and how to hand the throttle to the only reader whose appetite actually counts.

The Great Test, Part 10: The Day the Fuse Box Melted
The Great Test, Part 10: The Day the Fuse Box Melted
Autonomous SystemsIncident record for the breaker churn regression, the lifetime fuse audit, and the two-tier redesign, day seventeen.

The breaker redesign began as a five-minute repair and briefly became the worst incident since the original deadlock. The fix for the stale-singleton problem, make the breaker one upserted record, deployed at 12:32 UTC and did the opposite: evaluation began inserting a new breaker record roughly every fifteen seconds, unbounded, twenty-five and climbing within minutes, while the reads meant to enforce safety drank from a firehose of contradictory state. Mid-churn the machine also tripped on a reason that could never resolve: dead-letter depth 35, the stable pre-restart backlog that had never been baselined, a fuse that would fire on every evaluation forever. The audit that followed the stop-the-bleed was the experiment's first full accounting, and it found the uncomfortable arithmetic underneath eighteen days of trips: sixteen lifetime trips, eleven of them verification failure rate, steady state around 30 percent against a 25 percent threshold on a ten-article window. Three failures in ten articles trips the fuse. The breaker could mathematically never stay closed; every batch re-tripped it, not because the machine was broken but because the fuse was designed to detect a trend using a sample that could not contain one. The audit also caught the machine's proudest number hiding its most expensive bug: an article that had redrafted itself 82 times rather than quit, and a second at 81, together a quiet credit hemorrhage wearing the costume of persistence.

The safest-sounding fix in the world, make the safety record a singleton, was the one that melted down: two new breaker records every fifteen seconds, forever. And the machine's proudest number, an article that refused to give up and redrafted itself 82 times, was actually its most expensive bug. The morning's lesson was that reliability theater has two costumes: fuses that trip on arithmetic, and persistence that never checks the meter.

The Great Test, Part 11: First Contact
The Great Test, Part 11: First Contact
Autonomous SystemsThe first real traffic and first AI answer usage, day eighteen, and the measurement and authoring stack built in response.

For seventeen days the experiment was a broadcast. The machine published article after article into what its own dashboards showed as silence: zero indexed pages, zero impressions, zero clicks, seventeen days of a search console the monitoring agent could not even properly see. The whole design of the test was to find out how long that silence lasts when the content is agent-created, labeled, and honest about being both. On day eighteen the first crack in the silence arrived, and it arrived from the direction the experiment cared about most. Real traffic, a healthy amount, and more than that: the first recorded use of the site by an AI answer engine, the site's content used to answer a question somewhere in the machine layer. Owner-verified on the platform dashboard within hours. The same morning, Google Search Console finally recognized the property under the right account, ending the seventeen-day measurement blackout that had made the experiment run on faith. The honest ledger is part of the story: GSC still shows zero impressions, which is expected, the data lags by days and most URLs were only submitted to the index on day fifteen and sixteen, and the traffic observation lives in platform analytics rather than the instruments built for it. First contact was seen before it could be measured. That is a very human way to learn the machines had arrived.

Seventeen days of publishing into silence, and the first answer came not from a click but from a machine using the site to answer someone. The response to first contact was not celebration. It was instrumentation and then a rewrite: if the readers include answer engines, the article templates should stop writing for skimmers and start writing for citers.

The Great Test, Part 7: Teaching the Machine to Fail Small
The Great Test, Part 7: Teaching the Machine to Fail Small
Autonomous SystemsArchitecture redesign, cost engineering, and incident record for the split-pipeline overhaul

By the morning of day nine, the machine had a biography of failure that read like a warning label. In a single week it tripped its own circuit breaker roughly a dozen times, most recently at 3:34 in the morning because three of its last ten drafts had failed self-inspection. The trip was fair. What followed was not: the breaker governed publishing, but nothing governed drafting, so the machine spent the rest of the night writing articles nobody was allowed to publish, at ten to twenty metered calls every thirty minutes, around the clock. This was the monolith's original sin, and it had repeated three times in nine days: a burst of blocked-publish spend that added up to more than three thousand credits, roughly a third of the machine's entire monthly budget, exchanged for nothing. Against that ledger stood twenty-four live articles and a queue of a hundred and thirty-two more, written, verified, and shelved. The week also exposed how poorly the machine handled its own good work. A published article, The Red Room Curse, was pulled back from live status overnight by a stale redraft job that rewrote its title on the way down, the third variation of the same class of bug: work that should never have touched a finished piece. A phantom nightly production batch, scheduled for two in the morning, turned out never to have run at all; every publish the machine had ever made slipped out through the side door of the maintenance heartbeat. A cache built to make the machine smarter filled for two days with beautiful, diverse sources and was consulted zero times, because the part that was supposed to read it was never wired. And when a fresh draft called The Dybbuk Box failed verification twice, the machine discarded a piece that was, by its own measurements, most of the way done. The problem was never any single bug. The problem was that one circuit carried everything, so every failure was a system failure, and the machine's response to trouble was to keep its most expensive habit running while its only valuable output sat blocked.

Every instinct in systems work says reliability comes from adding: more checks, more retries, more intelligence. This overhaul came almost entirely from subtraction. The machine was made dumber on purpose, scripts where judgment was not needed, one dossier instead of five fetches, and it was given permission to do absolutely nothing when there was nothing worth doing. The deepest fix was not teaching it to write better; it was teaching it to fail in pieces small enough to afford.

The Great Test, Part 5: SEO That Cannot Drift
The Great Test, Part 5: SEO That Cannot Drift
Autonomous SystemsDeterministic SEO normalization and machine-readable publishing

Technical SEO is mostly tedious, and tedious is exactly what language models are bad at. An LLM can write a compelling meta description, but ask it to always produce one under 160 characters, always set the canonical URL, always emit valid JSON-LD with the right schema.org types, always size the social image at 1200x630, and always label its own authorship, and it will comply impressively for a while, then drift. Every AI content site eventually ships an article with a missing OG tag or a malformed schema graph. On an unattended site with no human ever glancing at the output, prompt-level SEO discipline is not a strategy; it is a countdown timer to silent breakage.

The trick was never getting the model to do SEO. It was refusing to let the model do SEO. Everything Google parses is written by deterministic code, and the only thing the model contributes is the one thing it is actually good at: the story.

The Great Test, Part 4: The Topic-to-Publish Machine
The Great Test, Part 4: The Topic-to-Publish Machine
Autonomous SystemsSystems architecture, job queue design, publish throttling

Anyone can wire four API calls in a row and call it a pipeline. The real problem with unattended publishing is that every stage depends on the last one, and the world is hostile: research APIs rate limit, extraction occasionally times out, a model call fails mid-article, and a domain you registered yesterday gets indexed at whatever pace Google feels like. A human editor absorbs all of that by checking in every day. A machine has to absorb it structurally, or it stops silently, and the worst failure of an autonomous site is not an error message, it is quiet. So the topic-to-publish chain had to be built as a system that survives its own failures: every step retryable, every state durable, and the whole thing metered to the call.

The queue's killer feature is boredom. A heartbeat that claims twelve jobs every thirty minutes, six of which politely decline, is not inefficiency; it is a machine breathing at a fixed rate so that any deviation shows up as a spike someone can read.

The Great Test, Part 6: Guarding the Machine Against Itself
The Great Test, Part 6: Guarding the Machine Against Itself
Autonomous SystemsModel strategy, verification design, drift containment

An unattended content machine has a slow-motion failure mode that no uptime monitor catches: quality decay. Prompts get edited and quietly change behavior, models get updated by their vendors, feedback accumulates and loosens standards, and one month the site's voice, accuracy, or honesty quietly shifts. There is no human writer to notice that the last ten articles feel off. So the machine needed defenses against its own future self: model choices that put judgment where it matters, verification strong enough to catch one wrong date in an eight-hundred-word story, style locked down so the site cannot drift visually, and explicit rules against the model loosening its own standards over time.

The best line in the whole codebase is a moderation rule: the model must apply the policy identically every time and must not loosen based on approval history. It is an AI being ordered, in writing, to refuse to learn from its own popularity.

The Great Test, Part 2: Anatomy of an Unattended Publishing Machine
The Great Test, Part 2: Anatomy of an Unattended Publishing Machine
Autonomous SystemsSystems architecture and honest automation design

Everyone talks about AI agents as if the hard part is generating the text. It isn't. Language models got good, and the writing is now the easy stage. The hard part is operations: a machine that publishes on its own also fails on its own, at three in the morning, while spending real metered money. If you automate only the writing, you get a factory that can produce confident nonsense at scale and stop silently the first time something upstream changes. So the design problem for crimesandmyths.com was never how do we generate articles. It was: how do you build a publishing machine whose failures are loud, whose spending is visible, and whose honesty does not depend on the mood of a model?

The design bet: honesty is a systems property, not a prompt. Verification limits, budget ceilings, and breaker trips live in code and database tables where a persuasive model cannot charm them away, and the machine's own spend ledger doubles as its lie detector.

The Great Test, Part 3: Week One, Everything That Broke
The Great Test, Part 3: Week One, Everything That Broke
Autonomous SystemsOperations log and failure analysis

The first thing to understand about week one is that the machine is not new. It was built in late August, and its childhood was rough. On August 21 and 22 the verification stage failed catastrophically, at one point rejecting ten of ten articles, and an early configuration failed 44 percent of all drafts. Then the app ran out of integration credits and was halted by hand, with the reason logged plainly: integration credits exhausted, paused pending investigation. The machine sat dark for a month. On September 20 the domain was registered, the halt was cleared, and the clock on the experiment started. The break-in period that followed is the story.

The most valuable thing an autonomous machine can publish is its own incident report. Week one's real product was not 36 articles; it was a field manual of exactly how unattended systems fail, written by the system.

The Great Test, Part 1: Can a Website Raise Itself?
The Great Test, Part 1: Can a Website Raise Itself?
Autonomous SystemsExperiment design, autonomous pipeline, daily machine monitoring

Series note, October 7, day eighteen: this piece was written on day six, when every number on the scoreboard was still zero. Keep that in mind as you read. The payoff arrives in Part 11, when an AI answer engine uses this site to answer a question for the first time. Everything between here and there, the deadlock, the flag, the meltdown, is the story of getting there, and if you want the shape of the whole arc before the details, start with the overview piece, Our Journey to First Contact. It has never been cheaper to flood the internet with machine-written text, and it has never been harder to know whether any of it earns attention on merit. Google has spent two years rolling out policies against scaled content abuse while quietly absorbing AI-written pages into results that pass its quality bars. Meanwhile the AI answer engines, ChatGPT, Perplexity, Copilot, and Google's own AI Overviews, are reorganizing search so that fewer people ever arrive at a website at all. Everyone has an opinion about what this means for AI content. Almost nobody has a clean, instrumented, fully published record of what actually happens when a brand-new website, written and operated entirely by AI agents, tries to earn relevance the honest way. So we built the instrument. crimesandmyths.com is a fully automated content site. An agent pipeline picks true crime and folklore stories, researches paid sources, writes 500-plus-word retellings in the spirit of Uncle John's Bathroom Reader, illustrates them, fact-checks itself, and publishes with no human writing a single word. There are no ads, no paid traffic, no clickbait, and every page is labeled as AI-authored. The question is blunt: how long does it take agent-created content on a brand-new domain to earn relevance in Google and in the AI answer engines?

Most AI content experiments measure whether a machine can fool a human reader. This one measures whether a machine can earn the respect of other machines: Google's index, Bing's crawler, and the AI answer engines, with every number published whether it flatters the experiment or not.

I'm selective.
You should be too.

The engagements I commit to get my full attention. Not a vendor — a collaborator.

Pitch your problem