Back to Hub
Notebook

The Data Dance

4,463 photos, one enrichment pass, thirty fields each

August 7, 2026
The Data Dance

The Unorthodox Angle

The hard problem isn't generating data with AI — it's designing a schema where every AI-generated field does double duty: powering search, filters, SEO, collections, visual similarity, and commerce from the same single enrichment pass. Most companies start with the AI and work backward. This pipeline starts with the downstream systems and works forward.

4,463 photos sat in Google Drive. Good photography — Venice canals, macro fungi, brutalist architecture, wildlife, florals. Decent folder organization. Zero metadata that a storefront could use.

No titles. No descriptions. No tags. No alt text. No color data. No searchable keywords. No SEO. No way to filter by mood, composition, color, or subject. Just files in folders — the way most photographers' archives look, and the way most enterprise data lakes look too.

What happened next is a case study in what AI is actually good at when you stop asking it to be creative and start asking it to do the work nobody wants to do.

The Scale of the Problem

A photography storefront needs dozens of data points per image to function. It needs titles for display. Descriptions for product pages. Alt text for accessibility. Tags for search. Keywords for filtering. Color data for visual matching. Camera data for technical credibility. Location data for GEO relevance. SEO metadata for discoverability. Orientation for layout. Pricing for commerce.

Multiply that by 4,463 images and you get a sense of the wall. One person doing this manually — looking at each photo, writing a title, crafting a description, picking tags, identifying colors, copying EXIF data — would need roughly five minutes per image. That's 372 hours. Nine weeks of full-time work. And it would be the worst nine weeks of their life.

The AI pipeline does it in minutes.

The Pipeline: Drive to Data

The pipeline starts with Google Drive. Eight folders of raw photographs — travel, wildlife, macro, florals, architecture. The sync system caches Drive file IDs for incremental processing, so it doesn't re-download files it's already seen. New files get pulled into Base44's file storage, and the enrichment begins.

What the sync system tracks tells you what the pipeline cares about:

  • Stage: What phase of processing each batch is in
  • Total: How many files are in the current batch
  • Created/Updated: Whether records were new or modified
  • Synced Drive IDs: Which Google Drive files were processed
  • Synced Slugs: The AI-generated URL slugs for each photo

That last one matters more than it looks. The slugs aren't filenames. They're things like velvet-bloom-in-shadow, ancient-guardians-of-the-arch, whispering-dome-spire, the-silent-sentinel-of-venice. SEO-friendly, human-readable, generated by the AI from the image content. Every photo gets a title and a slug before it gets anything else.

What the AI Actually Generates

Here's where it gets interesting. A single photo — one image file — goes through the pipeline and comes out with roughly 30 data points. Let me walk through a real example.

Take a photo that the AI titled "Dusky Garden Reverie." Here's what the pipeline produced:

Identity

  • Title: Dusky Garden Reverie (AI-generated, evocative, not a filename)
  • Slug: dusky-garden-reverie (SEO-friendly URL)
  • Description: A 100+ word rich description — "A close, painterly study of flowering stems and foliage rendered in muted purples and soft greens. The shallow depth of field melts background shapes into a gentle bokeh while a warm golden highlight punctuates the scene, creating an intimate, contemplative atmosphere of quiet texture and fading light."
  • AI Description: A concise one-sentence summary — "A moody close-up of purple blooms and sage-green leaves with soft bokeh and a warm gold accent."
  • Alt Text: Accessibility text — same as the concise description

That's five data points from a single image, and we haven't gotten to the technical analysis yet.

Content Classification

  • Tags: flower, plant, foliage, leaf, wildflower, garden, purple flower, greenery, texture, bokeh, close-up, soft light — subject-based tags
  • AI Keywords: Everything in tags plus colors, moods, and composition terms — muted purple, sage green, olive green, gold, shallow depth of field, soft focus, minimalist, serene, moody, nostalgic, dreamy, natural light
  • Mood: serene, moody, nostalgic, dreamy — emotional descriptors
  • Composition Style: close-up, shallow depth of field, bokeh, soft focus — technical composition analysis

The AI didn't just identify what was in the photo. It identified how it was composed, what mood it evokes, and what visual techniques were used. These aren't tags a photographer would manually add — they're analytical observations.

Color Science

This is where the pipeline gets genuinely sophisticated:

  • Primary Colors: muted purple, sage green, olive green, warm gold — named colors in plain English
  • Palette Hex: #9481a0, #c5b1d9, #647c5b, #8ba27d, #7d8258 — five dominant hex values
  • Palette LAB: Five LAB color values — perceptually uniform color space coordinates like {L: 56.6, a: 13.2, b: -13.8}
  • Palette Weight: 0.3, 0.2, 0.18, 0.17, 0.15 — the proportional representation of each color in the image
  • Palette Mean LAB: {L: 59.7, a: 0.19, b: 1.24} — the average color, for similarity matching
  • Palette Version: 1 — algorithm version tracking

Why LAB instead of just hex? Because LAB is perceptually uniform. Two colors with the same Euclidean distance in LAB space look equally different to the human eye. Hex distance doesn't work that way. If you want to build a "similar colors" feature — or match photos to a collection by palette — you need LAB. The hex is for display. The LAB is for computation.

The weights matter too. A photo that's 30% muted purple and 15% warm gold is different from one that's 15% purple and 30% gold, even if the palette contains the same colors. The weight vector is how you tell the difference.

Technical Data (EXIF)

  • Camera Model: X-T3
  • Lens: 65.0 mm
  • Aperture: f/1
  • Shutter Speed: 1/125
  • ISO: 800
  • Focal Length: 65mm
  • Date Taken: 2026-08-05T09:09:38
  • Orientation: square

All extracted from the image's EXIF data — the metadata embedded by the camera at capture time. No AI needed for this part, but the pipeline handles it in the same pass.

SEO and GEO

  • SEO Priority: 0 (defaults to zero, can be manually boosted for featured images)
  • Alt Text: Already generated as part of the identity pass
  • Location Name/Region/Country: Available when GPS data is present in EXIF — reverse-geocoded to human-readable place names

Commerce

  • Available Digital: true (can be sold as a digital download)
  • Price Digital: $10 (default, adjustable per photo)
  • Digital File URI: Path to the high-res file for fulfillment
  • Watermarked URL: Dynamically generated via wsrv.nl with parameters for width (1024), format (webp), and quality (80) — the watermark and optimization happen at request time, not at upload time

That last point is worth pausing on. The watermarked URL isn't a stored file — it's a dynamic request to an image processing service: https://wsrv.nl/?url=...&w=1024&output=webp&q=80. Every display-quality image is generated on demand from the original. No pre-processing. No storage of multiple sizes. The URL encodes the transformation.

The Schema as Architecture

Count the fields: title, slug, description, AI description, alt text, tags, AI keywords, mood, composition style, primary colors, palette hex, palette LAB, palette weight, palette mean LAB, palette version, camera model, lens, aperture, shutter speed, ISO, focal length, date taken, orientation, GPS lat, GPS lng, location name, location region, location country, SEO priority, available digital, price digital, digital file URI, watermarked URL, collection name, collection IDs.

That's 34 fields. Most of them are populated autonomously — by AI vision analysis, EXIF parsing, color extraction, or reverse geocoding. A few (SEO priority, price, availability) are configurable defaults that can be manually overridden.

The schema isn't just a list of fields. It's an architecture where each field exists because something downstream needs it.

The Data Dance: One Record, Many Systems

Here's where the "elegant data dance" happens. Every enriched photo record feeds at least a dozen downstream systems — all from a single AI pass:

Site Search runs against title, description, tags, and AI keywords. A search for "purple flower" hits the tags (purple flower) and the AI keywords (muted purple), and the description (mentions "purples"). One query, multiple matching fields.

Color Filters use primary_colors for named-color browsing ("show me purple photos") and palette_hex for visual swatches. But the actual similarity computation uses palette_lab and palette_mean_lab — because two photos can both be "purple" in completely different parts of the purple spectrum.

Mood Filters use the mood array. "Show me dreamy photos" is a filter on mood: ["dreamy"]. The AI generated the mood from the visual content — no human had to tag it.

Composition Filters use composition_style. "Show me close-ups with bokeh" filters on composition_style: ["close-up", "bokeh"]. Again, AI-generated from the image.

Smart Collections are the most elegant part. Each collection defines keywords and match_fields — which fields to match against. For example:

  • "Perfectly Purple" matches keywords: ["purple"] against match_fields: ["primary_colors"] — any photo where a named primary color contains "purple" auto-joins this collection.
  • "Nature's Anatomy" matches keywords like anatomy, vascular, veins, venation, cells, cellular, tissue, histology, microscopic against match_fields: ["tags", "ai_keywords"] — any photo the AI tagged with those terms auto-joins.
  • "Mycopia" matches keywords: ["mushroom", "fungus"] — self-organizing collection of all fungal photography.
  • "Travel" matches keywords: ["architecture"] — architectural photography from travel.

No manual assignment. The AI generates the tags, the collection rules match against the tags, and the photos land in the right galleries automatically. Add a new photo of a mushroom? It's in Mycopia before you think about it.

SEO runs on title, description, and alt text — all AI-generated and optimized for the actual content of the image. The slug is clean and descriptive. When GPS data is available, location name, region, and country add GEO signals for local search relevance.

Visual Similarity uses the LAB palette. Two photos with similar mean LAB values and similar palette weight distributions look similar to the human eye. You can compute Euclidean distance in LAB space and rank by proximity. "More like this" becomes a math problem, not a tagging problem.

Commerce uses available_digital, price_digital, and digital_file_uri for the checkout flow. The watermarked URL handles display. The original file handles fulfillment.

Image Optimization is handled at the URL level. The watermarked_url field stores a wsrv.nl transformation URL that encodes width, format, and quality. Need a different size? Change the URL parameters. No image processing pipeline to maintain.

Layout uses orientation — landscape, portrait, or square — to make grid layout decisions. CSS handles the rest.

The Versioning Question

The palette_version field is easy to miss but important. It tracks which version of the color extraction algorithm produced the palette. When the algorithm improves — better clustering, more accurate dominant colors, finer weight distribution — you can reprocess photos and know which ones have the old version. That's the difference between a pipeline and a script. A pipeline is designed to evolve.

The same principle applies to location_version and iptc_version. The pipeline doesn't assume it got everything right the first time. It versions its outputs so they can be selectively reprocessed.

Why This Matters Beyond Photography

Here's the thing that makes this more than a photography tool: the same pattern applies to any data lake where raw assets need structured enrichment.

Replace "photos" with "documents." Replace "Google Drive" with "SharePoint." Replace "AI vision analysis" with "AI document analysis." The architecture is the same:

  1. Ingest from an existing store (Drive, SharePoint, S3, whatever)
  2. Enrich with AI — generating the fields that downstream systems need
  3. Structure into a schema where every field has a purpose
  4. Deploy the enriched record across search, filters, categorization, SEO, and commerce

The hard part isn't the AI. Vision models can describe images. Language models can extract entities from text. The hard part is the schema design — figuring out which fields you need, which fields the AI can generate reliably, and how those fields connect to the systems that consume them.

Most companies trying to "leverage AI in their existing data lakes" start with the AI and work backward to the data. They run an LLM over their documents and get unstructured summaries. Then they try to figure out what to do with them. The summaries sit in a field nobody queries.

The photography pipeline works the other way. It starts with the downstream systems — search, filters, collections, SEO, commerce — and asks: what data does each system need? Then it asks: can the AI generate that data? Then it builds the schema around the answers.

The result is a record where every field is load-bearing. Nothing is generated "just in case." Nothing sits unused. The AI's output isn't a summary — it's the structural data that makes the system work.

The Sync Exclusion Layer

One more piece worth mentioning: the SyncExclusion entity. Not every file in Drive should end up on the storefront. The exclusion layer tracks files that were deliberately kept out — by title and Drive file ID. This is the filtering mechanism between the raw archive and the public-facing catalog.

It's a small thing, but it represents an important architectural decision: the sync is inclusive by default, and exclusion is explicit. That's the right direction. You'd rather have too much in the pipeline and filter down than too little and miss things.

What I'd Tell Someone Building This

Start with the schema. Not the AI, not the pipeline, not the sync. The schema. Map every downstream system that will consume the data. For each system, list the fields it needs. For each field, decide: can the AI generate this? Can it be extracted from the source? Does it need manual input? Then build the pipeline to populate the schema.

Version your AI outputs. The algorithm will improve. You'll want to know which records have the old version.

Use the right data format for the job. Hex for display. LAB for computation. Named colors for user-facing filters. Raw coordinates for computation. Human-readable place names for display. Don't try to make one format do everything.

And don't let the AI generate fields nobody uses. Every field in the schema should have a consumer. If it doesn't, delete it. The pipeline is elegant because every data point does work — not because there are a lot of them.

The dance isn't in the AI. It's in the architecture.

AIdata pipelinephotographyDrive syncschema designinteroperabilitySEOcolor scienceLAB color spaceDW8 Photographysmart collectionsdata lakes