
The Unorthodox Angle
The biggest AI cost savings don't come from switching models — they come from restructuring the call pattern. One image, one call, one structured output. The pipeline that asks the model three times is paying three times, and getting a slower result for the privilege.
The Data Dance covered what the pipeline generates — 30+ data points per photo from a single enrichment pass. This is the story of how it got there, and why the first version was burning tokens it didn't need to burn.
If you're building any system where AI processes data at scale, the architecture of your LLM calls matters more than which model you use. I learned this the hard way.
The Original Architecture
The first version of the Drive sync pipeline worked. It pulled photos from Google Drive, ran them through AI, and populated the schema. The output quality was good. The problem was the cost — in tokens, in time, and in rate limits.
Here's what the call pattern looked like:
- Visual analysis call — The model looks at the image and describes what's in it. This generates the foundational understanding that downstream calls depend on.
- Caption call — A separate LLM call takes the visual analysis output and generates a title, description, and alt text.
- Color call — Another separate LLM call analyzes the image for dominant colors and generates named color descriptors.
- Tag call — Another call generates subject tags, keywords, mood, and composition style.
These were running as parallel GPT model calls — which sounds efficient, but wasn't. Here's why.
Problem One: The Same Image, Sent Three Times
Each of those calls is a vision request. The model has to see the image to analyze it. When you send a photo to a vision model, the image is tokenized — typically at a cost of hundreds to low thousands of tokens per image, depending on resolution.
When you make three separate calls that each need to see the image, you're paying the image token cost three times. The model is looking at the same photo three times and generating three separate pieces of analysis when it could have generated all of them in a single pass.
At 4,463 photos, that's 4,463 × 3 = 13,389 redundant image tokenizations. The model saw the same images three times and was billed for each viewing.
Problem Two: The Wrong Order of Operations
The visual analysis call ran before color and tagging. That sounds logical — understand the image, then describe it. But it created a dependency chain that forced a second pass.
Here's what was happening: the visual analysis call would produce a description. Then the color call would run — but to generate meaningful color names (not just hex values, which are computed separately), it needed the context of what was in the image. So either it re-analyzed the image from scratch, or it consumed the visual analysis output and made another call.
Either way, you're paying for two passes over the same image when one would have been enough. The fix was counterintuitive: don't separate visual analysis from enrichment. Ask the model to do everything in one structured response — identify the subject, generate the description, name the colors, assign the mood, identify the composition style, generate the tags. One call. One image. One pass.
The ordering wasn't just inefficient — it was architecturally wrong. The pipeline was treating "understand the image" and "enrich the image" as separate operations when they're the same operation. The model doesn't need to describe an image before it can identify its colors. It sees everything at once. The sequence was a human mental model imposed on a system that doesn't work sequentially.
Problem Three: The Default Model Trap
The pipeline was using the "default" model setting, which auto-selects the "best available" model. On most platforms, that means the most capable — and most expensive — model in the lineup. For a pipeline processing 4,463 images, that's a significant cost multiplier.
The "best available" model is designed for complex reasoning, nuanced analysis, and high-stakes generation. Photo enrichment is none of those things. It's a structured extraction task: look at an image, generate a title, a description, some tags, some color names, a mood assessment. A mid-tier model does this just as well as the flagship — and costs a fraction per call.
The difference between a top-tier and mid-tier vision model for this task is marginal in quality and massive in cost. The top-tier model might write slightly more poetic descriptions. But the pipeline doesn't need poetry — it needs accuracy, consistency, and structured output. A mid-tier model delivers all three.
After the refactor, switching from "default" to a specific mid-tier model reduced per-call cost significantly while maintaining the same output quality. The descriptions are still rich and usable. The tags are still accurate. The color names still make sense. The only thing that changed was the bill.
The Refactor: One Call, One Pass, One Model
The refactored pipeline does what the original should have done from the start:
- Single LLM call per photo — The model sees the image once and generates everything in one structured response: title, slug, description, AI description, alt text, tags, AI keywords, mood, composition style, and primary colors.
- No ordering dependency — There's no "visual analysis first, then enrichment" sequence. The call does both simultaneously because the model processes the whole image at once.
- Mid-tier model — Specifically selected for cost efficiency, not auto-selected for maximum capability.
- Structured output — The response is JSON-structured, not prose. The model returns fields directly, which means no parsing step and no wasted tokens on transitional language.
The non-LLM work — EXIF extraction, LAB palette computation, GPS reverse geocoding, orientation detection — runs algorithmically and was never the cost problem. The cost problem was purely in the LLM call pattern.
The Math
I'm not going to publish exact dollar figures because pricing changes and your mileage will vary, but the structure of the savings is clear:
Before:
- 3-4 LLM calls per photo (visual analysis + caption + color + tags)
- Each call sends the image (image token cost paid 3-4× per photo)
- Each call uses the top-tier model (highest per-token rate)
- Sequential dependency on visual analysis creates latency
- Rate limits on the top-tier model throttle throughput
After:
- 1 LLM call per photo (all enrichment in one structured response)
- Image sent once (image token cost paid 1× per photo)
- Mid-tier model (fraction of the per-token cost)
- No sequential dependency (one call, done)
- Higher rate limits on mid-tier models = faster throughput
The multiplier is roughly 10-15× on token cost per photo — from 3-4 expensive calls to 1 cheaper call. At 4,463 photos, that's the difference between a pipeline that's economically viable and one that isn't.
The Generalizable Lessons
These optimizations aren't specific to photography. They apply to any system where AI processes data at scale:
1. Consolidate calls, don't parallelize them
Parallel calls sound efficient — "run them all at once!" — but if each call needs the same input (the same image, the same document), you're paying the input token cost multiple times. One call with a structured output that covers all your enrichment needs is cheaper than N calls that each do one thing.
The instinct to split work into specialized calls comes from a software engineering pattern: single responsibility. But LLMs aren't microservices. They're general-purpose processors that can handle multiple tasks in a single pass. Forcing them to work one task at a time doesn't make the output better — it makes the pipeline more expensive.
2. Order matters when there's a dependency
If call B depends on the output of call A, you can't parallelize them — you have to sequence them. But if B's dependency on A is just "context that the model would have generated anyway if you asked it to do both at once," then the dependency is artificial. You created it by splitting the work.
The original pipeline ran visual analysis before color tagging because it assumed the color analysis needed the visual context. But the model sees the image. It has the context. You don't need to generate a description, pass it to another call, and ask for colors based on that description. You just ask for both.
3. Use the cheapest model that does the job
"Default" or "best available" is a trap for production pipelines. It's great for prototyping — you want the best output while you're figuring out what you need. But once you know what your pipeline does, you know what capability level it requires. And structured extraction from an image does not require the flagship model.
The test is simple: run 100 images through both the top-tier and mid-tier models. Compare the output. If the quality difference doesn't change how your downstream systems work, switch to the cheaper model. The 100-image test costs you a few dollars. The savings across 4,000+ images pay for it many times over.
4. Structured output is a cost optimization
When you ask an LLM for prose, it generates prose — with all the connecting words, transitions, and filler that prose requires. When you ask for JSON, it generates only the fields you specified. The token count is lower because there's no narrative overhead.
A response like "A close, painterly study of flowering stems and foliage rendered in muted purples and soft greens..." costs more tokens than a structured response that puts that same text in a description field alongside title, tags, mood, and colors in a single JSON object. The model is doing the same work — looking at the image and generating content — but the output format is more compact.
5. Cache aggressively
The sync system caches Drive file IDs. If a file hasn't changed since the last sync, it doesn't get reprocessed. This sounds obvious, but it's easy to forget when you're focused on the enrichment pipeline. The cheapest LLM call is the one you don't make.
At scale, incremental processing is the difference between a pipeline that costs $50 on the first run and $5 on subsequent runs, and one that costs $50 every time.
What I'd Build Differently Next Time
I'd start with one call per item. Not because I know it'll work — because I know it's the first thing to test. If the quality is good enough, you're done. If it's not, then you split — but you split with evidence, not with assumptions about what the model can and can't do in a single pass.
I'd pick a specific model from day one. "Default" is for exploration. Production pipelines need a known quantity.
And I'd design the output schema before writing the prompt. The schema tells you what you need. The prompt asks for exactly that. Nothing more, nothing less. The pipeline that asks for a summary it never uses is paying for tokens it doesn't need.
The refactor wasn't about making the pipeline smarter. It was about making it cheaper by asking the model to do less — or more precisely, by asking it to do the same work with fewer calls. The AI didn't change. The architecture around it did.