Case Study
Natalya Rugs — shipment-to-storefront pipeline
Aug–Sep 2026
A Python pipeline that turns a supplier's invoice spreadsheet and a Drive folder of photos into published Shopify listings. Deterministic code handles size, price and SKU; AI handles colour, style and copy — and every AI answer has to clear a hard validator before it reaches the store.
The AI calls were the easy part. The hard part was that nothing coming in could be trusted. The invoice labels 1,014 of 1,100 rugs the same style, and reports centimetres in a column marked metres. Photo folders carry typos in the rug number. Some rugs have no photos at all. So the spreadsheet is the source of truth, photos are matched on the 4-digit rug number with an override list for mislabelled folders, and the invoice's weaving style reaches the model as a hint only — it publishes if the photos confirm it, otherwise the field is left blank and the rug is flagged.
The split is deliberate: anything with a right answer is plain Python under test, and the model is only asked to make judgment calls. Four LangGraph steps — vision, copywriter, SEO, reviewer — produce colour, style and copy, and that output has to pass hard validators (banned jargon, no words implying wear or age, a sentence-length cap, colour-naming rules) or the product lands in a needs_review queue instead of the live store. The vision step also has to discard photos of the rug's back, which carry a paper label and skew the colour read. Dropping them by position didn't work — they turn up mid-sequence — so the model identifies them by weave and reports what it ignored.
Every red rug is rust. The first version handed the model a fixed list of 64 shade names, and "rust" was the only red on it, so the entire red half of the catalogue collapsed onto one word while crimson and oxblood — the shades people actually search for — weren't available at all. Opening up the vocabulary traded that for a different drift: mood words like "warm" and "faded" in place of colours. The fix validates the shape of the answer rather than the words. "Muted blue" passes, "muted" alone doesn't, and "faded rose" is rejected because it describes wear, not colour. A shade-frequency report now prints how often each name appears, so the next collapse of this kind shows up as a number rather than as a customer complaint.
Re-runnable at full scale. A SQLite ledger plus a live SKU check lets an interrupted run resume without creating duplicates, and --refresh rewrites copy on products already in the store while never touching price, photos or stock — or putting a sold rug back on sale.
Stack
- Python
- LangGraph
- Claude
- Pydantic
- Typer
- SQLite
- Shopify GraphQL API
- pytest
Private client work. The LLM provider is swappable per step — Anthropic API, Claude Agent SDK, Hugging Face, or a local Ollama model.