Solutions / Catalog auto-tagging

Product catalog auto-tagging (vision)

Attributes are read straight from product photos and supplier documents — colour, material, pattern, shape, category — so filters and search stop failing on the half of your catalog nobody had time to tag.

E-commerce Vision Data
Whole catalog tagged
including what arrived untagged
Attributes from the image
not from the title
Confidence per tag
so review goes where it matters
Who it is for

E-commerce teams whose filters, search and recommendations underperform because most products are missing most attributes.

Short answer

A vision model reads your product images and supplier documents and fills the attributes your taxonomy expects, with a confidence score on each one. High-confidence tags go straight in, uncertain ones queue for a person, and new arrivals are tagged on ingest instead of waiting for someone to get to them.

The problem

Filters only work on the products someone had time to tag, which is never the new ones.

01

Untagged products are invisible

A product with no colour or material never appears under a filter, so it does not sell, so nobody prioritizes tagging it. The loop is self-sealing and quietly expensive.

02

Manual tagging cannot keep up with intake

Every new supplier drop arrives faster than the team can process it. The backlog only grows, and the freshest products are the worst covered.

03

Inconsistent values break the filters that exist

Navy, dark blue and blue-navy become three filter values, splitting the same products into three groups and making the filter look broken to a shopper.

How it works

From trigger to result, step by step.

01

Start from your taxonomy, not the model’s

Your attribute list and allowed values per category. Tags that do not match your controlled vocabulary are useless downstream, however accurate they look on their own.

02

Read the image and the documents together

Vision handles colour, pattern, shape and visual style; supplier PDFs and spec sheets carry material, dimensions and compliance data. Combining both fills far more fields than either alone.

03

Score confidence and route the uncertain

Every tag carries a score. Above your threshold it applies automatically, below it queues for review, and every correction feeds back so the threshold can be raised over time.

04

Normalize to controlled values

Free text is mapped to your vocabulary — navy becomes the one value you actually filter on — and new values are proposed for approval rather than created silently.

05

Run on ingest, then backfill

New products are tagged as they arrive so the backlog stops growing; the existing catalog is backfilled in batches, highest-traffic categories first.

Before / after

What changes on the ground.

Today, by hand
×Most of the catalog missing most attributes
×New arrivals waiting weeks to be tagged
×Three spellings of the same colour
×Filters and search quietly underperforming
With the automation running
Attributes filled across the whole catalog
New products tagged on ingest
Values normalized to your controlled vocabulary
Review time spent only on low-confidence tags
What you get

Delivered, not demoed.

A vision and document pipeline mapped to your existing taxonomy
Attribute extraction from images, PDFs and supplier spec sheets
Confidence scoring with a review queue for uncertain tags
Normalization to controlled values with proposals for new ones
Write-back to your PIM or storefront, on ingest and as a backfill
Built with

We build in your stack rather than moving you onto ours. The list below is what this solution most often connects to — other systems are a scoping question, not a blocker.

Claude OpenAI Vision Python PostgreSQL Akeneo Shopify n8n Elasticsearch
Time to production4–8 weeks
Build priceFixed price
First stepFree mini-audit
Honest limits

When this is not the right solution.

·If you have no taxonomy, build one first. Tags without controlled values create a second mess on top of the first, and we will say so before starting.
·Attributes that are simply not visible or documented — thread count, origin, certifications absent from the paperwork — cannot be inferred, and guessing them is a compliance risk.
·With a few hundred products, a person tags the catalog in a week and the project is not worth running.

Questions we get about this one

Visual attributes like colour, pattern and shape are strong; fine material distinctions are weaker and usually come from documents instead. We measure per attribute on your own catalog before rollout.

Bring us the process that hurts.

The mini-audit is free: we take your version of this process apart and tell you plainly whether automating it pays. If it does, you get a scope and a fixed price.