Product Data

Fashion Product Data Quality

Data quality is not a feeling that the catalogue looks clean; it is a set of dimensions you can measure and hold to a standard.

Updated 2026-07-01 · 8 min read

It is easy to say a catalogue has good data and hard to prove it. Data quality only becomes actionable when it is broken into dimensions that can each be defined, measured and improved independently — completeness, accuracy, freshness, consistency and coverage. Treated as a single vague virtue, quality drifts. Treated as measurable dimensions, it becomes something a discovery platform can manage like any other part of its infrastructure.

This page frames data quality as a discipline rather than an aspiration. It defines the dimensions that matter for fashion discovery, explains how each can be measured, and connects the practice to the problems it exists to solve, from structural messiness to offer quality and the AI systems that increasingly sit on top, such as shopping assistants.

Why quality has to be a discipline

Fashion product data degrades continuously and silently. Merchants change feed formats, brands rename models, prices and stock move by the hour, and every one of these events introduces error that no single check would catch. A platform that treats quality as a one-time clean-up will find its catalogue quietly rotting, because the inputs never stop changing.

Making quality a discipline means measuring it continuously against defined dimensions, setting thresholds, and treating a drop on any dimension as an incident to investigate rather than noise to ignore. It is the difference between hoping the data is good and knowing, at any moment, exactly how good it is and where it is weakest.

Completeness

Completeness asks how much of the data that should exist actually does. For footwear that means: do listings carry a style code, a resolvable model, a colourway, a size grid, a current price and a working image. Missing fields do not just leave gaps — they push downstream systems onto weaker signals, so an incomplete listing is harder to match and less trustworthy to show.

Completeness is measured as the share of records that carry each field, tracked per merchant so that a feed which stops sending a field is caught quickly. It is directional: some fields matter far more than others, so completeness is weighted toward the fields — style code, size-level stock, current price — that most affect matching and trust.

Accuracy

Accuracy asks whether the data that exists is correct. A field can be present and complete but wrong: a mismatched image, a mislabelled colour, a stale price, a size marked available that is not. Accuracy is the hardest dimension to measure because it requires a source of truth to compare against, and for much of fashion data there is no single authority.

In practice accuracy is estimated through cross-checks and consistency signals — does the image agree with the colourway, does the style code agree with the model, does the claimed list price agree with observed price history — and through sampling against reliable references. Perfect measurement is impossible, but relative accuracy per merchant and per field is both measurable and actionable.

Freshness and consistency

Freshness asks how recently the data was confirmed, and it dominates the volatile fields — price and per-size stock above all. A platform measures freshness as the age distribution of its most volatile fields and holds it to a threshold, suppressing or flagging data that has aged past the point of reliability. For fast-moving fields, fresh-but-approximate beats precise-but-stale.

Consistency asks whether the same thing is represented the same way everywhere. It is the dimension most directly attacked by normalization: are brands, models, colours and sizes expressed in the canonical vocabulary, or has an unmapped variant slipped through. Inconsistency fractures products, breaks filters and undermines matching, so it is measured by looking for un-normalised or conflicting values that should have been reconciled.

Coverage and measuring the whole

Coverage is the outward-facing dimension: does the catalogue actually contain the products and merchants shoppers are looking for. A catalogue can be complete, accurate, fresh and consistent yet still fail its users because it is missing the shoes or the shops they care about. Coverage is measured against demand — the models, colourways, sizes and regions shoppers search for — and against the set of merchants worth carrying.

Taken together these dimensions give a measurable picture of catalogue health:

  • Completeness — share of records carrying each important field.
  • Accuracy — share of fields that agree with cross-checks and references.
  • Freshness — age distribution of volatile fields against thresholds.
  • Consistency — rate of un-normalised or conflicting values.
  • Coverage — how well the catalogue meets actual demand.

Tracked continuously and per merchant, these turn data quality from a vague claim into an operational metric that a discovery platform can defend, improve and hold suppliers to.

Note. The dimensions trade off. Chasing maximum coverage by admitting low-quality feeds can lower accuracy and consistency, so quality management is about balancing the dimensions, not maximising any one.

Sources & further reading

  1. Google Search Central, “Product data quality guidelines” (2024)
  2. McKinsey, “Analysis of retail data and analytics” (2023)

Sources are attributed to their publishers and link to each publisher's own site. Figures reflect general market direction rather than point-in-time precision; consult the linked publishers for their current data.