All benchmarks

Distributor product data statistics, measured

Measured August 2026 · 37 distributors scored from a universe of 223 · edition 2026

The short answer

These figures come from scoring live product pages at the largest North American distributors against a published framework, not from a survey and not from vendor self-reporting. The pattern they describe is consistent: catalogs are online and machine-illegible. The median measured distributor publishes roughly a dozen structured attributes per product page, a minority expose any product identifier a machine could match on, and fewer still publish Product structured data. Every number below is computed directly from the underlying dataset, which is published in full with a named company list, so any figure here can be checked against the companies it came from.

Most product data statistics in circulation have no method behind them. They get cited, recycled, and eventually appear in AI answers with a confidence nothing in their origin supports.

These are measured. The sample is named, the framework is published, the raw dataset is downloadable, and the sample size sits next to every number. It is a specific sample — large North American distributors with a public catalog — and it should not be read as a statement about distribution generally.

What the measurement found

Shares are of the measured sample. Companies without a public catalog cannot be scored and are reported separately.

58/100
Median Digital Readiness Index
37 distributors measured

Scored against live product pages, not self-reported.

11
Median structured attributes per product page
37 distributors measured

The count a machine can actually read as fields — not bullet points in a description.

35%
Publish any GTIN on a product page
13 of 37

Without an identifier, a product cannot be matched to the same product anywhere else.

22%
Publish Product structured data (JSON-LD)
8 of 37

The single clearest signal to a retrieval system that a page describes a product.

70%
Publish a discoverable sitemap
26 of 37
14%
Block AI crawlers outright
5 of 37

A deliberate choice to be absent from answer engines.

45 pts
Largest gap between best and worst category on one site
widest observed

Catalog quality is rarely uniform — the average hides the categories that are failing.

17
Have no public catalog to measure at all
of 223 in the universe

Large distributors whose products are not visible to any search or answer engine.

Where the points are actually lost

Median score per pillar as a share of the points available. The shape matters more than the total.

Product Data Depthmedian 18.3 of 35 · 52%

Can a machine tell what this product is and match it to the same product elsewhere?

Buyer Answerabilitymedian 12 of 25 · 48%

Does the page answer what a buyer actually asks before they commit?

Commerce Transparencymedian 13 of 20 · 65%

Can a buyer find out what it costs and whether it ships, without asking a human?

Machine & Agent Readinessmedian 13 of 20 · 65%

Can a crawler, a marketplace, or an AI shopping agent actually consume any of it?

Why attribute count is the number to watch

Of everything measured, structured attribute count per product page is the figure that predicts the most downstream behaviour. It determines whether faceted search can narrow a catalog, whether a marketplace listing will pass validation, whether a comparison is possible, and whether a retrieval system can answer a question about the product without guessing.

The distinction that matters is between attributes and prose. A page can describe a product at length and expose almost nothing a machine can read as a field. Descriptions are counted by humans and ignored by facets.

The identifier problem underneath everything

A product without a GTIN, MPN or any published identifier cannot be reliably matched to the same product in a marketplace, a competitor's catalog, or a retrieval system's index. Every downstream capability — syndication, comparison, marketplace listing, AI citation — depends on that match being possible.

It is also the cheapest gap to close, because identifiers usually exist somewhere in the business. They sit in purchasing records and supplier files, and they were simply never published on the page.

How to compare yourself against this

Take your top five categories by revenue and count, on a sample of live product pages, how many attributes a machine could read as structured fields — not bullet points, not sentences. Then check whether any identifier is published, and whether the page carries Product structured data.

That gives you three numbers directly comparable to the ones above. Most teams find the exercise uncomfortable, because internal completeness reporting measures against an internal schema and this measures against what a buyer or a machine can actually use.

Method, and how to check this

Every figure on this page is computed from the published dataset rather than written into the text, so it cannot drift from the measurement behind it. Scores come from live product pages sampled at each company, graded against a published framework. Groups with fewer than 3 measured companies are suppressed rather than published thin.

Frequently asked questions

Where do these product data statistics come from?

From scoring live product pages at the largest North American distributors against a published Digital Readiness framework. The company list, the framework, the measurement date and the raw dataset are all published, and the sample size appears next to every figure.

How many distributors were measured?

The measured sample is smaller than the universe, because a company needs a public, sampleable catalog to be scored at all. Both numbers are shown above — the measured count and the number with no public catalog to measure.

Is this a representative sample of distribution?

No, and it should not be read that way. It is the largest North American distributors with a public catalog. Smaller distributors are not represented, and the figures would likely look different for them.

What is a good attribute count for a product page?

It depends on the category rather than on a universal number. The useful test is coverage of the attributes buyers actually filter and search on in that category, which is usually a smaller and more specific set than the one a manufacturer publishes.

Can we use these figures in our own material?

Yes, with attribution and the sample size stated. The dataset is published under an open licence with a citation line, precisely so the numbers can be checked rather than merely repeated.

Keep reading

See where your catalog lands

We'll score a sample of your live product pages against the same framework and show you the number next to your sector — including the categories your internal reporting averages away.

Book a measurement