Distributor product data statistics, measured
Measured August 2026 · 37 distributors scored from a universe of 223 · edition 2026
The short answer
These figures come from scoring live product pages at the largest North American distributors against a published framework, not from a survey and not from vendor self-reporting. The pattern they describe is consistent: catalogs are online and machine-illegible. The median measured distributor publishes roughly a dozen structured attributes per product page, a minority expose any product identifier a machine could match on, and fewer still publish Product structured data. Every number below is computed directly from the underlying dataset, which is published in full with a named company list, so any figure here can be checked against the companies it came from.
Most product data statistics in circulation have no method behind them. They get cited, recycled, and eventually appear in AI answers with a confidence nothing in their origin supports.
These are measured. The sample is named, the framework is published, the raw dataset is downloadable, and the sample size sits next to every number. It is a specific sample — large North American distributors with a public catalog — and it should not be read as a statement about distribution generally.
What the measurement found
Shares are of the measured sample. Companies without a public catalog cannot be scored and are reported separately.
Scored against live product pages, not self-reported.
The count a machine can actually read as fields — not bullet points in a description.
Without an identifier, a product cannot be matched to the same product anywhere else.
The single clearest signal to a retrieval system that a page describes a product.
A deliberate choice to be absent from answer engines.
Catalog quality is rarely uniform — the average hides the categories that are failing.
Large distributors whose products are not visible to any search or answer engine.
Where the points are actually lost
Median score per pillar as a share of the points available. The shape matters more than the total.
Can a machine tell what this product is and match it to the same product elsewhere?
Does the page answer what a buyer actually asks before they commit?
Can a buyer find out what it costs and whether it ships, without asking a human?
Can a crawler, a marketplace, or an AI shopping agent actually consume any of it?
Why attribute count is the number to watch
Of everything measured, structured attribute count per product page is the figure that predicts the most downstream behaviour. It determines whether faceted search can narrow a catalog, whether a marketplace listing will pass validation, whether a comparison is possible, and whether a retrieval system can answer a question about the product without guessing.
The distinction that matters is between attributes and prose. A page can describe a product at length and expose almost nothing a machine can read as a field. Descriptions are counted by humans and ignored by facets.
The identifier problem underneath everything
A product without a GTIN, MPN or any published identifier cannot be reliably matched to the same product in a marketplace, a competitor's catalog, or a retrieval system's index. Every downstream capability — syndication, comparison, marketplace listing, AI citation — depends on that match being possible.
It is also the cheapest gap to close, because identifiers usually exist somewhere in the business. They sit in purchasing records and supplier files, and they were simply never published on the page.
How to compare yourself against this
Take your top five categories by revenue and count, on a sample of live product pages, how many attributes a machine could read as structured fields — not bullet points, not sentences. Then check whether any identifier is published, and whether the page carries Product structured data.
That gives you three numbers directly comparable to the ones above. Most teams find the exercise uncomfortable, because internal completeness reporting measures against an internal schema and this measures against what a buyer or a machine can actually use.
Method, and how to check this
Every figure on this page is computed from the published dataset rather than written into the text, so it cannot drift from the measurement behind it. Scores come from live product pages sampled at each company, graded against a published framework. Groups with fewer than 3 measured companies are suppressed rather than published thin.
Frequently asked questions
Where do these product data statistics come from?
From scoring live product pages at the largest North American distributors against a published Digital Readiness framework. The company list, the framework, the measurement date and the raw dataset are all published, and the sample size appears next to every figure.
How many distributors were measured?
The measured sample is smaller than the universe, because a company needs a public, sampleable catalog to be scored at all. Both numbers are shown above — the measured count and the number with no public catalog to measure.
Is this a representative sample of distribution?
No, and it should not be read that way. It is the largest North American distributors with a public catalog. Smaller distributors are not represented, and the figures would likely look different for them.
What is a good attribute count for a product page?
It depends on the category rather than on a universal number. The useful test is coverage of the attributes buyers actually filter and search on in that category, which is usually a smaller and more specific set than the one a manufacturer publishes.
Can we use these figures in our own material?
Yes, with attribution and the sample size stated. The dataset is published under an open licence with a citation line, precisely so the numbers can be checked rather than merely repeated.