Does your commerce platform predict product data quality?
Measured August 2026 · 37 distributors scored from a universe of 223 · edition 2026
The short answer
Commerce platform is a much weaker predictor of product data quality than platform vendors imply. Grouping measured distributors by their detected commerce platform produces medians that sit close together, and the spread inside each group is wider than the gap between groups — meaning the same platform hosts both strong and weak catalogs. The one grouping that does separate cleanly is bot wall configuration: sites behind an aggressive CDN or bot wall are frequently unreadable to retrieval systems regardless of how good the underlying data is, which converts an infrastructure decision nobody framed as a content decision into exactly that. Platform detection is published only where storefront fingerprints and vendor corroboration agree, so these groups are small and deliberately conservative.
A recurring question in platform selection is whether one commerce platform produces better product data outcomes than another. The measurement says: not much, and less than the sales cycle suggests.
What follows are the detected groups with at least three measured companies. The samples are small — platform detection is only published where two independent signals agree — so treat these as directional.
By commerce platform
Detected from storefront fingerprints and published only where corroborated. Groups under three are suppressed.
| Group | n | Median DRI | Highest in group |
|---|---|---|---|
| Shopify | 4 | 62.5 | — |
| SAP Commerce Cloud (Hybris) | 5 | 61 | — |
| No known vendor (custom or in-house build) | 17 | 54 | — |
| Storefront never readable (bot wall) | 6 | 53.5 | — |
By storefront CMS
Directly observable from the rendered storefront.
| Group | n | Median DRI | Highest in group |
|---|---|---|---|
| Adobe Experience Manager | 3 | 59 | — |
| No CMS signature (commerce platform or custom front end) | 28 | 58 | — |
| WordPress | 5 | 50 | — |
By CDN and bot wall
The grouping with the clearest practical consequence for answer engines.
| Group | n | Median DRI | Highest in group |
|---|---|---|---|
| Cloudflare | 16 | 58 | — |
| Akamai | 3 | 57 | — |
| No bot wall observed | 13 | 49 | — |
The bot wall is a content decision nobody made deliberately
Of the three groupings here, bot wall configuration has the most direct practical consequence. A catalog behind an aggressive challenge is not readable by retrieval systems, which means the quality of the data underneath is irrelevant to whether it can ever be cited.
This is almost always an infrastructure decision, made for good reasons — scraping, competitive price harvesting, load. It is rarely evaluated as what it also is: a decision about whether your products appear in AI answers. Some distributors in the measured universe could not be sampled at all for this reason, and those companies are invisible to any answer engine for the same reason they were invisible to us.
What platform choice does and does not decide
Platforms differ in how easily they render structured data, how well they handle deep attribute sets, and how much engineering is needed to publish a clean product page. Those are real differences.
What they do not decide is whether the attributes exist. Every platform in these groups can publish a rich, machine-readable product page, and every one of them also hosts catalogs that publish almost nothing — because the data was never there to publish. Replatforming a thin catalog produces a thin catalog on newer software, which is the most expensive way to learn this.
Method, and how to check this
Every figure on this page is computed from the published dataset rather than written into the text, so it cannot drift from the measurement behind it. Scores come from live product pages sampled at each company, graded against a published framework. Groups with fewer than 3 measured companies are suppressed rather than published thin.
Frequently asked questions
Which commerce platform is best for product data?
The measurement does not support a clear answer, and that is the finding. Medians across detected platform groups sit close together and the spread inside each group is wider than the gaps between them.
How was the platform detected?
From storefront fingerprints, published only where a second independent signal corroborates them. Where the two disagree, or where only one exists, the company is reported as detected or unrecognised rather than confirmed.
Does a bot wall hurt AI visibility?
Yes, directly. If a retrieval system cannot fetch the page, the product cannot appear in an answer regardless of how complete the data is. Several large distributors are unreadable for this reason.
Will replatforming improve our product data?
Only the rendering of it. A replatform moves existing data onto new software; it does not create attributes that were never collected. Teams frequently discover this at go-live.