What causes bad product data in B2B ecommerce
Bad B2B product data comes from ERP item masters built for orders, supplier files in many formats, free-text specs, no data owner, and cleanups that decay.

Bad product data in B2B ecommerce is mostly caused upstream of the website: ERP item masters built to process orders rather than answer buyer questions, supplier data arriving in dozens of incompatible formats, specs stored as free text instead of structured attributes, and no one owning the data once it is loaded. The site only shows the symptoms (empty filters, 15W next to 15 Watts, a spec sheet two revisions out of date), so the fix has to land where data is received, structured, and maintained, not on the product page.
In the MasterB2B 2026 State of B2B eCommerce report, practitioners named data cleanliness and hygiene as their single biggest barrier to growth for the second year in a row. The more useful question is which cause produces which symptom, because each one has a different fix.
What causes bad product data in B2B ecommerce: six root causes
1. The ERP item master was built for transactions
An ERP record exists to buy, stock, price, and invoice an item: part number, unit of measure, cost, and a short description that fits on a pick ticket. It does not need voltage_rating or thread_size as typed fields, so it usually lacks them. When the ERP feeds the website directly, the site inherits that shape. Elastic Path describes the result as a recurring cycle to pull from ERP, scrub manually, reformat, load, and repeat for every catalog update. We cover the structural argument in why the ERP item master is not a catalog.
Symptom on the site: titles like BRKR 2P 20A QO PLUG and product pages with a description but no spec table.
2. Supplier files arrive in a dozen formats
A distributor does not have one data source. It has hundreds of suppliers, and each one sends data its own way. The NAED product data journey study, built on more than 60 interviews at 42 electrical industry companies, found Excel files from many manufacturers requiring manual review, QC, integration, and testing. It also found that when a manufacturer changes its Excel template, the distributor has to re-automate the import.
Wholesalers outside electrical report the same pattern. Novomind notes that many suppliers use their own catalogues and classifications, with different designations for the same attribute, and that data arrives as Excel, PDF, or CSV.
Symptom on the site: one category where half the SKUs filter cleanly and half show nothing, split along supplier lines.
3. Specs live in free-text fields
When there is no structured slot for an attribute, people type it into whatever field exists. The NAED study lists the results participants actually saw: Volts vs V, 15W vs 15 Watts, BLK vs Black, similar products filed in different categories, and a unit price supplied for a case quantity.
Each is a correct value in a format the next system cannot compare. Classification standards such as ETIM exist to solve this by giving each product class a uniform set of technical features with predefined values, but a standard only helps once someone maps each supplier's free text onto it.
Symptom on the site: a "Color" filter offering both BLK and Black, and a wattage facet with 40 options for what should be 8.
4. Nobody owns the data
The NAED research found product data roles commonly decentralized across product divisions, and many distributors reported not knowing who to contact at a manufacturer for product data. Most distributors reported no or only rudimentary measurement of data quality. GreyBath's drift audit lists ownership and process failures, such as nobody knowing who owns or approves a field, among the causes of values slowly diverging across ERP, website, PDFs, and partner portals.
Symptom on the site: a value that is wrong in the same way for months, even though a rep has corrected it in a quote a dozen times.
5. Cleanups are one-time projects, and data decays
The usual response is a cleanup project before a launch or replatform. It works, then decays from day one. NAED participants describe what decay looks like: products retired with no notice, spec sheets changed and never sent, price changes received after the new price took effect, and links to product data that no longer work.
For illustration only: a distributor carrying 50,000 SKUs where 5% of records change in a year has 2,500 records going stale annually with no process to catch them. Three years after a cleanup, roughly 15% of the catalog could be out of date without anyone noticing.
Symptom on the site: a downloadable PDF that contradicts the spec table above it.
6. Every channel gets its own rewrite
Marketplaces, punchout, print, and the website each want a different shape, so teams edit copies instead of the source. Akeneo's 2024 B2B survey found that 40% of companies say keeping product data consistent across channels is one of their biggest challenges. Every local rewrite is a fork that the next supplier update will not reach.
Symptom on the site: the same SKU with different dimensions on your site and on your marketplace listing.
B2B product data quality problems mapped to their fix
| Root cause | What buyers see | Where the fix lives |
|---|---|---|
| ERP item master shape | Cryptic titles, no spec table | A product-content layer separate from the ERP |
| Supplier format sprawl | Facets that work for some brands only | Normalization at intake, mapped to one attribute model |
| Free-text specs | BLK and Black as separate filter values | Typed attributes with controlled values and units |
| No owner | Errors that never get fixed | Named owners plus completeness and accuracy KPIs |
| One-time cleanup | Stale PDFs, retired items still live | Ongoing re-sourcing against current manufacturer docs |
| Channel rewrites | Different specs on different channels | Edit once at source, transform per channel on output |
The three classic quality complaints, incomplete, inaccurate, and inconsistent product information, are not separate problems. Incomplete data usually traces to causes 1 and 5, inaccuracy to 4 and 5, and inconsistency to 2, 3, and 6.
Supplier product data quality: why distributors inherit the mess
Manufacturers often write product copy once and send the same file to many distributors, so those distributors inherit the same gaps, and the distributor's own data model rarely matches. NAED participants describe a distributor data model that still has gaps after manufacturer data is populated, because the two models were never synchronized.
Industry syndicators and data pools reduce the format problem. In the GS1 world, the Global Data Synchronization Network connects trading partners through certified data pools, with the GS1 Global Registry matching subscriptions to registered items identified by GTIN, and AHRMM describes it as a way to synchronize core product data to standard specifications. Where your suppliers publish through GDSN or a channel data pool, use it. But a feed carries what the manufacturer put into it. If the filterable attributes your category needs were never in the source, synchronization delivers the gap faster.
How distributors handle bad supplier product data
The distributors who keep data usable tend to do four things:
- Define the attribute model per category first. Decide which fields a buyer filters on for, say, a ball valve (
size,end_connection,body_material,pressure_rating) and make those required. - Accept any format, normalize at intake. Forcing every supplier into a template fails when the template changes. Mapping inbound data to your model, with units and controlled values, holds up better.
- Source missing values from documents, not guesses. The spec sheet, catalog page, or manufacturer site usually has the value. Record which document it came from, so a conflict can be reviewed instead of silently overwritten.
- Measure and maintain. Track completeness by category and supplier, re-check against current manufacturer sources on a schedule, and feed the gaps back to the supplier. Our guide to product data governance covers owners and KPIs in detail, and the distributor product data statistics page collects published figures on how large the gap tends to be.
The labor is the hard part. NAED interviewees reported a minimum of 0.5 to 1.5 full-time staff dedicated to product data improvement, with most of that time going to manually received data. That is where the "who does the work" question gets decided: offshore data entry, a generic AI description writer, or an ongoing, source-grounded enrichment practice.
Your PIM stores the data; Anglera does the work. It runs alongside whatever PIM, ERP, or syndication platform you already have, or from a flat CSV export, and extracts attribute values from real spec sheets and catalogs, attaches a confidence score and source to each one, and flags conflicts for review rather than inventing values. A first enrichment pass is typically live in 30 days or less.
How to decide where to start
Pull one high-revenue category and audit 50 SKUs against the six causes above. If the gaps cluster by supplier, start at intake normalization. If they cluster by attribute, fix the model. If the data was clean two years ago and is not now, you need maintenance, not another project. The last-mile ownership argument explains why distributors should own that layer rather than wait for suppliers.
Bad product data is a supply-chain problem that surfaces on a web page. Anglera does the ongoing work in the middle: turning supplier documents into structured, sourced attributes in the systems you already run, and keeping them current as suppliers change.
