How MRO distributors manage product data for millions of SKUs
MRO distributors run millions of SKUs by classifying each item, applying a per-class attribute template, automating supplier intake and enriching continuously.

MRO distributors manage product data for millions of SKUs by classifying every item into a standard class (UNSPSC, ECLASS, ETIM or an internal taxonomy mapped to them), then filling a per-class attribute template from manufacturer documents and supplier feeds. The ones that keep up treat it as a standing operation: automated supplier intake, validation rules at the door, and continuous enrichment of the long tail, not a one-time catalog cleanup.
The scale MRO distributors are actually managing
The numbers in public filings explain why nobody does this by hand. Grainger's 2025 annual report, as summarized from the 10-K, puts its High-Touch business at approximately 2 million products, Zoro at roughly 13 million and MonotaRO at roughly 29 million, and says more than 5,000 primary suppliers provide more than 1.5 million products. The filing notes products are regularly added and removed.
The churn matters more than the count. In a Databricks customer story, Grainger is described as handling 400,000 product changes a day across about 2.5 million products. RS Group told suppliers in its May 2022 supplier newsletter that it aims to host millions of supplier products across its e-commerce channels. HD Supply, which Home Depot reacquired in December 2020, is listed through the OMNIA Partners cooperative as offering over 150,000 MRO, janitorial and property management products.
At that size the question is not how to fix the catalog but how to keep a catalog that changes daily from decaying.
How MRO distributors manage product data: the operating model
Strip away vendor names and the large catalogs share five layers. Public sources describe the layers; they do not describe each company's internals, so treat this as the pattern, not a blueprint of any one distributor.
| Layer | What it decides | Typical reference |
|---|---|---|
| Classification | Which class a SKU belongs to | UNSPSC, ECLASS, ETIM, internal tree |
| Attribute template | Which fields that class must carry, in which units | ECLASS properties, ETIM features |
| Supplier intake | How manufacturer data enters | Feeds, APIs, BMEcat, spreadsheets |
| Validation | What gets rejected or flagged | Picklists, unit rules, required fields |
| Continuous enrichment | How gaps and changes get filled | Spec sheets, catalogs, manufacturer sites |
Classification: UNSPSC, ECLASS and ETIM do different jobs
Taxonomy normalization starts here, and the three standards are not interchangeable.
UNSPSC is a code tree. It has four levels: segment, family, class and commodity, each two digits, producing an 8-digit code. It is common in procurement; GS1 US managed it through the end of 2024, and the UNDP has handled revisions and change requests since. It tells you that a part is a ball bearing. It does not tell you which bore diameter field to fill.
ECLASS classifies and describes. Its standard overview cites about 48,000 product classes and more than 23,000 unique properties across four classification levels, with an eight-digit class code, and it is built for catalog exchange formats such as BMEcat.
ETIM is a classification model for technical products. Its model documentation defines each class by predefined features of four types: A alphanumeric with a fixed value list, L logical yes/no, N numeric, and R range, with units attached to numeric and range features.
In practice, a distributor keeps one internal browse taxonomy and maps each node to whichever standards its customers and suppliers use: UNSPSC for a procurement customer's invoice lines, ETIM for a manufacturer's feed. One SKU record then serves both.
Attribute templates per class
The template is where most MRO data quality is won or lost. A template for a ball bearing class might require bore_diameter, outside_diameter, width, seal_type, material, dynamic_load_rating and max_speed, with fixed units and a picklist for seal_type so 2RS, rubber sealed and double sealed collapse into one value.
ETIM's model makes the principle explicit: every element is predefined. That is what makes faceted search and cross-reference work. Grainger's product information lead frames the goal as accuracy, consistency and completeness, and the same piece cites gray versus grey as the kind of inconsistency that confuses buyers when validation rules do not enforce a standard. Multiply that across every color, thread and voltage field.
We go deeper on which fields matter for fasteners, bearings, motors and safety products in MRO and industrial attributes.
Supplier onboarding: how Grainger, RS Group and HD Supply take in data
The public record shows the big distributors moving away from the supplier spreadsheet.
- Grainger historically relied on manual spreadsheet submissions and is shifting toward automation, API integrations and validation rules, according to the same Salsify interview, with what its team calls a "supplier tool belt" to meet suppliers at different levels of digital maturity. That interview cites 1.5 million products from more than 3,000 suppliers; counts differ across sources and dates, so read them as orders of magnitude.
- RS Group asks suppliers to send product data in whatever format their system provides, says automated sharing gets products live four times faster than its spreadsheet process, and targets launch within 24 hours.
If you searched for a "Grainger product data API": the public sources above describe API-driven integration on the supplier intake side, not the terms of any catalog access for third parties. Suppliers should confirm the current requirements in each distributor's supplier portal, because those change.
Content providers and the long tail
No distributor has the people to read every spec sheet for 2 million SKUs. Many buy a baseline from content providers. Distributor Data Solutions, for example, describes connecting to manufacturer data sources and delivering structured, normalized content into the ERP, PIM or e-commerce systems a distributor already runs.
Syndicated content covers the brands that publish good data. The long tail is what is left: small manufacturers, private-label items, discontinued-but-still-ordered parts, and SKUs whose only source is a scanned PDF. That tail is where completeness drops. The economics are covered in long-tail SKU economics.
MRO product data enrichment in practice
Enrichment for MRO is extraction plus verification. A useful sequence:
- Anchor on the manufacturer part number. Normalize the manufacturer name and MPN first (strip spaces, dashes and prefixes consistently) so every downstream source can be matched to the right SKU.
- Pull from the primary document. Spec sheet, catalog page, manufacturer product page. Record which document each value came from.
- Map to the template. Convert units, apply picklists, put
1/2"-13 UNCinto the thread field rather than the description. - Score and flag. When two sources disagree on a load rating, flag it for a person instead of picking one.
- Re-run on change. New SKUs, new required attributes and supersessions all trigger the same loop.
ISO 8000 is the standard reference for the quality side. It defines requirements for exchanging master data between business partners, including portability (an agreed syntax and explicit semantic encoding), with part 110 covering characteristic data exchange, part 115 quality identifiers and part 120 provenance. Recording the source per value is the practical version of provenance.
For illustration: a 1,000,000-SKU catalog where 30 percent of items are missing three template attributes is 900,000 empty fields. At two minutes of manual lookup per field, that is 30,000 hours, before a single new SKU arrives. That arithmetic is why the work has to be automated and source-grounded rather than staffed.
Where large MRO catalogs go wrong
- Classifying without templates. A UNSPSC code on every SKU looks complete in a dashboard and still leaves search filters empty.
- Project thinking. A cleanup that ends leaves the catalog decaying at the rate of daily changes.
- Invented values. Generative tools that write plausible specs without a source create errors that are harder to find than blanks.
Who does the enrichment work
The PIM, ERP or MDM is where the record lives. Someone still has to read the spec sheet and fill the field. Options are internal staff, an offshore data-entry team, syndicated content where it exists, or a maintained enrichment layer.
That last option is what Anglera is. Your PIM stores the data; Anglera does the work. It extracts values from real manufacturer documents, quality-scores each one, flags conflicts for review, and writes back into whatever system you already run, starting from a flat CSV or ERP export if that is all you have. You can see the mechanism on how it works, and how it fits industrial catalogs on Anglera for industrial and MRO distributors. If you are still choosing the system of record itself, our PIM guide for industrial and MRO distributors compares the options.
How to decide where to start
Pick the 20 classes that drive the most search traffic or revenue. Lock a template for each, with units and picklists. Measure fill rate per required attribute, not per SKU. Then point enrichment at the gaps, starting with the long tail that syndicated content does not reach.
Large MRO catalogs stay usable because classification, templates and enrichment run as one ongoing practice. Anglera works alongside your existing PIM or ERP to do the extraction and backfill, with implementation measured in weeks rather than quarters, so new attributes and new SKUs get filled from source documents instead of waiting in a queue.
