All posts
Amay Aggarwal
Amay Aggarwal
Co-founder, Anglera

What team do you need to run a PIM? Roles and headcount for 100,000 SKUs

Running a PIM takes five roles: a product data owner, data stewards, catalog specialists, a PIM admin and an integration owner. Headcount hinges on enrichment.

What team do you need to run a PIM? Roles and headcount for 100,000 SKUs

Running a PIM takes five roles: a product data owner who sets the standards, one or more product data stewards who enforce them, catalog or content specialists who fill in the values, a PIM administrator who runs the system, and an integration owner who keeps data flowing in and out. For a 100,000-SKU catalog, the governance core is a handful of people whatever you do; the headcount swing comes from who does the enrichment, and in the worked example below in-house keying needs about 21 FTE-years for the initial backfill while automated extraction with human review needs well under one FTE-year of review time.

What team you need to run a PIM, role by role

Product data owner. A business leader, usually from merchandising, category management or product management, who decides what "complete" means for each category and approves exceptions. Hypotenuse describes the owner as the person who sets the standards for an attribute group and approves exceptions. At 100,000 SKUs you usually have one per category family (fasteners, electrical, janitorial), as a part-time duty rather than a hire.

Product data steward. The full-time operator of the governance rules. Covered in detail below.

Catalog or content specialists. The people who find the spec sheet, read it, and enter voltage, thread_pitch, material and pack_quantity into the PIM. In Akeneo's collaboration model, contributors draft changes and send proposals that product owners approve or reject, and a reviewer can only approve changes in the attribute groups and locales they have edit rights on. This is the role whose size scales with SKU count, and the one most operating-model decisions are really about.

PIM administrator. Owns the system itself: users, permissions, families, attribute definitions, validation rules, workflows. WisePIM's role glossary defines the administrator as the person who manages user accounts, settings, and how data moves in and out of the system. At 100,000 SKUs this is typically one person, sometimes combined with the steward.

Integration owner. Keeps the ERP, supplier feeds, ecommerce platform and marketplace syndication connected. Hypotenuse calls this the technical owner who runs imports, integrations and the connections to each channel. Usually sits in IT or ecommerce engineering and is shared with other systems.

During implementation you add two temporary roles. Start With Data's implementation team list includes an executive sponsor with budget authority and a project manager. Those step back after go-live; the five above do not.

What a product data steward does day to day

The owner decides that every motor SKU needs frame_size, hp and enclosure_type. The steward makes the PIM enforce it and works the queue when it does not.

Public job postings are the clearest description of the job. Regal Rexnord's posting for a product data steward asks the role to set standards for product hierarchies, material types, units of measure, commodity codes, build scorecards for completeness, accuracy, uniqueness and timeliness, and keep product data synchronized across ERP, MDM and PIM systems. TVH's technical product data steward posting adds two duties that matter for headcount: enforcing data standards, policies, and guidelines established by the Data Owner, and coordinating the allocation of data tasks among data maintenance staff.

In practice, the steward's week looks like this:

  • Maintains picklists and units: stainless steel vs SS vs 304 SS, in vs inch.
  • Writes and tunes validation rules so a voltage of 1200 on a light bulb gets caught at import.
  • Works the exception queue: conflicting values between the ERP and the supplier sheet, duplicates, missing mandatory fields.
  • Reports fill rate by category to the data owners and decides which gaps get worked first.
  • Directs the specialists, or the vendor, or the automation, and samples their output.

That last point is the data quality job in one line. The steward does not type every value; the steward knows which values are wrong. Start With Data notes stewards are generally pulled from merchandising, product management, or marketing, which is right: they need to know that a 3/8-16 bolt description is a thread spec, not a typo. AtroCore notes stewardship is often distributed across people who also hold other titles, which works until the exception queue outgrows side-of-desk time.

For the policy layer that sits above the steward (who owns which attribute group, what the thresholds are, how exceptions escalate), see our product data governance guide.

How many people it takes to manage product data for 100,000 SKUs

There is no reliable public benchmark for FTE per SKU, so here is the arithmetic. Replace every assumption with your own.

Assumptions for illustration:

  • 100,000 active SKUs, about 25 category-specific attributes each that need real values.
  • Manual enrichment from a spec sheet or manufacturer page: 20 minutes per SKU.
  • Each year: 15,000 new SKUs plus 10,000 SKUs needing a revision at 10 minutes each.
  • 1,600 productive hours per FTE per year after meetings, training and leave.

Initial backfill, fully manual: 100,000 SKUs at 20 minutes is 2,000,000 minutes, or about 33,300 hours. At 1,600 hours per FTE that is roughly 21 FTE-years.

Ongoing, fully manual: 15,000 new SKUs at 20 minutes is 5,000 hours. 10,000 revisions at 10 minutes is about 1,670 hours. Together roughly 6,670 hours, or about 4 FTE, every year, indefinitely.

One public reference point: Codesoltech suggests that implementation projects for catalogs over 100,000 SKUs staff a project team of 4 to 6 roles, including 2 to 3 data enrichment specialists. Run that through the same math: three specialists at 1,600 hours is 4,800 hours, which covers about 14,400 SKUs a year at 20 minutes each. That staffing level works if most SKUs arrive already enriched or something else does the extraction. It does not cover a 100,000-SKU backfill by hand.

Product data team headcount under three operating models

The governance core stays roughly constant across models. What changes is who fills the 25 attributes, and how much review time that creates.

Operating modelWho fills attributesIn-house effort, backfillIn-house effort, ongoing per yearCore team
Fully manualYour specialists~33,300 hrs (~21 FTE-years)~6,670 hrs (~4 FTE)owner time, 1 steward, 1 admin, part-time integration
Outsourced data entryOffshore or BPO team~1,670 hrs QA plus vendor management~420 hrs QA plus vendor managementsame, steward spends more time on spec writing and sampling
Automated extraction with human reviewSource-grounded extraction, people review flagged values~1,000 hrs (~0.6 FTE-years)~250 hrssame

How the last two rows are built:

  • Outsourced: assume your steward QA-samples 10% of delivered SKUs at 10 minutes each. That is 10,000 SKUs and about 1,670 hours for the backfill, and 2,500 SKUs or about 420 hours a year ongoing, plus spec writing and rework. Pivotree frames its outsourced service around taking the SKU-by-SKU grind off the plate while strategy stays in-house, which matches this split.
  • Automated with review: assume 15% of SKUs carry a low-confidence or conflicting value that a person checks, at 4 minutes each. That is 15,000 SKUs and 1,000 hours for the backfill, and about 3,750 of the 25,000 yearly events, or 250 hours, ongoing.

Change any assumption and the totals move, but the shape holds: manual scales headcount with SKU count, the other two scale review time with error rate.

What goes wrong with each staffing plan

Manual teams fall behind on the long tail. New SKUs get worked first, so revisions and older SKUs never reach the top of the queue. Knowledge about which supplier mislabels amperage lives in two people's heads, a problem we cover in tribal knowledge in catalog attributes.

Outsourced teams deliver what the spec says, not what the buyer needs. Ambiguous attribute definitions produce consistent wrong answers at scale.

Automation without review invents values. A description generator can call a cable plenum-rated with nothing to back it. Extraction is only useful if every value traces to a source document and conflicts get flagged. We cover this in what to check before trusting your PIM's AI button.

Nobody owns the steward role. The PIM goes live, the project team disbands, and the exception queue has no owner.

How to decide on your PIM team structure

  1. Count SKUs that need attributes filled today, and SKUs added or changed per year.
  2. Time 50 real SKUs end to end across a few categories. Use that number, not 20 minutes.
  3. Decide who fills values: your team, a vendor, or extraction with review.
  4. Staff the governance core regardless, and make the steward a named, full-time role once the catalog passes a few tens of thousands of SKUs.
  5. Pick the PIM last. It stores whatever the team produces; our PIM software comparison covers the options.

This is where Anglera fits. Your PIM stores the data; Anglera does the enrichment work, extracting values from spec sheets, catalogs and manufacturer pages, quality-scoring each one and flagging conflicts for your steward, working alongside whatever PIM, ERP or flat file you already run. Implementation is typically about 30 days, and the team you keep is the governance core described above. See how Anglera works for the review flow.

Frequently asked questions

What is the difference between a product data owner and a product data steward?

The product data owner is a business leader, usually in merchandising or category management, who sets the standard for what complete data looks like in a category and approves exceptions. The product data steward enforces those standards day to day by maintaining picklists and validation rules, working the exception queue, and reporting data quality metrics such as completeness and accuracy. Owners decide; stewards make the PIM follow the decision.

How many catalog specialists do I need for 100,000 SKUs?

It depends on minutes per SKU and who does the extraction. As an illustration, at 20 minutes per SKU and 1,600 productive hours per FTE, a manual backfill of 100,000 SKUs is about 21 FTE-years, and 15,000 new plus 10,000 revised SKUs a year is about 4 FTE ongoing. If extraction is automated and people only review flagged values, the same work can drop to well under one FTE-year of review time.

Can the PIM administrator and data steward be the same person?

In smaller catalogs they often are, since both roles work on attribute definitions and validation rules. As the catalog grows, the exception queue and supplier conflicts usually take a steward's full time, while the administrator handles users, permissions, workflows and system configuration. One practical signal to split them is data quality work waiting behind system tickets.

Who should be on a PIM implementation team?

Beyond the ongoing roles, implementation adds an executive sponsor with budget authority and a project manager who turns the plan into milestones. The data steward defines the attribute model and taxonomy, the technical lead or integration owner designs data flows to the ERP and channels, and a user advocate checks that workflows match how people work. The temporary roles step back after go-live.

Amay Aggarwal

About the author

Amay Aggarwal — Co-founder, Anglera

Amay is a co-founder of Anglera, where he's building the AI pipeline that turns messy supplier catalogs into structured, AI-readable product data for distributors and answer engines. He built the catalog AI systems at Uber Eats on top of research from Stanford's AI lab.

See it on your own SKUs.

A 30-minute walkthrough on your categories and your supplier data.

Book a demo