All posts
Ray Iyer
Ray Iyer
Co-founder, Anglera

How to measure product data accuracy (and audit outsourced data entry)

Measure product data accuracy by checking a random sample of attribute values against a source of truth, then report the error rate with a confidence interval.

How to measure product data accuracy (and audit outsourced data entry)

You measure product data accuracy by checking a random sample of attribute values against a defined source of truth for each attribute, then dividing correct values by values checked. Report it as a field-level error rate with a confidence interval, broken out by supplier, vendor, and category, because one catalog-wide percentage hides where the wrong values live.

Accuracy is not completeness

IBM defines accuracy as how well data represents real-world entities and whether it can be validated against trusted sources, and completeness as whether all required values are present (IBM). The two move independently. Fill rate is cheap to measure: count the non-empty cells. Accuracy needs what Apollon calls an external truth, a reference the value can be checked against.

That gap is where catalogs get hurt. A voltage column that is 95% filled looks finished on a dashboard. If a chunk of those values were copied from a sibling SKU's spec sheet, the dashboard cannot tell you, and the buyer finds out at install.

A two-by-two quadrant of fill rate versus accuracy. High fill with low accuracy is the dangerous quadrant: the catalog looks 95 percent complete but the values are quietly wrong.

We go deeper on that split in fill rate vs accuracy.

Data quality dimensions, and where accuracy sits among them

IBM lists six dimensions: accuracy, completeness, consistency, timeliness, validity, and uniqueness (IBM). ISO 8000-8 frames it differently, as syntactic quality (data conforms to its specified syntax), semantic quality (data corresponds to what it represents), and pragmatic quality (data is suitable for a particular purpose). The standard says syntactic and semantic quality are measured through verification and pragmatic quality through validation (ISO 8000-8 preview; summary at arc42).

The practical translation for a product catalog:

  • Validity, or syntactic checks, run on 100% of records automatically: GTIN check-digit logic, value ranges, unit plausibility (Apollon). A gtin with a bad check digit or a weight of 0 fails without anyone opening a PDF.
  • Accuracy, or semantic checks, ask whether max_operating_pressure: 150 psi is actually true of that valve. A rule cannot answer that. ISO 8000-8 notes that a complete check of correspondence is in many cases close to impossible, and that sampling and statistical methods can give a probabilistic view instead (ISO 8000-8 preview).

That note is the case for a sampling audit. Valid is not the same as true: 24 VDC passes every format rule on a 120 VAC contactor.

How to measure product data accuracy in five steps

1. Define the truth source for each attribute

Decide what "correct" means before anyone scores, including the match rule. Ruling on "close enough" after the fact is how audits drift.

AttributeTruth sourceMatch rule
gtin, mpnBrand owner's record or the packageExact
height, width, depth, gross_weightManufacturer spec sheet or physical measurementWithin tolerance
voltage, amperageSpec sheet or nameplateExact after unit normalization
material, finishSpec sheetMatch to your controlled vocabulary
costCurrent supplier price listExact

For package dimensions, GS1 Canada flags discrepancies against a global standard tolerance of plus or minus 4% (GS1 Canada). For prices, stock, and identifiers, Inventory Source recommends checking against the latest supplier price list, inventory system, and source records.

2. Draw a random, stratified sample

Pull the sample at random from what was delivered, not the first 100 rows of the file. Stratify by supplier and category so each slice you plan to report on gets enough checks. Inventory Source suggests mixing categories, high-value items, and frequently updated SKUs into a controlled sample before approving a full feed.

Include blanks. If you only sample filled fields, you never catch a value that exists on the spec sheet but was left empty.

3. Score each field, then roll up

Mark every checked value as correct, wrong, wrongly blank, or wrongly filled (a value where the attribute does not apply). Keep wrong and missing as separate counts so accuracy and completeness stay separate numbers.

Then report two rates. ARDEM's SLA framing is useful here: field-level accuracy as the share of fields captured correctly, by criticality, and record-level accuracy as the share of records with zero critical-field errors (ARDEM). Apollon also advises stating which share was checked automatically and which by manual sampling.

4. Compute the error rate with a confidence interval

A point estimate from a small sample is a guess. For a proportion, NIST's handbook gives the Wilson interval and notes that Agresti and Coull recommended it for virtually all combinations of sample size and proportion (NIST).

For illustration: you sample 150 SKUs from a vendor batch and check 8 critical attributes each, which is 1,200 field checks. You find 36 wrong values, a 3.0% error rate. The 95% Wilson interval is roughly 2.2% to 4.1%. Now run the same 3% on only 100 checks (3 errors): the interval widens to roughly 1.0% to 8.5%. That sample cannot tell a good vendor from a poor one.

One caveat. Errors inside a single SKU are correlated (same keyer, same spec sheet), so a field-level interval is somewhat optimistic. Report the SKU-level rate next to it.

5. Track by supplier, vendor, and category over time

Report the rate per source, not just per catalog. Tag each error with a cause, because each one points to a different fix:

  • Transcription: 3/4 in keyed as 3/8 in.
  • Unit: millimetres stored in an inches field.
  • Wrong source: a value taken from a sibling SKU or the wrong revision of the sheet.
  • Stale: the manufacturer revised the spec and nobody re-checked.
  • Unsupported: no source document backs the value at all.

How to check the quality of outsourced product data entry

Treat each delivered batch as a lot you accept or reject. You draw the sample, not the vendor. An independent reviewer checks it against the same source document the keyer used, which means the vendor has to record the source for every value. If nobody can say which page of which PDF a value came from, you can audit the format but not the truth.

ExtendedDesk recommends splitting the work so one person enters and another validates, and logging errors by category to find training or form-design problems (ExtendedDesk). Both fit inside the audit above. Our comparison of outsourced data entry covers the wider trade-offs.

How to write accuracy into an outsourcing SLA

A clause that says "99% accurate" with no method attached is hard to enforce. Spell out:

  • Definition. Field-level accuracy by criticality tier, plus record-level accuracy on critical fields.
  • Truth sources and match rules, as an appendix (the table from step 1).
  • Sampling. Who draws it (you), sample size per batch and per stratum, and a logged random seed.
  • Acceptance rule. Judge the batch on the upper bound of the 95% interval, not the point estimate, so a lucky small sample cannot pass a bad batch.
  • Rework. A failed batch is corrected and re-sampled at the vendor's cost.
  • Traceability. ARDEM lists user-level change logs, reviewer identity, reason codes for overrides, and before/after correction logs.

On targets, ARDEM publishes directional ranges of 99.2% to 99.8% for structured forms and 97.0% to 99.2% for multi-source packets. Attribute work from spec sheets and catalogs usually resembles the second. Set your number from a baseline audit of your own data, not a vendor brochure.

What goes wrong in accuracy audits

  • Circular checks. The reviewer compares the value to the same spreadsheet the keyer produced instead of the source document.
  • One audit, then never again. Spec sheets get revised, so last year's correct value can be wrong today.
  • One number for the whole catalog. A 2% average can hide a single supplier at 12%.
  • No cause codes. You know the rate but not whether to retrain, fix a form, or replace a source.

For a scorecard that combines accuracy with the other dimensions, see how to measure product data quality and our post on scoring product data quality.

Who does the work, and how that changes the audit

The audit is only as cheap as the provenance behind each value. This is where Anglera fits: your PIM stores the data, and Anglera does the work of extracting attribute values from spec sheets, catalogs, and manufacturer sites, quality-scoring them, and flagging conflicts for review instead of inventing a value. It works alongside your PIM, ERP, or a flat CSV export, and typical implementation is about 30 days.

When each value carries a pointer to its source document, the sampling audit above becomes closer to a lookup than an investigation. Whether the work is done in-house, by a BPO, or by Anglera, ask for the same thing: a random sample, a named truth source, and an error rate with its interval.

Hero photograph by EqualStock on Unsplash

Frequently asked questions

What is the difference between product data accuracy and completeness?

Completeness asks whether a required value is present; accuracy asks whether that value is true for the product. A field can be filled and wrong, so fill rate cannot stand in for accuracy. Accuracy has to be checked against an outside reference such as the manufacturer spec sheet or the supplier price list.

What are the main data quality dimensions for a product catalog?

IBM lists six: accuracy, completeness, consistency, timeliness, validity, and uniqueness. ISO 8000-8 groups quality as syntactic, semantic, and pragmatic, measuring the first two by verification and the third by validation. In practice, validity checks can run automatically on every record, while accuracy needs a sampled comparison against source documents.

How big should a data accuracy audit sample be?

Big enough that the confidence interval is narrow enough to make a decision. For illustration, a 3 percent error rate on 1,200 field checks has a 95 percent Wilson interval of roughly 2.2 to 4.1 percent, while the same rate on 100 checks spans roughly 1 to 8.5 percent. Size each supplier or category stratum you intend to report on separately.

What accuracy should an outsourced data entry SLA require?

Set the target from a baseline audit of your own data, and define it as field-level accuracy by criticality plus record-level accuracy on critical fields. ARDEM publishes directional ranges of 99.2 to 99.8 percent for structured forms and 97.0 to 99.2 percent for multi-source packets. Judge each batch on the upper bound of the error-rate interval rather than the point estimate.

Ray Iyer

About the author

Ray Iyer — Co-founder, Anglera

Ray is a co-founder of Anglera, building the product-data infrastructure for agentic commerce — turning messy catalogs into structured, AI-readable data that buyers and answer engines can find. Previously product at Uber; Stanford CS.

See it on your own SKUs.

A 30-minute walkthrough on your categories and your supplier data.

Book a demo