All posts
Ray Iyer
Ray Iyer
Co-founder, Anglera

Stop Cleaning Your Product Data to Get Ready for AI. AI Is the Cleanup Crew.

Why waiting for clean product data before deploying AI gets the sequence backwards: for catalogs, AI enrichment is the cleanup, not the reward for it.

Stop Cleaning Your Product Data to Get Ready for AI. AI Is the Cleanup Crew.

For three years the trade press has run the same sequencing argument on a loop: clean your data first, or start with AI and accept the risk. Both camps are answering the wrong question when the domain is the product catalog. AI isn't the thing you earn the right to use after a cleanup project. It is the cleanup project, and unlike most AI deployments, you can audit every row it touches before it ships.

The debate as the trade press has framed it

Distribution Strategy Group has taken both sides of this within the same year, which tells you something about how unresolved the question actually is. In "AI Is an Amplifier, Not a Shortcut," the argument is that AI is a mirror. It magnifies whatever mess already exists in your CRM, so you fix the fields, kill the duplicates, and standardize the hierarchy before you let a model near it. Three months later, in "The Data Perfection Myth," the position flips: cleanup has no finish line, competitors who started six months ago are already banking value, and a good chunk of modern AI, large language models specifically, was built to work through inconsistency rather than choke on it.

Both pieces are right about their own subject matter. CRM records feed a decision system: a forecast, a route, a comp calculation. Feed that system garbage and the AI doesn't fix the garbage, it launders it into a confident-looking wrong answer faster than a human would have. That's a real risk, and "clean first" is the correct call for anything where the AI's output is a number you're about to act on blind.

But the catalog isn't that kind of system, and neither piece grapples with why.

Why the catalog inverts the sequence

A CRM record or a forecast input is consumed downstream by another automated process. Garbage in one field corrupts a calculation nobody re-checks. A catalog attribute (a dimension, a material spec, a compliance flag) is consumed by a human or an AI shopping agent making a single, visible decision on a single, visible product page. The stakes and the review surface are not the same at all.

That difference is what makes enrichment auditable in a way forecasting and pricing AI simply are not. When an LLM extracts a voltage rating from a spec sheet PDF and writes it into an empty attribute field, you can check that one cell against that one source document in seconds. When a forecasting model adjusts thirty thousand SKUs' safety stock based on a demand signal nobody can point to, there's no comparable row-level trace to check. Distribution Strategy Group's own framing in "AI Is an Amplifier," cited above, actually concedes this point by reaching for CRM and pricing as its examples, never the catalog. The amplifier metaphor holds for systems where AI acts on aggregated, structure-dependent data. It breaks down for systems where AI is filling in one missing fact at a time against a source you can point to.

Put differently: the risk of "AI on messy data" scales with how much the AI's output is trusted without a human ever looking at it again. Catalog enrichment is close to the opposite end of that spectrum from algorithmic pricing.

Who benefits from the prerequisite

It's worth asking why "clean your data before AI" became gospel in the first place, given how selectively it actually applies. Most of the vendors repeating it sell something that sits downstream of the catalog: pricing engines, CRM platforms, forecasting tools. For those tools to work, somebody upstream has to have already fixed the product master, the customer master, the sales history. That somebody was never going to be the pricing vendor. Telling a distributor "get your data clean, then we'll talk" is a rational thing for a downstream AI vendor to say. It's also, not coincidentally, a way to push the hard part of the deployment back onto the customer indefinitely.

Call it an incentive rather than a conspiracy. But it means the clean-data prerequisite gets applied as a blanket rule to a domain, the catalog itself, where the actual mechanics don't support it. Nobody selling pricing AI benefits from pointing out that catalog cleanup is a solvable, auditable, AI-native problem you could knock out in weeks instead of a governance initiative you fund for years.

What the numbers around this actually look like

The evidence for urgency sits outside any one vendor's pitch. A 2026 survey of more than 400 distribution leaders by NAW and Modern Distribution Management found 73% of leaders expected measurable AI results, and only 16% had actually gotten them, with item masters cluttered by duplicates and dead SKUs called out as a specific drag. The same research found practitioners at a NAW symposium ranking organizational resistance above data readiness as the real blocker, which undercuts the idea that distributors need years of cleanup before they can start anything. IBM's 2025 Institute for Business Value research put the annual cost of poor data quality above $5 million for more than a quarter of organizations surveyed. And on the customer-facing side, a 2026 report from Salsify and 1Point1 attributes roughly a fifth of online returns directly to inaccurate product descriptions. That's the downstream cost of the exact attribute gaps that sit unfilled while a distributor debates sequencing.

Anglera's own Top Distributors 2026 index measures this directly. The Digital Readiness Index scores 200+ distributors against 14 live-site signals, and the fill-rate gaps it surfaces are not spread evenly by archetype. They cluster exactly where you'd expect: in the long-tail SKUs nobody has had headcount to touch.

Where this leaves the operator

Don't wait to deploy AI on the catalog until the catalog is clean. Deploy AI because the catalog is the cleanup. Run it against your existing PIM, and check its work row by row against source documents the way you'd check any new hire's first week. Not because the technology is infallible, but because catalog enrichment is one of the rare AI use cases where "trust but verify" is cheap and fast instead of theoretical.

That's the wedge Anglera is built around: your PIM keeps storing the data, and the enrichment layer does the work of filling it in, attribute by attribute, with every change traceable back to a source. No twelve-month cleanup phase required before the AI touches anything. The cleanup phase is the AI, running now, audited as it goes.

Ray Iyer

About the author

Ray IyerCo-founder, Anglera

Ray is a co-founder of Anglera, building the product-data infrastructure for agentic commerce — turning messy catalogs into structured, AI-readable data that buyers and answer engines can find. Previously product at Uber; Stanford CS.

See it on your own SKUs.

A 30-minute walkthrough on your categories and your supplier data.

Book a demo