All posts
Ray Iyer
Ray Iyer
Co-founder, Anglera

AI Pilots Don't Die From Bad Models. They Die From Uncountable Output.

Pilots don't die from bad models, they die from output nobody can count. A selection test for which AI project actually ships first, and why.

AI Pilots Don't Die From Bad Models. They Die From Uncountable Output.

The keynote advice going around distribution conferences this year is "stop experimenting, move to production." It's not wrong, but it's incomplete, and the gap is expensive. It skips the one question that actually predicts whether a pilot ships: can you count what the model produces? Pilots don't die from bad models. They die from output nobody can audit, price, or defend in a budget meeting.

What the keynote got right, and what it left out

At the Applied AI for Distributors keynote in June, Graybar executives urged distributors to move AI projects into production, arguing that the real blockers are organizational, not technical — "your data, your processes and your people," not model quality. Distribution Strategy Group's takeaways piece framed it as execution replacing experimentation, a clean rhetorical turn that's been repeated at half the conferences since.

Both pieces are right that change management, not model capability, is usually the binding constraint. Neither says how to pick which pilot gets the change-management budget in the first place. "Move to production" is an instruction with no selection criterion attached, and distributors have taken it as license to promote whatever pilot had the best demo. That's the wrong filter. The right one is narrower and less exciting: does this AI application produce a discrete, checkable unit of output, or does it produce a vibe?

The divide is real, and it isn't about model quality

MIT's widely cited "State of AI in Business 2025" research — reviewing 300+ enterprise initiatives and 153 executive interviews — found that 95% of generative AI pilots show zero measurable return, despite $30–40 billion in enterprise spend. The researchers called it the "GenAI Divide": over 80% of companies have piloted something, nearly 40% report some deployment, and almost none of it moves an enterprise P&L. The tools that stall are general-purpose ones — chat assistants, copilots — because they hand a flexible instrument to a worker and ask the organization to somehow instrument the resulting judgment calls.

Separate 2026 research on enterprise AI agents found a similar pattern split by use case, not by industry: back-office automation delivered the biggest, most dependable wins — the same connected agent trimming the same hours every day, consistently — while sales- and marketing-facing bets, where most of the budget actually goes, produced the least reliable returns. Forrester's read on Copilot adoption lands on the same fault line: only 20–30% of licensed seats get weekly use, because the tool doesn't change anyone's workflow, it just sits next to it.

Distribution's own numbers confirm the pattern from the demand side. DSG's own inventory survey found 81% of warehouse and operations professionals want to implement AI, but only 11% currently use it day to day — and tellingly, chatbots and general-purpose applications ranked lowest in stated interest, well behind demand forecasting and replenishment. Operators are already voting, with their attention if not their budgets, for AI that produces a number they can check against a shelf.

Countability is the selection criterion, not ambition

Here's the mechanism. A sales copilot's output is a suggestion embedded inside a human conversation. To know if it worked, you'd have to instrument the whole call, the whole quote cycle, the whole relationship — separate the AI's contribution from the rep's skill, the customer's mood, the competitor's price that week. Nobody actually builds that instrumentation, so six months in, the project has a Net Promoter Score anecdote and no board-ready ROI line. That's not a data problem. It's a shape-of-the-work problem: the output was never a discrete unit to begin with.

Compare that to a product-content pipeline: a spec pulled off a manufacturer PDF, a UNSPSC code assigned, a duplicate SKU matched across two supplier feeds, a missing attribute filled from a datasheet. Each of those is a row. Each row is right or wrong — check it against the source document, the manufacturer's own spec sheet, the competing distributor's listing. You can sample 200 rows, measure the error rate, and know by Friday whether the model is production-grade. You can price it, because "cost per enriched SKU" is a number that existed before the AI project and still means something after. Anglera's Top Distributors 2026 index measured this gap directly: the distributors furthest ahead on the Digital Readiness Index aren't the ones with the flashiest chat interface, they're the ones whose catalogs simply have fewer holes in them — a boring, countable outcome that compounds.

This is why product-content operations keeps quietly reaching production while sales copilots sit in demo purgatory. It was never about which use case is more strategically important — inventory forecasting and quote assist are both defensible bets on paper. It's about which one produces an artifact a human can audit in isolation, without first solving the much harder problem of measuring an entire workflow.

The test to apply before you fund the next pilot

Before greenlighting an AI project this quarter, ask three questions instead of one. Does the model's output land as a discrete row — a field, a match, a code — or as advice buried in a conversation? Can a reviewer check that row against an independent source in under a minute, without reconstructing the context that produced it? And can you name the unit cost today, before the project starts, in a way that still means something after?

If the answers are yes, fund it and move fast — Graybar's two-month quote-extraction build is the right cadence for that kind of project. If the answers are no, the pilot isn't dying because the model is bad or the org isn't ready for AI. It's dying because nobody could ever have counted what it produced. Fix the selection criterion before the next keynote tells you to fix your culture.

That's the same discipline behind how we built Anglera: your PIM stores the data, we do the enrichment work as auditable, per-row output — the kind you can QA and price before you ever have to defend it in a board deck.

Ray Iyer

About the author

Ray IyerCo-founder, Anglera

Ray is a co-founder of Anglera, building the product-data infrastructure for agentic commerce — turning messy catalogs into structured, AI-readable data that buyers and answer engines can find. Previously product at Uber; Stanford CS.

See it on your own SKUs.

A 30-minute walkthrough on your categories and your supplier data.

Book a demo