Plain-English definitions of product data, PIM, enrichment, syndication, and AI-search terms — for distributors, retailers, and brands.
Agentic checkout is a purchase completed by an AI agent on a buyer's behalf, with the agent handling cart construction and payment authorization while the merchant remains the merchant of record. It is implemented through protocols such as ACP and UCP and through platform-specific integrations, and it depends on a product record accurate enough that the agent's assumptions about price, unit, and availability hold at the moment of the order.
Agentic Commerce Protocol (ACP)The Agentic Commerce Protocol (ACP) is an open specification for checkout between a buyer's AI agent and a merchant, covering cart, payment delegation, and order handling so the agent can complete a purchase without becoming the merchant of record. It is jointly governed by OpenAI and Stripe as founding maintainers and is still labelled beta. The repository describes neutral foundation stewardship as a future path, not a current state.
AI catalog enrichmentAI catalog enrichment is the use of machine learning and language models to complete an entire product catalog at once — extracting attribute values from supplier documents, classifying SKUs into a taxonomy, normalising units and vocabulary across suppliers, and generating structured content — rather than enriching records one at a time by hand. Its defining requirement is consistency at volume: being correct across hundreds of thousands of records and being able to show where each value came from.
AI citationAn AI citation is a link or source attribution an AI answer engine attaches to a claim in its response, pointing back at the page it retrieved that claim from. Citations are the visible evidence that retrieval happened, and for commerce they are the closest equivalent to a ranking position: being cited is how a product enters the answer a buyer actually reads.
AI crawlerAn AI crawler is an automated agent that fetches web content on behalf of an AI system, for one of three distinct purposes: building model training data, building a retrieval index for an AI search product, or fetching a page live because a user just asked about it. Those three purposes are governed by different robots.txt tokens from the same vendor, which is why blanket "block AI bots" rules so often produce an outcome nobody intended.
AI product attribute generationAI product attribute generation is the automated production of structured attribute values — dimensions, materials, ratings, certifications, compatibility references — for product records that lack them, using extraction from source documents and images together with inference from category norms. It is distinct from description generation, which produces prose, and it is the step that determines whether a product can be filtered, compared and retrieved at all.
AI product description generatorAn AI product description generator is a tool that produces product titles, descriptions, feature bullets and metadata from structured input using a language model. Output quality is governed almost entirely by input quality: given complete specifications the results are strong and cheap, while given a bare part number the model will produce fluent, confident text containing facts it inferred rather than read.
AI referral trafficAI referral traffic is the sessions arriving on your site from links inside AI assistant answers — chatgpt.com, perplexity.ai, gemini.google.com, copilot.microsoft.com and similar. It is the one AI-visibility signal that lands in your own analytics, which makes it the easiest to report and the easiest to misread, because it counts only the answers someone clicked out of.
Amazon A+ ContentAmazon A+ Content is Amazon's enhanced product description module, available to brand-registered sellers and vendors, that replaces the plain text description on a detail page with formatted images and comparison charts. It sits below the fold, applies at the ASIN or parent level, and renders mostly as images. It does not populate structured attributes, and Amazon does not index it for search.
Amazon flat fileAn Amazon flat file is a spreadsheet template downloaded from Seller Central and used to create or update listings in bulk. Each template is category-specific: the columns required for Hardware are not the columns required for Lighting or Industrial & Scientific. Sellers fill the Template tab, upload it, and Amazon returns a processing report with row-level errors. A category listing report is the reverse trip — an export of your live listings in a similar column shape.
Answer engine optimization (AEO)Answer engine optimization (AEO) is the practice of structuring product content and attributes so that AI-powered systems — including Google AI Overviews, ChatGPT Shopping, and Perplexity — can read, extract, and cite a specific SKU as the direct answer to a buyer's query. Where traditional SEO targets a ranked list of links, AEO targets the single synthesized recommendation a model returns.
Answer engine optimization (AEO)Answer engine optimization (AEO) is the practice of structuring content so that AI answer engines — ChatGPT, Perplexity, Google AI Overviews, Gemini, Copilot — can retrieve it, understand it, and cite it inside a generated answer. For product catalogs, AEO is mostly a data problem rather than a copywriting one: an engine can only cite a specification it can parse, so completeness, structured markup and machine-readable identifiers determine eligibility long before wording does.
ApplebotApplebot is Apple's web crawler, gathering the page content that powers Siri web answers, Spotlight Suggestions, and Safari's search features. It identifies itself with "Applebot" in its user-agent string, respects robots.txt, and adjusts its crawl rate automatically rather than honoring a Crawl-delay directive. Applebot-Extended, added in 2024, is a separate, independent user-agent that controls whether content already gathered can be used to train Apple's generative AI models, including Apple Intelligence — disallowing it doesn't remove a page from Siri, Spotlight, or Safari search, and allowing standard Applebot doesn't automatically allow Applebot-Extended.
ASIN (Amazon Standard Identification Number)An ASIN is a 10-character alphanumeric identifier that Amazon assigns to each product page in its catalog. Unlike a GTIN or MPN, Amazon owns and issues it - you do not, and neither does a standards body. Amazon mints a new ASIN when a submitted listing matches nothing in the catalog, and attaches your offer to an existing ASIN when it does match. One SKU maps to one ASIN per marketplace.
Attribute fill rateAttribute fill rate is the percentage of required attribute fields that contain a usable value across a set of SKUs. It is calculated as filled cells divided by expected cells, and is usually cut per attribute, per category, or per channel. Data teams report it to executives as the headline measure of catalog completeness. On its own it says nothing about whether the values are correct.
BMEcatBMEcat is an open XML standard for exchanging structured product catalog data between a supplier and a buyer's ERP, PIM, or e-procurement system as a single file rather than a live connection. It carries item numbers, descriptions, classification codes (commonly ETIM or eCl@ss), prices, units of measure, and links to images or documents in a fixed schema both sides already agree on. It is maintained by BME e.V., the German association for materials management, purchasing, and logistics, and is most entrenched in DACH-region industrial and electrical distribution.
Buyer signalsBuyer signals are the behavioral and contextual data points — search queries, facet selections, comparison patterns, and purchase criteria — that reveal how a specific buyer discovers, evaluates, and selects a product. In B2B product data, these signals define which attributes, terminology, and use-case framing belong in a listing to match how that buyer actually shops, not just what the supplier documented.
Catalog managementCatalog management is the ongoing process of collecting, structuring, enriching, and distributing product data so that every SKU in a company's assortment is accurate, complete, and consistent across every channel where buyers encounter it. In B2B commerce, it spans supplier data ingestion, attribute normalization, content enrichment, and syndication to e-commerce sites, distributor portals, and procurement platforms.
ClaudeBotClaudeBot is Anthropic's training crawler, described in Anthropic's crawler documentation as collecting web content that could contribute to training its generative AI models. It sits alongside two other agents: Claude-User, which fetches pages in response to a Claude user's question, and Claude-SearchBot, which navigates the web to improve search result quality. As with OpenAI, the training crawler and the search crawler are separate controls.
Content chunkingChunking is the step in a retrieval pipeline that splits source content into passages small enough to embed and retrieve independently. The chunk, not the page, is what gets matched against a query and handed to the model, so a fact's retrievability depends on whether it survived the split with enough context attached to make sense on its own.
Content health scoreA content health score is a composite grade, usually 0-100, that rates how complete and compliant a product listing is against a rubric. Typical checks cover required attributes, image specs, description length, and identifiers. Retailers, PIM vendors, and analytics tools each write their own rubric, so the same SKU can pass in one system and fail in another. The score measures conformance to a checklist, not whether a buyer can choose the product.
Content-to-ConversionContent-to-conversion is the measurable relationship between product content quality—completeness, accuracy, and buyer relevance—and the rate at which that content moves a B2B buyer from product discovery to a confirmed purchase signal such as an order, quote request, or catalog add. In B2B e-commerce, it is the primary lens for evaluating whether product data is doing commercial work or simply occupying database rows.
Controlled vocabularyA controlled vocabulary is a fixed, approved list of the values an attribute is allowed to hold: the only entries a field like Finish or Drive Type may contain. Anything a supplier writes outside that list is either mapped back into it or rejected. Controlled vocabularies are what make attributes filterable, comparable across brands, and readable by search engines and marketplace feeds.
DAM (Digital Asset Management)A DAM (digital asset management system) is the system of record for product media files — packshots, lifestyle photos, 360 spins, video, CAD models, spec sheet PDFs, safety data sheets. It stores the master file, generates channel-specific renditions, tracks usage rights, versions replacements, and serves stable CDN URLs. A PIM stores attributes about the SKU; a DAM stores the files attached to it. The two link by SKU.
Data cleansing vs data enrichmentData cleansing corrects errors, removes duplicates, and standardizes inconsistent values already in a product record; data enrichment adds attributes, descriptions, and context that were never captured in the first place. Cleansing makes existing data accurate; enrichment makes a record complete enough to be found, compared, and bought.
Data governanceData governance is the set of rules, roles, and approvals that decide who may create, change, or approve a product data value, and what a valid value looks like. For product data it answers four questions per attribute: who owns it, which source wins, what a valid value looks like, and what requires review. Without it, enriched catalogs drift back toward blanks and conflicts as new SKUs and supplier feeds arrive ungoverned.
Data normalizationData normalization in B2B product data is the process of transforming product attributes — units of measure, naming conventions, and value formats — into a consistent schema so every SKU in a catalog can be accurately compared, searched, and evaluated on equal footing. Without it, identical specifications stored in different formats are invisible to faceted search, comparison engines, and downstream channel feeds.
Data validation rulesData validation rules are the machine-checkable conditions a product record must satisfy before it counts as complete and publishable. They cover required fields, allowed values, numeric ranges, formats, and conditional logic that changes by category or sales channel. A PIM stores and enforces them, and they act as the acceptance criteria that decide whether a SKU passes review or gets sent back for rework.
Digital Product Passport (DPP)A Digital Product Passport (DPP) is a standardized, machine-readable record — reached by scanning a data carrier such as a QR code on the product or packaging — that holds a product's sustainability, compliance, and lifecycle data and stays attached to it after sale. It's a regulatory requirement under the EU's Ecodesign for Sustainable Products Regulation (ESPR, Regulation (EU) 2024/1781), rolling out category by category rather than all at once.
DUNS numberA DUNS number (Data Universal Numbering System) is a nine-digit business identifier issued by Dun & Bradstreet, assigned per physical business location and used to verify and track a company for credit, trade, and vendor-onboarding purposes. It is not a GS1 identifier and doesn't convert into a GLN or GTIN. Since April 2022, the U.S. federal government no longer uses DUNS for contractor and grant registration — SAM.gov issues its own Unique Entity Identifier (UEI) instead — but D&B still issues and maintains DUNS commercially, and retailers, GPOs, and trading partners still ask for it during vendor setup.
E-E-A-T (Experience, Expertise, Authoritativeness, Trustworthiness)E-E-A-T is the framework Google's search quality raters use to assess content quality, standing for Experience, Expertise, Authoritativeness and Trustworthiness. Google states plainly that "E-E-A-T itself isn't a specific ranking factor," but that its systems use a mix of signals that identify content with good E-E-A-T, and that of the four, trust is the most important. It is a description of what Google is trying to reward, not a checklist that produces rankings.
EAN (European Article Number)An EAN (European Article Number, now formally the International Article Number) is a 13-digit GS1 product identifier encoded in the barcode on retail packaging. It is the same thing as a GTIN-13. EAN-13 is the global default; the 12-digit UPC used in the US and Canada is a subset that becomes an EAN by adding a leading zero. Every EAN ends in a calculated check digit.
eCl@sseCl@ss is a cross-industry product classification and attribute standard maintained by the eCl@ss e.V. association in Germany. It pairs a four-level, eight-digit class hierarchy with standardized property dictionaries, so each class carries a defined list of attributes with their units and permitted values. It is widely used for industrial, MRO, electrical, and chemical catalogs in DACH markets, typically exchanged inside BMEcat files.
EDI (Electronic Data Interchange)EDI (Electronic Data Interchange) is the structured, computer-to-computer exchange of business documents — purchase orders, invoices, shipment notices — between trading partners in a standardized format, replacing manual entry from email, fax, or a portal. In North America the dominant standard is ANSI X12, organized into numbered transaction sets like the 850 (purchase order) and 810 (invoice); most of the rest of the world uses EDIFACT. EDI moves transactions, not catalog content — that's a separate job handled by formats like BMEcat or a PunchOut session.
ERP (Enterprise Resource Planning)An ERP is the system of record for a company's transactional operations: purchasing, inventory, pricing, order management, and finance. In product data terms, the ERP holds the item master, the record that has to exist before a part can be bought, stocked, priced, or shipped. It is authoritative for commercial and logistics fields, but it was never built to store the descriptive attributes, images, or copy a buyer reads.
ETIMETIM is an open, international classification standard for technical products such as electrical, HVAC, plumbing, and building supplies. It assigns each product to a numbered class, then defines the exact features, units, and permitted values that class requires. Manufacturers publish ETIM-classified data, and distributors use it to build comparable catalogs and filterable search across brands.
Faceted / Attribute-Based SearchFaceted search (also called attribute-based search) is a catalog navigation method that lets buyers filter results simultaneously across multiple independent attribute dimensions — such as voltage rating, material, thread size, or certification — so only the SKUs satisfying every selected criterion are returned. It depends entirely on structured, consistent attribute data stored at the field level; specifications buried in descriptions or PDFs are invisible to it.
GDSN (Global Data Synchronization Network)GDSN is a GS1-governed network of certified data pools that enables suppliers and retailers to continuously exchange structured product information through a standardized publish-and-subscribe model. It is a delivery mechanism for product data — not a source of enriched or buyer-ready content.
Generative engine optimization (GEO)Generative engine optimization (GEO) is the practice of shaping content so that AI systems which generate answers — ChatGPT, Google AI Mode, Perplexity, Gemini — retrieve it, use it, and cite it. The term comes from a 2024 KDD paper by Pranjal Aggarwal and coauthors that framed it as a black-box optimization problem: content owners cannot see inside the model, so they optimize the inputs it can reach. In practice GEO and answer engine optimization describe the same work under different names.
GLN (Global Location Number)A GLN (Global Location Number) is a 13-digit GS1 identifier for a legal entity, a function, or a physical location: the company you are, the warehouse you ship from, the department that receives the invoice. Trading partners use it to address you and your places in EDI documents and GDSN publications. GDSN will not register an item without one, and retailer EDI documents route on it.
Golden recordA golden record is the single trusted version of a product's data, assembled attribute by attribute from every system that holds a claim about it. When a supplier feed, your ERP, and a manufacturer's spec sheet disagree on the thread pitch of a 3/8-16 hex bolt, survivorship rules decide which value wins. The result is not a copy of any one source. It is the best available value for each field, with its origin recorded.
Google AI ModeAI Mode is Google's conversational search experience, where a query is answered by a generated response with follow-up questions rather than by a page of links. Like AI Overviews, it is a Search surface: Google states that appearing in it requires only that a page be indexed and eligible to be shown in Search with a snippet, and that it may use query fan-out to gather a wider set of sources than a single query would.
Google AI OverviewsAI Overviews are AI-generated summaries Google shows above traditional results for some queries, with links out to the sources they drew on. Google states there are no additional requirements or special optimizations to appear in them: a page must simply be indexed and eligible to be shown in Google Search with a snippet. They are a Search feature, not a separate product, which is why Google-Extended does not opt you out of them.
Google product categoryGoogle product category is Google's own product taxonomy: a fixed, hierarchical list of category paths like "Hardware > Hardware Accessories > Hardware Fasteners > Nuts & Bolts", each carrying a numeric ID, that you assign to items in a Merchant Center feed. It is optional for most items — Google auto-assigns a category when you leave it blank — but it is required for categories such as apparel, alcohol, and mobile devices, and it drives tax and shipping rules.
Google Shopping feed specificationThe Google Shopping feed specification is Google's published contract for the attributes a product feed must contain to be accepted by Merchant Center and shown in Shopping ads and free listings. It defines required attributes such as id, title, description, link, image_link, availability, and price, plus conditionally required ones like gtin, mpn, brand, and item_group_id. Fail the contract and the item is disapproved, not just ranked lower.
Google Shopping GraphThe Shopping Graph is Google's product data layer: a continuously updated model of products, sellers, prices, reviews and availability, assembled from Merchant Center feeds, crawled retailer pages, manufacturer data and other sources. It is what shopping results, Shopping in AI Mode, and Google's agentic shopping features draw on, which makes feed accuracy and page accuracy two inputs to the same system.
Google-ExtendedGoogle-Extended is a robots.txt token that controls whether content Google crawls can be used to train and ground its generative AI products — Gemini models in Gemini Apps, the Vertex AI API for Gemini, and grounding in both. Google states plainly that it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search." It is not a crawler; Googlebot does the fetching and this token governs the downstream use.
GPC (GS1 Global Product Classification)GPC (GS1 Global Product Classification) is GS1's standard scheme for sorting products into a four-level hierarchy: segment, family, class, and brick. The brick is the working level — an eight-digit code that says what a product fundamentally is, independent of brand, packaging, or supplier. In GDSN, the brick you assign is what tells a data pool and its recipients which attributes and validation rules apply to that item.
GPTBotGPTBot is OpenAI's training crawler. OpenAI's own bot documentation describes it as "used to crawl content that may be used in training our generative AI foundation models," and it is a separate agent from OAI-SearchBot, which is what surfaces sites in ChatGPT's search features. Blocking GPTBot does not remove you from ChatGPT search, and allowing it does not get you cited — the two are constantly confused.
GroundingGrounding is the practice of tying an AI model's output to retrieved source material rather than to its training memory, so each claim can be traced back to a document. In commerce, a grounded answer about a product is one where the price, spec, and availability came from a live source at answer time instead of from whatever the model absorbed months earlier.
GS1 Digital LinkGS1 Digital Link is a GS1 standard that encodes a product's GTIN and other identifiers as a structured web URL, carried inside a 2D barcode such as a QR code, so one scan can resolve to both point-of-sale checkout and a webpage of product information. It replaces the old split between a UPC for scanning and a separate marketing QR code for content with a single, standardized code that any GS1-compliant scanner can parse.
GTIN (Global Trade Item Number)A GTIN (Global Trade Item Number) is a GS1-standardized numeric identifier — 8, 12, 13, or 14 digits — that uniquely identifies a trade item at a specific packaging level anywhere in the world. It is the universal key that links a physical product to its data record across every system in the supply chain, from manufacturer to distributor to retailer.
HS code (Harmonized System code)An HS code is a numeric product classification code from the World Customs Organization's Harmonized System, used by customs authorities to identify goods and assess duty. The first six digits are standardized across roughly 200 countries; individual countries extend them to eight or ten digits for their own tariff schedules. Formal customs entries require one on the commercial invoice and entry filing.
Human-in-the-loop reviewHuman-in-the-loop review is a governance pattern where AI-generated product data is scored for confidence and routed to a human reviewer before it reaches your PIM. High-confidence values publish automatically; low-confidence, high-risk, or conflicting values escalate to a person who approves, corrects, or rejects them. Every decision is logged, so each attribute value carries a source, a reviewer, and a timestamp.
Intelligent enrichmentIntelligent enrichment is the practice of augmenting product data not just with missing attributes, but with the right attributes—framed in the language and detail level that buyers actually use when searching, comparing, and purchasing. It goes beyond reformatting supplier content by reading buyer behavior signals to determine what to add, how to phrase it, and how to prioritize SKUs by commercial impact.
Item masterAn item master (called the material master in SAP) is the core ERP record for a SKU: the internal item number, unit of measure, standard cost, GL account coding, and operational flags — buy vs. make, lot- or serial-tracked, active or discontinued — that purchasing, warehouse, and finance systems key off of. It's built for internal operations, not for what a buyer sees. A PIM holds the customer-facing counterpart: marketing copy, images, spec attributes, and taxonomy meant for syndication. The two records share the same item number but serve different audiences and are usually owned by different teams.
JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a format for embedding structured data in a web page as a self-contained block of JSON, placed inside a script tag. Search engines parse it to learn what a page is about — that a PDP describes a 3/8-16 Grade 8 hex cap screw with a specific GTIN, price, and availability — instead of inferring it from visible HTML. It is the structured data format Google recommends.
Kit and bundle SKUA kit or bundle SKU is a single salable part number that ships as a set of separate items rather than one manufactured piece, unlike an assembly, whose parts are consumed into one unit. A kit is picked from stocked components at order time; a bundle is a merchandising grouping of finished goods. Either way the set carries its own GTIN and MPN, and its attributes have to be derived from its components rather than typed in.
LLM optimization (LLMO) and AI search optimization (AISO)LLM optimization (LLMO) and AI search optimization (AISO) are two of several competing names for the same discipline: making content that language-model-powered search systems retrieve and cite. Neither has an authoritative definition or standards body behind it, and neither describes work that differs materially from answer engine optimization or generative engine optimization. The proliferation of acronyms is a marketing artifact, not four distinct practices.
llms-full.txtllms-full.txt is a community convention for publishing an expanded, single-file version of a site's content for language models to consume, alongside the shorter index-style llms.txt. It is not part of the original llms.txt proposal, which describes optional expanded files generated by tooling rather than a standard filename, and like llms.txt it is honoured by no major AI provider as a matter of policy.
llms.txtllms.txt is a proposed convention: a markdown file at the root of your domain that gives large language models a curated map of your site's most useful content, in plain text instead of HTML. It is a hint, not a rule — no major AI provider has committed to reading it, and it grants or blocks nothing. For ecommerce, it is cheap to publish and worth doing only after your product pages are actually machine-readable.
Long-tail SKUA long-tail SKU is a low-volume, infrequently ordered item that sits in the bottom band of a catalog by sales — individually small, collectively most of the assortment. Because merchandising attention follows revenue, long-tail SKUs rarely get funded enrichment, so their attributes, images, and descriptions stay thin. The result is a catalog where a minority of items are complete and the majority are effectively unsearchable.
MAP (Minimum Advertised Price)MAP (Minimum Advertised Price) is the lowest price a brand allows a reseller to display publicly — in an ad, on a product page, or in a search result. It governs the advertised number, not the price a buyer actually pays at checkout. Brands publish a MAP policy plus a per-SKU price list, monitor the digital shelf for violations, and enforce commercially by pulling rebates, co-op funds, or supply.
Master data management (MDM)Master data management (MDM) is the set of policies, processes, and technology a company uses to create and maintain a single, authoritative record — the "golden record" — for its core business entities (products, customers, suppliers, locations) and distribute that record consistently across every downstream system. It is a governance discipline, not a data-quality or enrichment tool.
Model Context Protocol (MCP)The Model Context Protocol (MCP) is an open standard for connecting AI applications to external tools and data sources, so an agent can call a function or read a record instead of guessing from training memory. Anthropic introduced it in November 2024 and donated it in December 2025 to the Agentic AI Foundation, a directed fund under the Linux Foundation, which moved it out of single-vendor governance. It is the plumbing layer beneath commerce protocols, not a commerce protocol itself.
MPN (Manufacturer Part Number)An MPN (Manufacturer Part Number) is the identifier a manufacturer assigns to its own product, such as 3M's 8210 for an N95 respirator. It is unique only within that manufacturer's catalog, so an MPN alone is ambiguous; brand plus MPN is what identifies a product. Unlike a GTIN, an MPN is not centrally registered and follows no format rules.
New item setup (NIS)New item setup (NIS) is the process a retailer, distributor, or marketplace uses to admit a new product into its catalog. The supplier files the item's identifiers, attributes, images, packaging data, and commercial terms; the buyer validates them against its category rules. Until the item passes, it cannot be ordered, listed, or shipped. Items are rejected when required attributes are missing, malformed, or contradict the manufacturer's spec sheet.
OAI-SearchBotOAI-SearchBot is OpenAI's retrieval crawler, described in OpenAI's bot documentation as "used to surface websites in search results in ChatGPT's search features." It is the agent that determines whether your pages can be found and cited when a ChatGPT user asks a question, and it is separate from GPTBot, which crawls for model training. If you want to be citable in ChatGPT, this is the one to allow.
OCI (Open Catalog Interface)OCI (Open Catalog Interface) is SAP's protocol for PunchOut catalog sessions, connecting a buyer's SAP-based procurement system to a supplier's hosted catalog over plain HTTP. Instead of exchanging XML documents like cXML, an OCI session opens a catalog with session details passed as URL parameters, and the buyer's cart is returned as HTTP-posted name-value pairs. It covers the shopping and cart-return steps only — order and invoice automation aren't part of the standard — and its adoption concentrates around SAP SRM, S/4HANA sourcing, and Ariba deployments, more heavily in Europe than North America.
OpenAI product feedThe OpenAI product feed is a structured product data specification that merchants submit so ChatGPT can index and display their products with current price and availability. It carries required fields including item_id, title, description, url, brand, image_url, price, availability, and seller details, plus eligibility flags controlling whether items can appear in search and in checkout. Its most consequential requirement is that product and variant identifiers stay stable over time.
Parent-child product variantsParent-child product variants are a data structure that groups sellable child SKUs — each with its own GTIN, price, and stock — under a non-sellable parent record holding the shared content. The children differ only along declared variation axes such as size, color, or length. Marketplaces use the relationship to render one product page with a picker instead of dozens of separate listings.
Part number cross-referenceA part number cross-reference (also written as a cross-reference part number) is a mapping that links one manufacturer's part number to an equivalent or replacement part from another manufacturer. Distributors publish these as interchange tables so a buyer holding a competitor's part number can find the equivalent they stock. The mapping is only as good as the attribute data behind it: two parts are interchangeable when their dimensions, ratings, and materials match closely enough for the application.
PDP (product detail page)A PDP is the page dedicated to a single purchasable product — one SKU, or one parent with its variants — carrying the identifiers, attributes, images, documents, and price a buyer needs to decide. Every category page, facet, feed row, and marketplace listing is assembled from what PDPs declare. A PDP is only as good as the product data behind it.
PerplexityBotPerplexityBot is Perplexity's retrieval crawler, described in Perplexity's documentation as designed to surface and link websites in search results on Perplexity, and it respects robots.txt. Perplexity separately runs Perplexity-User, which fetches a page when a user's question requires it and which the documentation says generally ignores robots.txt because the request was initiated by a person. Neither crawls for model training.
Product attributesProduct attributes are the individual data fields — dimensions, materials, certifications, compatibility specs, electrical ratings, and similar properties — that formally describe what a product is and how it performs. In B2B e-commerce, they are the primary mechanism by which buyers filter search results, compare competing SKUs, and validate products against procurement or engineering requirements.
Product content intelligenceProduct content intelligence is the measurement layer over a catalog: analytics that assess how complete, accurate and competitive product content is, where it is failing across channels, and what the commercial cost of each gap is. It converts a diffuse quality problem into a ranked, quantified work queue — and it is diagnostic rather than corrective, since identifying a gap does not fill it.
Product content localizationProduct content localization is the work of adapting product data for a specific market. It converts units of measure, swaps compliance and certification attributes, remaps taxonomy and classification codes, and uses the terms local buyers search. A UL 486A listing means little to a buyer looking for a VDE mark. Translation changes the words; localization changes the record.
Product content managementProduct content management is the discipline of creating, maintaining, governing and distributing everything a buyer sees about a product — attributes, descriptions, imagery, documents, video and channel-specific variants — across every place the product is sold. It is broader than product information management, which centres on structured data, and broader than digital asset management, which centres on media.
Product content syndicationProduct content syndication is the process of distributing standardized product data — titles, descriptions, specifications, and digital assets — from a single authoritative source to multiple downstream channels such as distributor catalogs, retail platforms, and B2B marketplaces. It is a delivery mechanism: the quality and completeness of what arrives at each channel depend entirely on the quality of what left the source.
Product data enrichmentProduct data enrichment is the process of supplementing raw or incomplete product records with additional attributes, corrected values, structured descriptions, and contextual metadata required for accurate search, comparison, and purchase decisions. In B2B contexts, it typically means transforming sparse supplier-exported data into complete, buyer-ready records that meet the attribute depth demanded by procurement professionals, digital commerce channels, and AI-assisted discovery engines.
Product data qualityProduct data quality is the degree to which product records are complete, accurate, consistent, valid, timely and unique enough to support the decisions and systems that depend on them. In commerce it is measured against a defined attribute standard per category rather than in the abstract — a record can be flawless as data and still be unfit for purpose if it lacks the specifications buyers filter and choose on.
Product feedA product feed is a structured, machine-readable file or data stream — typically formatted as CSV, XML, or JSON — that transmits product attributes, pricing, and availability from a source system to a downstream channel such as a marketplace, distributor portal, search engine, or procurement platform. In B2B commerce, the completeness and accuracy of that feed directly determines whether a buyer can find, evaluate, and purchase a product without calling a sales rep.
Product feed managementProduct feed management is the practice of transforming a catalog into the specific format each sales and advertising channel requires, then keeping those outputs correct as products, prices and channel requirements change. It covers field mapping, value transformation, category matching, validation, scheduled submission and error remediation across destinations such as Google Shopping, Meta, Amazon and marketplaces.
Product information management (PIM)Product information management (PIM) is the practice of centralizing product content — specifications, descriptions, digital assets, and categorization — in a single governed repository so it can be distributed consistently across every sales channel. A PIM system serves as the authoritative source of record for a product catalog, but does not itself gather, clean, or enrich the data it holds.
Product knowledge graphA product knowledge graph is a machine-readable model of a catalog in which every product, attribute value, category, manufacturer, and standard is an explicit entity connected by named relationships. Instead of storing "1/2 in. NPT brass ball valve" as a description, it records the product as a node linked to brass, to NPT, to a nominal size, and to the fittings it mates with. That structure lets software answer questions keyword search cannot.
Product matchingProduct matching is the process of deciding whether two product records from different sources describe the same physical item. It applies entity resolution — identifier lookup, attribute comparison, and text similarity — to link a supplier's spreadsheet row, a competitor's listing, and your own SKU to one entity. It is the backbone of deduplication, cross-referencing, and any enrichment that pulls data from a source you do not control.
Product schema markupProduct schema markup is structured data — typically written in JSON-LD using the Schema.org Product vocabulary — embedded in a web page's HTML to expose product attributes (name, SKU, GTIN, price, availability) as typed, machine-readable facts rather than prose. Search engines use it to generate rich results; AI systems use it to evaluate and recommend products without inferring meaning from unstructured copy.
Product taxonomyProduct taxonomy is the hierarchical classification system that organizes a catalog into categories and subcategories, determining where each product lives in a browse tree, which attribute schema applies to it, and whether buyers and search engines can find it at all. Every downstream function — faceted navigation, attribute completeness, marketplace syndication — depends on a product landing in the right node.
Prompt testing (AI visibility tracking)Prompt testing is the practice of running a fixed set of buyer questions through AI assistants on a repeating schedule and recording whether your brand, products, and URLs appear in the answers. It is the standard method behind AI visibility tools, and it is sampling rather than reporting: the underlying systems are non-deterministic and none of them publish the data directly.
PunchOut catalogA PunchOut catalog is a supplier-hosted catalog that a buyer reaches from inside their own procurement system, such as SAP Ariba, Coupa, or Jaggaer, instead of loading a static price file. The procurement system opens an authenticated session on the distributor's site. The buyer builds a cart, and the cart returns to the procurement system as a requisition line for approval and PO issue.
PXM (product experience management)PXM (product experience management) is the practice, and the software category, of managing product content so it reads well and converts on every channel a buyer sees, not just so it is stored correctly. It covers a PIM's attributes and taxonomy plus assets, channel-specific copy, syndication, and performance measurement. Vendors use the label to draw a wider boundary around product content, so treat it as a claim about scope and ask what sits inside that boundary.
Query fan-outQuery fan-out is the technique of issuing several related searches behind a single user question, then assembling an answer from the combined results. Google names it explicitly in its guidance on AI features in Search, saying AI Overviews and AI Mode may use it to display a wider and more diverse set of helpful links than one query would return.
Retrieval-augmented generation (RAG)Retrieval-augmented generation (RAG) is an AI architecture that grounds a language model's answer by first retrieving relevant passages from an external source — a product catalog, a document index — and feeding them into the model's context, instead of relying only on what the model memorized during training. The term comes from a 2020 paper by Patrick Lewis and coauthors at Facebook AI Research, UCL, and NYU, and the pattern now underlies most AI shopping assistants and enterprise search over a company's own content.
Rufus (Alexa for Shopping)Rufus is Amazon's AI shopping assistant, launched in beta in early 2024 and described by Amazon as "an expert shopping assistant trained on Amazon's product catalog and information from across the web to answer customer questions on shopping needs, products, and comparisons." Amazon renamed it Alexa for Shopping on 13 May 2026, so both names refer to the same assistant and much published guidance still uses the old one.
sameAs (schema.org property)sameAs is a schema.org property defined as the "URL of a reference Web page that unambiguously indicates the item's identity," such as a Wikipedia page, a Wikidata entry, or an official website. It is available on any Thing, and its practical job is entity disambiguation: telling a machine that the organization or brand on your page is the same entity as the one described elsewhere.
Semantic completenessSemantic completeness is an informal term for whether a piece of content answers a question fully enough to stand on its own, without the reader needing to go elsewhere. It is a useful writing principle and not a measurable quantity: there is no standard definition, no published measurement method, and no peer-reviewed correlation between it and citation rates, despite figures to that effect circulating in AI-search marketing content.
Semantic searchSemantic search is a retrieval method that matches on meaning rather than exact words. It converts the query and each product record into vectors, numeric representations of meaning, and returns the records whose vectors sit closest to the query's. A shopper searching "bolt that won't rust outdoors" can surface a 316 stainless hex cap screw even though the listing never uses the words "rust" or "outdoors."
Server-side rendering (SSR) for AI crawlersServer-side rendering means the server returns fully-formed HTML containing your content, rather than a JavaScript shell the browser fills in afterwards. It matters for AI search because most AI crawlers do not execute JavaScript: analysis of large-scale crawl logs found no evidence that OpenAI's GPTBot does. A product page whose specifications and JSON-LD are injected client-side is, to those crawlers, a page with no product on it.
Share of modelShare of model is the proportion of AI answers in a category that mention or cite your brand or products, measured by running a fixed set of buyer prompts across assistants on a schedule and counting appearances. It is an industry-coined metric with no standard definition and no first-party reporting behind it, so two vendors' numbers for the same brand are rarely comparable.
Share of searchShare of search is the percentage of search activity in a category that belongs to your brand: either your slice of total search volume for brand terms, or the share of visible results your SKUs win for category queries. Brands and distributors track it category by category on Amazon, Google, and distributor sites. The same measure extends to AI answers, where the unit counted is citations instead of ranked links.
SKU EnrichmentSKU enrichment is the process of systematically augmenting a product record's existing data — typically sparse supplier copy — with accurate, structured attributes that help buyers find, evaluate, and choose the product. It goes beyond reformatting to add net-new information: specifications, categorized values, search-optimized copy, and buyer-relevant comparisons that the original supplier data rarely includes.
SKU rationalizationSKU rationalization is the periodic review of every active SKU in a catalog against sales velocity, margin, and carrying cost, ending in a decision to keep, consolidate, or discontinue each one. The goal is concentrating merchandising, inventory, and data-quality effort on the SKUs that earn it, instead of spreading it evenly across an assortment where a small share of items drives most of the revenue.
Structured data validationStructured data validation is checking that your markup parses, conforms to schema.org, and contains the properties consuming systems expect. Two tools do different jobs: Google's Rich Results Test reports eligibility for Google's own features, while the Schema Markup Validator checks schema.org conformance generally. A page can fail the first and still be perfectly good input for an AI retrieval system.
Structured vs. unstructured product dataStructured product data is stored in discrete, machine-readable fields — part numbers, voltage ratings, dimensions, certifications — that systems can query, filter, and compare without human interpretation. Unstructured product data is everything else: PDFs, prose descriptions, images, and supplier documents that contain useful information but require parsing before any system can act on it.
Taxonomy MappingTaxonomy mapping is the process of translating each product's classification from one category hierarchy — such as a supplier's internal schema — to the equivalent node in a target system, such as a distributor's PIM, a marketplace browse tree, or an industry standard like UNSPSC or eCl@ss. The target node determines which attribute template governs the product, and therefore which specifications buyers can search and filter on.
The digital shelfThe digital shelf is the collective set of online touchpoints — search engines, e-commerce websites, B2B marketplaces, punchout catalogs, and AI answer engines — where a buyer can discover, evaluate, and purchase a product. Unlike a physical shelf where placement is a retail decision, digital shelf presence is determined algorithmically by the quality, completeness, and structure of the underlying product data.
Universal Commerce Protocol (UCP)The Universal Commerce Protocol (UCP) is a Google-led open protocol, published on 11 January 2026, that lets a business advertise commerce capabilities to AI agents at a /.well-known/ucp discovery endpoint. It is broader than a checkout spec, covering catalog search and lookup, cart building, identity linking, checkout, and order management, and it runs over REST and JSON-RPC with AP2, A2A and MCP support built in.
UNSPSCUNSPSC (United Nations Standard Products and Services Code) is a global classification system that assigns every product and service an eight-digit code across four levels: Segment, Family, Class, and Commodity. It is maintained by GS1 US and used mainly by procurement teams for spend analysis, catalog search, and supplier reporting. It classifies what a thing is. The code carries no attributes, and unlike a GTIN or MPN it does not identify a specific item.
UOM (unit of measure)UOM (unit of measure) is the unit a product is counted, measured, priced, or sold in: each, case, foot, pound, milliliter. In distributor catalogs one SKU usually carries several UOMs at once: a stocking UOM, a selling UOM, and a pricing UOM can all differ on the same record, alongside units attached to physical attributes like length and weight. A bolt stocked in EA can sell as a box of 50 and be priced per hundred.
UPC (Universal Product Code)A UPC (Universal Product Code) is the 12-digit numeric identifier encoded in the barcode on retail packaging in the US and Canada. It is assigned by the brand owner from a licensed GS1 company prefix and identifies one specific sellable item: a single size, color, and pack count. A UPC is not a separate standard from GTIN. It is a GTIN-12, the 12-digit member of the GTIN family.
Vector database (vector index)A vector database is a store built to hold embeddings and answer nearest-neighbour queries over them quickly — given a query vector, return the closest records out of millions. It is the retrieval layer underneath semantic search, RAG pipelines, and most AI shopping assistants, and it is only as good as the product text that was embedded into it.
Vector embeddingA vector embedding is a list of numbers that represents the meaning of a piece of text, an image, or a product record, positioned so that semantically similar items sit close together in the same space. Embeddings are what make semantic retrieval possible: a query and a product description can be compared numerically even when they share no words.
WikidataWikidata is a free, collaboratively edited knowledge base run by the Wikimedia Foundation, where every entity has a stable Q-identifier and structured statements about it. It is one of the most widely reused public sources of entity data, which makes a Wikidata item the closest thing to a canonical machine-readable reference for an organization or brand, and the value schema.org names first as an example for sameAs.