← Legibility Index 2026

How it is measured

An index that publishes a number next to a company's name has to be arguable with. That means the rubric, the weights, the sample and the seed all have to be public — and it means being just as explicit about what we refuse to measure as about what we do.

The instrument

Five pillars, nineteen signals, one hundred points. Every signal is observed from the retailer's own live public site. Nothing is surveyed, nothing is self-reported, and there is no model judgement anywhere in the scoring — re-running the scorer over the same saved inputs produces the same numbers, which is the property that lets the index survive being disputed.

Product Data Depth

25

Can a machine tell what this product is, and match it to the same product elsewhere?

Identity6
Identifiers12
Attribute depth7

Buyer Answerability

20

Does the page answer what a shopper actually asks before committing?

Descriptive content6
Imagery5
Ratings and reviews6
Catalog consistency3

Commerce Transparency

20

Can a shopper — or an agent — find out what it costs, whether it ships, and how to send it back?

Public pricing5
Stock visibility4
Shipping in markup4
Return policy in markup4
Un-gated access3

Machine Readability

20

Is the structured data actually valid and complete, not merely present?

Product markup6
Required-field completeness5
Recommended-field depth5
Variant expression4

Agent Interface

15

Can a crawler, a marketplace, or an AI shopping agent discover and act on this catalog?

Crawler access to products4
AI crawler stance4
Product sitemap4
Agent surfaceplatform-derived3

Two tracks, and the gap between them

Most benchmarks read a page one way. We read it two ways, because the two fail in opposite directions.

Track A — what you publish

A plain, honest HTTPS client fetches the product page and we parse its JSON-LD graph properly — not with a regular expression. This is what a crawler gets.

Track B — what a machine recovers

The same page through Anglera's extraction service, which renders and normalises past bot walls. This is what an extractor can actually get back.

A retailer can publish a rich GTIN and MPN that an extractor loses, or refuse a crawler entirely while its page is full of data. Those are completely different problems and a single number flattens them into the same word. The distance between the tracks is reported per company as the Legibility Gap, and is never folded into the score.

How pages are chosen

Sampling is where a benchmark is most easily rigged, so the judgement and the selection are split apart.

  1. A researcher identifies the public files. Which sitemap holds the retailer's products, and what a product URL looks like. That is a checkable claim about a public file, not a discretionary pick.
  2. A script draws the sample. Seeded pseudo-random selection over every product URL in that sitemap, bucketed by top-level path so the sample spans the catalog rather than landing in one category.
  3. Anyone can reproduce it. The seed is anglera-legibility-2026 and the sample size is 8. Same seed, same sitemap, same eight pages.

Where a retailer publishes no usable product sitemap we fall back to a verified category sample — products taken from the middle of a category listing, never featured or best-seller placements, spread across at least five categories. The scorecard says which method was used, every time.

The seed is frozen for this edition and published here. Changing it after scores are out would let us reroll a sample we did not like.

Four rules that keep the number honest

Each of these is a place where a careless rubric publishes a confidently wrong number about a real company.

1

Platform-derived is separated from merchant-authored

Some commerce platforms serve an agent surface — a protocol manifest, an llms.txt — free on every storefront they host. Scored flat, signals like those measure which platform a retailer bought, not what the retailer did. They sit in their own bucket, and we publish a platform-adjusted score that counts only what the merchant authored. A retailer whose platform hands it nothing is lifted by that adjustment; one that got it free is nudged down.

2

Blocked is unknown, never zero

Roughly a third of large U.S. retail refuses a plain standards-compliant client, robots.txt included. Every signal carries a fetch status, and a signal we could not observe is removed from both the numerator and the denominator. It never silently reads as a failure. Where too little of the instrument applied — below 85% coverage — we publish the pillars we measured and no composite at all, because a score from half an instrument is not comparable to one from a whole one.

3

Assert on content, never on status

One major retailer serves a bot interstitial at HTTP 200; another serves an 800KB HTML page under a 404. Every file probe checks the content type and the shape of the body before it believes a status code. Without that, roughly half of all llms.txt findings across this industry are false positives.

4

Absence is not evidence of absence

Several of the largest retailers are named launch partners for agentic-commerce protocols and publish no manifest at all, because they integrate through a private channel instead. Presence of an artifact scores. Absence scores nothing and is never described as a failing.

The measurement behind rule 1

Eleven of the companies in this index publish an agentic-commerce protocol manifest at /.well-known/ucp. They split cleanly in two, and the split is not about effort.

StorefrontVersionCapabilitiesBuilt by
Six storefronts on one commerce platform2026-08-258the platform
Academy Sports2026-01-234the retailer
Ace Hardware2026-01-233the retailer
Chewy2026-01-233the retailer
Scheels2026-04-083the retailer
Lululemon2026-04-081the retailer

Every storefront on the same commerce platform reports exactly eight capabilities on exactly the same protocol version. Every retailer that built its own reports between one and four. Fingerprinted on version, capability keys, transports and payment handlers, the platform-hosted profiles collapse into identical shapes — three unrelated retailers did not independently arrive at a byte-identical manifest.

So the signal a naive index rewards most — capability breadth — runs backwards to merchant effort. A retailer that stood up its own integration scores three. A tenant that did nothing scores eight. That is why protocol surface sits in its own bucket here, why every row carries a platform-adjusted score, and why the build fails if the adjustment ever stops lifting the retailers whose platform gave them nothing.

Rendering mode, and a decision we reversed

Track A reads a rendered DOM, captured in a real browser, uniformly for every company — so a retailer that injects its markup client-side is measured on what it actually publishes, and one that refuses a plain crawler is still measured rather than excluded.

The first build of this index scored the raw server HTML instead — what a plain crawler receives. That is a defensible thing to measure and we abandoned it, because of what it cost: 31 of the largest retailers in the country became unscorable, not because their catalogs were thin but because their edge refuses an ordinary fetch. Our own extractor read most of them perfectly well. “We could not score Costco” was a statement about our transport, not about Costco’s data.

Rendering also separates two things the first design conflated: “this retailer publishes no product markup” is a real finding that earns a real score, while “we could not see whether they do” is not a finding at all. A benchmark should be able to tell those apart.

We got this wrong once, in this direction, and it is worth stating. An earlier draft of this page used Costco as the example: rendered on our own hardware its pages returned no structured data, so we described it as a retailer publishing none. Through a better browser the same pages return a valid Product node, and Costco scores 67. The measurement was at fault, not the retailer — which is exactly the failure this instrument is built to avoid, and the reason the bottom of the ranking gets checked by hand before anything is published.

We did not lose the crawler question, we promoted it. Whether a plain, standards-compliant client can read the same page is now measured for every company and published on its scorecard and in the dataset. It moved from a coverage gap that silently dropped companies out of the ranking to a stated fact about each of them — which is where its force was all along.

The mode is applied identically to all of them. Measuring blocked retailers one way and open ones another would make the two populations incomparable, which is a worse defect than measuring fewer companies.

What we refuse to score

A benchmark is defined as much by what it leaves out. These are real things that matter, and we do not score any of them, because we cannot observe them fairly from a retailer's own public site:

  • Merchant feeds. The product feeds retailers send to Google, OpenAI and Perplexity are private by construction, and for most answer engines they are the actual selection substrate. Anyone claiming to measure them from a crawl is not measuring them.
  • AI referral traffic. Estimating it requires a third-party panel. Panels describe what happened to a retailer; we measure what a retailer ships. Only one of those is checkable by the retailer, and only one tells them what to fix.
  • Whether a specific AI crawler is actually admitted. robots.txt states intent; the edge decides reality, per IP and per fingerprint. Testing it properly would mean originating from a vendor's own address space.
  • Protocol endpoints behind authentication — checkout sessions, carts, orders. No credentials, no measurement, no score.
  • Anything inside a walled marketplace. It is not on the retailer's own domain and is not comparable across the list.

One consequence worth stating plainly: structured data here is scored as merchant-listing eligibility and machine legibility, not as a ranking factor in any assistant. Google's own guidance says no special markup or file is required for its generative search features. We are not going to claim otherwise to make the score sound more valuable than it is.

The six statuses

Unscored is never zero, and the unscored states say completely different things. Conflating them would misrepresent most of them.

Scored

Enough of the instrument applied to publish a comparable score.

Partly observable

We read the storefront, but too little of the instrument applied to publish a score that is comparable to the rest of the list. The pillars we did observe are shown.

Not observable

The storefront returned no readable product page to a standards-compliant client. Some of these do serve a full browser — what we report is that an ordinary crawler cannot read them, not that the pages are empty or that the retailer publishes nothing.

Crawling declined by robots.txt

This retailer's robots.txt disallows crawling its product paths for every user-agent, and we honoured it. Nothing here is a judgement — an index that scored retailers on crawler policy while ignoring that policy itself would be worthless.

No public catalog

No product detail pages reachable without a login, membership, or store selection. This is a finding about the retailer, not a measurement failure.

Not yet sampled

Our pipeline has not measured this company in this edition. This is a gap on our side.

Calibration, and changing the rubric

The attribute-depth scale was calibrated once against the observed distribution and then frozen. Any future change to a threshold ships in the same commit as a change to this page, and is versioned to the edition it applies to. Re-crawling and re-scoring at the same time would move a score for two reasons at once, which makes it impossible to say what actually changed — so the scorer can replay every score from its saved audit trail without refetching anything.

Every published-data change is logged in public on the corrections page. Previous editions stay live at their own URLs. If we get something wrong about your company, that is the fastest way to get it fixed.

Where the rubric comes from

The signals are grounded in public specifications: schema.org's Product, Offer, ProductGroup and MerchantReturnPolicy vocabularies; Google's published requirements for merchant listing experiences; the GS1 GTIN standard, including check-digit validation, so a field that says "gtin" but is not one does not score as one; RFC 9309 for robots.txt; and the published crawler documentation of the major AI vendors. The universe is the NRF/Kantar Top 100 Retailers 2026 list, ranked on 2025 U.S. retail sales.

Beyond that, the weights reflect Anglera's own operating experience with product data at scale. They are a point of view, published so it can be argued with. No third party has endorsed this index, and no company in it was consulted before publication.

Take the data

The full dataset is published under CC BY 4.0 as JSON and CSV. Cite it, recompute it, or check our arithmetic.