Prompt testing (AI visibility tracking)
Prompt testing is the practice of running a fixed set of buyer questions through AI assistants on a repeating schedule and recording whether your brand, products, and URLs appear in the answers. It is the standard method behind AI visibility tools, and it is sampling rather than reporting: the underlying systems are non-deterministic and none of them publish the data directly.
How a panel is built
Four decisions define it, and all four should be written down before the first run.
The prompt set. Real buyer questions, frozen so results stay comparable. Changing the set resets the baseline whether you intend it to or not.
The surfaces. Which assistants and which models, recorded by version where available, since providers update models continuously.
The cadence. Weekly or monthly, with several repeat runs per prompt so you can see within-prompt variance rather than mistaking it for change.
The scoring. At minimum: brand mentioned, specific product named, your URL cited, competitor set present. Keep those separate. They move independently and collapsing them into one visibility score destroys the diagnostic value.
Writing prompts a buyer would actually type
The most common mistake is testing prompts that flatter the brand. "Is Acme Supply a good electrical distributor" tells you almost nothing, because nobody asks it.
Buyers ask constrained, specification-shaped questions, and the good sources for them are already inside the business: quote requests, site-search logs, support tickets, and the unbranded queries in Search Console. In distribution the productive shapes are compatibility ("what breaker fits a Square D QO panel"), supersession ("what replaced part number X"), selection under constraint ("150W high bay for a 30 ft ceiling, DLC listed"), and comparison ("difference between 316 and 304 stainless for coastal use").
Include prompts you expect to lose. A panel that only contains winnable questions produces a chart that goes up and teaches you nothing.
Reading the results without fooling yourself
Non-determinism is the whole methodological problem. Two runs of the same prompt on the same day return different sources. A week-over-week move on a 50-prompt panel run once each is well inside the noise, and reporting it as a change is how these programs lose credibility internally.
The useful outputs are not the headline percentage. They are the patterns: which categories you never appear in, which competitors keep showing up, and which of your own pages get cited. That last one is the actionable list, because it tells you what kind of content the retrieval step is finding on your domain.
When a category never appears, check coverage before touching content. A category where forty percent of SKUs lack the attribute the prompt filters on cannot be won with copywriting.
Where it fits alongside other measurement
Prompt testing sees answers; server logs see retrieval; analytics sees clicks. Each is blind to what the others catch, which is why a serious program runs all three.
Crawler logs are the underrated one and the cheapest to start. If OAI-SearchBot or PerplexityBot is not fetching your product pages, or is fetching them and receiving a JavaScript shell, no prompt panel result is going to be interpretable — you are measuring the outcome of a pipeline that is broken at step one.
Frequently asked questions
What is prompt testing in AI SEO?
Running a fixed set of questions through AI assistants on a schedule and recording whether your brand, products, or URLs appear in the answers. It is how AI visibility tools generate their numbers, since no assistant publishes visibility data directly.
How often should we run a prompt panel?
Monthly is enough for most catalogs, with several repeat runs per prompt so within-prompt variance is visible. Weekly runs on a small panel mostly produce noise, and the temptation to react to it is the main failure mode of these programs.
How many prompts should a panel contain?
Enough to cover your major categories and question types, including ones you expect to lose. Coverage across categories matters more than raw count, because a blended score hides the category-level gaps that are the actual output of the exercise.
Why do two AI visibility tools report different numbers for us?
Because they use different prompt sets, different models, different run counts, and different rules for what counts as an appearance. There is no standard, so the numbers are not comparable across vendors. Pick one method and track it against itself.