Blog

How to Build a Reliable Prompt Set for AI Visibility Measurement

AI visibility data is only as useful as the prompts behind it. A repeatable prompt set makes results less sensitive to wording, location, product names, and one-off model behavior.

A prompt set is the measurement instrument behind an AI visibility report. If the set overuses your brand name, asks only one type of question, or changes wording each month, the resulting percentages can look precise while measuring very little. The goal is not to find a handful of prompts that produce favorable answers. It is to create a balanced, documented sample that reflects how real people ask for information and compare options.

Start by defining what “visibility” means

AI visibility is not a single event. A site can be accessible to an AI crawler without appearing in a live answer. A page can be present in an open-web or training corpus without being cited today. An answer can mention a company without linking to it, or cite a page without recommending the company. Before writing prompts, decide which outcome you are measuring.

A practical measurement model separates at least four layers:

  • Access: Can relevant crawlers reach the site, and do robots.txt rules permit the intended access?
  • Content readiness: Are important pages available as crawlable HTML with usable metadata, structured data, internal links, and a sitemap?
  • Presence: Is the brand or its content represented in an indexed or training-related web corpus?
  • Live answer visibility: Does an AI system mention, recommend, compare, or cite the brand for a particular question?

This distinction matters because prompt testing primarily measures the final layer. It can reveal live mentions, recommendations, citations, and competitor share of voice, but it cannot by itself prove why an answer was generated. BatSignal’s measurement methodology treats these signals separately rather than collapsing them into a single score.

Build intent buckets before writing individual prompts

Intent buckets prevent the prompt set from becoming a collection of questions someone happened to think of. Each bucket should represent a recognizable user job. The exact buckets depend on the business, but most commercial sites can begin with the following structure.

Intent bucketWhat the user is trying to doExample prompt patternUseful visibility signals
Category discoveryUnderstand the available optionsWhat are the main ways to solve [problem]?Mention and category share of voice
EducationLearn how a process, product, or issue worksHow does [category] work for [use case]?Accuracy, explanation, and source citation
ShortlistingCreate an initial list of providers or productsWhich [providers/products] should I consider for [need]?Recommendation rate and position in list
ComparisonChoose between named or unnamed alternativesCompare [brand/product] with [alternative] for [criterion].Inclusion, strengths, weaknesses, and citations
Buyer intentMove toward a purchase or vendor decisionWhat is the best [product/service] for [role] in [location]?Recommendation, qualification, and commercial relevance
Problem recoveryFind help after a poor result or failed attemptWhat should I do if [problem] happens?Trust signals and helpful citations

Do not assume every bucket deserves equal weight. If the business sells to enterprise buyers, high-consideration comparison and implementation prompts may matter more than broad educational questions. Document the weighting before collecting results. Otherwise, a change in the mix can be mistaken for a change in visibility.

Represent users, locations, and products explicitly

Real prompts contain context. “What is the best project management software?” is not equivalent to “What is the best project management software for a 20-person remote agency in the UK?” The second question gives the model more constraints and may produce a different set of recommendations.

Create small, deliberate dimensions rather than trying to model every possible user. For example:

  • User roles: founder, procurement manager, developer, marketer, operations lead, or consumer.
  • Experience level: beginner, experienced practitioner, or specialist buyer.
  • Company or household context: small business, mid-market, enterprise, regulated organization, or individual.
  • Location: country, region, or city where availability, pricing, regulation, or service coverage changes the answer.
  • Product scope: core product, product tier, service line, use case, or adjacent category.
  • Decision criteria: price, reliability, integration, support, security, speed, accessibility, or sustainability.

Use a matrix to keep the combinations manageable. A six-role by three-location by four-intent design creates 72 combinations before product and comparison variations. That may be excessive for a first baseline. Select representative combinations and record why they were chosen.

Locations deserve special care. A location in the prompt is not the same as the location of the model session or the location implied by the account. Record both where possible. If a test is run from different regions, treat that as a test variable, not as a hidden source of variation.

Balance branded, unbranded, and competitor prompts

A reliable set needs both prompts that name the brand and prompts that do not. Branded prompts test recognition and entity understanding; unbranded prompts test whether the brand appears when the user has not already supplied it.

Prompt typeExampleWhat it can tell youMain limitation
BrandedWhat are the strengths and weaknesses of Brand A?How the system describes the brand and whether it cites relevant pagesDoes not test discovery
Unbranded categoryWhat are the best tools for [use case]?Whether the brand enters a consideration setResults can vary with broad wording
Competitor comparisonBrand A or Brand B for [criterion]?Relative positioning and comparative claimsMay encourage a narrow, artificial frame
Problem-ledHow can a team solve [problem]?Whether the brand appears as a relevant solutionThe brand may not be expected in every answer
Source-seekingWhich sources explain [topic] accurately?Citation and source-selection behaviorMeasures publisher authority as well as brand visibility

A common mistake is to report branded visibility as if it were category visibility. A brand that appears in nearly every “tell me about Brand A” prompt may still be absent from unbranded buyer questions. Report the two populations separately.

Design comparison frames instead of asking for generic winners

“What is the best tool?” is often too vague to produce a stable or interpretable result. It also invites the model to choose criteria that may not match the buyer’s needs. Comparison frames make the decision rule visible.

Useful frames include:

  1. Best for a specified role or organization size.
  2. Best when a particular integration, workflow, or compliance requirement matters.
  3. Best for a defined budget or purchasing model.
  4. Best for a location with relevant availability or regulatory constraints.
  5. Best for a stated trade-off, such as simplicity versus customization.
  6. Best alternative when a user has rejected a familiar option.

Vary the direction of the comparison. Test “Brand A versus Brand B,” “Brand B versus Brand A,” and an unbranded request where both are eligible. If only one direction is tested, ordering effects can be mistaken for a product preference.

Also record what counts as a favorable result. A mention in a list is not necessarily a recommendation. A recommendation without a citation is not the same as a cited recommendation. Define separate fields for mention, recommendation, citation, citation accuracy, relative position, and any material qualification.

Use controlled variation and repetition rules

AI outputs are stochastic and can change with model updates, conversation history, location, tools, and other session conditions. Repeating the same prompt does not eliminate that variability, but it helps estimate it.

Use three prompt groups:

  • Core prompts: unchanged across reporting periods so trends remain comparable.
  • Paraphrase prompts: equivalent questions with different natural wording to test wording sensitivity.
  • Exploratory prompts: new use cases, products, competitors, or market language that may be promoted into the core set later.

For each core prompt, define a repetition rule before running the scan. For example, run each prompt three times in the same model and session conditions, or run a smaller number of repetitions when the test set is large. Do not quietly remove inconvenient outputs. Store every response, timestamp, model or product name, locale, and relevant settings.

If resources are limited, prioritize balanced coverage over excessive repetition. Ten repetitions of a narrow branded set are usually less informative than three repetitions of a broader, well-designed set. Report the number of observations behind every rate.

Write prompts that are specific without becoming leading

A prompt should provide enough context to represent a real question, but not so much that it forces the answer. Leading prompts such as “Why is Brand A the best option?” measure compliance with your premise more than visibility.

A useful prompt template is:

  1. User situation: who is asking and what are they trying to accomplish?
  2. Category or problem: what decision or task is in scope?
  3. Constraints: where, for whom, at what scale, or under what budget?
  4. Evaluation frame: which criteria should the answer consider?
  5. Output request: shortlist, comparison, explanation, steps, or sources.

For example: “I run a 15-person design agency in the UK and need a project management tool that integrates with our existing workflow. Which options should I shortlist, and what trade-offs should I investigate?” This is more useful than inserting a brand into the question, because it tests whether the brand is surfaced for a plausible buyer need.

Keep prompts concise enough to reflect real use. Long briefs can be useful for a specialized test, but they should be labeled as such. The more assumptions a prompt contains, the harder it becomes to generalize the result.

Score answers with a consistent annotation scheme

Collecting outputs is not enough. Two reviewers can read the same response and disagree about whether it recommended a brand or merely mentioned it. Create annotation rules before reviewing results.

FAQ

How many prompts do I need to measure AI visibility?

There is no universal number, but a useful starting set usually contains 30 to 60 prompts across several intent buckets, audience roles, locations, products, and comparison frames. Add repetitions only after the core set is balanced.

Should I use the exact same prompts every time?

Use a stable core set for trend measurement, then maintain a smaller exploratory set for new wording, products, or market questions. Changing the entire set makes month-to-month results difficult to interpret.

Should prompts include my brand name?

Use both branded and unbranded prompts. Branded prompts test whether an AI system recognizes and describes your entity. Unbranded prompts are more useful for measuring discovery, recommendation, and category-level visibility.

Do AI answers prove that a website was crawled or used for training?

No. A live answer citation is different from crawler access, inclusion in a web corpus, or training presence. Measure these as separate signals rather than treating a citation as proof of all three.