Blog
What Should Agencies Report When AI Visibility Results Are Not Stable?
AI answers fluctuate by prompt, model, location, date, and source set. Agencies should report what was observed, how it was sampled, which technical conditions were checked, how confident the result is, and what happens next—not promise a fixed ranking position.
Start with the right definition of AI visibility
AI visibility is not one stable ranking metric. It is an observed set of outcomes from a defined sample of prompts, systems, dates, and user contexts. An answer may mention a company, recommend it, link to its site, quote its content, or omit it entirely. Those are different events and should not be collapsed into a single score without explanation.
This distinction matters because agencies are often asked for a familiar SEO-style answer: “What position are we in?” In many AI interfaces there is no fixed position, universal result set, or reproducible public index. The more defensible question is: “Under which test conditions did the brand appear, how did it appear, and what evidence supports that observation?”
- Mention: the brand appears in the answer, whether or not it is recommended.
- Recommendation: the system suggests the brand as a provider, product, or option.
- Citation: the answer links to, names, or visibly relies on a source from the brand’s site or another page about it.
- Share of voice: the brand’s observed presence compared with named competitors within the same prompt sample.
- Technical readiness: measurable conditions that make content accessible and interpretable, such as crawl access, HTML availability, metadata, structured data, and sitemaps.
A useful AI visibility measurement guide should therefore be treated as a measurement protocol, not a promise that a site will rank or be cited.
Report five layers instead of one unstable score
A client report should separate outcomes from explanations. A concise framework uses five layers: observed outcomes, sampled evidence, technical state, uncertainty, and next actions. Keeping these layers distinct makes the report useful even when results move between reporting periods.
| Layer | What to report | What it does not prove |
|---|---|---|
| Observed outcomes | Mentions, recommendations, citations, linked sources, competitor appearances, and changes from the prior sample | A permanent ranking, guaranteed traffic, or causal attribution |
| Sampled evidence | Exact prompts, systems, dates, locations, screenshots or exports, and repeat tests | Performance across every user, model, or query |
| Technical state | Crawler access, robots rules, HTML, metadata, JSON-LD, sitemap, llms.txt, and indexable content | That technical fixes will produce citations |
| Uncertainty | Sample size, reproducibility, known variables, confidence level, and missing data | That an unstable result is a confirmed trend |
| Next actions | Specific content, technical, measurement, or research tasks with owners and dates | A promise that the action will change an answer |
This structure also prevents a common reporting error: treating a technical audit finding as an outcome. For example, an accessible sitemap is a verifiable condition. It is not evidence that an AI system used the sitemap or will cite a page.
Define the sample before collecting results
Unstable AI results become more interpretable when the sample is fixed in advance. Agencies should document the prompt set and avoid silently replacing prompts that produce inconvenient results. A small, stable panel is usually more useful than a large set that changes every month.
- Group prompts by intent, such as category discovery, comparison, problem-solving, local selection, and buyer intent.
- Record the exact wording, including punctuation and qualifiers such as location, budget, industry, or audience.
- Specify the product or model tested, account state where relevant, date, time window, and location.
- Run the same prompts on a repeat schedule, then add a clearly labelled exploratory sample for new questions.
- Capture the complete answer, citations, linked sources, and named competitors—not just whether the brand appeared.
For buyer-intent work, separate prompts such as “best accounting software for a small nonprofit” from informational prompts such as “how does fund accounting work?” The former may reveal recommendation behavior; the latter may reveal source or content coverage. Combining both into one rate can conceal useful differences. BatSignal’s buyer-intent content guide provides a practical way to organize these categories.
Use a comparison window
For each reporting period, show a current sample against a prior sample collected with the same protocol. If the prompt set, model, or location changed, mark the comparison as non-equivalent. This simple label is more honest than presenting two percentages in a trend chart when the underlying tests are different.
Use outcome metrics with denominators and definitions
Every percentage in an agency report should answer three questions: what was counted, out of how many tests, and during what period? “AI visibility increased 20%” is incomplete. “The brand was mentioned in 8 of 20 fixed prompts this month, compared with 5 of 20 last month” is auditable, even if the sample is limited.
| Metric | Suggested calculation | Reporting note |
|---|---|---|
| Mention rate | Prompts with a brand mention ÷ prompts tested | Count one prompt once unless repeated runs are intentionally analyzed separately |
| Recommendation rate | Prompts where the brand was recommended ÷ eligible prompts | Define what qualifies as a recommendation before reviewing results |
| Citation rate | Prompts with a brand-owned source cited or linked ÷ prompts tested | Separate first-party citations from third-party citations |
| Competitor share of voice | Brand appearances ÷ total brand appearances among the defined competitor set | A relative sample measure, not market share |
| Repeat agreement | Repeated runs with the same coded outcome ÷ total repeated runs | Useful for showing stability or instability in a controlled test |
Keep the raw counts beside the percentages. A change from 1 of 5 prompts to 2 of 5 is not equivalent to a change from 20 of 100 to 40 of 100, even though both can be described as a 20 percentage-point increase. Small samples should be presented as directional observations, not precise estimates.
Show evidence, not just dashboard numbers
A client should be able to understand why a metric changed. Include a compact evidence appendix or expandable record containing the prompt, answer excerpt, cited URLs, timestamp, system, location, and coding decision. Screenshots can help preserve what the interface displayed, but they should not replace structured notes.
- Keep the exact answer text or an export where permitted by the product’s terms.
- Record whether the brand was mentioned, recommended, compared, or cited.
- List each cited URL and identify whether it is first-party, publisher, directory, review, or another source type.
- Note whether the answer used current web links, unattributed knowledge, or an unclear source path.
- Flag ambiguous cases for a second reviewer rather than forcing every answer into a positive or negative category.
Evidence should also include negative findings. If a brand was mentioned but no first-party page was cited, that is materially different from a cited recommendation. If a competitor appeared in a comparison answer, record the basis shown by the system rather than assuming the competitor has a stronger overall position.
Separate access, training presence, and live citations
AI visibility reports often confuse three different pathways by which a site may be present in an ecosystem.
| Pathway | What can be checked | Appropriate conclusion |
|---|---|---|
| Crawl access | Robots directives, server responses, blocked paths, rendered HTML, and crawler-specific rules | The site is more or less accessible to a named crawler or retrieval process |
| Training or open-web presence | Historical crawl or corpus records, such as Common Crawl coverage where available | The domain or pages were observed in a dataset; this does not show current answer use |
| Live answer citation | A defined prompt test that returns a linked or named source | The source appeared in the tested answer under the recorded conditions |
A robots file that permits a crawler does not force a model to use the content. An entry in Common Crawl training presence does not establish that a current answer was generated from that crawl. Likewise, a live citation does not prove that the site is broadly visible across all prompts.
Technical checks still matter. Agencies can report findings from robots.txt and AI crawler checks, crawlable HTML versus a SPA, JSON-LD and AI discovery, metadata, sitemaps, and llms.txt. The wording should remain precise: these are conditions, signals, or potential barriers—not ranking levers.
Assign uncertainty levels and explain volatility
A useful report does not hide volatility; it classifies it. Agencies can use a simple confidence label based on repeatability and sample quality. The label is a reporting aid, not a statistical claim unless the agency has a formal sampling design.
| Label | Use when | Recommended wording |
|---|---|---|
| High confidence | The same protocol was repeated, the outcome is clear, and the pattern appears across enough observations for the account | “Observed consistently in the defined sample.” |
| Medium confidence | The result is clear in one sample but has limited repeats, or some conditions changed | “Observed, but the comparison has important limitations.” |
| Low confidence | The result appeared once, was ambiguous, or could not be reproduced | “A directional observation requiring confirmation.” |
When results fluctuate, list the variables that may explain the movement: prompt wording, geography, language, logged-in state, model version, search freshness, cited-source changes, answer length, personalization, and test timing. Do not select one cause simply because it is plausible. If no controlled test supports the explanation, describe it as a hypothesis.
A practical volatility note might say: “The brand appeared in 7 of 20 prompts this period versus 6 of 20 previously. Three prompts changed model behavior on repeat testing, so the movement is not treated as a confirmed gain. Technical access remained unchanged.” That is more useful than “visibility is up 16.7%.”
Connect technical changes to measurable checks
Technical recommendations should have their own acceptance criteria. If an agency recommends allowing a crawler, improving server-rendered HTML, adding JSON-LD, or clarifying a sitemap, report whether the implementation was completed and whether the relevant check now passes. Then measure answer outcomes separately over a defined observation window.
- State the issue in observable terms, such as a blocked path, missing title, absent canonical, empty initial HTML, or invalid structured data.
- Describe the proposed change and the pages or templates affected.
- Verify deployment with a repeatable scan, fetch, or validation step.
- Record the date the change became available to crawlers and users.
- Continue the fixed prompt sample without claiming that any later answer change was caused by the fix unless a stronger controlled design supports that conclusion.
This approach makes technical work accountable without overstating causality. The AI visibility guide and AEO checklist can help teams turn broad recommendations into checkable implementation tasks.
Build the report around decisions and next actions
The final section should tell the client what to do with the evidence. Keep actions limited, specific, and prioritized. A report that lists every possible GEO tactic without an owner or test plan creates activity rather than learning.
- Fix confirmed access or rendering barriers before interpreting content-level results.
- Improve pages that directly answer recurring buyer-intent questions, then add those prompts to the monitoring panel.
- Review cited third-party sources when the brand is recommended but the first-party site is absent from the evidence.
- Test competitor and category language where the brand is omitted, rather than assuming more generic content will help.
- Schedule a repeat sample after a defined interval and preserve the same protocol for comparison.
FAQ
Why do AI visibility results fluctuate so much?
AI answers can change when the prompt wording, model, location, user context, retrieval sources, index state, or date changes. Some systems also generate answers from a mixture of live web results and previously learned information. A change in one answer does not necessarily mean that the site gained or lost a durable visibility position.
Should agencies report an AI visibility percentage?
They can, but only when the percentage has a clear denominator and sampling method. For example, report the percentage of a defined prompt set in which a brand was mentioned, recommended, or cited during a specified test window. Do not present that figure as a universal share of all AI answers.
How should an agency explain a drop that cannot be reproduced?
Label it as an observed change with low or medium confidence, show the original and repeat samples, and state what was held constant. If the result does not reproduce under the same conditions, avoid attributing it to a site change. Recommend monitoring or a controlled follow-up test instead.
Does passing an AI crawler check mean a site will be cited?
No. Crawler access is an eligibility condition, not a citation outcome. A site may be accessible to crawlers and still be absent from a particular answer because of relevance, authority, source selection, freshness, competition, or model behavior.
What should be included in an AI visibility dashboard?
Include the test date, prompts, model or product, location where relevant, brand and competitor mentions, recommendation and citation rates, linked sources, sample size, confidence or limitations, technical findings, changes since the last period, and clearly assigned next actions.