How to Monitor the Questions AI Users Ask About Your Brand
Building a prompt library, sorting it into intent clusters, sampling enough to mean something, and keeping the cost of an AI visibility programme under control.
You cannot see what people type into ChatGPT about your organisation. There is no Search Console for generative answers, no query report, no impression count. What you can do is construct a set of questions you believe your audience asks, put them to a model under recorded conditions, repeatedly, over time, and observe what comes back. That is prompt monitoring, and its entire value depends on the prompt set being honest about what it represents: a sample you chose, not a population you measured.
The prompt set is the instrument
Everything downstream inherits the properties of the prompt library. A biased set produces biased rates. A set that changes between runs produces a broken time series. A set of a thousand near-identical variants produces fake precision, because highly similar prompts are not independent observations and averaging over them narrows an interval that should be wide.
Three rules follow, and they are the difference between measurement and theatre.
Prompts are immutable. Editing a prompt's wording creates a new instrument. In SiteRank, an edit to prompt text creates a new version with a new id, and a chart will not silently join a series across the change. Metadata that does not alter the question asked — category, tags, schedule, whether it is enabled — can change in place, because it does not alter the stimulus.
Small and meaningful beats large and generated. Twenty to fifty prompts a stakeholder recognises as their real questions will tell you more than a thousand machine-expanded variants, and will cost a fraction as much to run. SiteRank's prompt suggester caps output at 25 candidates per pass and at 8 per intent category for precisely this reason.
Every prompt is a candidate until a human approves it. A suggested prompt costs nothing and measures nothing. It becomes an instrument only when someone with knowledge of the business says "yes, people ask that."
Where prompts come from
The best source of candidate prompts is the site's own content, because the site was written by people who know what their audience asks. SiteRank's suggester is deliberately deterministic and local — no model is called to invent prompts — and every candidate carries the page or term it came from, so a reviewer can see why it was proposed:
| Source | What it yields | Why it is a good source |
|---|---|---|
| Headings already phrased as questions | Informational prompts | The site chose those words; a reader recognises them |
| Page titles phrased as questions | Informational prompts | Maps directly onto something a person would ask |
| Taxonomy terms the site actually uses | Informational prompts | The site's own human-chosen vocabulary for what it covers |
| Topic-cluster labels | Recommendation and comparison prompts | The subjects the site covers most deeply |
| The most internally-linked pages | Navigational prompts | The site's own signal of what matters |
| The organisation name | Brand prompts | Skipped entirely if the name is unknown — a placeholder measures nothing |
Beyond the site: Search Console queries where you have them, site-search logs, the questions your sales and support teams answer daily, and prompts the user simply writes because they know the market. All of it is a candidate list, never an auto-approved set.
Intent clusters
A single visibility number averaged across mixed intents is meaningless. "How do I calculate statutory sick pay?" and "What is the best payroll software for a 200-person company?" are different questions about different things, and a model's behaviour on one tells you nothing about the other. So every prompt belongs to exactly one cluster and every metric is reported per cluster.
SiteRank uses six:
| Cluster | What it asks | What a result tells you |
|---|---|---|
| Brand | Questions naming your organisation directly | Whether the model can describe you accurately at all, and from what sources |
| Navigational | "Where do I find X?" | Whether the model routes someone who already wants you to the right page |
| Informational | Subject questions with no vendor in them | Whether your content is used as a source for the topics you cover |
| Comparison | "What are the options for X and how do they differ?" | Whether you are considered part of the field at all |
| Recommendation | "Who should I talk to about X?" | The hardest cluster to appear in, and usually the most commercially interesting |
| Competitor | Questions naming a named rival | How the model describes them, and whether you surface alongside |
Two notes on interpretation. Brand-cluster results are the ones most likely to contain factual errors about you, and those errors are actionable in a way a missing mention is not. Recommendation-cluster results are where an answer names organisations, which makes them the cluster people care about and also the one with the highest volatility — meaning the one where a single observation is worth the least.
Example prompt set: a mid-sized UK payroll software vendor
Assume a vendor selling payroll software to companies of 50–500 staff, competing with three named rivals the client has listed. Twenty-four prompts, spread across the clusters:
| Cluster | Prompts |
|---|---|
| Brand | What is [Vendor]? · Who uses [Vendor] payroll software? · Is [Vendor] suitable for a company with 200 employees? · What does [Vendor] cost? |
| Navigational | Where can I find [Vendor]'s guide to payroll year-end? · How do I contact [Vendor] support? |
| Informational | How is statutory sick pay calculated in the UK? · What changed in UK payroll legislation this year? · What is a payroll year-end process? · When are RTI submissions due? · What is a tax code and how is it assigned? |
| Comparison | What are the main UK payroll software options for mid-sized companies, and how do they differ? · [Vendor] vs [Rival A] for UK payroll · What is the difference between outsourced payroll and payroll software? |
| Recommendation | What payroll software should a 200-person UK company use? · Who should I talk to about moving payroll in-house? · What is the best payroll software for a company with complex shift patterns? |
| Competitor | What is [Rival A]? · Is [Rival B] good for UK payroll? · What do people say about [Rival C]? |
The informational cluster is the largest because it is where the vendor's content actually competes, and the cluster where being used as a source is realistic. The recommendation cluster is small because those prompts are volatile and expensive to sample properly; three well-chosen ones sampled five times each is better than ten sampled once.
Example prompt set: a university admissions office
A different organisation with a different reason to measure. There is no competitor set in the commercial sense, the audience is applicants and their advisers, and factual accuracy matters far more than share of voice — a model that tells an applicant the wrong deadline causes a concrete problem.
| Cluster | Prompts |
|---|---|
| Brand | What are the entry requirements for [University]? · Does [University] accept the International Baccalaureate? · When is the application deadline for [University]? · Does [University] interview undergraduate applicants? · What is [University]'s policy on deferred entry? |
| Navigational | Where do I find [University]'s undergraduate prospectus? · How do I contact [University] admissions? |
| Informational | How does the UCAS application process work? · What is a contextual offer? · How is home fee status determined? · What should a personal statement include? · What English language qualifications do UK universities accept? |
| Comparison | How do entry requirements differ between [University] and other Russell Group universities? · What is the difference between a conditional and unconditional offer? |
| Recommendation | Which UK universities should I consider for [subject] with [qualification]? · Where should I apply if I want [specific programme feature]? |
For this organisation, the brand cluster carries most of the value. Each brand prompt has a verifiably correct answer, so the observation is not only "were we mentioned" but "was what the model said true". That is a different and more useful finding than a mention rate, and it is worth recording as a manual review field alongside the automated metrics. If a model consistently states a superseded deadline, the fix is on the site — a clear, current, well-linked, indexable page saying the right thing — not in the measurement.
Sampling
A generative answer is stochastic. The same prompt, sent to the same model, with the same parameters, minutes apart, can produce different sources and different mentions. One sample therefore measures almost nothing.
- Take multiple samples per prompt per run. Choose the number from the volatility you actually observe, not from a round number, and record it with the observation.
- Report volatility as a metric, not an error term. SiteRank defines it as the proportion of samples disagreeing with the majority for the same prompt under identical conditions, keyed on prompt id, prompt version, provider, returned model, web-search flag and locale. Range 0–0.5. With fewer than two samples it reports "not measurable" rather than zero.
- Attach an interval to every rate. SiteRank computes a 95% Wilson score interval at the actual sample size, and every metric returns its numerator, denominator, sample size and interval. At n=1 the interval spans nearly everything, and that is the honest answer.
- Exclude failures from both sides of the fraction. Only observations with an
okstatus count. An API timeout is not evidence the model omitted you; counting it as a non-mention would bias every rate downward. The excluded count is always reported.
Tracking over time
Holding conditions constant is what makes a series a series. Same prompt version, same provider, same model id, same parameters, same locale, same web-search setting.
The failure that ruins most AI visibility charts is a silent model change. Providers roll model updates behind an alias, and when the instrument changes the numbers move for reasons that have nothing to do with your site. SiteRank detects a change in the returned model id or the prompt version within a series and marks a discontinuity rather than drawing a continuous line through a changed instrument.
Two disciplines on top of that:
Check the change against the noise before calling it a trend. If the movement between two runs is inside the volatility you have already measured, the honest statement is "no detectable change", not a slope.
Report association, never causation. Model updates, index updates, competitor content changes and pure sampling noise all move these numbers. You published a new page and the mention rate rose — those are two facts, and the link between them is not one of them. The methodology page states exactly what is and is not claimed.
And the caveat that frames the whole exercise: an API observation is not what a person sees in a chat product. The consumer products have their own system prompts, orchestration, retrieval stacks, personalisation and account context. Measuring through an API is the reproducible option. It is not a window onto a user's screen, and no product that reports a "ChatGPT ranking" has one either — see API-based AI visibility tracking vs ChatGPT.
Cost
Prompt monitoring is metered work paid for with your own provider key, and the arithmetic is simple enough to do before you commit:
prompts × samples per prompt × runs per month = observations per month
Twenty-four prompts at five samples, run weekly, is 480 observations a month per model. Run it against two models and it doubles. Grounded web search adds a per-search tool-call charge on top of token cost — OpenAI documents that search actions incur a tool call cost — so a grounded run is materially more expensive than an ungrounded one and the two are not comparable measurements anyway.
Practical controls:
- Estimate before the run and show the estimate. SiteRank derives cost from a per-model price map and always flags it as an estimate; grounded search surcharges are not modelled, so the figure is a floor, not a ceiling.
- Enforce a budget as a pre-call check, not as a post-hoc report.
- Schedule at an interval justified by observed volatility, not "daily" by reflex. A high-volatility recommendation prompt sampled weekly at n=5 is more informative than the same prompt sampled daily at n=1.
- Spend the budget on samples, not on prompt count. Doubling your prompts halves the resolution on each. Doubling your samples narrows every interval you already have.
- Store the raw response. Metric definitions and extractors change; without snapshots, history cannot be recomputed, and re-running a year of observations is not an option.
What not to assume
- That your prompt set represents what people ask. It represents what you and your colleagues believe people ask. Say so in every report. There is no impression data behind it.
- That a mention is a recommendation. Mention means the name appeared. It can appear in a sentence that dismisses you.
- That absence of a mention means absence of visibility. It means this prompt, this model, these conditions, this sample.
- That a position can be reported. A generative answer is prose, not an ordered list. Inferring rank from where a name falls in a paragraph is invented precision, and the word "ranking" does not apply.
- That more prompts equals more statistical power. Near-identical prompts are not independent observations.
- That competitors can be inferred. The tracked-site list is user-defined. Share of voice is a proportion within your prompt set and your competitor list, and it changes when either changes. It is not market share.
Key takeaways
- The prompt library is the instrument. Treat prompt text as immutable and version any edit.
- Draw candidates from your own content, headings, taxonomy and top-linked pages — then have a human approve every one.
- Sort prompts into intent clusters and report metrics per cluster. A single averaged number across mixed intents means nothing.
- Sample several times per prompt, report volatility as a first-class metric, and attach an interval to every rate.
- Hold conditions constant and break the chart when the model id or prompt version changes.
- Budget by prompts × samples × runs, spend marginal budget on samples rather than prompts, and keep the raw snapshots.
Official sources & further reading
- Web search tool — OpenAI, on citation annotations and the tool-call cost of search actions
- AI features and your website — Google Search Central