Home/Services/AI Search Visibility

Grow · Pillar 04

AI search visibility: measurement, not guesswork.

You cannot manage a channel you cannot see. This is the reporting discipline that makes generative search a measured channel rather than an anxiety.

How do you measure visibility in AI search?

AI search visibility is measured by running a fixed set of realistic buyer prompts across the major assistants on a schedule, then recording for each one whether the business is mentioned, how it is described, which sources are cited and which competitors appear. Tracked over time, that produces a share-of-answer metric comparable to share of voice in traditional search.

  • A fixed prompt set, so month-on-month comparison is valid.
  • Records citations and competitors, not just whether you appeared.
  • Separates your improvement from the model simply changing.

Why it matters

Generative search is the only channel nobody reports on.

Every other channel has a number. Organic has Search Console, paid has a platform, email has open and click rates. Generative search — which is increasingly where supplier research starts — typically has nothing, because it leaves almost no trace in analytics and because the answers are not public.

So it gets discussed anecdotally. Someone tries a prompt, gets a bad answer, and it becomes a board topic with no baseline to compare against. Measurement fixes that: a defined prompt set, run consistently, turns a worry into a trend you can act on.

What the report contains

Why anecdote is not data

A fixed prompt set run the same way every month, so the figures mean something when compared. Each report covers:

  • Nobody can say whether you appear in assistant answers or not.
  • A single bad answer someone noticed is driving strategy.
  • You do not know which sources the models cite about your market.
  • No baseline exists, so no change can be demonstrated.
  • Competitor presence in AI answers is completely unknown.
  • Referrals from assistants are invisible in your analytics.

Scope

What the monitoring covers.

This is a retained measurement service. It pairs with GEO, which is the work that changes the numbers.

  1. Prompt set design

    Twenty to forty prompts reflecting how your buyers genuinely ask — by problem, by service, by location, by comparison — agreed with you and then held stable.

  2. Multi-model testing

    The same prompts across ChatGPT, Gemini, Perplexity, Claude and Copilot, because presence varies substantially between them and an average hides that.

  3. Share of answer

    How often you are mentioned, how prominently, and in what terms — tracked as a trend rather than reported as a snapshot.

  4. Competitor comparison

    Who else appears for your prompts and how they are described. Frequently the most commercially useful part of the report.

  5. Citation source analysis

    Which pages and third-party sources the models draw on for your market, which tells you precisely where corpus work should be aimed.

  6. Referral attribution

    Identifying assistant-referred sessions in your analytics where the referrer is detectable, so the channel has a traffic figure as well as a visibility one.

Deliverables

What the measurement gives you.

The point is a decision, not a dashboard. Each report ends with what we would do next and why.

  1. A baseline, then a trend

    The first report establishes where you stand; subsequent ones show movement against the same prompts. Without a stable baseline, nothing in this channel can be claimed honestly.

  2. Knowing which sources matter

    The specific pages and third-party sites the models cite for your market. This converts GEO from guesswork into a target list.

  3. Competitive intelligence

    How rivals are described and positioned by the assistants your buyers use. Occasionally uncomfortable, consistently useful.

  4. Early warning on inaccuracy

    Wrong descriptions caught while they are a monitoring item rather than after a prospect has acted on one.

Is this the right answer?

When ai search visibility is worth doing — and when it is not.

We would rather lose a project at this stage than six weeks in. If the right-hand column describes you, say so and we will tell you what we would do instead.

Worth doing when

  • Nobody can say whether you appear in assistant answers.
  • One bad answer somebody noticed is driving strategy.
  • You are investing in GEO and want to know if it is working.
  • You need a number for this channel to take to a board.

Probably not when

  • You want the measurement to also fix the problem — that is the GEO work.
  • You will not act on the findings, in which case it is a report nobody reads.
  • You expect a complete census rather than a measured sample.
  • Weekly measurement, which mostly produces noise.

How we deliver

How the service runs.

Deliberately boring and repeatable. Changing the prompt set is what destroys the comparability that makes the data worth having.

  1. Month zero

    Design & baseline

    Agree the prompt set and the competitor list, run everything, and produce the baseline report with verbatim answers attached.

  2. Monthly

    Re-run & report

    The same prompts, the same models, the same method. Movement, citation changes and competitor shifts, with the raw answers included so nothing rests on our summary.

  3. Quarterly

    Review & recommend

    A session on what the trend means, what to fix in the corpus, and whether the prompt set needs extending as your market changes.

  4. Ongoing

    Separate signal from noise

    Model updates move results independently of anything you do. Tracking several models and holding the prompts constant is what lets us tell the difference.

What you receive

The things that actually land.

Artefacts, not adjectives. Everything below is listed in the scope document before a phase starts, so “done” is a defined state rather than an opinion.

  1. An agreed prompt set

    Twenty to forty prompts reflecting how your buyers genuinely ask, including location-qualified ones, held stable so comparison means something.

  2. A baseline report

    Where you stand across the major assistants before any work starts, with the raw answers attached rather than only our summary.

  3. Monthly share-of-answer tracking

    How often you are mentioned, how prominently and in what terms, as a trend rather than a snapshot.

  4. Named competitor benchmarking

    Who else appears for your prompts and how they are described. Frequently the most commercially useful part.

  5. Citation source mapping

    Which pages and third-party sources the models draw on for your market — the target list for any corpus work.

  6. A recommended action each month

    What to do next and why, or an explicit statement that nothing needs doing. The report exists to inform a decision, not to justify a retainer.

Golden Triangle

Why this matters here specifically.

In B2B markets with long cycles and few suppliers, a single assistant answer can shape a shortlist you never see.

  1. Small supplier sets Three names When an assistant names three suppliers for a specialist Midlands niche, being one of them is close to decisive. The stakes per answer are far higher than in a consumer market.
  2. Invisible research No referrer Assistant research mostly leaves no analytics trace, so without deliberate measurement this channel is genuinely invisible. That is precisely why it goes unmanaged.
  3. Regional prompts Location-qualified Buyers ask for suppliers near Birmingham, in the Midlands, or close to a named site. We include location-qualified prompts because that is how the research actually happens.

We monitor AI search visibility for businesses across the Golden Triangle, reporting monthly with the verbatim answers and cited sources attached.

Technology

How the measurement is done.

Consistent method, raw answers retained. You get the evidence, not just our conclusion about it.

  • Multi-model prompt runs
  • ChatGPT & Copilot
  • Gemini & Perplexity
  • Share-of-answer tracking
  • Citation source mapping
  • Competitor benchmarking
  • Referral detection
  • Verbatim answer archive
  • Change detection
  • Monthly reporting

Where we deliver this

AI Search Visibility across the Golden Triangle.

35 locations, each with a page written for it — the sectors it is actually built on, and what that means for this work. See all areas we serve.

Questions

AI search visibility, answered plainly.

Mostly asked by people who need a number to take to a board.

Can you actually track visibility in ChatGPT and Gemini?

Yes, by systematic sampling rather than by a reporting API, because none of the assistants publishes one. We run a fixed prompt set across the models on a schedule and record the outcomes, which gives a reliable trend provided the prompts stay constant. It is a measured sample, not a complete census, and we are explicit about that distinction.

What is share of answer?

The proportion of your target prompts where your business is mentioned, weighted by how prominently and how accurately. It is the generative-search equivalent of share of voice, and its value is comparative — against last month and against named competitors — rather than as an absolute figure.

How is this different from GEO?

This service measures; GEO changes the result. The monitoring tells you where you stand, which sources the models cite and who else appears, and the GEO work acts on that. They are sold separately because the measurement is genuinely useful on its own, including as a way to judge whether anyone’s GEO work is achieving anything.

Why can we not see AI referrals in Google Analytics?

Many assistant interactions never produce a click, and those that do often arrive without a usable referrer or are grouped as direct traffic. Some sources are now detectable and we identify those, but the honest position is that analytics alone cannot measure this channel, which is why prompt-set monitoring exists.

How often should we measure?

Monthly suits most businesses: frequent enough to catch an inaccuracy or a competitor shift, infrequent enough that normal model variation does not look like a trend. Weekly measurement mostly produces noise, and quarterly leaves a wrong description circulating for too long.

What do we do with the report?

Each one ends with a short list of recommended actions and the reason for each, normally aimed at the specific sources the models cited. If the answer is that nothing needs doing this month, we say that rather than inventing work — the report exists to inform a decision, not to justify a retainer.

How do you make sure the measurement is consistent?

By holding the prompt set, the models and the method constant, and by running them on a schedule rather than ad hoc. Changing the prompts is what destroys comparability, so we add to the set deliberately and report on the original subset separately when we do.

Can you track specific competitors?

Yes, and it is part of the standard report. We agree a named list at the start and record when each appears, how they are described and which sources are cited for them. That often reveals which third-party sources matter in your market faster than any other method.

What is a good share of answer?

The absolute figure means little; the comparison does. What matters is your trend month on month and your position relative to the named competitors. In a narrow B2B niche, appearing in a third of relevant prompts can be a strong position; in a crowded one it may not be.

Do we need this if we are not doing GEO?

It is useful on its own as an early warning system. An out-of-date or wrong description circulating in assistant answers is worth knowing about whether or not you have a programme to improve visibility, and it is usually cheaper to correct than to leave.

Related services

Works with these.

  1. Grow

    Generative Engine Optimisation

    Being present, and described correctly, when an AI assistant recommends suppliers.

    Explore
  2. Grow

    Answer Engine Optimisation

    Being the answer that gets extracted, not the tenth link nobody clicks.

    Explore
  3. Grow

    Conversion & Analytics

    Measuring what happens after the click, then fixing where it leaks.

    Explore

All 19 GTX Digital services

Next step

Get the baseline before you do anything else.

One prompt set, five assistants, verbatim answers and a named competitor comparison. It is the cheapest way to find out whether this channel needs your attention.