Blog · AI Search

Tracking brand mentions in AI search, honestly

Almost every page on this subject is selling a monitoring tool. The measurement is genuinely harder than those pages admit, and knowing why changes what you should bother tracking.

Why this is harder than it sounds

Conventional rank tracking works because a query has an answer that is stable enough to sample. Ask Google the same thing twice today and you get materially the same page of results.

Assistants do not behave that way. The same question, asked twice, can produce different sources. Rephrase it slightly and the set changes again. Ask from a different account, or a week later, and it changes once more. The systems differ from each other too, which the LLM SEO overview sets out.

The reason sits in the mechanism: assistants retrieve passages rather than ranking pages, as what AI search actually retrieves from a page sets out, and a passage that clears the relevance bar for one phrasing may not for another.

So there is no position to record, and the thing you are sampling is a distribution rather than a value. That is not a reason to skip measurement, but it does mean any single check tells you almost nothing, and any tool reporting a precise score is reporting a sample it has chosen not to describe.

Three sources of variance, and what each does to your numbers
Phrasing
“Best CRM for startups” and “what CRM should a startup use” are the same question to a person and can retrieve completely different sources. Any measurement is a measurement of the phrasings you chose.
Time
Retrieval indexes update, models get replaced, and the systems behind them change without announcement. A drop between two months may be your content or may be a version change you were never told about.
Personalisation and context
Some assistants carry conversation history and account context. The same question inside a longer chat is not the same query, which makes clean-room testing the only comparable kind.

The practical consequence: measure the same questions the same way every time, and treat the trend as the signal rather than any individual result. That is the whole method, and everything below is detail on doing it properly.

The question set is the measurement

Since you cannot sample everything, what you choose to sample defines what you learn. This is the part most monitoring setups get wrong, because the temptation is to track the questions you want to win rather than the ones buyers ask.

A useful set has three kinds of question in it, and the proportions matter.

What belongs in a twenty-question monitoring set
Type How many Example shape
Category questions 8-10 “What is the best tool for X”, high value, high competition, rarely won early
Problem questions 6-8 “How do I solve Y”, where a good explainer genuinely competes
Named comparisons 3-5 “X versus Y”, where you appear or you do not, very clear signal
Your own brand 1-2 “What is X”, a control. If you lose this, something is badly wrong
Twenty is enough to see movement and small enough to run monthly by hand

The brand control row earns its place. If an assistant cannot describe your own product accurately, no amount of category-level work matters, and that failure is both common and fixable in a way the others are not. The three product questions worth checking every month are set out in the ChatGPT visibility piece.

Write the set once, then leave it alone. Changing questions between runs is the fastest way to produce a chart that measures your editing rather than your visibility.

What the monitoring tools actually do

There is a growing category of products offering AI visibility scores. Some are useful. None can do more than sample, and understanding that is the difference between using one well and being misled by it.

Underneath, almost all of them run a question set against one or more assistants on a schedule and record whether a brand appears. That is the same method you can run yourself. What you are buying is the scheduling, the storage and the presentation.

The three questions worth asking any vendor
Which questions
and can you change them
Which assistants
and which model version
How many runs
per question, per period
If a vendor will not answer the third, the score is a single sample dressed as a metric

The third question is the one that separates a real measurement from a number. A single run per question per month produces a figure that will move ten points on noise alone. Several runs, averaged, produces something you can act on.

What none of them can tell you is how often a real person asked that question and saw that answer. There is no impression count. Anyone presenting share of voice as though it were a traffic figure is converting a sample into a volume that does not exist.

The signals that survive when direct measurement fails

Direct citation tracking is noisy and partial, so it is worth pairing with signals that do not depend on catching an assistant in the act.

  • Branded search volume. The most reliable indirect signal available. If people who did not know you existed start searching your name, something upstream created that demand, and assistant answers are increasingly part of that upstream.
  • Referral traffic from assistant domains. Real but heavily undercounted, because most answers are read without a click. Useful as a floor, useless as a total.
  • Direct traffic from unusual geographies or at unusual times. Weak, and worth watching when the other two move together.
  • What prospects say on calls. Unscientific and often the earliest signal you get. “I asked ChatGPT and it mentioned you” is a data point no dashboard will ever show you.

The last one is genuinely worth formalising. One question on a discovery call, asked consistently, produces a record that is closer to the truth than most of the tooling. It has the same weakness as any self-report and none of the sampling problems.

The wider attribution problem behind all of this, and why last-touch reporting undercounts anything that happens early, is covered in measuring organic when attribution breaks.

Running it yourself, in forty minutes a month

The manual version is genuinely competitive with the tooling, and it has one advantage no product offers: you see the raw answers rather than a score derived from them.

The setup is a spreadsheet with the questions down the side and the run dates across the top. Each cell records whether you appeared, and a second sheet records what was cited instead. That second sheet turns out to be the valuable one.

The monthly routine, start to finish
01
Run the set in a clean session
No conversation history, no account context. The same conditions every time, or you are comparing different measurements.
02
Record appearance, not position
There is no position. Appeared or did not, plus which competitors did. Binary is the honest resolution here.
03
Capture what was cited instead
Paste the sources. This is the column that tells you what to change, and it is the one every dashboard omits.
04
Compare against the previous three runs
Never against the previous one. Single-month comparisons are noise at this sample size.

Forty minutes is a real estimate for twenty questions across two assistants, and it is a task that delegates well. The judgement is in writing the question set, which happens once.

The place it stops scaling is beyond about thirty questions or three assistants, which is where the tooling starts to earn its price. Below that, the manual version is better because it is transparent.

Reading the results without fooling yourself

Once you have a few months of data, the interpretation is where the errors happen.

How much a month-on-month change actually tells you
A single question changing noise
Three or four questions moving together worth investigating
The whole set moving in one direction a real change, yours or theirs
Brand control question failing act immediately
Directional. The point is that small movements are not findings

The failure mode is treating every movement as a result and reacting to it. A question dropping out one month and returning the next has told you nothing except that the system is stochastic, which you already knew.

The opposite failure is assuming a whole-set decline is your fault. Model versions change, retrieval indexes get rebuilt, and a broad drop across every question at once is more likely to be a platform change than anything you did. Check whether competitors moved too before rewriting anything.

What justifies action: a sustained decline over three or more runs, concentrated in questions that share a topic. That pattern points at content rather than at the platform, and it is specific enough to act on.

What being cited is actually worth

A question worth asking before investing much here: what does an assistant citation get you, in terms anyone would recognise as a result.

The honest answer is that it varies enormously and nobody has good numbers. A citation inside an answer someone reads without clicking produces no session, no attribution, and possibly a purchase months later that lands as direct traffic. That is real value and it is invisible to every reporting system you have.

What can be said with more confidence is where it matters most, and it is not evenly spread.

Where a citation is worth more, and where it is worth less
Situation Value of a citation Why
Early research, unfamiliar category High The assistant is shaping which vendors get considered at all
Named comparison against a competitor High The answer arrives at someone already deciding
Definitional question about your own product Critical If this is wrong, it is wrong at scale and you may not know
Broad how-to with no purchase intent Low Cited, read, and forgotten. Pleasant and not commercial
The third row is the one companies discover late and expensively

The third row deserves the emphasis. An assistant describing your product incorrectly does that consistently, to everyone who asks, until whatever it learned from changes. That is a materially different risk from a bad search result, and it is the strongest single argument for running a brand control question every month.

The fourth row is worth stating plainly because it cuts against the enthusiasm: being cited on broad informational questions is genuinely pleasant and produces very little. If your citation tracking is improving only on that row, the programme is working and not yet paying.

What to do when you are not appearing

The measurement is only worth running if there is a response to a bad result, and the response is not usually more content.

The first check is whether the pages that should be cited are retrievable at all: self-contained sections, subjects named rather than referred to, claims stated before caveats. Those changes are covered in the AI search visibility guide, and they are cheaper than writing anything new.

The second is whether you have anything genuinely citable. A model choosing between six syntheses of the same received wisdom has no reason to pick yours. A model answering a question about a number has very few sources, which is why original data outperforms a better-written guide.

The third, and the one people skip: check what is being cited instead. The sources appearing in place of yours tell you what the system considers a good answer to that question, and that is a more useful brief than any keyword tool. Sometimes the answer is that a forum thread is winning, which tells you the question is being asked in a register your content does not match.

The Entity Extractor makes that comparison concrete: paste the page being cited alongside yours and it lists the entities theirs names that yours does not.

One last thing worth saying about the whole exercise. Tracking is not the work, and it is easy to let it become the work. A monitoring setup that consumes an afternoon a month and produces a chart nobody acts on is worse than no tracking, because it converts genuine uncertainty into a number that feels like knowledge. If the measurement is not changing what you write, stop measuring and spend the time writing instead. The point of the question set is to tell you which pages are failing and why, and if it has not done that in three months either the set is wrong or the answer is that nothing needs changing yet.

The short version

There is no rank and no impression count, so the only direct measurement is running a fixed question set on a schedule and recording whether you appear. Sample the same questions every time or the variance swamps the signal. Treat any vendor score as a sampled appearance rate, because that is all it can be.

Operator note

I ran a twenty-question set monthly for a client through most of last year, by hand, in a spreadsheet. It took about forty minutes a month and told me more than any of the tools I trialled, mostly because I knew exactly what it sampled and could see the raw answers rather than a score.

The thing that surprised me was how often the reason for not appearing was visible in the answer itself. The cited sources were usually more specific, not better written. That is a fixable problem and it is not the one most people assume they have.

Frequently asked

How do I track brand mentions in AI search?
Write a fixed set of about twenty questions a buyer would ask, run them on a schedule across the assistants you care about, and record whether you appear and what was cited instead. Keep the questions identical between runs, because changing them measures your editing rather than your visibility. Several runs per question beats one, since the same question asked twice can return different sources.
Are AI visibility tracking tools worth paying for?
They automate a method you can run yourself, so the value is in scheduling and presentation rather than in any capability you lack. Before buying, ask which questions are sampled and whether you can change them, which assistants and model versions are used, and how many runs happen per question per period. If the last question goes unanswered, the score is a single sample presented as a metric.
Can you measure impressions or share of voice in AI search?
No. There is no impression count, because nobody reports how many people asked a given question or saw a given answer. Any share-of-voice figure is an appearance rate across a sample of questions the vendor chose, which is a useful trend and not a volume. Treating it as traffic is the most common error in this area.
What should I do if my brand is not being cited?
Check retrievability first, since self-contained sections and specifically named subjects cost editing time rather than new content. Then check whether you publish anything genuinely citable, because a synthesis of received wisdom gives a model no reason to choose you. Then look at what is cited instead: those sources show you what the system considers a good answer, which is a better brief than any keyword tool.
Related reading
Getting cited by AI search: what actually moves the needle What AI search actually retrieves from a page Measuring organic when attribution breaks
Working on this yourself?

I take on a small number of engagements at a time. If you are past product-market fit and want to talk through what this would look like for your situation, the calendar is open.

Book a call