Tracking brand mentions in AI search, honestly
Almost every page on this subject is selling a monitoring tool. The measurement is genuinely harder than those pages admit, and knowing why changes what you should bother tracking.
Why this is harder than it sounds
Conventional rank tracking works because a query has an answer that is stable enough to sample. Ask Google the same thing twice today and you get materially the same page of results.
Assistants do not behave that way. The same question, asked twice, can produce different sources. Rephrase it slightly and the set changes again. Ask from a different account, or a week later, and it changes once more. The systems differ from each other too, which the LLM SEO overview sets out.
The reason sits in the mechanism: assistants retrieve passages rather than ranking pages, as what AI search actually retrieves from a page sets out, and a passage that clears the relevance bar for one phrasing may not for another.
So there is no position to record, and the thing you are sampling is a distribution rather than a value. That is not a reason to skip measurement, but it does mean any single check tells you almost nothing, and any tool reporting a precise score is reporting a sample it has chosen not to describe.
Phrasing
Time
Personalisation and context
The practical consequence: measure the same questions the same way every time, and treat the trend as the signal rather than any individual result. That is the whole method, and everything below is detail on doing it properly.
The question set is the measurement
Since you cannot sample everything, what you choose to sample defines what you learn. This is the part most monitoring setups get wrong, because the temptation is to track the questions you want to win rather than the ones buyers ask.
A useful set has three kinds of question in it, and the proportions matter.
| Type | How many | Example shape |
|---|---|---|
| Category questions | 8-10 | “What is the best tool for X”, high value, high competition, rarely won early |
| Problem questions | 6-8 | “How do I solve Y”, where a good explainer genuinely competes |
| Named comparisons | 3-5 | “X versus Y”, where you appear or you do not, very clear signal |
| Your own brand | 1-2 | “What is X”, a control. If you lose this, something is badly wrong |
The brand control row earns its place. If an assistant cannot describe your own product accurately, no amount of category-level work matters, and that failure is both common and fixable in a way the others are not. The three product questions worth checking every month are set out in the ChatGPT visibility piece.
Write the set once, then leave it alone. Changing questions between runs is the fastest way to produce a chart that measures your editing rather than your visibility.
What the monitoring tools actually do
There is a growing category of products offering AI visibility scores. Some are useful. None can do more than sample, and understanding that is the difference between using one well and being misled by it.
Underneath, almost all of them run a question set against one or more assistants on a schedule and record whether a brand appears. That is the same method you can run yourself. What you are buying is the scheduling, the storage and the presentation.
The third question is the one that separates a real measurement from a number. A single run per question per month produces a figure that will move ten points on noise alone. Several runs, averaged, produces something you can act on.
What none of them can tell you is how often a real person asked that question and saw that answer. There is no impression count. Anyone presenting share of voice as though it were a traffic figure is converting a sample into a volume that does not exist.
The signals that survive when direct measurement fails
Direct citation tracking is noisy and partial, so it is worth pairing with signals that do not depend on catching an assistant in the act.
- Branded search volume. The most reliable indirect signal available. If people who did not know you existed start searching your name, something upstream created that demand, and assistant answers are increasingly part of that upstream.
- Referral traffic from assistant domains. Real but heavily undercounted, because most answers are read without a click. Useful as a floor, useless as a total.
- Direct traffic from unusual geographies or at unusual times. Weak, and worth watching when the other two move together.
- What prospects say on calls. Unscientific and often the earliest signal you get. “I asked ChatGPT and it mentioned you” is a data point no dashboard will ever show you.
The last one is genuinely worth formalising. One question on a discovery call, asked consistently, produces a record that is closer to the truth than most of the tooling. It has the same weakness as any self-report and none of the sampling problems.
The wider attribution problem behind all of this, and why last-touch reporting undercounts anything that happens early, is covered in measuring organic when attribution breaks.
Running it yourself, in forty minutes a month
The manual version is genuinely competitive with the tooling, and it has one advantage no product offers: you see the raw answers rather than a score derived from them.
The setup is a spreadsheet with the questions down the side and the run dates across the top. Each cell records whether you appeared, and a second sheet records what was cited instead. That second sheet turns out to be the valuable one.
Forty minutes is a real estimate for twenty questions across two assistants, and it is a task that delegates well. The judgement is in writing the question set, which happens once.
The place it stops scaling is beyond about thirty questions or three assistants, which is where the tooling starts to earn its price. Below that, the manual version is better because it is transparent.
Reading the results without fooling yourself
Once you have a few months of data, the interpretation is where the errors happen.
The failure mode is treating every movement as a result and reacting to it. A question dropping out one month and returning the next has told you nothing except that the system is stochastic, which you already knew.
The opposite failure is assuming a whole-set decline is your fault. Model versions change, retrieval indexes get rebuilt, and a broad drop across every question at once is more likely to be a platform change than anything you did. Check whether competitors moved too before rewriting anything.
What justifies action: a sustained decline over three or more runs, concentrated in questions that share a topic. That pattern points at content rather than at the platform, and it is specific enough to act on.
What being cited is actually worth
A question worth asking before investing much here: what does an assistant citation get you, in terms anyone would recognise as a result.
The honest answer is that it varies enormously and nobody has good numbers. A citation inside an answer someone reads without clicking produces no session, no attribution, and possibly a purchase months later that lands as direct traffic. That is real value and it is invisible to every reporting system you have.
What can be said with more confidence is where it matters most, and it is not evenly spread.
| Situation | Value of a citation | Why |
|---|---|---|
| Early research, unfamiliar category | High | The assistant is shaping which vendors get considered at all |
| Named comparison against a competitor | High | The answer arrives at someone already deciding |
| Definitional question about your own product | Critical | If this is wrong, it is wrong at scale and you may not know |
| Broad how-to with no purchase intent | Low | Cited, read, and forgotten. Pleasant and not commercial |
The third row deserves the emphasis. An assistant describing your product incorrectly does that consistently, to everyone who asks, until whatever it learned from changes. That is a materially different risk from a bad search result, and it is the strongest single argument for running a brand control question every month.
The fourth row is worth stating plainly because it cuts against the enthusiasm: being cited on broad informational questions is genuinely pleasant and produces very little. If your citation tracking is improving only on that row, the programme is working and not yet paying.
What to do when you are not appearing
The measurement is only worth running if there is a response to a bad result, and the response is not usually more content.
The first check is whether the pages that should be cited are retrievable at all: self-contained sections, subjects named rather than referred to, claims stated before caveats. Those changes are covered in the AI search visibility guide, and they are cheaper than writing anything new.
The second is whether you have anything genuinely citable. A model choosing between six syntheses of the same received wisdom has no reason to pick yours. A model answering a question about a number has very few sources, which is why original data outperforms a better-written guide.
The third, and the one people skip: check what is being cited instead. The sources appearing in place of yours tell you what the system considers a good answer to that question, and that is a more useful brief than any keyword tool. Sometimes the answer is that a forum thread is winning, which tells you the question is being asked in a register your content does not match.
The Entity Extractor makes that comparison concrete: paste the page being cited alongside yours and it lists the entities theirs names that yours does not.
One last thing worth saying about the whole exercise. Tracking is not the work, and it is easy to let it become the work. A monitoring setup that consumes an afternoon a month and produces a chart nobody acts on is worse than no tracking, because it converts genuine uncertainty into a number that feels like knowledge. If the measurement is not changing what you write, stop measuring and spend the time writing instead. The point of the question set is to tell you which pages are failing and why, and if it has not done that in three months either the set is wrong or the answer is that nothing needs changing yet.
There is no rank and no impression count, so the only direct measurement is running a fixed question set on a schedule and recording whether you appear. Sample the same questions every time or the variance swamps the signal. Treat any vendor score as a sampled appearance rate, because that is all it can be.
I ran a twenty-question set monthly for a client through most of last year, by hand, in a spreadsheet. It took about forty minutes a month and told me more than any of the tools I trialled, mostly because I knew exactly what it sampled and could see the raw answers rather than a score.
The thing that surprised me was how often the reason for not appearing was visible in the answer itself. The cited sources were usually more specific, not better written. That is a fixable problem and it is not the one most people assume they have.
Frequently asked
How do I track brand mentions in AI search?
Are AI visibility tracking tools worth paying for?
Can you measure impressions or share of voice in AI search?
What should I do if my brand is not being cited?
I take on a small number of engagements at a time. If you are past product-market fit and want to talk through what this would look like for your situation, the calendar is open.
Book a call