Blog · AI Search

ChatGPT visibility: how it picks sources, and what you can influence

ChatGPT sources answers in two quite different ways depending on whether it searches, and the distinction decides which of your problems is fixable this quarter.

Two mechanisms, and only one you can move

ChatGPT answers a question in one of two ways, and almost all confusion about visibility here comes from treating them as one thing.

It either answers from what the model already contains, or it runs a search and reads results. Which happens depends on the question, the settings, and how current the model judges the answer needs to be.

The two paths, and what each means for you
Answered from training
No search happens. The answer reflects what the model absorbed during training, which is a snapshot from before it shipped. You cannot influence this on any useful timescale, and neither can anyone selling you a service that claims otherwise.
Answered with search
The model runs a query, reads a handful of results and synthesises. This is influenceable, and the route is conventional: rank for the query, then be extractable when read.
A mix, which is common
It searches, then blends what it finds with what it already believed. This is why a page can be cited and the surrounding claims still be wrong.

The practical consequence is that the same question can produce a stale answer one day and a current one the next, depending on whether a search was triggered. Testing without noting which happened produces results that look random.

There is a simple tell. If the answer cites sources with links, a search ran. If it does not, you are reading the training snapshot, and whatever it says reflects the state of the web at some point before the model shipped. Recording which happened turns an incoherent set of test results into two coherent ones, and it costs nothing beyond noticing.

The distinction also decides who should care about a bad answer. A wrong searched answer is a content problem and yours to fix. A wrong trained answer is closer to a reputation problem, it will persist for months, and the work is spread across places you mostly do not control.

The searched path, which is mostly conventional SEO

When ChatGPT searches, it behaves close enough to a search engine that the existing playbook applies with one addition.

You need to rank for the query it constructs, which is often not the query the user typed. Someone asking a long conversational question produces a shorter, more conventional search behind the scenes, so the terms you already target are usually closer to what matters than the phrasing of the question.

Then the passage has to survive being read out of context. That requirement follows from how retrieval works at all, which what AI search actually retrieves from a page sets out, and the practical version is in the AI search visibility guide.

What to do for the searched path
01
Rank conventionally for the underlying query
Not the conversational phrasing. The search it runs is usually a compressed version of the question.
02
Answer in the opening lines of a matching section
It reads a few results quickly. An answer in paragraph five is an answer it does not reach.
03
Be specific enough to be quoted rather than paraphrased
A number or a named method survives synthesis. Generic phrasing gets absorbed into the summary with no attribution.

None of this is new work if the conventional programme is sound. That is the reassuring part of an area that mostly generates anxiety.

The trained path, and why it matters more than it seems

The answers you cannot influence quickly are the ones worth worrying about, because they are the ones repeated most consistently.

If the model absorbed something incorrect about your product, it will say that thing to everyone who asks, in every session, until a version trained on different data replaces it. There is no correction mechanism, no support ticket, and usually no notification that it is happening.

The leverage, such as it is, is indirect: what gets absorbed comes from what is widely published about you. That includes your own site, and also documentation, directories, forum threads, comparison sites and anywhere else your product is described by someone other than you.

  • Make your own documentation unambiguous. Vague self-description is what gets replaced by someone else’s confident summary.
  • Correct third-party descriptions where you can. Directory entries and comparison sites carry more weight here than their traffic suggests, because they are widely duplicated.
  • Publish the specific facts you want repeated. Pricing model, what the product does not do, who it is not for. Precise statements survive summarisation better than positioning language.

That last one is the most useful and the least intuitive. A page stating plainly what your product does not do is more likely to be reflected accurately than three pages of benefits, because it contains a fact rather than a claim.

Memory, custom instructions and why nobody sees the same answer

ChatGPT carries context that a search engine does not, and it changes what different people see enough to matter.

An account with memory enabled may retain that a user works in a particular industry, prefers certain tools, or asked about a competitor last week. Custom instructions add a standing layer on top. Both shift which sources feel relevant to a question, which means two people asking identical words can receive materially different answers.

What varies between two people asking the same thing
Factor Effect Can you influence it
Account memory Prior context colours what is surfaced No
Custom instructions Standing preferences reshape the answer No
Whether a search ran Current sources or a training snapshot No
What ranks for the underlying query Which sources get read Yes
Only the bottom row is work. The other three are reasons to distrust a single test

The bottom row being the only influenceable one is the honest summary of this whole subject, and it explains why the useful advice keeps collapsing back into conventional SEO.

It also explains why anecdotes here are close to worthless. Someone reporting that they appear for a query has told you about one account, one moment, and one path through the two mechanisms. It is not evidence about what a prospect sees.

Testing it without misleading yourself

Checking your own visibility here is easy to do badly, and there are three specific traps.

Three ways a self-test misleads
Trap What happens Fix
Testing in your own account Memory and custom instructions bias the answer Use a logged-out or temporary session every time
Asking a leading question Naming your product guarantees it appears Ask the question a buyer would ask, without your name in it
Testing once The same question returns different sources on different runs Several runs, and compare against previous months rather than yesterday
The first trap is the commonest and the most flattering

The first row deserves emphasis. If you have discussed your own company in that account before, the model may carry that context, and you will get a reassuring answer that no prospect would ever see. Almost every founder who tells me they appear in ChatGPT tested it while logged in.

The disciplined version of this, as a monthly routine rather than a one-off check, is set out in tracking brand mentions in AI search.

What the training path actually rewards

Since the trained path is the one you cannot move quickly, it is worth being precise about what does eventually move it, because the honest answer is unglamorous and takes a long time.

What gets absorbed is what is widely and consistently said. Not what ranks, not what is well written, and not what you would prefer. A description repeated across twenty sites in similar words is more likely to persist than a better description on one.

What plausibly influences how a model describes you, over years
Consistent description across many third-party sites strongest
Your own documentation, stated as plain fact strong
Widely cited coverage or research moderate
Your marketing pages and positioning copy weak
Anything published in the last few months negligible at this timescale
Reasoning from how training corpora are assembled rather than from any published detail

The bottom row is the one that frustrates people. Work done this quarter has essentially no effect on the trained path, and anyone promising otherwise is describing the searched path and calling it the other thing.

The fourth row is the one worth internalising. Positioning copy is written to be distinctive and is therefore inconsistent with how everyone else describes your category, which makes it less likely to be reproduced than a plain factual sentence. The description that survives is the boring one.

The practical implication is a slow, cheap habit rather than a project: whenever your product is described somewhere you do not control, check that the description is accurate and correct it if not. Directory entries, integration listings, comparison pages. Individually trivial and collectively the thing that shapes the trained path.

What being wrong actually costs

The framing of this whole area as a marketing opportunity understates the more immediate issue, which is accuracy.

A prospect asking about your pricing model, your integrations or your compliance posture gets one answer, delivered confidently, with no competing results beside it to prompt scepticism. If that answer is wrong, it is wrong at scale and quietly.

The three questions worth checking about your own product
Pricing
the most commonly wrong, and the most costly
What it does not do
misattributed capability creates bad-fit leads
Who it is for
wrong segment framing filters out real buyers
Run these monthly in a clean session. It takes five minutes

The second row produces a failure people rarely trace back: if a model confidently says your product does something it does not, you get demos from people who wanted that thing, and a sales team wondering why qualification has degraded.

This is why I would treat the brand control question as the highest-priority check in any monitoring set. Category visibility is an opportunity. Being described incorrectly is a live problem, and one you would want to know about whether or not you cared about AI search at all.

A proportionate response

Given the volume of noise on this subject, worth ending with what is proportionate for a company that is not an AI-search specialist.

Check the three product questions monthly, in a clean session, which takes five minutes. Make sure your documentation states facts plainly rather than positioning. Keep doing conventional SEO, since that is what feeds the searched path. That is the whole of it.

What is not proportionate: a dedicated programme, a monitoring subscription before you have established whether you have a problem, or rewriting content on the assumption that a specific mechanism works a specific way. The broader case for treating this as an editing standard rather than an initiative is in the LLM SEO overview.

The Entity Extractor is useful for one narrow part of this: checking whether your own pages actually establish the facts you want repeated, or merely gesture at them. A page that never states its subject specifically gives a model nothing precise to carry forward.

There is a version of this subject that deserves more attention than it gets, and it is not the visibility one. As assistants become a normal way to research a purchase, the accuracy of what they say about you becomes part of your product surface in the way that documentation or a status page is. Nobody owns it, it is not in anyone’s remit, and it degrades silently. The five-minute monthly check is worth doing for that reason alone, independent of whether being cited in category answers ever produces a single lead.

The short version

ChatGPT answers either from training or from a live search, and only the second is influenceable on any useful timescale. For searched answers the route is conventional ranking plus an extractable passage. For trained answers, the leverage is being described correctly wherever the training data comes from, which is a slower and mostly indirect job.

Operator note

The most useful five minutes I spend on this for any client is asking a clean session what the product costs and what it does not do. It has been wrong more often than right, and it is the finding that gets acted on fastest, because it is unambiguous.

Category visibility conversations tend to go in circles because nobody can size the opportunity. A model telling prospects your pricing works in a way it does not is a different kind of conversation entirely.

Frequently asked

How do I get my brand to show up in ChatGPT?
It depends which mechanism answers the question. When ChatGPT searches, the route is conventional: rank for the underlying query, then make the relevant passage answer the question in its opening lines. When it answers from training instead, there is no fast route, and the indirect leverage is having your product described accurately and specifically wherever it is widely published.
Can I optimise for ChatGPT specifically?
For its searched answers, yes, and the work is close to ordinary SEO because a search is what happens underneath. For answers drawn from training data there is no direct optimisation, and any service claiming to place you inside a model’s knowledge is selling something that does not work that way. The distinction is worth insisting on when evaluating vendors.
Why does ChatGPT say something wrong about my product?
Most likely it absorbed an inaccurate or outdated description during training, and repeats it consistently because there is no correction mechanism. The indirect fix is making your own documentation state facts plainly, correcting third-party descriptions in directories and comparison sites, and publishing precise statements about pricing, limitations and fit, which survive summarisation better than positioning language.
How do I test my ChatGPT visibility properly?
Use a logged-out or temporary session, because memory and custom instructions in your own account will bias the answer toward a flattering result. Ask the question a buyer would ask without naming your product, since naming it guarantees it appears. Run each question several times, because the same question returns different sources on different runs, and compare month to month rather than day to day.
Related reading
LLM SEO: what the term means, and what is actually different Getting cited by AI search: what actually moves the needle Tracking brand mentions in AI search, honestly
Working on this yourself?

I take on a small number of engagements at a time. If you are past product-market fit and want to talk through what this would look like for your situation, the calendar is open.

Book a call