What AI search actually retrieves from your page
Assistants do not cite pages. They cite passages. Once you internalise that, a lot of advice about optimising for AI search turns out to be optimising the wrong unit.
The retrieved unit is a chunk
A retrieval system does not hold your page in memory as a page. It splits documents into chunks, embeds them, and retrieves the chunks that match a query. What gets surfaced and attributed is a passage, with your URL attached as provenance.
The consequence is direct. A page can be excellent overall and still never be retrieved, because its useful claims are distributed across paragraphs that each individually lack the context to stand alone. A chunk that begins “this approach has three main drawbacks” is nearly useless in isolation, because the retrieval system has no idea what “this approach” refers to.
In B2B SEO the passages that matter most are the ones answering a specific buying question rather than a general one.
Writing passages that survive extraction
| Retrieves poorly | Retrieves well | Why |
|---|---|---|
| This approach has three main drawbacks | Utility-led SEO has three main drawbacks | No pronoun pointing outside the passage |
| Impressions rose 12x | Organic impressions rose 12x over eighteen months | Number carries its unit and window |
| …building to a conclusion at the end | Conclusion first, elaboration after | A chunk cut early still contains the claim |
| The framework recommends… | The MEDDIC framework recommends… | Named entity is recognisable |
The adjustment is small and slightly unnatural at first: make each section self-contained enough that a reader dropped into it cold would understand what it is about.
- Restate the subject at the start of a section rather than relying on a pronoun that points back three paragraphs.
- Put the claim before the elaboration. A section that opens with its conclusion retrieves well; one that builds to a conclusion in the final line retrieves as a preamble.
- Keep a definition near the term it defines, in the same passage, rather than in a glossary elsewhere on the page.
- Attach numbers to their units and their context in the same sentence. “Impressions rose 12x” is not extractable. “Organic impressions rose 12x over eighteen months” is.
This reads as slightly more repetitive prose than you would write for a linear reader. That is the trade, and it is a mild one.
Why entities do the heavy lifting
Retrieval matches meaning rather than strings, which is why keyword density stopped being interesting some time ago. What does still matter is whether the concepts on your page are named specifically enough to be recognised as the concepts they are.
Writing “the framework” twelve times where you could have written the framework’s actual name costs you the association. Naming the tools, standards, companies, and methods you are discussing gives the retrieval layer recognisable anchors and gives your page a clearer position in the graph of things it is about.
This is not an instruction to stuff proper nouns. It is an instruction to stop hedging them out. Technical writers do this instinctively. Marketing writers have usually been trained out of it, because specificity narrows the audience, which is exactly why it works here.
This matters doubly for generated page families, where the template decides how every passage reads at once, as the piece on programmatic pages covers.
What you can and cannot measure
Be honest about the measurement situation. There is no clean equivalent of Search Console for assistant citations. What you can do is run your priority queries against the major assistants on a schedule, record whether you appear and in what form, and treat that as a coarse panel rather than a metric.
It is genuinely noisy. Answers vary between sessions and shift when models update. What it does reliably tell you is whether you are in the considered set at all, which is the difference that matters most, and which is usually stable enough to read.
Doing this deliberately rather than hoping for it is what the AI-search visibility work is, and the measurement problem above is why it has to be run as a programme rather than a one-off fix.
Free tools are unusually good at earning citations for this reason, and choosing which tool to build first covers how to pick one.
Optimise the passage, not the page. Self-contained sections, claims stated before elaboration, and specifically named entities are what make a chunk worth retrieving and attributable back to you. The practical version of all of it, in the order worth doing it, is set out in the guide to getting cited by AI search. Knowing whether any of it worked is a separate problem, covered in tracking brand mentions in AI search.
Writing for retrieval feels slightly wrong at first, because you repeat the subject more than a linear reader needs. I have made peace with it. The cost is prose that reads marginally more redundant. The benefit is passages that survive being cut out of context, which is the only form in which an assistant will ever see them.
I do not trust any single assistant check. I run the same queries monthly and read the trend, because individual answers vary between sessions and shift when models update. What stays stable is whether you are in the considered set at all, and that is the thing worth knowing.
Frequently asked
How do I get my content cited by ChatGPT or Perplexity?
Does schema markup help with AI search visibility?
Can I track whether AI assistants cite my site?
Is AI search optimisation different from SEO?
I take on a small number of engagements at a time. If you are past product-market fit and want to talk through what this would look like for your situation, the calendar is open.
Book a call