Getting cited by AI search: what actually moves the needle
Most advice here is repackaged SEO with the word AI in front. Some of it transfers and some does not. The parts that do not are where the work is.
What transfers from SEO, and what does not
The honest starting point: a large part of conventional SEO does carry over. Assistants retrieve from an index, and the things that made a page findable and trustworthy still make it retrievable. Crawlability, clear structure, real expertise on the page, and links from places that matter.
So the advice to throw out your SEO playbook is wrong. But three things genuinely do not transfer, and they are the ones worth spending effort on. A fuller account of what the term covers and how the platforms differ is in the LLM SEO overview.
The retrieved unit is a passage, not a page
There is no position to hold
Volume stops being the planning input
The mechanics of how retrieval actually works, and what a chunk is in practice, are set out in what AI search actually retrieves from a page. This piece is about what to do with that.
Write sections that survive being lifted out
This is the highest-return change available, and it costs nothing but attention.
A passage gets retrieved and inserted into a context window on its own. Everything around it is gone. So a section that opens with “as we saw above” or “this approach” is a section that arrives at the model as a fragment referring to something it cannot see.
The test is simple: copy any section out of your page, paste it somewhere with no surrounding text, and ask whether it still answers a question. If it does not, it will not be cited.
| Written for a reader | Written to survive extraction |
|---|---|
| “This makes the second approach preferable in most cases.” | “Cluster-based architecture is preferable to flat structure in most cases, because…” |
| “As mentioned, the lag is significant.” | “Organic SEO typically shows impressions at four months and pipeline at twelve.” |
| “It depends on your stage.” | “Before product-market fit, the right SEO spend is close to zero.” |
The pattern in the right column: name the subject in the sentence rather than referring back to it, and put the claim before the qualification. A model reading a chunk needs the answer in it, not a pointer to where the answer was.
Name entities specifically
Retrieval works on meaning rather than exact strings, but specificity still decides whether a passage is judged relevant to a question.
A page that says “leading analytics platforms” is less retrievable for a question about a named tool than one that says the name. This is not keyword stuffing revived; it is the difference between a passage that demonstrably concerns a subject and one that gestures at a category.
The practical check is to read a section and count how many of its claims could apply to any company in your category. If most could, the section is generic in the specific way that makes it uncitable, because there is no reason to retrieve yours rather than anyone else’s.
The Entity Extractor reports which entities a page actually establishes and which a competing page names that yours does not, weighted by where they appear rather than how often. That gap list is a usable brief.
Answer the question in the first two sentences
Assistants attend to the beginning and end of a retrieved passage more reliably than the middle. Burying the answer in paragraph four of a section is a structural mistake even when the answer is excellent.
This runs against a habit that serves human readers well. Building to a conclusion, establishing context first, holding the reveal. All of that is good writing and all of it reduces the chance of being cited.
The resolution is not to abandon the craft. It is to lead each section with its claim and then develop it, which reads perfectly well to a person and puts the answer where a model will find it.
Structured data does a smaller job than people claim
Schema is worth having and it is not the lever it is often sold as.
What it reliably does: helps a system parse what a page is, who wrote it, when it was updated, and what questions it answers. FAQPage markup in particular gives a question-and-answer pair in a form that needs no extraction at all.
What it does not do: make a weak passage citable. No amount of markup rescues a section that does not answer anything on its own.
The genuine risk is worse than absence. Contradictory or duplicated structured data, especially two blocks claiming the same identifier, produces a page whose own description of itself does not agree. The Schema Validator checks a whole page rather than one block at a time, which is the only way to catch that.
The pages worth optimising, and the ones that are not
Not every page benefits from this, and treating it as a site-wide project wastes most of the effort.
Assistants get asked questions. A page that answers a question is a candidate for citation; a page that exists to convert someone already convinced is not, and rewriting it for retrievability improves nothing.
| Page type | Worth the effort | Why |
|---|---|---|
| Explainers and guides | Yes, first priority | They answer questions in the form assistants receive them |
| Comparison and alternatives | Yes | Heavily asked, and specific enough to be citable |
| Original research or data | Yes, highest value per page | A number nobody else has is the strongest possible reason to cite you |
| Pricing and product pages | Marginal | Asked about, but the assistant usually wants a fact you can supply in one line |
| Case studies | Rarely | Specific to a client, and the question that would surface them is rare |
The third row deserves emphasis. Original data is the most citable thing you can publish, because a model answering a question about a number has very few sources to choose from. A small survey of your own customers, published with its method, will be cited more reliably than a well-written guide competing with a hundred others.
The corollary is uncomfortable: if everything you publish is a synthesis of what already exists, there is no structural reason for any assistant to prefer you. That is true in conventional search too, and retrieval makes it sharper because there is no second position to occupy.
Chunking, and why your headings matter more than they did
Retrieval systems split documents before indexing them. Where those splits fall is not something you control, but you influence it heavily through structure.
A page with clear headings and self-contained sections tends to split along sensible boundaries. A wall of text splits arbitrarily, which produces chunks that start mid-argument and end mid-sentence. Neither is retrievable for anything.
This is why headings do more work than they did in conventional SEO. A heading was a scanning aid and a mild relevance signal. Now it is also the label on a retrievable unit, and a heading like "The next step" labels a chunk with nothing a retrieval system can use.
A practical consequence for long pieces: a four-thousand-word article with six headings produces enormous chunks that dilute whatever is specific in them. The same content under fifteen headings produces units that each concern one thing. The writing does not change; the retrievability does.
The same argument applies to lists and tables, which chunk badly when they run long. A table of forty rows arrives as a fragment with no header row attached. Splitting it under sub-headings, with a sentence stating what each part shows, survives extraction far better.
Measuring something that has no rank
The measurement problem is real and most reporting has not caught up with it.
There is no position to track. A page is either used to answer a question or it is not, and that varies by phrasing, by assistant, and between one week and the next. Rank tracking has nothing to measure.
Three things that can actually be tracked, in descending order of reliability.
- Manual citation checks on a fixed question set. Write twenty questions a buyer would ask, run them monthly across the assistants you care about, and record whether you appear. Tedious, unglamorous, and the only direct measurement available.
- Referral traffic from assistant domains. Small numbers, and they undercount badly because most assistant answers are read without a click. Directionally useful, not a total.
- Branded search volume. If people who did not know you existed start searching your name, something upstream created that. The most reliable signal and the least specific.
What not to do: buy a tool promising an AI visibility score without understanding what it samples. Most run a question set and report appearance rate, which you can do yourself and understand properly. The broader attribution problem is covered in measuring organic when attribution breaks. The method, its failure modes and what to do with a bad result are set out in tracking brand mentions in AI search.
What to do first, in order
The work sorts cleanly by return against effort, and the order matters because the early items make the later ones worth doing.
Start with the pages you already have that rank well. They are already retrievable and already trusted; making their sections self-contained is a few hours of editing against pages with existing authority. New content written correctly from the start is the second job, not the first.
The thing to resist is treating this as a separate programme. It is an editing standard applied to work you were doing anyway, and companies that spin it out as an AI visibility initiative tend to produce a document rather than a change.
One honest caveat on all of it. The retrieval systems behind these assistants change without notice and without documentation, and anyone presenting a fixed playbook here is describing a snapshot. What makes the advice above worth following is that none of it is a trick: sections that answer one question completely, subjects named rather than referred to, and claims stated before caveats are better writing regardless of what any model does next. If the mechanism changes, the pages are still good pages, which is not something you could say about most tactics aimed at a specific system.
Assistants retrieve passages, not pages, so the unit of optimisation is the section rather than the article. Make each section answer one question completely, name entities specifically, and put the answer before the context. Then measure citation appearance rather than rank, because rank is not the mechanism any more.
The change that produced the clearest difference for me was unglamorous: going through existing pages and rewriting the first sentence of every section so it named its subject instead of referring back. A day of work, no new content, and the pages started appearing in assistant answers they had not before.
I would still not describe that as proven. Citation behaviour is noisy enough that a month of observation is not evidence, and anyone claiming a precise causal number here is overselling what the measurement supports.
Frequently asked
How do I improve my brand’s visibility in AI search engines?
Is AI search optimisation different from SEO?
Does schema markup help with AI search visibility?
How do I measure whether AI search visibility is improving?
I take on a small number of engagements at a time. If you are past product-market fit and want to talk through what this would look like for your situation, the calendar is open.
Book a call