Blog · AI Search

Getting cited by AI search: what actually moves the needle

Most advice here is repackaged SEO with the word AI in front. Some of it transfers and some does not. The parts that do not are where the work is.

What transfers from SEO, and what does not

The honest starting point: a large part of conventional SEO does carry over. Assistants retrieve from an index, and the things that made a page findable and trustworthy still make it retrievable. Crawlability, clear structure, real expertise on the page, and links from places that matter.

So the advice to throw out your SEO playbook is wrong. But three things genuinely do not transfer, and they are the ones worth spending effort on. A fuller account of what the term covers and how the platforms differ is in the LLM SEO overview.

Three things that behave differently, and why
The retrieved unit is a passage, not a page
A ranking system returns documents and a person scans them. A retrieval system returns chunks and a model reads them. A section that only makes sense after reading the three above it is a section that cannot be retrieved usefully.
There is no position to hold
You are either in the context window for a given question or you are not. There is no page two, no gradual climb, and no partial credit for being fourth.
Volume stops being the planning input
Nobody types a keyword at an assistant. They ask a question in their own words, so the useful unit of planning is the question, and one well-answered question covers a spread of phrasings a keyword tool would list as separate terms.

The mechanics of how retrieval actually works, and what a chunk is in practice, are set out in what AI search actually retrieves from a page. This piece is about what to do with that.

Write sections that survive being lifted out

This is the highest-return change available, and it costs nothing but attention.

A passage gets retrieved and inserted into a context window on its own. Everything around it is gone. So a section that opens with “as we saw above” or “this approach” is a section that arrives at the model as a fragment referring to something it cannot see.

The test is simple: copy any section out of your page, paste it somewhere with no surrounding text, and ask whether it still answers a question. If it does not, it will not be cited.

The same content, written two ways
Written for a reader Written to survive extraction
“This makes the second approach preferable in most cases.” “Cluster-based architecture is preferable to flat structure in most cases, because…”
“As mentioned, the lag is significant.” “Organic SEO typically shows impressions at four months and pipeline at twelve.”
“It depends on your stage.” “Before product-market fit, the right SEO spend is close to zero.”
The right column costs a few more words and is the entire difference between cited and not

The pattern in the right column: name the subject in the sentence rather than referring back to it, and put the claim before the qualification. A model reading a chunk needs the answer in it, not a pointer to where the answer was.

Name entities specifically

Retrieval works on meaning rather than exact strings, but specificity still decides whether a passage is judged relevant to a question.

A page that says “leading analytics platforms” is less retrievable for a question about a named tool than one that says the name. This is not keyword stuffing revived; it is the difference between a passage that demonstrably concerns a subject and one that gestures at a category.

The practical check is to read a section and count how many of its claims could apply to any company in your category. If most could, the section is generic in the specific way that makes it uncitable, because there is no reason to retrieve yours rather than anyone else’s.

The Entity Extractor reports which entities a page actually establishes and which a competing page names that yours does not, weighted by where they appear rather than how often. That gap list is a usable brief.

Answer the question in the first two sentences

Assistants attend to the beginning and end of a retrieved passage more reliably than the middle. Burying the answer in paragraph four of a section is a structural mistake even when the answer is excellent.

This runs against a habit that serves human readers well. Building to a conclusion, establishing context first, holding the reveal. All of that is good writing and all of it reduces the chance of being cited.

The resolution is not to abandon the craft. It is to lead each section with its claim and then develop it, which reads perfectly well to a person and puts the answer where a model will find it.

The shape of a section that gets cited
01
The claim, in one sentence
Stated plainly, with the subject named rather than referred to.
02
The reasoning or evidence
Why it is true. This is where the depth goes, and where a human reader who wants more finds it.
03
The qualification
When it does not hold. Placed last on purpose, because a model reading the opening should get the claim, not the caveat.

Structured data does a smaller job than people claim

Schema is worth having and it is not the lever it is often sold as.

What it reliably does: helps a system parse what a page is, who wrote it, when it was updated, and what questions it answers. FAQPage markup in particular gives a question-and-answer pair in a form that needs no extraction at all.

What it does not do: make a weak passage citable. No amount of markup rescues a section that does not answer anything on its own.

The genuine risk is worse than absence. Contradictory or duplicated structured data, especially two blocks claiming the same identifier, produces a page whose own description of itself does not agree. The Schema Validator checks a whole page rather than one block at a time, which is the only way to catch that.

The pages worth optimising, and the ones that are not

Not every page benefits from this, and treating it as a site-wide project wastes most of the effort.

Assistants get asked questions. A page that answers a question is a candidate for citation; a page that exists to convert someone already convinced is not, and rewriting it for retrievability improves nothing.

Which pages this work applies to
Page type Worth the effort Why
Explainers and guides Yes, first priority They answer questions in the form assistants receive them
Comparison and alternatives Yes Heavily asked, and specific enough to be citable
Original research or data Yes, highest value per page A number nobody else has is the strongest possible reason to cite you
Pricing and product pages Marginal Asked about, but the assistant usually wants a fact you can supply in one line
Case studies Rarely Specific to a client, and the question that would surface them is rare
The third row is the one most content programmes never invest in

The third row deserves emphasis. Original data is the most citable thing you can publish, because a model answering a question about a number has very few sources to choose from. A small survey of your own customers, published with its method, will be cited more reliably than a well-written guide competing with a hundred others.

The corollary is uncomfortable: if everything you publish is a synthesis of what already exists, there is no structural reason for any assistant to prefer you. That is true in conventional search too, and retrieval makes it sharper because there is no second position to occupy.

Chunking, and why your headings matter more than they did

Retrieval systems split documents before indexing them. Where those splits fall is not something you control, but you influence it heavily through structure.

A page with clear headings and self-contained sections tends to split along sensible boundaries. A wall of text splits arbitrarily, which produces chunks that start mid-argument and end mid-sentence. Neither is retrievable for anything.

What a well-structured section looks like in practice
150-400
words per section, a workable chunk
1 question
answered per section, completely
Named subject
in the heading and the first sentence
Guidance rather than rules. The principle is that a section should stand alone

This is why headings do more work than they did in conventional SEO. A heading was a scanning aid and a mild relevance signal. Now it is also the label on a retrievable unit, and a heading like "The next step" labels a chunk with nothing a retrieval system can use.

A practical consequence for long pieces: a four-thousand-word article with six headings produces enormous chunks that dilute whatever is specific in them. The same content under fifteen headings produces units that each concern one thing. The writing does not change; the retrievability does.

The same argument applies to lists and tables, which chunk badly when they run long. A table of forty rows arrives as a fragment with no header row attached. Splitting it under sub-headings, with a sentence stating what each part shows, survives extraction far better.

Measuring something that has no rank

The measurement problem is real and most reporting has not caught up with it.

There is no position to track. A page is either used to answer a question or it is not, and that varies by phrasing, by assistant, and between one week and the next. Rank tracking has nothing to measure.

Three things that can actually be tracked, in descending order of reliability.

  • Manual citation checks on a fixed question set. Write twenty questions a buyer would ask, run them monthly across the assistants you care about, and record whether you appear. Tedious, unglamorous, and the only direct measurement available.
  • Referral traffic from assistant domains. Small numbers, and they undercount badly because most assistant answers are read without a click. Directionally useful, not a total.
  • Branded search volume. If people who did not know you existed start searching your name, something upstream created that. The most reliable signal and the least specific.

What not to do: buy a tool promising an AI visibility score without understanding what it samples. Most run a question set and report appearance rate, which you can do yourself and understand properly. The broader attribution problem is covered in measuring organic when attribution breaks. The method, its failure modes and what to do with a bad result are set out in tracking brand mentions in AI search.

What to do first, in order

The work sorts cleanly by return against effort, and the order matters because the early items make the later ones worth doing.

Return against effort for the common interventions
Rewrite section openings to lead with the claim highest
Make sections self-contained high
Name entities specifically moderate
Add or fix FAQPage schema small but cheap
General schema expansion low
Effort is roughly inverse to this order, which is unusual and worth exploiting

Start with the pages you already have that rank well. They are already retrievable and already trusted; making their sections self-contained is a few hours of editing against pages with existing authority. New content written correctly from the start is the second job, not the first.

The thing to resist is treating this as a separate programme. It is an editing standard applied to work you were doing anyway, and companies that spin it out as an AI visibility initiative tend to produce a document rather than a change.

One honest caveat on all of it. The retrieval systems behind these assistants change without notice and without documentation, and anyone presenting a fixed playbook here is describing a snapshot. What makes the advice above worth following is that none of it is a trick: sections that answer one question completely, subjects named rather than referred to, and claims stated before caveats are better writing regardless of what any model does next. If the mechanism changes, the pages are still good pages, which is not something you could say about most tactics aimed at a specific system.

The short version

Assistants retrieve passages, not pages, so the unit of optimisation is the section rather than the article. Make each section answer one question completely, name entities specifically, and put the answer before the context. Then measure citation appearance rather than rank, because rank is not the mechanism any more.

Operator note

The change that produced the clearest difference for me was unglamorous: going through existing pages and rewriting the first sentence of every section so it named its subject instead of referring back. A day of work, no new content, and the pages started appearing in assistant answers they had not before.

I would still not describe that as proven. Citation behaviour is noisy enough that a month of observation is not evidence, and anyone claiming a precise causal number here is overselling what the measurement supports.

Frequently asked

How do I improve my brand’s visibility in AI search engines?
Make each section of your pages answer one question completely without depending on the surrounding text, because assistants retrieve passages rather than whole pages. Then name entities specifically rather than referring to categories, and lead each section with its claim rather than building to it. Those three changes cost editing time rather than new content, and they apply to pages you already have.
Is AI search optimisation different from SEO?
Partly. Crawlability, structure, genuine expertise and links still matter, so most of an existing SEO foundation carries over. Three things differ: the retrieved unit is a passage rather than a page, there is no ranking position to hold since you are either in the context window or absent, and planning starts from questions rather than keyword volume because nobody types keywords at an assistant.
Does schema markup help with AI search visibility?
It helps a system parse what a page is and what it answers, and FAQPage markup in particular supplies a question-answer pair needing no extraction. It does not make a weak passage citable. The larger risk is contradictory structured data, especially two blocks claiming the same identifier, which produces a page whose self-description disagrees with itself.
How do I measure whether AI search visibility is improving?
There is no rank to track, so direct measurement means running a fixed set of twenty buyer questions monthly across the assistants you care about and recording whether you appear. Referral traffic from assistant domains undercounts heavily because most answers are read without a click. Branded search volume is the most reliable indirect signal and the least specific.
Related reading
What AI search actually retrieves from a page Measuring organic when attribution breaks Programmatic pages without the thin-content penalty
Working on this yourself?

I take on a small number of engagements at a time. If you are past product-market fit and want to talk through what this would look like for your situation, the calendar is open.

Book a call