Programmatic pages without the thin-content penalty
Scaled page generation is not inherently spam. It becomes spam at the exact point where the template stops carrying information the reader could not have assembled themselves.
The line is utility per page, not page count
Publishing ten thousand pages from a database is not the offence. Publishing ten thousand pages where the only difference between them is a substituted noun is. The distinction matters because the first is how several genuinely useful properties are built, and teams avoid the whole category out of a vague fear of penalty.
The test to apply per page: could a reasonably motivated reader have produced this page themselves in under a minute using a search box? If yes, the page is a wrapper around a query and adds nothing to the index. If no, because the page contains a comparison, a calculation, a normalisation across sources, or genuine editorial judgement, then it is a real page that happens to have been generated.
Where the information actually comes from
Every defensible programmatic page has a source of information beyond the template. Usually one of four:
- Proprietary data. Numbers you collected and nobody else has. The strongest position and the most expensive.
- Computation. Public inputs, non-obvious output. A calculator page that runs the arithmetic a reader would otherwise do in a spreadsheet.
- Normalisation. Several messy public sources reconciled into one comparable format. The work is the reconciliation, and it is real work.
- Editorial layer. A generated skeleton with a human-written judgement on top. Scales worse, defends best.
| Source | Example | Defensibility | Scales |
|---|---|---|---|
| Proprietary data | Numbers only you collected | Highest | Poorly, expensive |
| Computation | A calculator over public inputs | High | Well |
| Normalisation | Messy sources reconciled | Medium | Well |
| Editorial layer | Generated skeleton, human judgement | Highest | Worst |
If you cannot name which of those four a proposed page family rests on, the family does not have a reason to exist yet. That conversation is much cheaper before the build than after.
Generated pages also have to survive being read by machines rather than people, which is a separate constraint covered in what AI search actually retrieves.
Governance, because these families rot quietly
The specific risk with generated families is that quality degrades in the tail without anyone noticing. The first hundred pages have rich data because they cover popular entities. Page four thousand covers an entity with three fields populated and renders as a mostly-empty shell.
Handle it with a coverage threshold enforced at build time rather than at audit time. Define the minimum fields a page needs to be worth indexing. Pages that clear it publish. Pages that do not are held back, not published with gaps, and not published with a placeholder apologising for the gaps.
The second control is a sampling review. Pull twenty pages at random every month, weighted toward the tail rather than the head, and read them as a reader would. Tail rot is invisible in aggregate metrics and obvious within thirty seconds of reading.
Crawl budget is a real constraint here
A large generated family competes with the rest of your site for crawl attention. If four thousand thin pages absorb the crawl, your commercial pages get visited less often and your genuinely good content takes longer to be reflected.
Two practical levers. Gate the family behind the coverage threshold so fewer, better pages exist in the first place. And give the family a clean hub structure so a crawler can reach any member in two or three hops rather than paginating through a hundred index pages.
Neither is exotic. Both get skipped, because the build gets scoped as a data problem and the discoverability layer gets treated as something to sort out afterwards.
Hub structure here means the same thing it means everywhere else on a site, which is covered properly in the piece on internal link graphs.
Name the information source before you build, enforce a coverage threshold at build time, and sample the tail monthly. Generated pages fail on thinness in the tail, not on volume at the head.
I ask one question before any programmatic build: which of the four information sources is this resting on. If the answer takes more than a sentence, the family is not ready, and building it anyway just means discovering that at page four thousand instead of at the whiteboard.
The coverage threshold is the control that actually saves these projects, and it is almost always the one that gets cut for launch. Hold pages back rather than shipping them with gaps. Nobody has ever regretted publishing fewer, better pages.
Frequently asked
Is programmatic SEO against Google’s guidelines?
How many programmatic pages is too many?
What is a coverage threshold?
Do programmatic pages hurt crawl budget?
I take on a small number of engagements at a time. If you are past product-market fit and want to talk through what this would look like for your situation, the calendar is open.
Book a call