A content audit is four decisions, not a spreadsheet
Most audits produce a tab with three hundred rows and a column called action, filled in optimistically. Six months later the tab is stale and the site is unchanged. The problem is not the data. It is that nobody decided anything.
Why most content audits change nothing
The standard method is to export every URL, pull metrics against each, colour the rows, and hand over a spreadsheet. It is thorough and it usually fails, for a reason that has nothing to do with the analysis.
A spreadsheet with three hundred rows and an action column is not a plan. It is three hundred small decisions nobody has authority or time to make, presented as though the hard part is already done. The hard part has not started.
An SEO audit has the same failure mode, and the defect list regrows for the same reason: cataloguing is not deciding. The difference with content is that the decisions are more consequential, because removing a page is irreversible in a way that fixing a title tag is not.
The version that works inverts the order. Start from the four decisions you are willing to make, then collect only the data that separates them. Everything else is interesting and does not change what you do on Monday.
The data that actually separates those four
Four decisions need surprisingly little data. Most audit templates collect twenty columns and use three.
- Clicks and impressions, ninety days. Clicks say whether anyone arrives. Impressions without clicks say the page is visible and unconvincing, which is a rewrite, not a removal.
- The queries it actually ranks for. The single most useful column, and the one most often missing. A page that misses its target and ranks for something else is not a failure; it is a page whose brief was wrong.
- Inbound internal links. A page with real editorial links has been endorsed by your own site. Removing it breaks something. A page with none is cheap to remove and was probably never supported in the first place.
- Last meaningful update. Not the modified date, which reflects plugin churn. When did the content last change in a way a reader would notice.
Two columns people collect and rarely use: word count, which tells you nothing on its own, and bounce rate, which for an informational page usually means the reader got their answer and left.
| Clicks | Impressions | Internal links | Decision |
|---|---|---|---|
| Some | Some | Any | Keep |
| None | Many | Any | Rewrite, it is visible and unconvincing |
| None | Few | Several | Investigate, something links here for a reason |
| None | None | None | Remove |
| Some | Some | Any, and a sibling ranks for the same query | Merge |
The page that ranks for nothing it was written for
This is the most interesting case in any audit and the one a spreadsheet handles worst.
A page was briefed against a keyword. It does not rank for that keyword. It does rank, quietly, at position nine for four queries nobody planned, and those queries convert better than the original target would have. On a colour-coded sheet it is red, because the sheet compares performance against intent rather than against value.
The action is not rewrite and it is certainly not remove. It is to re-target the page at what it already earns: change the title, lead with the question it actually answers, and let the original brief go. This is the cheapest win available in most audits and it is invisible unless you pull actual query data rather than target keywords.
The reverse case is worth naming too. A page that ranks exactly as briefed and earns nothing is a page whose brief was aimed at a query with no commercial value. That is a planning fault, not a content fault, and rewriting it will not help.
Ranks for nothing, earns nothing, nothing links to it
Misses its target, ranks for four unplanned queries
Hits its target exactly, earns nothing
What to do about pages that are simply old
Age gets treated as a defect in most audits, and it is not one. A page from 2021 that still answers its question correctly is a page that has had four years to accumulate links and ranking history. Refreshing it because of its date is how sites reset assets that were working. A related judgement, deciding which queries are now answered on the results page and worth conceding, is covered in the AI Overviews piece.
The distinction worth making is between pages that are old and pages that are wrong. Only the second needs anything.
| Condition | Signal | Action |
|---|---|---|
| Old and still correct | Stable rankings, steady clicks | Leave it. Update the date only if something changed. |
| Old and factually stale | Names a deprecated tool, a superseded figure, a dead link | Fix the specific claim, not the whole page. |
| Old and structurally wrong | Answers a question people stopped asking | Re-target or remove, depending on what it earns. |
The one genuine exception is anything with a year in the title or the URL. Those pages carry an implicit expiry date, and a page called “best tools 2024” is actively telling readers it is out of date. Either commit to updating them annually or stop putting years in titles.
There is a quieter version of staleness worth checking: outbound links. A page full of links to sites that have moved or closed reads as abandoned even when the argument is still sound, and it is a ten-minute fix rather than a rewrite.
Removal is the decision that frees capacity
Every other decision costs something. Keep costs nothing and changes nothing. Merge and rewrite both cost writing time. Removal is the only one that gives capacity back, and it is the one people put off indefinitely.
The hesitation is understandable and usually misplaced. The fear is losing traffic from a page that might be earning something invisible. The check for that is two minutes: does it have clicks, does anything link to it, does it rank for anything. If all three are no, the page is costing you and returning nothing.
What removal actually gives back, in rough order of value.
That last bar matters. Removing dead pages does not directly lift the pages that remain, and anyone promising that is overselling. What it does is stop you spending crawl budget, review time and internal links on pages that were never going to work, which is a real return measured over quarters rather than weeks.
One caution: removing and redirecting are different decisions. A page with no value and no inbound links can simply go. A page with inbound links, even bad ones, should redirect somewhere sensible, because the alternative is manufacturing your own broken links.
The merge decisions are where the risk sits
Merging is where an audit either earns its cost or destroys something. Two pages competing for one query is the obvious candidate, and the mechanics of consolidating them well are worth reading separately before you start.
The audit-specific point is about selection: a spreadsheet identifies merge candidates by title similarity, which is the wrong signal. Two pages with near-identical titles can serve completely different intents, and two pages with unrelated titles can compete directly.
The signal that actually works is whether both pages rank for an overlapping set of queries. That is data you already have and it takes one pivot to see.
Auditing a site that publishes programmatically
Everything above assumes pages a person wrote. A site with generated page families needs a different unit of analysis, because auditing two hundred location pages one at a time is neither possible nor useful.
The unit becomes the family, not the page. The questions change with it.
- Does the family as a whole earn anything? If the top ten pages in a two-hundred-page family carry ninety-five percent of the clicks, you do not have a family of two hundred useful pages. You have ten, and a hundred and ninety pages of overhead.
- What differs between two adjacent members? If the answer is a city name and a number, a reader cannot tell them apart and neither can a search engine. That is a template problem, not two hundred content problems.
- Where does the tail flatten? Most generated families have a point past which members earn nothing at all. That point is your removal boundary, and it is usually much earlier than anyone expects.
The action on a failing family is almost never to improve every member. It is to cut the family to the members that earn, and to fix the source data so the next generation does not reproduce the same problem at the same scale.
This is also where the crawl argument stops being theoretical. On a forty-page site, dead pages cost you nothing measurable. On a site with three generated families, they are most of what a crawler sees.
Running one without the spreadsheet trap
The practical version, in the order that keeps it from stalling.
- Removals first. They are unambiguous, they free capacity, and finishing something builds the momentum the rest of the audit needs. An audit that opens with the hardest decisions never finishes.
- Re-targets second. Cheapest wins available, and they need no new writing. Title and opening paragraph, informed by the queries the page already earns.
- Merges third, one at a time, with a full crawl cycle between each so you can see what each one did.
- Rewrites last, and only where the query is worth it. Most rewrites in most audits are not worth doing and get done anyway because rewriting feels productive.
The mechanical parts are worth automating rather than eyeballing. A link map shows which pages have real editorial support and which are orphaned, which is the input the removal decision needs and the one most audits guess at. Two pages competing for the same query is the other automatable check, and it is the one a manual scan misses most often.
One practical constraint that keeps audits honest: cap the number of decisions you will act on in a quarter before you start looking. Thirty is a realistic number for one person. Knowing the cap changes how you triage, because you stop trying to grade every page and start looking for the thirty that matter most.
The alternative, which is what most audits do, is to grade all three hundred equally and then discover there is no capacity to act on any of it. The grading was never the constraint.
The measure of a good audit is not how many rows it has. It is how many URLs are different a month later.
Audit to make four decisions: keep, merge, rewrite, remove. Judge a page on what it earns rather than what it was written to target, because pages that miss their brief often rank for something better. Do the removals first, because they are the only ones that free up anything.
The first content audit I ran produced a beautiful spreadsheet, four hundred rows, every column filled. Nothing happened. Six months later I reran it and the rows had barely changed, because nobody had ever been asked to decide anything, only to look at it.
The version that works now starts with removals, because they are unambiguous and finishing something is what keeps the rest moving. The audits that stall are the ones that open with the merges.
Frequently asked
What is a content audit in SEO?
How often should you do a content audit?
What should I remove in a content audit?
Does deleting pages improve SEO?
I take on a small number of engagements at a time. If you are past product-market fit and want to talk through what this would look like for your situation, the calendar is open.
Book a call