Skip to content
Have a project in mind?
The Gloria JournalWebsite management

Auditing old blog posts: keep, rewrite, merge or delete

Old newspapers stacked and annotated in red pen

For six years, one of the oldest posts on this site was a perfect example of what not to leave online. In March 2020, with schools shut and fixtures cancelled, I published a call for players to form a team for an online game tournament: a short text, a screenshot, a username. No connection to search, no value for a business owner. It stayed indexed anyway.

Almost every business blog carries pages like it. The press release from 2017. The seasonal greetings page. The article written for a keyword the organisation abandoned two strategies ago. The question is not whether to have a clear-out — that is a reflex, not a method — but how to decide, page by page, on criteria you can measure.

What follows is the editorial audit I run for clients, and eventually turned on my own archive. Three steps: inventory, measurement, decision. The third matters most, because deletion is one of five outcomes, and rarely the best.

Why an old, useless post is a genuine problem

The first effect is human rather than algorithmic. A prospect who lands on an off-topic, dated page full of dead links forms an opinion about how seriously the business takes its own work. That matters most when the page ranks on the brand name — exactly the query a buyer runs before making contact.

The second effect is statistical, and it cuts the other way. In a study published by Ahrefs in December 2023, covering roughly 14 billion pages from its Content Explorer index, 96.55% of pages received no organic traffic from Google. Those estimates come from Ahrefs' own keyword database, which probably understates the long tail. Either way, owning pages that attract nothing is the norm. It is not the count of silent pages that should trigger an audit, but their nature.

The third effect concerns how the site reads as a whole. The self-assessment questions in Google's guidance on helpful, reliable, people-first content are about the reader's experience, demonstrated expertise, and whether the content is worth someone's time. A site where a visible share of pages fails those questions sends a consistent signal.

Step 1: inventory before you judge

An editorial audit starts with an exhaustive list of indexable URLs, not with an impression. Three sources cross-check each other: the CMS export (under WordPress, posts and pages with date, category and author), the XML sitemaps, and a full crawl. The gaps between them are instructive on their own — orphan pages missing from the sitemap, indexable tag archives nobody knew existed, attachments published as standalone pages.

For each URL, build one row: title, publication date, last modified date, word count, category, HTTP status, indexing status, and incoming internal links. You are documenting, not judging.

Step 2: measure, with four indicators only

A reliable sort does not need twenty metrics. It needs four, read over at least twelve months so that seasonality is absorbed, not mistaken for a trend.

The four measures that decide a page's fate
Indicator Source What it tells you
Clicks and impressions Search Console, Performance report Whether the page is seen, and for what
Queries captured Search Console, filtered by page The subject actually covered, often not the original intent
Incoming links Link tool, Search Console links report Earned value a blunt deletion throws away
Assisted conversions Analytics, goals or events Its role in a journey, even without direct traffic

These columns consolidate easily in a spreadsheet. Past a hundred URLs, a Looker Studio report connected to Search Console saves re-exporting at every review. The tool is not the point; freezing the figures at a date and reading them cold is.

If you have no dependable tracking yet, that is the job to do before the editorial audit: our page on measurement and traffic management covers that groundwork. On the links column, a backlink audit is worth running in parallel, because one page with genuine earned links changes the decision entirely.

Step 3: decide, with five possible outcomes

Every URL leaves the audit with one dated decision. Five outcomes are enough.

Decision grid for an editorial audit
Decision When to choose it How it is implemented
Keep Steady traffic or conversions, subject still accurate No action, annual review
Rewrite Relevant subject, but dated, thin or misaligned with queries captured Rewrite in place, same URL, modified date updated
Merge Several pages cover the same ground and compete Consolidated on the strongest URL, 301 redirects from the rest
De-index Useful to visitors, of no interest in search (legal notices, thank-you pages) noindex rule, URL left open to crawling
Delete No traffic, no incoming links, no link to the business, beyond repair 404 or 410 response, no convenience redirect

Merging: the 301 redirect as a signal

When two or three articles tread on each other, consolidation beats deletion. Google's documentation on permanent redirects explains that Googlebot follows the redirect and that the indexing pipeline treats it as a signal designating the target as canonical. Incoming link value and page history are not lost — on one condition. The redirect must point at content that covers the same subject. Sending retired posts to the home page is not consolidation; it is tidying up in public.

De-indexing: use noindex, never robots.txt

To keep a page reachable but out of results, the rule goes in a meta tag or an HTTP header:

<meta name="robots" content="noindex">

X-Robots-Tag: noindex

Google is explicit in its documentation on blocking indexing: the URL must not be blocked in robots.txt. If the crawler cannot load the page, it never sees the rule, and the URL can stay in the index. This is the most common mistake I find in audits, usually made by someone being careful rather than careless. Afterwards, a systematic check on indexing status is worth the ten minutes.

Deleting: 404, 410 and the removals tool

As far as Google is concerned, 404 and 410 lead to the same outcome. The documentation on HTTP and network errors puts both codes in one category and states that the crawler then reports to the next stage of processing that the content does not exist. A 410 marks a deliberate, permanent removal and is the better choice for that reason, but it will not change the timetable. What matters more is not producing a "deleted" page that answers 200 with an error message on it: the textbook soft 404.

The URL removals tool in Search Console deletes nothing. It hides the page within about a day, but the request lasts only around six months, according to Google's documentation on removing information. Treat it as an emergency dressing, always paired with a permanent fix: deletion, a noindex rule, or password protection.

What to expect: modest, slow, never guaranteed

This is where I temper expectations. Deleting fifty dead pages does not lift the survivors through some law of communicating vessels. Google is direct about this in its page on core updates: removal is a last resort, for content beyond saving, and mass deletion mostly reveals that those pages were made for search engines rather than readers. The same page notes that some changes take effect within days, while others may take months.

The real benefits sit elsewhere. A site a stranger understands in two minutes. Internal links pointing at the pages that carry the business. Merged content that stops competing with itself. When ranking gains arrive, they almost always come from the rewrites and merges, not the deletions.

What I decided about that post

My 2020 call for players went through the grid like everything else. Organic traffic: none. Queries captured: nothing connected to the work I do. Incoming links: nothing usable. Conversions: zero. On those criteria alone, the logical outcome was deletion with a 410.

I chose rewriting instead, for a reason I will defend. The URL is short and old, and the post illustrated the question clients actually ask during an audit better than any client example could. A worthless old page can become useful again if the subject it illustrates is useful. That is the nuance the grid exists to allow: decide case by case, with the figures in front of you.

If you run this exercise on your own blog, start with the oldest pages. They are rarely the most numerous, but almost always the most embarrassing. Then cross the editorial sort with a technical audit: a page kept for good editorial reasons is of little use if it is slow, orphaned or blocked from crawling.

Common questions

Should I delete blog posts that get no traffic at all?

Not automatically. Pages with no traffic are very common, and that alone is not a reason to remove anything. Check the incoming links, the queries the page captures and its role in conversion journeys first. No traffic but a relevant subject usually means rewriting or merging.

Is 404 or 410 better when removing a page?

Google treats both the same way: in each case its crawlers record that the content no longer exists. A 410 marks a deliberate, permanent removal; a 404 simply reports an absence. The real risk is elsewhere — a removed page must never return 200 with an error message, which Google flags as a soft 404.

What is the difference between de-indexing a page and deleting it?

A noindex rule keeps the page reachable for visitors and crawlers but takes it out of search results. Deleting removes the page, which then answers 404 or 410. De-index what is useful to visitors but of no interest in search; delete what serves no purpose at all.

Is blocking a page in robots.txt enough to de-index it?

No, and it is counterproductive. If the URL is blocked, the crawler never loads the page and never sees the noindex rule it contains, so the URL can remain in the index. Let the crawl through and serve the rule in a meta tag or an HTTP header.

How long does a Search Console removal request last?

It hides the page within about a day, but lasts only around six months according to Google's documentation. It is an emergency measure, not a solution, and should always be paired with something permanent: deletion, a noindex rule, or password protection.

How long before a content clear-out has any effect?

Google states that some changes take effect within days, while its systems can take several months to account for others. Measure over two or three months at minimum, and expect no mechanical gain: improvements come mainly from rewrites and merges.

This link opens in a new tab.