Affinero

Blog

Deleting a post is a last resort, not archive hygiene

GENERATED · 947 words · EN→FI ✓

Editorial illustration: a pair of closed secateurs resting on a workbench beside one small cut branch, with a dense untr

Somewhere around post sixty, a blog acquires a graveyard. Pieces that were true the day they shipped, get almost no traffic now, and make the archive feel heavier than it actually is. The standard advice is to prune — cull the dead weight so the surviving pages can breathe. That advice is usually wrong, and Google's own documentation is closer to our position than to the position of most people citing it.

Google puts deletion last, in writing

Google's guidance on core updates has a section on what to do after a site loses ground, and it is unusually blunt about removal. Deleting content is "a last resort", to be considered only when a page genuinely cannot be salvaged. The page goes further, and that part is the interesting one: if you find yourself wanting to delete whole sections of your site, that impulse is diagnostic. Sections built for search engines rather than readers are the ones that feel disposable, because they always were.

That is not an instruction to keep everything forever. It is a claim about order of operations. Salvage first. Remove only what fails the salvage attempt.

The confusion has a traceable origin. In August 2023, Gizmodo reported that CNET had deleted thousands of older articles, and CNET's own explanation was that leaving old content live was costing it in search. Google's Danny Sullivan replied in public that deleting pages because Google supposedly dislikes old ones is, in his phrase, "not a thing". John Mueller added that routine clean-up is fine — the error is treating age by itself as evidence of harm. Search Engine Land covered the whole exchange, including the 2011 Panda-era advice that the modern version of this belief is descended from.

This is the same argument we made in Google doesn't penalize AI content, pointed at a different target. The systems are asking whether a page was worth publishing. Not when it was published, and not who typed it.

Four moves, and three of them keep the page

A post that has stopped working gives you four options, and deletion is the smallest of them.

You can leave it. This is the correct answer far more often than it feels, because a page that costs nothing to host and helps twelve people a year is not a problem you have. You can refresh it: same URL, same slug, rewritten wherever it went stale. Refreshing is the highest-return move on this list, because the page has already accumulated whatever links and history it has, and you keep all of it. You can consolidate: two thin posts that half-answer the same question become one that answers it properly, with the weaker URL redirected to the survivor. And you can remove it, which is worth doing when a page has no reason to exist for anyone — a stale announcement, a product you no longer sell, an event that happened.

Notice that three of the four keep the URL alive. That is not sentimentality. A URL that has been indexed for two years carries information — internal links pointing at it, the occasional external one, a search history — and deletion throws that away in exchange for a page count you were never being scored on.

Traffic is the worst available first filter

The usual pruning workflow starts by sorting the archive by sessions and drawing a line. It is the wrong first question, and Mueller made the point better than we can: hardly anyone reads your "About us" page, it probably has not changed in years, and nobody sane deletes it. Low traffic is a fact about demand, not about worth.

The better first question is whether the page still has a reason someone would land on it. A post answering a question people have stopped asking is genuinely finished. A post answering a question people still ask, badly, is a refresh. A post answering the same question as another post is a merge. Only the first is a deletion, and it is rarer than a sessions-sorted spreadsheet suggests.

An engine that only publishes is accruing a liability

Here is the part that cuts against us. A blog that publishes every morning generates the exact condition this post is about, faster than a human blog does. Volume is the cheap part now. Re-reading is not.

We have said before that an automated blog cannot do everything, and this belongs on that list in a specific form: a publishing engine with no memory of its own archive will eventually write the post it already wrote, and neither of them will be the good one. The fix is the same mechanism that makes internal linking work — the engine has to hold the whole archive in view, not just today's subject. Where that is missing, the honest description is not "automated content" but "accumulating content", and the difference shows up in about a year.

Prune the reason, not the page count

If your archive feels bloated, the useful audit is not a traffic export. It is reading the twenty oldest posts and asking, one at a time, what each was for and whether it still is. Most will need a paragraph rewritten. Some will need to be folded into a neighbour. One or two will genuinely be finished, and those you can delete without ceremony.

What changes if you accept this is mainly what you stop doing. You stop treating the archive as inventory to be trimmed to a number, and start treating it as a body of work with maintenance costs. That is slower, it makes a worse case study, and it is the version that survives a core update.