Tom McSherry
AI & Search

The cannibalisation trap in AI content buildouts, and how to avoid it

Tom McSherry

Tom McSherry

20 July 2026 · 8 min read

The short version: the most common way an AI content buildout goes wrong is not bad writing. It is that you finish with four pages all quietly aiming at the same search, and Google ranks none of them properly. That is cannibalisation. Once it is baked into sixty published pages it is slow and expensive to unpick, and it is almost entirely avoidable. The avoiding happens before you generate a single word.

I have written separately about why publishing at volume creates a mess in the first place, indexing problems and all - that is the hidden cost of AI articles at scale, and it is the piece to read for the why. This one is the other half: what you actually do so it does not happen to your site. If you are about to point a tool at your website and let it write fifty pages, this is the bit to sort out first.

What cannibalisation actually is, in plain English

Cannibalisation is when two or more of your own pages are chasing the same search. Google has to decide which one of yours is the answer, it is not confident, so it hedges. It half-ranks both. And half of two is worse than all of one - you finish below a competitor who has a single strong page on the topic.

Say you run a physio clinic in Hamilton. You have a service page built around 'sports physio Hamilton'. Six weeks into a content buildout you also have 'How to choose a sports physio in Hamilton' and 'Sports physiotherapy in Hamilton: what to expect'. All three answer the same thing for the same searcher. Google now has three candidates from one site and no clear signal about which one you meant to put forward, so your authority for that phrase gets split three ways. The competitor down the road with one solid page and nothing else sits above all of them.

Worth being clear: this is not a penalty. Google is not punishing you for having similar pages, it is doing its job with the signals it has, and you have given it ambiguous ones. I have covered how Google chooses which page to rank separately - cannibalisation is just that decision made hard, by accident.

Half-ranking two pages is worse than fully ranking one. That is the whole problem in a sentence.

Why AI buildouts cause this specifically

Plenty of sites cannibalise themselves without any AI involved. What is different about a buildout is the speed and the mechanism. Three things stack up.

The tool has no memory of what it already wrote

Every generation starts from nothing. Unless you have deliberately built a system that feeds it what already exists on your site, and most people have not, the tool does not know page 12 covered the same ground as page 41. A person would remember. They would get halfway in and say 'hang on, did I not already write this in June?' The tool has no such moment. It writes what you asked for, competently, every time, including the fourth time.

Near-duplicates are its default failure mode

Give a model two similar prompts and you get two similar articles. Not identical - it varies the phrasing, reorders the sections, picks a different opening. That is what makes it dangerous. The pages look different enough to a human skimming the list that nobody flags them, while being close enough in intent that a search engine reads them as competing for the same query. Obvious duplicates are easy to catch. It is the near-misses that get published.

The briefing habit makes it worse

Most buildouts start with a keyword export - two hundred phrases in a spreadsheet - and the plan becomes 'a page for each'. It feels systematic. It is the fastest way to build a site that fights itself, because a keyword list is not a list of topics. It is one topic listed twenty ways, plus plurals and 'near me' variants. Same trap as the older, slower version in why more pages can mean worse rankings. AI has not created the problem. It has made it possible to create the problem in a weekend.

The symptoms you can spot without any tools

You do not need software to catch this. If you check your own rankings by hand, or you just watch your traffic, the signs are recognisable:

  • A keyword's position swings between URLs. One week your service page shows for the phrase, the next week it is a blog post, then it is back. That flip-flopping is the clearest tell there is - Google is genuinely unsure which one you meant.
  • A new page never gains any traction. You publish something good, it gets indexed, then it sits there doing nothing for months. Often an older page of yours already owns the ground and neither is strong enough to take it outright.
  • Traffic is flat despite publishing steadily. Thirty new pages and the line on the chart has not moved. Flat while publishing usually means new pages are taking from old ones rather than adding.
  • Google keeps showing the wrong page. You search your own term and the article turns up instead of the page with the enquiry form on it.
  • Steady rankings got wobbly right after a publishing push. Nothing else changed. That timing is not a coincidence.

The cheapest check takes two minutes. Type your main service phrase into Google with 'site:yourdomain.co.nz' in front of it, and see how many of your own pages come back looking like they are trying to answer it. More than one and you have your answer.

How to avoid it: four things, all before you publish

1. Make a keyword map first

A keyword map is a boring document and it is the whole game. One row per page you intend to have. Each row gets the URL, the one phrase that page owns, the handful of closely related phrases it also covers, and a one-line note on who is searching for it and what they want. Nothing else. It fits in a spreadsheet.

The value is not the document, it is the rule it creates: no phrase appears on two rows. If you find yourself wanting to put 'emergency plumber Geelong' on two pages, you have caught a cannibalisation problem for free, before it cost you anything. Doing this after you publish is an audit. Doing it before is just planning. Same work, wildly different price. More on the sequencing in plan the site first, generate second.

2. One page per pool of demand, not one page per keyword

This is the rule that does most of the work. Group your keywords by what the person actually wants, then build one page per group. 'Sports physio Hamilton', 'sports physiotherapist Hamilton', 'sports injury physio near me' and 'running injury physio Hamilton' are not four pages. They are one page, written well, that happens to cover all four. Same person, same need, same thing they want to see when they land.

A new page is justified when the person behind the search wants a genuinely different thing. 'Sports physio Hamilton' and 'ACC physio claim' are different pages: one is choosing a clinic, the other is working out how funding works. 'Sports physio Hamilton' and 'sports physiotherapist Hamilton' are the same page with two names. Before you approve a page for generation, answer this in one sentence: what does this page own that nothing else on my site owns? If you cannot answer it cleanly, do not build it - add the material to the page that already exists.

If you have branches or multiple locations, the rule applies with more force, because location pages are near-identical by nature and a buildout will happily produce twelve of them. That has its own piece: stopping branch pages competing with each other.

3. Check before you publish, not after

Put one gate between generation and publishing, and make it a human one. It does not have to be heavy: someone reads the target phrase, checks it against the keyword map, and does the site: search to see what already exists. Two minutes a page. On a fifty-page buildout that is under two hours, and it is the difference between a site that works and a cleanup project.

The instinct is to review after publishing, because publishing is the satisfying bit and you can always fix things later. You can, but the cost is not the same. Fixing before publication is deleting a row from a spreadsheet. Fixing after means redirects, internal links to rebuild, and a period where the site looks unstable to Google. Same decision, two very different prices.

4. Publish slower than the tool can write

There is no prize for volume. Ten pages a month that each own something beats a hundred that overlap, and it gives you time to see what is working before committing to more. The tool can produce faster than any site can absorb, and that gap is where the mess accumulates.

If it has already happened, consolidate rather than delete

Deleting the losing pages is the tempting move and usually the wrong first one. Those pages may be weak, but some have links pointing at them, some have a trickle of traffic, and most have something worth keeping. Deleting throws that away and leaves you dead URLs to clean up as well.

The order I would work in:

  • List every page that touches the contested phrase. All of them, not just the obvious two.
  • Pick the winner. Usually the page with the strongest existing rankings, the most links, or the clearest commercial job - the one with the enquiry form beats the article almost every time.
  • Move anything genuinely useful onto the winner. Not a copy-paste dump - the good sections, folded in so the page still reads as one thing.
  • 301 redirect the losing URLs to the winner, so the links and history carry across instead of evaporating.
  • Fix the internal links. Anything that pointed at a retired page now points at the winner, with anchor text that says what the winner is about. This is the step people skip, and it is the one that tells Google which page you chose.
  • Then wait. Consolidation is not instant. Give it weeks, not days, before judging whether it worked.

If you are sitting on a large buildout and do not know how bad it is, do not start deleting to find out. Map what you have, decide what each page should own, then merge towards that. The wider version of this is in using AI for your website content without wrecking it.

Where AI genuinely earns its keep here

None of this is an argument against using AI to build content. I use it constantly, including for the posts on this site. It is very good at the mechanical work: drafting, restructuring, grouping a keyword export into rough clusters for you to correct, spotting where two pieces of text overlap once you point it at both of them.

What it is not good at is deciding which pages should exist. That is a judgement about your business and what you already have, and it is exactly the decision that, made wrong at volume, produces the mess. So I use AI for the absolute minimum amount of thinking possible and keep the page-level calls with a human. That is not a limitation I am apologising for. It is the part that makes the rest safe to automate.

The tools will get better, and when one can reliably hold a whole site in its head and tell me a planned page would undercut an existing one, I will say so. I would be happy to be proven wrong. Right now, a spreadsheet with one row per page and a rule that no phrase appears twice will save you more money than any feature on the market.

Want this done for you?

See how a profit-first SEO strategy could work for your business - no obligation.

See the case studies