← Deep Thoughts
SEO·2026 take·8 min read

The honeymoon curve.

Add thousands of pages and traffic climbs for a month or two. Then the whole domain sinks, not just the new pages. The spike was never a reward. It was a window before the corpus got scored.

Kasey Cox
Founder & Director, Foxz Creative

The most common panicked question we field about search is some version of this: “We added a few thousand pages, traffic went up for a month, and then it fell off a cliff. What did we break?” The uncomfortable answer is that nothing broke. The curve did exactly what that strategy always makes it do. The early climb was the bait. The cliff was the bill.

We call it the honeymoon curve, and once you have watched it run its full length a few times you can see it coming from the first week. This is the mechanism, drawn properly, with the telemetry to match.

The shape of the curve

Google indexes new pages aggressively. When a domain suddenly emits thousands of fresh URLs, the crawler takes them, indexes a large fraction of them fast, and ranks them provisionally while it gathers behavioural data — clicks, dwell, bounce, the signals it cannot know until real users touch the pages. During that window the pages rank. Some of them rank well. Traffic climbs, the dashboard turns green, and everyone involved concludes the strategy is working.

It is not working. It is being measured. Four to twelve weeks in, once the corpus has accumulated enough behavioural signal to be scored, Google applies a quality judgement — and it applies it to the site, not the page. That is the part that surprises people. The demotion is not surgical. The thin pages do not quietly drop while the good pages hold. The whole domain gets re-rated downward, including the honest pages that were there before the experiment started. The homepage that used to rank for the company name softens. The one genuinely good service page slides four positions. Everything wearing the same domain pays for the corpus it now lives inside.

The early traffic is not a reward for the volume. It is a window before the volume is scored.

What the classifier is actually scoring

The mistake underneath the whole strategy is imagining that ranking is decided one page at a time. It is not. Google’s helpful-content and spam systems increasingly judge the corpus — the pattern across your pages — and a site assembled by mail-merging one template across a list of cities or industries has a pattern that is trivial to detect. Cross-page similarity is the giveaway. When two hundred pages differ only by the noun that got swapped into the same skeleton, a similarity score does not need to be clever to see it. Neither does a reader.

So the thing being scored is not “is this page about cannabis branding in Colorado.” It is “does this domain contain a genuine entity behind each page, or does it contain one template wearing three thousand costumes.” Real local presence, real client work, real differentiation between one page and the next — those are the signals that a page deserves to exist. Generate pages without that entity underneath and you have not built three thousand pages. You have built one page and photocopied it, and the classifier grades it accordingly.

The telemetry

Numbers make this concrete. The pattern is consistent enough across the sites we have diagnosed that its shape barely changes from one to the next. Picture a site scaled past three thousand pages through programmatic generation, in the usual proportions:

  • A few hundred location pages, one per city on a list.
  • Several hundred service sub-pages, one template per service.
  • A couple of thousand industry-by-state pages, the classic /industries/[thing]/[state]/ mail-merge.
  • A handful of comparison pages to round it out.

Traffic during the early phase is real. Calls happen. Clicks happen. Then the classifier finishes scoring the corpus, and the readouts turn:

  • Average position slides into the mid-40s — page four or five of the results, on aggregate.
  • Click-through rate collapses under a quarter of one percent — the pages are impressed and ignored.
  • Hundreds of URLs settle into “Discovered – currently not indexed,” Google’s way of saying it found them and decided they were not worth the crawl budget.
  • More sit in “Crawled – currently not indexed” — looked at once, then declined.

“Discovered and declined” at that scale is not a technical bug. It is an editorial verdict rendered by a machine, and it is remarkably close to the verdict a person would render skimming the same pages: there is nothing here that is not somewhere else.

Why the obvious fix does not work either

The instinct once the cliff arrives is to cut. Take the site from thousands of pages back down to a few hundred and the problem should shrink with it. It does not, and this is the trap that costs the most time. A rebuild that keeps the structure — fifty state pages sharing an identical heading pattern, service hubs still cut from one mould, comparison pages still mail-merged — still reads to the classifier as one template across N geographies. Rebuilds like that sit at cross-page similarity in the high teens and fail to recover for months, because the thing that triggered the demotion was never the page count. It was the pattern. Shrinking the pattern is not the same as removing it.

There is a second, quieter defect in most rebuilds: the removal is done with the wrong signal. Legacy URLs get left returning a plain 404, which puts them on a slow multi-week retry loop and keeps them counted in the domain’s footprint the whole time. A 410 Gone tells Google the page is deliberately and permanently removed, and it deindexes on a much faster clock. And the removal signal only works if Google is allowed to fetch the URL to see it — block the path in robots.txt at the same time and the crawler never receives the instruction it was supposed to act on. Get either of those wrong and the junk lingers in the index for months while everyone wonders why the recovery has stalled.

What actually turns it around

Recovery is not a trick and it is not fast, but the direction is unambiguous: raise the domain’s signal-to-noise ratio and let Google see you doing it. In practice that means removing the templated corpus outright rather than trimming it, with a permanent-removal signal the crawler is free to fetch. It means rebuilding only the pages that have a real entity behind them — work you actually did, places you actually operate, a point of view no competitor could paste onto their own site. And it means adding new pages at a rate that keeps the ratio improving every week rather than diluting it, which is a ceiling to respect, not a quota to hit. A domain recovers when it becomes, page for page, more credible than it was — not when it becomes bigger.

We have written the companion piece to this one on the wider tells of dishonest search work in cutting through the SEO nonsense, and on how this same pattern reads once AI engines are the ones doing the citing in SEO in the age of AI. The through-line across all three is the same: scale without a real thing underneath it is a liability that markets itself as an asset.

The one thing to take from this

If a search partner pitches you a plan whose engine is “we will generate a page for every service in every city,” you now know exactly what you are being sold. Not growth. The first half of the honeymoon curve, priced as if the second half will not arrive. It always arrives. The teams that understand why are the ones you want holding the domain when it does.

Search work that survives its own second month?

We run honest retainers built on real pages with a real reason to exist, and reporting that tells you what we did versus what merely happened.

New Vulpin AI-Assisted SEO, powered by SemRush data. Learn More