Verified against ChatGPT · 2026-07-28
Turn a raw keyword export into a cannibalization-proof cluster map
Groups a messy keyword list into distinct content clusters by search intent rather than shared words, names the primary target per cluster, and flags every pair of keywords that would fight each other for the same page instead of building two.
The prompt
Ready to copy — highlighted parts are example details you can swap.
You are an SEO information architect deciding how many pages a keyword list actually justifies — not how many keywords it contains. SITE CONTEXT: we sell CRM software aimed at small teams under 20 people, self-serve signup with no sales calls RAW KEYWORD LIST (with monthly search volume if available): best crm for small business (2400) crm software comparison (1900) free crm for startups (1600) crm pricing (880) how to choose a crm (720) crm implementation checklist (390) EXISTING PAGES ALREADY LIVE, IF ANY: /pricing — current pricing page; /blog/crm-vs-spreadsheets — comparison post TASK 1. GROUP BY INTENT, NOT WORDS. Cluster these keywords into groups a single page could realistically satisfy without competing against another page on this same site. Two keywords sharing several words but serving different intents — "crm pricing" and "crm software comparison" is the canonical example — must NOT be forced into one cluster unless you can name the specific reason one page genuinely covers both. 2. NAME EACH CLUSTER by the page title it implies, outcome-first, not the seed keyword restated. 3. FOR EACH CLUSTER, state the dominant intent (informational / commercial-investigation / transactional) and the single keyword that becomes the page's primary target. Every other keyword in that cluster becomes a secondary target the page earns through topical depth — never a second H1 competing for the same click. 4. FLAG CANNIBALIZATION RISK explicitly: any two keywords across the whole list, not just within one cluster, that look like they'd want the same page. For each pair, recommend which one becomes primary and which folds in as supporting content, rather than leaving both to become competing pages later. 5. FLAG WHAT YOU CANNOT CONFIDENTLY CLUSTER — ambiguous keywords whose intent genuinely depends on the searcher's stage or on business context you don't have — and state exactly what additional information would resolve it. Do not force a confident-sounding cluster onto an ambiguous keyword just to avoid an "unclear" answer. 6. IDENTIFY THE PILLAR-CLUSTER HIERARCHY: if any clusters are naturally subtopics feeding a broader pillar page, name the pillar and which clusters should link up to it rather than sit as unrelated siblings. 7. If /pricing — current pricing page; /blog/crm-vs-spreadsheets — comparison post is provided, check whether any new cluster would compete with a page that already exists rather than filling a genuine gap, and say so before recommending a new page be built. OUTPUT A table: Cluster name | Primary keyword | Secondary keywords | Intent | Suggested page type. Follow it with the cannibalization-risk list, the unclustered/ambiguous list, and the pillar hierarchy, each as its own labeled section.
Customize
Optional — swap in your own details for the highlighted parts above.
Why this works
Clustering by intent instead of shared words is the actual difference between designing a site architecture and producing a pile of similar pages, because two pages chasing the same intent split ranking signals between them instead of compounding into one authoritative result — a real and measurable cost search engines don't forgive just because the pages use different exact-match phrasing. The cross-list cannibalization check matters separately from the within-cluster grouping step, because the riskiest pairs are rarely the ones that look similar on the surface inside one cluster; they're the ones that surface only when the entire list is compared against itself, which is why the task explicitly requires scanning across all clusters, not just within each one. Explicitly flagging ambiguous keywords rather than forcing a confident-sounding grouping targets the actual failure mode that makes DIY keyword clustering expensive after the fact: an architecture decision made on a guess doesn't surface as a problem in a spreadsheet, it surfaces months later as two live pages quietly competing for the same ranking, at which point merging them costs a redirect and lost link equity instead of five minutes of honest uncertainty up front. Checking new clusters against pages that already exist closes the same gap from the other direction — a keyword list analyzed in isolation from the live site will happily recommend building a page that duplicates one already ranking, which is cannibalization the model could have caught for free if it had been given the context to look.
What you get back
Cluster: "Best CRM for Small Business" | Primary: best crm for small business | Secondary: free crm for startups | Commercial-investigation | Comparison listicle Cluster: "CRM Pricing Explained" | Primary: crm pricing | Secondary: — | Commercial-investigation | Pricing breakdown page Cannibalization risk: "crm software comparison" and "best crm for small business" want the same page — merge, keep "best crm for small business" as primary, fold the comparison angle into that page's structure instead of building a second one. Unclear: "how to choose a crm" could be a standalone guide or an intro section on the pillar page — depends on whether you want a separate top-of-funnel asset; resolve by checking if it has enough distinct search volume to justify its own page. Pillar: "CRM Pricing Explained" is a natural subtopic of "Best CRM for Small Business" — link it as supporting content, not a competing target.
Verified against
ChatGPT GPT-5.1 · 2026-07-28
Gemini Gemini 3 Pro · 2026-07-21
Changelog
- 2026-07-28 — Initial publish, verified against ChatGPT and Gemini 3 Pro.
Building this for real?
This is a free starting point. If you'd rather have SEO built and running for your business, that's Scult's day job.
EXPLORE SEO
