Cleaning Keyword Lists
Paste a keyword list, one per line or comma-separated, and this tool removes duplicates and reports how many it stripped. Comparison ignores case and collapses extra whitespace, so “Running Shoes”, “running shoes” and “running shoes” are recognised as one entry.
Duplicate keywords are inevitable once you combine sources. A list assembled from a research tool, a competitor export and your own Search Console data will overlap heavily, and the duplicates distort everything you do with it afterwards.
Why duplicates distort your analysis
The most direct problem is counting. If you are sizing a content plan by how many keywords fall into each topic, duplicates inflate some topics more than others depending on which sources overlapped. A topic that looks twice as big as another may simply have appeared in two of your exports.
Clustering is affected the same way. The Keyword Clustering Tool processes a limited number of keywords per run, so duplicates consume capacity that unique keywords needed, and they inflate the apparent size of whichever cluster they land in.
There is a practical cost too. A list with duplicates means someone eventually writes two pages for the same keyword because it appeared twice in the spreadsheet under different rows.
What counts as a duplicate
- Exact repeats, anywhere in the list.
- Case differences — “Running Shoes” and “running shoes”.
- Extra whitespace — leading, trailing, or doubled spaces between words.
- Empty lines, which are removed rather than kept as blank entries.
What is deliberately kept
Word order matters. “Shoes for running” and “running shoes” stay as separate entries, because they are genuinely different queries with potentially different search volumes and intents.
Singular and plural forms are kept separate for the same reason. “Running shoe” and “running shoes” often have meaningfully different volumes, and merging them would lose information you may need.
Punctuation differences are preserved. Some keyword sets legitimately distinguish “men's running shoes” from “mens running shoes”, since both are searched and may have different volumes.
If you want those variations collapsed, that is a judgement call requiring your knowledge of the market — automated merging would remove data you might need.
Where duplicates come from
Combining exports is the main source. Research tools, competitor analysis and Search Console all surface overlapping keyword sets by design, since they are describing the same search landscape.
Modifier expansion is the second. Generating variations around several seed keywords produces overlap wherever two seeds share modifiers — expanding both “running shoes” and “running trainers” yields near-identical modified lists.
Manual list-building over time is the third. A keyword spreadsheet maintained across months accumulates entries nobody remembers adding.
Where this fits in the workflow
Deduplicate before doing anything analytical with a list. Cleaning first means clustering, grouping and intent classification all work on accurate data rather than inflated counts.
The usual order is: combine your sources, deduplicate here, sort alphabetically with the Sort Keywords Alphabetically tool if you want to eyeball near-duplicates the tool deliberately kept, then cluster or group.
For grouping into pages afterwards, the Keyword Grouping Tool maps the cleaned list to a content plan with a suggested title per page.
Spotting near-duplicates the tool keeps
Sorting the deduplicated list alphabetically puts variations next to each other, which makes it easy to scan for singular/plural pairs and word-order variants you may want to merge manually.
Whether to merge them is a real decision rather than a cleanup step. Two variants with meaningfully different search volumes may deserve separate treatment; two that are clearly the same query with a typo do not. That judgement needs your knowledge of the market, which is why the tool leaves them alone.
Frequently asked questions
Does it treat singular and plural as duplicates?
No, they are kept separate. “Running shoe” and “running shoes” often have meaningfully different search volumes, so merging them automatically would lose information you may need.
Is the comparison case-sensitive?
No. “Running Shoes” and “running shoes” are recognised as the same keyword and only one is kept. Extra whitespace is also collapsed before comparing.
Does word order matter?
Yes. “Shoes for running” and “running shoes” are kept as separate entries because they are genuinely different queries with potentially different volumes and intents.
Why deduplicate before clustering?
Clustering tools process a limited number of keywords per run, so duplicates consume capacity unique keywords needed. They also inflate the apparent size of whichever cluster they fall into, distorting your content plan.
How do I find near-duplicates it kept?
Sort the cleaned list alphabetically, which puts variations next to each other. Whether to merge them is a judgement call about your market rather than something a tool should decide.