Find repeated phrases and word combinations in articles, essays, and other texts. See how often each phrase appears and where it occurs, with all processing kept locally in your browser.
Overlapping occurrences count: “go go go” contains “go go” twice.
Phrase boundaries are . ! ? 。 ! ? … and line breaks (CR, LF, U+0085, U+2028, U+2029).
Commas, colons, semicolons and other punctuation separate words without ending a phrase.
This deterministic rule also splits abbreviations and decimals; it is not perfect linguistic segmentation.
Limits: 50,000 words and 2,000,000 UTF-16 text units. No partial analysis.
Sorted by occurrences descending, word count descending, then alphabetically (Unicode code-point order). Phrase labels use normalized words from the first occurrence; the preview preserves the original text.
| Phrase | Words | Occurrences | Action |
|---|
All occurrences are highlighted across preview pages; overlapping ranges are merged. The current occurrence is outlined. Text is displayed in pages of 12,000 UTF-16 units to limit rendering.
Repeated Phrase Finder is a free duplicate phrase checker for articles, essays, reports and other texts. It finds consecutive combinations of two to twelve words, counts their occurrences and shows their original locations. Repetition can be useful for emphasis or consistent terminology; it is not automatically a writing mistake.
The overview counts analyzed words, unique repeated phrases meeting your length and frequency settings before hiding, and visible phrases after hiding and result search. Input or analysis-setting changes mark previous results as out of date and disable viewing and exports until you analyze again.
Words begin with a Unicode letter or number and may contain letters, combining marks and numbers. Internal straight or curly apostrophes (U+0027, U+2019, U+02BC) and hyphens (U+002D, U+2010, U+2011) are retained, so don't and well-known each count as one word. Text without spaces, such as a continuous Chinese string, is not split into dictionary words.
Comparison uses Unicode NFC normalization without modifying your original text. Case-insensitive matching is enabled by default and uses JavaScript Unicode lowercase conversion; it is not locale-specific case folding. Visually similar letters, different apostrophe styles and different hyphen styles are not equated. Phrase labels join the first occurrence's normalized words with spaces; they omit separators such as commas.
The deterministic boundary rule ends a phrase at any . ! ? 。 ! ? … or line break (CR, LF, U+0085, U+2028, U+2029) between words. Commas, colons, semicolons and other punctuation can separate words but do not end a phrase. This rule deliberately also splits abbreviations, URLs and decimal numbers at periods. It is predictable, not perfect linguistic sentence segmentation.
red blue green. red blue yellow. contains red blue twice.Red blue. red blue. matches twice by default, but not with case-sensitive comparison.alpha beta. gamma delta. never creates beta gamma; neither does a line break between these word groups.go go go contains two overlapping occurrences of go go. Every valid starting word counts, even when ranges overlap.red, blue. red blue. matches red blue twice; the preview retains the comma.The optional English stopword filter hides only phrases consisting entirely of words in the tool's fixed English list, such as in the or of the. It is off by default, is not appropriate for other languages, and is not an exhaustive list of English function words. Stopwords are never removed before analysis: red and blue must not create a false red blue match.
The compact filter is on by default. It hides a shorter phrase only when every occurrence corresponds to the same fixed-position subphrase of a longer repeated phrase in the selected length range, with no extra occurrences. In red blue green. red blue green., the longer phrase remains while red blue and blue green are hidden. Add red blue. and red blue remains visible because it now occurs three times.
Result search is a case-insensitive substring filter on phrase labels. The interface states how many phrases the analysis filters and result search hide. Sorting is occurrences descending, word count descending, then alphabetical Unicode code-point order.
Select a phrase to highlight its occurrences without changing the input. Original case, punctuation and whitespace are retained, including decomposed accents. Overlapping highlighting ranges are merged, and the current occurrence is outlined. Previous and Next navigate individual occurrences with a counter such as 2 of 7.
The preview is paginated into 12,000 UTF-16-unit windows. All occurrences are marked on their respective pages; occurrence navigation opens the appropriate text page automatically. Long occurrences can continue onto the next page. User content is rendered as text nodes, never interpreted as HTML or executed as code.
TXT and Copy produce the same tab-separated report. CSV uses the columns phrase,word_count,occurrence_count with quoted, escaped phrase fields. JSON exports an array of objects with these fields; word and occurrence counts are numbers. All formats use all currently filtered and sorted results, not just the displayed table page.
Files must be readable UTF-8 .txt files; invalid UTF-8 produces an explicit error. Copy requires browser clipboard permission and HTTPS or localhost. Clear removes input, results, previews and messages while keeping your analysis settings.
All text processing and file reading stay in your browser. No text is uploaded, no AI or API is used, and the tool does not save your input. The safety limits are 50,000 words, 2,000,000 UTF-16 text units and 8 MB per imported file. Exceeding a limit rejects the whole input rather than silently processing a portion.
Maps group n-grams efficiently in a Web Worker. Large, varied texts can still require substantial browser memory, especially with the full 2–12-word range. Worker processing time is shown after analysis; rendering and exports take additional time. Performance depends on your device and browser, and no universal speed is promised. This is an exact word-combination finder, not plagiarism detection, grammar checking, semantic analysis or a writing-quality verdict.