Web Analytics

 Repeated Phrase Finder

Find repeated phrases and word combinations in articles, essays, and other texts. See how often each phrase appears and where it occurs, with all processing kept locally in your browser.

Analysis settings

Overlapping occurrences count: “go go go” contains “go go” twice. Phrase boundaries are . ! ? 。 ! ? … and line breaks (CR, LF, U+0085, U+2028, U+2029). Commas, colons, semicolons and other punctuation separate words without ending a phrase. This deterministic rule also splits abbreviations and decimals; it is not perfect linguistic segmentation. Limits: 50,000 words and 2,000,000 UTF-16 text units. No partial analysis.

Find repeated phrases in text

Repeated Phrase Finder is a free duplicate phrase checker for articles, essays, reports and other texts. It finds consecutive combinations of two to twelve words, counts their occurrences and shows their original locations. Repetition can be useful for emphasis or consistent terminology; it is not automatically a writing mistake.

⭐ How to use Repeated Phrase Finder

  1. Paste or type text, import a local UTF-8 .txt file, or load the example.
  2. Choose minimum and maximum phrase lengths and minimum occurrences. Defaults are 2–6 words and at least two occurrences.
  3. Click Find repeated phrases. Processing runs locally in a Web Worker, not on every keystroke. Use Cancel to stop it.
  4. Search the results, browse pages of 50 rows, and select View occurrences for an original-text preview.
  5. Use Copy or Download as TXT, CSV or JSON to export all filtered results in the displayed sort order, including results on other pages.

The overview counts analyzed words, unique repeated phrases meeting your length and frequency settings before hiding, and visible phrases after hiding and result search. Input or analysis-setting changes mark previous results as out of date and disable viewing and exports until you analyze again.

🔎 Word matching and explicit phrase boundaries

Words begin with a Unicode letter or number and may contain letters, combining marks and numbers. Internal straight or curly apostrophes (U+0027, U+2019, U+02BC) and hyphens (U+002D, U+2010, U+2011) are retained, so don't and well-known each count as one word. Text without spaces, such as a continuous Chinese string, is not split into dictionary words.

Comparison uses Unicode NFC normalization without modifying your original text. Case-insensitive matching is enabled by default and uses JavaScript Unicode lowercase conversion; it is not locale-specific case folding. Visually similar letters, different apostrophe styles and different hyphen styles are not equated. Phrase labels join the first occurrence's normalized words with spaces; they omit separators such as commas.

The deterministic boundary rule ends a phrase at any . ! ? 。 ! ? … or line break (CR, LF, U+0085, U+2028, U+2029) between words. Commas, colons, semicolons and other punctuation can separate words but do not end a phrase. This rule deliberately also splits abbreviations, URLs and decimal numbers at periods. It is predictable, not perfect linguistic sentence segmentation.

💡 Practical examples and overlapping occurrences

  • red blue green. red blue yellow. contains red blue twice.
  • Red blue. red blue. matches twice by default, but not with case-sensitive comparison.
  • alpha beta. gamma delta. never creates beta gamma; neither does a line break between these word groups.
  • go go go contains two overlapping occurrences of go go. Every valid starting word counts, even when ranges overlap.
  • red, blue. red blue. matches red blue twice; the preview retains the comma.

⚙️ Hiding filters and compact results

The optional English stopword filter hides only phrases consisting entirely of words in the tool's fixed English list, such as in the or of the. It is off by default, is not appropriate for other languages, and is not an exhaustive list of English function words. Stopwords are never removed before analysis: red and blue must not create a false red blue match.

The compact filter is on by default. It hides a shorter phrase only when every occurrence corresponds to the same fixed-position subphrase of a longer repeated phrase in the selected length range, with no extra occurrences. In red blue green. red blue green., the longer phrase remains while red blue and blue green are hidden. Add red blue. and red blue remains visible because it now occurs three times.

Result search is a case-insensitive substring filter on phrase labels. The interface states how many phrases the analysis filters and result search hide. Sorting is occurrences descending, word count descending, then alphabetical Unicode code-point order.

📍 Safe original-text preview and navigation

Select a phrase to highlight its occurrences without changing the input. Original case, punctuation and whitespace are retained, including decomposed accents. Overlapping highlighting ranges are merged, and the current occurrence is outlined. Previous and Next navigate individual occurrences with a counter such as 2 of 7.

The preview is paginated into 12,000 UTF-16-unit windows. All occurrences are marked on their respective pages; occurrence navigation opens the appropriate text page automatically. Long occurrences can continue onto the next page. User content is rendered as text nodes, never interpreted as HTML or executed as code.

📥 Local import, copy and downloads

TXT and Copy produce the same tab-separated report. CSV uses the columns phrase,word_count,occurrence_count with quoted, escaped phrase fields. JSON exports an array of objects with these fields; word and occurrence counts are numbers. All formats use all currently filtered and sorted results, not just the displayed table page.

Files must be readable UTF-8 .txt files; invalid UTF-8 produces an explicit error. Copy requires browser clipboard permission and HTTPS or localhost. Clear removes input, results, previews and messages while keeping your analysis settings.

⚠️ Privacy, limits and what this tool cannot do

All text processing and file reading stay in your browser. No text is uploaded, no AI or API is used, and the tool does not save your input. The safety limits are 50,000 words, 2,000,000 UTF-16 text units and 8 MB per imported file. Exceeding a limit rejects the whole input rather than silently processing a portion.

Maps group n-grams efficiently in a Web Worker. Large, varied texts can still require substantial browser memory, especially with the full 2–12-word range. Worker processing time is shown after analysis; rendering and exports take additional time. Performance depends on your device and browser, and no universal speed is promised. This is an exact word-combination finder, not plagiarism detection, grammar checking, semantic analysis or a writing-quality verdict.

🔗 Related text tools


This tool is also known as

  • repeated phrase finder
  • find repeated phrases in text
  • duplicate phrase checker
  • repeated word combinations finder
  • text repetition checker

Frequently Asked Questions

It finds repeated consecutive combinations of two to twelve words in one text, counts their occurrences and highlights their original locations. It is free and runs locally in your browser.

Paste your text or import a UTF-8 .txt file, choose your phrase lengths and minimum occurrences, then click Find repeated phrases. Search or browse results and select View occurrences to see the original text.

No. JavaScript and a local Web Worker process your text in your browser. There is no text upload, backend processing, AI or analysis API, and the tool does not store your input.

Choose whole-number minimum and maximum lengths between 2 and 12 words. The defaults are 2 to 6 words with at least 2 occurrences. The minimum length must not exceed the maximum.

Yes, by default comparison uses JavaScript Unicode lowercase conversion after NFC normalization. Enable case-sensitive comparison to distinguish Red blue from red blue. This is not locale-specific case folding.

No. The explicit rule stops phrases at . ! ? 。 ! ? … and CR, LF, U+0085, U+2028 or U+2029 line breaks. Commas, colons and semicolons separate words without stopping a phrase. Periods also split abbreviations and decimals; this is not perfect linguistic segmentation.

Yes. Go go go contains go go twice, starting at the first and second words. Overlapping preview ranges are merged visually, but each starting position counts and can be navigated separately.

The compact filter, enabled by default, hides shorter phrases when a longer phrase in the selected length range contains exactly the same aligned occurrences. If the shorter phrase has an additional occurrence, it remains visible. Turn off this setting to see all lengths.

It optionally hides phrases made entirely of words in a fixed English stopword list, such as in the. It is off by default and is intended only for English. Words are never removed before analysis, so non-consecutive words never become false phrases.

Yes. Unicode letters, combining marks and numbers are supported. Internal straight and curly apostrophes and common hyphens stay inside words. NFC-equivalent accents match without changing original positions. Different apostrophe or hyphen styles are not treated as identical. Continuous unspaced scripts are not segmented into dictionary words.

The limits are 50,000 words and 2,000,000 UTF-16 text units, with an 8 MB import-file limit. Inputs exceeding a limit are rejected with an explicit message; no partial analysis is performed. Large varied texts may use substantial browser memory.

Yes. Copy and TXT produce the same tab-separated report; CSV and JSON contain phrase, word_count and occurrence_count. All currently filtered and sorted results are exported, including results beyond the 50-row table page. JSON counts are numeric and CSV phrase fields are escaped.

Import requires a readable UTF-8 .txt file within the size limits. Invalid UTF-8 and read failures show an error. Clipboard access needs HTTPS or localhost, browser support and permission. If copying is blocked, use Download as TXT instead.

Yes. Cancel terminates the Worker. Changing the input or analysis settings also cancels pending analysis and marks previous results as out of date. Run the analysis again; an older run cannot replace newer results.

No. Repetition can provide emphasis, structure and consistent terminology. Use results as a review aid, not an automatic quality judgment. This tool does not perform plagiarism detection, grammar checking or semantic analysis.

Repeated Phrase Finder searches combinations of multiple consecutive words. Find duplicate words works with individual words, Word density measures word frequency, Text difference compares texts, and Text Counter & Analyzer provides general text statistics.

Our text tools

General
Text counters
Find in text
Text cleaning