Both directions, not one average
How much of A appears in B is a different number from how much of B appears in A, and the gap is the whole story when one text is longer. Both are reported in plain words.
You have two pieces of writing and you need to know how much of one is in the other. Paste a text into each box — or a web address, and the page is fetched for you — and get a percentage, the sentences that match, and the passages the two share word for word. Nothing is uploaded. No sign-up.
How much of A appears in B is a different number from how much of B appears in A, and the gap is the whole story when one text is longer. Both are reported in plain words.
Every sentence is paired with its closest match on the other side, so you can read the overlap for yourself instead of arguing with a percentage.
The longest runs of identical words the two texts share, with a word count on each. A forty-word identical run is more convincing evidence than any percentage.
The sentences with nothing close to them on the other side — the part that is genuinely original, listed so you can see what survived.
The comparison is JavaScript running on your machine. Nothing is uploaded, nothing is stored and nothing is added to anybody's archive of submitted work.
Every matched pair with its score, ready for Excel, Sheets or Numbers — so an editor can work down the list without opening the tool.
Paste the writing itself, or paste a URL and the page is fetched and reduced to its readable text. One side can be text and the other a URL.
The whole comparison happens on your own machine, in under a second. Nothing is uploaded and nothing is kept.
Shaded sentences in the side-by-side view jump to their partner when clicked. Work down the matched pairs, then export the list to CSV.
Most free similarity checkers give you one percentage and no way to argue with it. These are the five figures this one reports and what each is actually measuring.
The same percentage answers two different questions depending on who is asking, and the searches that land on a similarity checker come from both.
Two of your own pages saying the same thing. Nobody has done anything wrong; the pages are just competing. Google picks one and filters the other out of the results, so the duplicate quietly stops earning traffic. The fix is a canonical tag, a redirect or a merge — not a rewrite. Most of it is created by templates and URL parameters rather than by writers.
Somebody else’s words presented as yours. Here the percentage matters far less than the passages: one uncited paragraph is serious at 4% overall, and a page of properly quoted and referenced material is fine at 40%. Read the matched sentences and the longest shared runs, and ignore the headline.
This tool measures the overlap. Which of the two problems you have is your call, and it changes what you do about it.
A lot of people searching for a Turnitin similarity checker — or a Turnitin similarity checker free of charge — want the Turnitin similarity report without paying for Turnitin. That is not something any free tool can give you, and the sites promising it are either selling something else or keeping your paper. Here is the honest version.
“AI similarity checker” gets searched for two different things, and the free AI similarity checker results that come back rarely say which one they are. This is an AI similarity checker free of any model at all, so it is worth being specific.
Text similarity is a measure of how much two pieces of writing have in common, usually given as a percentage. Lexical similarity counts shared words and phrases, so it catches copying. Semantic similarity compares meaning, so it catches the same idea written differently. Most checkers report the lexical figure, because it is the one you can point at and verify.
Duplicate content is the same or near-identical text appearing at more than one address, either within a site or across sites. It happens most often by accident: printer-friendly pages, URL parameters, filtered category listings, product descriptions copied from a supplier, or the same article republished under two paths. It is a housekeeping problem rather than a moral one, but it costs traffic all the same.
There is no universal figure, and any tool that gives you one is guessing. Academic institutions typically accept somewhere between 10% and 25% including quotations and references, and set their own bar. For two pages on the same website, anything above about 60% means one of them should be merged or canonicalised. Read the matched passages rather than the headline number: 40% made of properly cited quotes is fine, 10% made of one lifted paragraph is not.
Semantic similarity measures whether two texts mean the same thing, regardless of the words used. "Cut costs" and "reduced expenditure" share no wording but carry the same claim, and a semantic comparison scores that as a match. It is normally computed by turning each text into a vector with a language model and measuring the angle between them. Lexical methods cannot do this, which is why heavy paraphrasing slips past them.
The common method breaks both texts into overlapping runs of words, called shingles, and counts how many runs they share. A second pass pairs each sentence with its closest match on the other side and scores the pair on shared vocabulary, usually with TF-IDF weighting and cosine similarity so that common words count for less. The two numbers answer different questions: verbatim overlap and reworded overlap.
Light paraphrasing, yes. Swapping a few words, reordering a clause or changing the tense leaves enough shared vocabulary for a sentence to still register as a match. A genuine rewrite in different words will not be caught by any method that compares wording alone, because there is nothing left to compare. Detecting that needs a semantic model, and even then the result is a judgement rather than proof.
Plagiarism is using someone else's work without credit, which is an academic and ethical question. Duplicate content is the same text existing at two addresses, which is a technical and SEO question. Self-duplication is common and is not plagiarism. A checker can measure overlap; only a person can decide whether that overlap was dishonest.
Not as a penalty. Search engines filter rather than punish: when two pages say the same thing, one is chosen for the results and the other is dropped, so the duplicate quietly stops earning traffic. The real costs are crawl budget spent on pages that will never rank and link signal split between two URLs that should have been one. The fix is a canonical tag, a redirect or a merge.
If two of your own pages came back heavily similar, the next questions are which one Google has picked, whether the canonical tags agree with you and how many other pages are in the same state. The free SEO audit tool crawls the site and answers all three, including duplicate titles and descriptions across every page it reads. The broken link checker covers the other half of a content clean-up: the links that stopped working while nobody was looking.
All three are pieces of Meelu, a desktop app where an AI marketing agent runs your marketing on your own machine — no word cap, no rate limit, and nothing leaves the laptop. Join the waitlist.