Why Arbitrary Similarity Scores Fail to Prove Plagiarism
Text-matching algorithms measure verbatim word overlap rather than intellectual theft, routinely flagging properly cited quotes and standard methods language. Relying on arbitrary percentage thresholds penalizes rigorous scholarship while failing to catch heavily paraphrased stolen ideas.
By Nabil Faris
In short
- Text-matching algorithms detect verbatim character strings, not intellectual theft, meaning properly cited quotes and standard methods language trigger the exact same algorithmic penalty as stolen text.
- Enforcing arbitrary similarity thresholds forces students to unnaturally paraphrase standard academic terminology merely to lower a mathematical score, actively degrading the clarity of their writing.
- Because the software cannot evaluate intent or conceptual overlap, a document can return a zero percent similarity score while consisting entirely of heavily reworded stolen ideas.
In this article
Inside a university disciplinary committee room, a student faces expulsion because a software dashboard flagged 32 percent of their thesis in bright orange. The panel stares at the number, treating the mathematical output as undeniable proof of academic misconduct. Yet the highlighted text consists entirely of properly formatted citations, standard methodological boilerplate, and a bibliography.
The algorithm functioned exactly as designed, successfully matching strings of text against a global database. The humans in the room, however, fundamentally misunderstood what that measurement actually means. Text-matching platforms like Turnitin and iThenticate are engineered to detect verbatim word overlap, not intellectual theft.
Plagiarism is the deliberate appropriation of another person's ideas, concepts, or expressions without proper attribution. It is a behavioral and ethical violation. Text-matching software, conversely, performs a purely statistical operation, identifying identical character sequences regardless of context, intent, or citation formatting.
"A Similarity Checker is a text-matching software that compares submitted text against a database and generates a similarity report," the Nanyang Technological University academic integrity policy states. "It does not check for plagiarism directly, but is used as a detector."
The Mechanics of Textual Overlap
When a document is uploaded to a platform like Turnitin, the system breaks the text into discrete strings and cross-references them against billions of web pages, academic journals, and previously submitted student papers. The resulting similarity index is simply the percentage of words that match an existing source.[3]
This mathematical output is entirely agnostic to academic conventions. If a researcher writes, "informed consent was obtained from all participants," the software will flag the sentence because thousands of other papers use that exact phrasing. There is only one correct way to describe many standard scientific procedures.
Properly quoted material triggers the exact same algorithmic response. A student who meticulously formats a block quote from a 2023 National Institutes of Health study, complete with quotation marks and a footnote, will see that entire paragraph highlighted in red. The software successfully found the match; it simply does not care about the quotation marks.[2]
Bibliographies and reference lists are virtually guaranteed to generate high similarity scores. A correctly formatted APA citation for a foundational textbook will perfectly match every other paper that has ever cited that same textbook, driving the overall percentage higher with every added source.[3]
The Trap of Arbitrary Thresholds
Despite explicit warnings from software developers, academic institutions routinely establish rigid similarity thresholds. Many universities and academic journals mandate that submissions must score below 15 to 20 percent to be accepted. Any paper exceeding this arbitrary cap is automatically flagged for review or outright rejected.[3]
This reliance on a hard numerical cutoff creates a severe structural penalty for rigorous scholarship. Because standard academic boilerplate and properly cited references trigger the algorithm, these legitimate elements consume a massive portion of the allowable threshold before any original analysis is even evaluated.[4]
According to HasteWire's 2025 analysis of detection systems, false positives occur in roughly 5 to 10 percent of all scans. When an institution enforces a strict 15 percent limit, a student whose bibliography and methods section account for 10 percent of the document has almost no margin for error remaining.[3]
"Given the persistent use of simplistic similarity score thresholds at some academic journals and educational institutions... a method is arguably needed that encourages examining the Similarity Reports," researchers noted in a 2023 study published by the National Institutes of Health. Evaluators routinely use the score as a proxy to avoid reading the detailed reports.[2]
Degrading the Quality of Writing
The enforcement of arbitrary percentage caps actively degrades the quality of academic writing. When students realize that standard terminology will inflate their similarity index, they begin actively avoiding the precise language required by their discipline.
Instead of using the universally understood phrase "two-tailed t-test," a panicked student might write "a dual-directional statistical evaluation of means" merely to evade the text-matching algorithm. This practice replaces clear, standardized communication with convoluted synonyms simply to satisfy a machine.
The anxiety surrounding these scores is profound. Students frequently submit multiple drafts to institutional portals, obsessively tweaking their prose to lower the number. The focus shifts entirely from mastering the subject matter to gaming a statistical threshold.
"A high similarity score from a plagiarism checker feels like an accusation, but it isn't proof of anything on its own," notes a 2026 report from PubLearnHub. "Understanding why false positives happen... matters more as more journals run automated checks before a human ever reads the manuscript."
The Illusion of the Zero Percent Score
If a high similarity score does not prove plagiarism, a low score is equally useless for proving originality. A document can return a zero percent similarity index while consisting entirely of stolen intellectual property.[3]
A sophisticated plagiarist does not copy and paste text verbatim. Instead, they take an existing paper, absorb its core arguments, structural flow, and unique conclusions, and rewrite the entire document using different vocabulary. Because the specific character strings no longer match, traditional text-matching software will declare the document entirely original.[1]
"For instance, two sentences expressing the same idea using different vocabulary yield semantically similar embeddings, even if surface similarity scores fail to flag them," researchers at the Universidad Evangélica del Paraguay observed in a September 2026 analysis of detection tools.[1]
This exposes the fundamental flaw in treating text-matching software as a definitive judge. The algorithm catches lazy students who copy and paste Wikipedia articles, but it is completely blind to the actual theft of novel ideas, which is the most severe form of academic misconduct.[1][4]
The Impact on Academic Writing
The reliance on automated similarity metrics disproportionately harms students who follow structural rules too closely. Scholars frequently rely on standardized academic phrasing and formulaic sentence structures to ensure their writing is professionally acceptable.
Because their prose adheres closely to established templates, text-matching algorithms flag their work at significantly higher rates. "A student may have used a lot of content from one specific source, or may have used too many direct quotations," notes the University of Namibia's academic guidance. "That would not count as plagiarism, just poor academic writing."
When a student is accused of plagiarism based solely on a 25 percent similarity score, the burden of proof is entirely inverted. The author is forced to defend their integrity by explaining that their matched text consists of common transitional phrases and standard definitions.
The psychological toll of defending oneself against an algorithm can derail a student's entire semester. "A 25% score sounds alarming until you see it's entirely methods boilerplate and correctly quoted material," the PubLearnHub guidance reminds authors, urging them to contest false flags.
Restoring Human Judgment
The solution is not to abandon text-matching software, but to demote it from a judge to a diagnostic tool. A similarity report should serve exclusively as a map, directing a human evaluator to specific passages that require closer inspection.[2][4]
If a report highlights a 30 percent overlap, the instructor must open the document and look at the colors. If the red highlights cover the bibliography and a series of properly formatted block quotes, the score is meaningless and the paper should be accepted without penalty.[3]
Conversely, if a 12 percent score consists of three paragraphs copied verbatim from an uncited blog post, the student has committed plagiarism, regardless of how low the overall number appears. The context of the match is the only metric that actually matters.[3]
Institutions must abolish hard percentage thresholds. A policy that accepts 19 percent similarity but triggers a disciplinary hearing at 21 percent is mathematically illiterate. It assumes that academic integrity can be quantified by a string-matching algorithm that cannot read quotation marks.[4]
Ultimately, proving intellectual theft requires intellectual effort. Software can scan a billion web pages in three seconds, but only a human reader can determine whether a student has genuinely engaged with the material or simply stolen someone else's thoughts.[4]
How we did this
- Method
- A derivation of the false-positive penalty inherent in text-matching algorithms by isolating the baseline percentage of unavoidable overlap in standard academic writing.
- What we found
- Because standard academic boilerplate and properly cited references consume up to half of an institution's allowable similarity threshold before any actual content is evaluated, students are mathematically penalized for adhering to rigid structural conventions, forcing them to actively degrade the clarity of their methods sections to remain under the arbitrary cap.
- What we worked from
- False positive rate in standard text scans: 5–10%
- Common institutional similarity threshold: 15–20% — PM Proofreading
- Limits of this analysis
- The exact baseline of unavoidable overlap varies heavily by academic discipline; STEM papers require more standardized methodological phrasing than humanities essays.
Analysis by camp
Academic Administrators
Rely on hard percentage thresholds to process thousands of submissions efficiently.
For university departments processing thousands of dissertations and assignments every semester, manually reading every similarity report is logistically impossible. Administrators frequently rely on arbitrary thresholds—such as a strict 20 percent cap—as a scalable triage mechanism. They view the software as an objective, impartial standard that removes human bias from the initial screening process, allowing them to flag potentially problematic papers automatically.
Pedagogical Researchers
Argue that plagiarism is a behavioral issue of intellectual theft, not a mathematical overlap.
Academic integrity experts stress that intellectual theft requires intent and the uncredited appropriation of ideas, neither of which a string-matching algorithm can detect. They argue that using a similarity score as a punitive threshold fundamentally misunderstands the technology. Instead, researchers advocate for using the software strictly as a diagnostic mapping tool, where the percentage is ignored entirely and human evaluators focus only on the context of the highlighted passages.
International Students
Experience high anxiety over false positives due to their reliance on standardized academic phrasing.
Non-native English speakers and international scholars frequently rely on formulaic sentence structures and standardized transitional phrases to ensure their writing meets professional academic standards. Because their prose adheres so closely to established templates, text-matching algorithms flag their work at significantly higher rates than native speakers. This creates a structural disadvantage, forcing these students to spend excessive time defending legitimate methodological descriptions against automated accusations of misconduct.
- Pedagogical Researchers
- Argue that plagiarism is a behavioral issue of intellectual theft, not a mathematical overlap, and that software should be strictly limited to a diagnostic mapping tool.
- Academic Administrators
- Rely on hard percentage thresholds to process thousands of submissions efficiently, viewing the similarity score as an objective, scalable proxy for academic integrity.
- Students and Authors
- Experience high anxiety over false positives, as their reliance on standardized academic phrasing triggers text-matching algorithms at significantly higher rates.
Perspectives this story doesn't cover
- Software Developers
Sources
[1]Universidad Evangélica del ParaguayPedagogical ResearchersTools and Techniques for Detecting Plagiarism
Read on Universidad Evangélica del Paraguay →
[2]National Institutes of HealthPedagogical ResearchersA four-band method for evaluating similarity scores
Read on National Institutes of Health →
[3]PM ProofreadingUnderstanding the Turnitin Similarity Report
Read on PM Proofreading →
[4]Factlen Editorial TeamPedagogical ResearchersSynthesis by Factlen editorial team
Read on Factlen Editorial Team →
More in Education
See all →Cognitive Development
The Formal Operational Stage: How the Capacity for Abstract Thought Defines the Final Level of Piaget's Cognitive Development
8 sources
Global Education Funding
Global Leaders Pledge $3.5 Billion at UN General Assembly to Address Education Crisis
5 sources
Assessment Science
Content, Criterion, and Construct: How the Three Types of Validity Define the Quality of an Educational Assessment
7 sources
Science of Reading
The Five Pillars of Reading Instruction: How Phonemic Awareness, Phonics, Fluency, Vocabulary, and Comprehension Define Literacy
6 sources
Comments
Every angle. Every day.
Get Education stories with full source coverage and perspective breakdowns, free every day.




