Updated July 2026
The claimed numbers vs. the tested numbers
Every detector publishes an impressive accuracy figure, and every one of those figures comes from the vendor’s own testing, on text conditions the vendor chose. Independent testing consistently lands lower. A few reference points worth keeping in mind:
| Data point | What it showed |
|---|---|
| Turnitin’s own target | Under 1% false positives at the document level, with roughly 85% of unmodified AI text caught |
| Independent real-world reviews | False-positive estimates of several percent, and much worse in small adversarial tests run by journalists |
| Multi-tool comparisons | Average accuracy around 60% across detectors; the best premium tools reached the mid-80s on unedited AI text |
| OpenAI’s own detector | Discontinued in 2023 for low accuracy — the maker of ChatGPT couldn’t reliably detect ChatGPT |
That last row is the honest headline of this entire topic. If any organization had the data advantage to build a reliable AI detector, it was OpenAI — and they shut theirs down rather than keep publishing unreliable scores.
How AI detectors actually work (and why that caps their accuracy)
AI detectors don’t “recognize” ChatGPT the way antivirus recognizes a known file. They’re classifiers trained to score how statistically typical a passage is of machine-generated prose. Two signals dominate:
- Predictability. Language models tend to pick the statistically likely next word. Text where every word choice is the expected one reads as machine-like to a detector.
- Uniformity. Human writing varies — sentence lengths lurch, tone drifts, small oddities accumulate. AI output is smooth. Detectors measure that smoothness.
Both signals are real, and both are shared by plenty of human writing: formulaic lab reports, careful writing by non-native English speakers, heavily templated business prose, and text polished by grammar tools. That overlap is the structural reason false positives can’t be engineered away — the detector isn’t detecting AI, it’s detecting AI-typical statistics, and some humans write that way.
The false-positive problem is not evenly distributed
The most important accuracy finding of the last few years isn’t the average error rate — it’s who absorbs it. Published research found that detectors flag writing by non-native English speakers at several times the rate of native speakers, because careful, simpler, rule-following prose pattern-matches what models produce. Other reliable false-positive magnets: very short texts, formulaic genres (five-paragraph essays, standard intros), and text that’s been through aggressive grammar polishing.
This is why several universities publicly stepped back from AI detection — most famously Vanderbilt, which disabled Turnitin’s AI detector in 2023 — and why nearly every vendor now ships a disclaimer that scores shouldn’t be the sole basis for academic action. The tools are lopsided: an 85% catch rate sounds strong until you realize the cost of each false alarm is an integrity accusation against a student who did the work.
What reliably breaks detection
Accuracy numbers are measured on clean conditions that barely exist in the wild. In practice, detection degrades fast when:
- The text is edited. Even moderate human revision of AI output shifts the statistics detectors rely on. Hybrid human-AI text is the hardest case and increasingly the most common one.
- The text is paraphrased. Detectors have added layers specifically trained on paraphraser output — Turnitin added one in 2024 — but vendors themselves acknowledge rewritten AI remains the weak spot.
- The model is new. Detectors train on the output of existing models, so each new generation of writing models temporarily outruns them. The race structurally favors the generator.
- The sample is short. Most tools need a few hundred words before their statistics stabilize; scores on short answers are close to noise.
“AI plagiarism” is the wrong frame — and the confusion matters
A lot of accuracy confusion comes from mixing up two completely different checks:
- Plagiarism checkers (the classic Turnitin similarity report, SafeAssign, Google Classroom’s originality reports) compare your text against existing sources — the web, journals, other students’ papers. They’re mature, precise, and largely uncontroversial: a match either exists at a URL or it doesn’t.
- AI detectors classify the writing style itself, with no source to point to. There’s nothing to “look up” — just a statistical opinion about how the prose reads.
AI-generated text sails through the first check, because it’s newly generated — it matches nothing. It only trips the second. That’s why a paper can be simultaneously “0% plagiarized” and “92% AI,” and why the two scores carry completely different evidentiary weight. A similarity match is checkable evidence; an AI score is an unverifiable estimate. Any school policy that treats the two numbers as equally solid is misreading its own tools.
The detector landscape, honestly sketched
Tool rankings churn constantly, but the categories are stable:
- Turnitin — the institutional default, wired into Canvas, Moodle, and Blackboard. Its accuracy story is covered in depth in our Turnitin guide: strong on unedited AI text, weaker on rewritten text, instructor-facing only.
- GPTZero — the best-known education-marketed standalone. Popular with individual teachers precisely because it’s free to start; scores frequently disagree with Turnitin’s on the same document, which tells you something about both.
- Copyleaks and similar API-first vendors — sold to institutions and platforms, often powering “AI detection” features inside other products. Same statistical approach, same structural limits.
- Grammarly and the authorship-tracking approach — a newer direction that sidesteps detection entirely: instead of guessing after the fact, it records how a document was written (typed, pasted, AI-assisted) as it happens. Process evidence beats statistical guessing, which is exactly why this direction is growing.
- Free web checkers — wildly variable quality, no accountability, and the source of most “I checked it and it was clean!” surprises. Useful for curiosity, not decisions.
How to sanity-check any AI score
Whether you’re a student disputing a flag or a teacher deciding whether to trust one, the same three tests expose how soft these numbers are:
- Run the same text through two or three detectors. Disagreement is the norm, not the exception — and if three tools give three verdicts, none of them is “the” answer.
- Look at what got highlighted. Detectors flag segments. If the flagged passages are the generic connective tissue (intros, summaries, definitions) while the substantive argument reads as human, that’s the classic false-positive signature.
- Check the length. Scores on anything under a few hundred words are statistical noise, and every vendor’s own documentation says so — including the fine print most policies never read.
If you’re a student staring at a false flag
The accuracy picture above is your defense material, but process evidence is what wins: version history in Google Docs or Word showing the paper being built over time, notes and sources, and your ability to discuss your own argument. Detector limitations are documented enough that a calm, evidence-backed response usually holds. We walk through the full playbook — including the non-native-speaker research worth citing — in our Turnitin guide.
If you’re a teacher deciding how much to trust a score
The research supports using detectors the way vendors now describe them: as one screening signal that prompts a closer look, weighed alongside things software can’t fake — the student’s writing history, draft evidence, source quality, and a short conversation about the work. A score plus corroborating signals is a reasonable basis for a process. A score alone, at documented error rates, is not — and policies built on scores alone tend to collapse on exactly the students least equipped to fight back.
Frequently asked questions
What is the most accurate AI detector?
Rankings shift with every model release, and every vendor’s own study crowns itself. In independent comparisons, the best premium tools have scored in the mid-80s percent range on unedited AI text, with the average tool closer to 60%. Treat any single accuracy number as a snapshot, not a spec.
Can an AI detector prove I used ChatGPT?
No. Detectors output a probability, not evidence. They can’t identify which tool wrote the text, can’t link it to an account, and are wrong often enough that vendors themselves — including Turnitin — say scores shouldn’t be the sole basis for punishing a student.
Do free AI checkers give the same result as Turnitin?
No. Different detectors use different models and training data, so the same essay routinely scores 5% AI on one tool and 60% on another. A clean result on a free checker doesn’t predict Turnitin’s score, and a flag on one doesn’t mean the others will agree.
Does using Grammarly make my writing look AI-generated?
Grammar and spelling corrections generally don’t. Heavy use of generative rewriting features can — because at that point portions of the text genuinely are machine-generated. Detectors also sometimes flag very polished, uniform prose regardless of how it got that way.
Will AI detectors get more accurate?
They improve, but they’re structurally behind: detectors are trained on the output of models that already exist, so every new writing model resets part of the race. The realistic future is detectors as one screening signal among several — not a reliable verdict machine.
Is using AI to write actually plagiarism?
Technically no — plagiarism is presenting someone else’s existing work as yours, and AI output is newly generated text. That’s why AI essays pass similarity checkers with clean scores. Most schools now treat undisclosed AI use as a separate academic-integrity violation (unauthorized assistance), which is why detectors and plagiarism checkers are two different reports.
Do AI detectors work on code, images, or other languages?
Barely, or not at all. Text detectors are trained mostly on English prose — support beyond a few major languages is thin, code detection is a different and less mature problem, and AI-image detection is a separate tool category entirely. Scores outside a detector’s training lane are the least trustworthy of all.
Josh Hutcheson — Editor, PriorityLearn
Josh researches, writes, and updates the answers on PriorityLearn, checking each one against current tools, official sources, and real school policies — and flagging what varies by state or district. About PriorityLearn →
