The verified data

Are AI Detectors Accurate? The Verified Data on False Positives.

The short answer

No AI detector is accurate enough to prove cheating on its own. Independent academic studies report false-positive rates of 5 to 20 percent. Stanford researchers found a 61.3 percent false-positive rate on essays by non-native English writers. Turnitin's own documentation admits a 4 percent sentence-level false-positive rate. OpenAI shut down its own classifier in July 2023, citing "low rate of accuracy."

Top-down view of printed pages with simple charts and a fountain pen on a wooden desk.

This is the citation page. Every number below carries a primary source. If you are preparing an appeal or a conversation with an integrity panel, this is the evidence base. Cite generously.

How accurate are AI detectors? The published rates at a glance

Two kinds of number circulate in this debate, and they do not agree. Vendors publish rates measured on test sets they selected themselves. Independent academics publish rates measured on real student writing. The independent numbers are consistently higher — in one population, dramatically so.

Detector or study False-positive rate Sample Source type
Stanford HAI (Zou et al., 2023)61.3%91 TOEFL essays, 7 detectorsPeer-reviewed
Weber-Wulff et al. (2023, IJEI)5–20% generalMultiple commercial detectorsPeer-reviewed
Vanderbilt scale estimate~750 per 75,000 papersTurnitin, institution-widePrimary source
Turnitin (own documentation)<1% document / 4% sentenceVendor benchmarkVendor self-report
Copyleaks (marketed)0.02%Vendor benchmark, unverifiedVendor self-report
GPTZero (own guidance)No figure — "not perfectly accurate"Vendor statementVendor self-report
OpenAI AI ClassifierDiscontinued July 2023Primary source

Independent academic studies (high confidence)

· Stanford HAI, peer-reviewed Stanford HAI (Zou et al., 2023) — 61.3% on non-native English writers

Seven AI detectors evaluated 91 TOEFL essays (written by non-native English speakers) and 88 US 8th-grade essays (written by native English speakers). The result on TOEFL essays:

  • 61.3 percent average false-positive rate across the seven detectors.
  • ~20 percent of essays were unanimously misflagged by all seven detectors.
  • Native-English essays were almost never misclassified.

Source: Stanford HAI; cross-verification at The Markup.

· Peer-reviewed, IJEI Weber-Wulff et al. (2023) — "neither accurate nor reliable"

Weber-Wulff and colleagues, publishing in the International Journal for Educational Integrity in 2023, tested multiple commercial AI-text detectors and concluded the technology is "neither accurate nor reliable." General false-positive rates across the literature run 5 to 20 percent in independent samples — far above vendor self-reported figures.

· Primary source — Vanderbilt Vanderbilt scale estimate — 750 false positives per 75,000 papers

Vanderbilt's published reasoning for disabling Turnitin's AI detector (August 16, 2023):

"Vanderbilt submitted 75,000 papers to Turnitin in 2022. If this AI detection tool was available then, around 750 student papers could have been incorrectly labeled as having some of it written by AI."

Source: Vanderbilt Brightspace.

Detector by detector — what each vendor admits

· Vendor self-report — Turnitin Turnitin AI detection accuracy — <1% document / 4% sentence

From Turnitin's own published blog posts:

  • Document-level false positive rate: less than 1 percent (for documents flagged 20% or more AI).
  • Sentence-level false positive rate: about 4 percent.

Turnitin's chief product officer has stated publicly that the tool is "advisory" — faculty decide whether to act on a flag. Sources: Turnitin — document level; Turnitin — sentence level.

Note: Turnitin admitted in June 2023 — two months after launch — that real-world false-positive rates were higher than the initial public claims. (K-12 Dive) If Turnitin is the only basis of your accusation, the Turnitin-specific page covers what to do next.

· Vendor self-report — GPTZero GPTZero accuracy — "not perfectly accurate," in its own words

GPTZero's own published guidance:

"No AI detector is perfectly accurate; detectors are especially error-prone on short, edited, or mixed (human+AI) writing, and results should be treated as clues paired with human judgment rather than definitive proof."

Source: GPTZero blog. GPTZero publishes no independently verified false-positive rate.

· Vendor self-report — Copyleaks Copyleaks accuracy — a marketed 0.02%

Copyleaks markets a 0.02 percent false-positive rate. This figure is vendor-self-reported and has not been independently verified. Even at that rate, in a 20,000-student university the absolute number of false accusations per year is on the order of dozens.

· Vendor self-report — Originality.ai Originality.ai accuracy — high claims, no independent validation

Originality.ai publishes high accuracy figures. Independent academic studies have not validated them.

· Primary source — OpenAI OpenAI — shut down its own classifier (July 2023)

OpenAI launched its AI Classifier in January 2023 and announced its discontinuation in July 2023, citing "low rate of accuracy." The company that builds ChatGPT could not reliably detect ChatGPT's output.

The scale arithmetic — why 1% is a lot

A 1 percent false-positive rate sounds reassuring. Applied at scale, it produces:

  • Vanderbilt 2022: 75,000 papers × 1% = ~750 students wrongly flagged.
  • A mid-size US R1 university: ~50,000 papers per year × 1% = ~500 students per year.
  • A large public state system (e.g., UC system): hundreds of thousands of papers × 1% = thousands of false accusations annually.

Every one of those numbers is a real student opening a real email at 11pm. Argued at the population level, the percentage rate is the point. Argued at the individual level — your level — the percentage is the floor, not the ceiling, of harm.

The "I Have a Dream" test

Multiple educators have demonstrated that AI detectors will flag famous human-written texts as AI-generated. The Martin Luther King Jr. "I Have a Dream" speech (1963) and passages from the Bible have been flagged by AI detectors at high "AI probability" scores. This is not a stress test of a fringe edge case — it is a demonstration that the underlying signal (low perplexity, predictability) is what these detectors see, and famously well-written prose has those qualities.

Do AI detectors actually work? The science in two paragraphs

AI detectors look at two main signals: perplexity (how surprising the next word is, given the previous words — low perplexity means predictable writing) and burstiness (variation in sentence length and complexity within a document). They flag writing as "AI" when perplexity is low and burstiness is low. They cannot directly observe whether a human or a machine wrote the text. They observe statistical patterns and infer.

ChatGPT produces low-perplexity, low-burstiness writing because it picks the most likely next word at every step. But so does any writer who uses common words and consistent sentence structure: non-native English speakers (smaller vocabulary), formal academic writers (consistent register), people who write under stress (less linguistic variety), writers with certain disabilities (formal, structured prose). The detectors cannot tell these populations apart from ChatGPT. That is the structural reason they fail, and it is the reason that no algorithmic fix will solve the bias.

What this means for your case

You are not arguing against a reliable instrument. You are arguing against a flawed one with documented bias against writers who happen to share statistical properties with ChatGPT. Cite the numbers above. Reference Vanderbilt's decision and OpenAI's shutdown. If English is not your first language, reference the Stanford study explicitly. The procedural-concerns section (Section III in the four-element appeal letter) is where these numbers belong — alongside your own writing-process evidence, which is what actually clears students.

So: are AI detectors accurate? Not accurately enough for any single flag to stand as proof, and every vendor on this page says a version of that in its own documentation. That sentence, with the sources attached, is the spine of your defense.

Frequently asked

How accurate are AI detectors?

Not accurate enough to prove cheating on their own.

No AI detector is accurate enough to prove cheating on its own. Independent academic studies report false-positive rates of 5 to 20 percent. Stanford researchers (Zou et al., 2023) found a 61.3 percent false-positive rate on TOEFL essays written by non-native English speakers. Turnitin's own documentation admits about 4 percent at the sentence level. OpenAI discontinued its own AI Classifier in July 2023, citing "low rate of accuracy."

What's the average AI-detector false-positive rate?

5 to 20 percent in independent studies. 61.3 percent on essays by non-native English writers.

Independent academic studies report false-positive rates of 5 to 20 percent across detectors. Weber-Wulff et al. (2023, peer-reviewed) reached a peer-reviewed conclusion that AI text detectors are 'neither accurate nor reliable.' Vendor self-reported rates are lower but have not been independently verified. The Stanford Zou et al. (2023) study on non-native English writers found a 61.3 percent false-positive rate.

Are AI detectors reliable enough to prove cheating?

No — the peer-reviewed verdict is "neither accurate nor reliable."

Weber-Wulff et al. (2023), publishing in the International Journal for Educational Integrity, tested multiple commercial detectors and concluded that AI text detectors are "neither accurate nor reliable." The vendors say much the same thing in their own documentation: GPTZero's published guidance states that detector results should be treated as clues paired with human judgment "rather than definitive proof," and Turnitin's chief product officer has described the tool publicly as advisory — faculty decide whether to act on a flag. A tool its own makers call advisory is not a tool that proves misconduct.

Why are stated false-positive rates different from real-world rates?

Vendor benchmarks vs. real student populations.

Vendor rates come from internal benchmarks on test sets curated by the vendor. Real-world rates emerge from diverse student populations — ESL students, students with disabilities, students writing in technical or formal disciplines — that vendor test sets often under-represent. Independent academic studies sampling real student writing repeatedly find higher rates than vendor self-reports.

Why did Vanderbilt disable Turnitin's AI detector?

Lack of transparency, bias, and an estimated 750 false positives per 75,000 papers.

On August 16, 2023, Vanderbilt University disabled Turnitin's AI detection tool. Their statement cites lack of transparency (Turnitin would not explain how decisions were made), documented bias against non-native English writers, and the scale arithmetic: 'Vanderbilt submitted 75,000 papers to Turnitin in 2022. If this AI detection tool was available then, around 750 student papers could have been incorrectly labeled as having some of it written by AI.' The official statement is on the Vanderbilt Brightspace Support site.

Are AI detectors biased against non-native English speakers?

Yes. Stanford found 61.3 percent false positives on TOEFL essays.

Stanford researchers (Zou et al., 2023, Stanford HAI) had seven popular AI detectors evaluate essays written by non-native English speakers (TOEFL exam essays) and native English speakers (US 8th-grade essays). The detectors flagged the TOEFL essays as AI-generated 61.3 percent of the time on average. On about 20 percent of TOEFL papers, the misclassification was unanimous across all seven detectors. The native-English US essays were almost never misflagged. The mechanism: both non-native writers and ChatGPT use simpler vocabulary and shorter sentences, which is what the detectors learn to flag.

Can two detectors agreeing prove I used AI?

No — they share the same blind spot.

Detectors are not independent. They all rely on the same general signals (perplexity, burstiness, predictability of word choice). When two detectors both flag a piece of writing, they are agreeing on the same signal — not independently confirming the verdict. If your writing is simpler and more predictable (because English is your second language, because you write formally due to a disability, because you write in a discipline that favors clear prose), every detector will flag you. That's correlation in their bias, not corroboration of the truth.

Why did OpenAI shut down its own AI detector?

Low accuracy. The maker of ChatGPT could not detect ChatGPT.

In July 2023, OpenAI announced it was discontinuing the AI Classifier it had launched only six months earlier. The official reason: 'low rate of accuracy.' The makers of ChatGPT could not reliably distinguish ChatGPT's output from human writing. This is one of the strongest pieces of evidence you can cite — the people with the most knowledge of how the model works could not build a reliable detector for it.

Can Copyleaks's claimed 0.02% rate be trusted?

It's vendor self-reported. Even if true, the absolute harm is material.

The 0.02 percent figure has not been independently verified. Independent academic studies (Weber-Wulff et al., 2023) report general false-positive rates of 5 to 20 percent across detectors. Even taking 0.02 percent at face value, in a 20,000-student university with four courses and five assignments each per term, the expected number of false accusations is on the order of dozens per year — and each of those is a real student in real distress.

Why is the absolute number of false positives so high even at 1%?

Tens of thousands of papers × small percentage = hundreds of students.

Vanderbilt's published scale arithmetic makes this concrete: 75,000 papers submitted in 2022 × 1 percent false-positive rate = about 750 students wrongly flagged. This was Vanderbilt's primary reason for disabling Turnitin's AI detector on August 16, 2023. Percentage numbers can sound reassuring; absolute numbers tell the real story.