Review · AI detector

GPTZero Review 2026: The Most-Used AI Detector, and the Most Divisive

GPTZero is the AI detector most teachers reach for, the one most students fear, and since June 2026 a subsidiary of the company that also sells an AI humanizer. It has the best free tier in the category and a 2.2 Trustpilot score. Both of those facts are true at once, and this review explains how.

By Jordan Hale Updated October 6, 2026 Pricing checked October 6, 2026

Meln is reader-supported. Some links on this page are affiliate links and we may earn a commission at no extra cost to you. Our rankings are not for sale. Read our disclosure.

3.7out of 5

The right detector for educators, the wrong tool for a verdict. GPTZero's sentence-level highlighting, LMS integrations, compliance certifications and 10,000-word free tier make it the most practical choice for teaching. Its accuracy swings wildly between independent tests, its false positives have real victims, and its billing record is poor. Use it as evidence, never as proof.

Best for
Educators, students checking their own work, publishers on a budget
Skip if
You need audited accuracy, team seats at scale, or paraphrase detection first
Price
Free 10,000 words/mo; $14.99/mo or $8.33 annual (150K words)
Owner
Superhuman (formerly Grammarly), since June 2026
Check the free tierAffiliate link.
Detection on plain AI text
4.0
False-positive handling
2.5
Features and integrations
4.5
Free tier
4.7
Billing and support
2.2
Value on paid plans
3.5

Pros and cons at a glance

What it gets right

  • Most generous credible free tier: 10,000 words a month
  • Sentence-level highlighting shows what was flagged and why
  • Writing Replay reconstructs how a document was composed
  • Chrome extension, Google Docs, Canvas and LMS integrations
  • SOC 2 Type II, GDPR, CCPA, FERPA compliance
  • Financially secure after the Superhuman acquisition

What it gets wrong

  • Independent accuracy results range from 52 to 95 percent
  • False positives on formal and non-native English writing
  • Trustpilot 2.2, with billing and cancellation complaints
  • Company does not respond to negative reviews
  • New parent company sells an AI humanizer
  • Pricing page hides dollar figures behind a toggle

What GPTZero is and how it detects

GPTZero was built in January 2023 by Edward Tian, then a Princeton senior, and Alex Cui, and it was the first detector most of the public heard of. It raised $3.5 million in 2023 and a $10 million Series A in 2024, was profitable with about 30 staff by the end of that year, and reported roughly $30 million in annual revenue and 20 million registered users by the time Superhuman acquired it in June 2026.

A printed essay with alternating sentences highlighted in two colours
Sentence highlighting is the feature that makes a detector discussable rather than just accusatory.

Under the hood it scores text on perplexity (how predictable each next word is to a language model) and burstiness (how much sentence length and structure vary), then layers newer classifiers on top, including what it calls a paraphraser shield. The output is a document-level probability plus per-sentence highlighting, which is the feature that makes it useful in a classroom: you can see which sentences drove the score instead of staring at a single percentage.

Where the numbers in this review come from

We did not run a private benchmark and we think you should be sceptical of reviewers who claim to have, because a few hundred samples tells you about those samples. We gathered every independent accuracy figure we could trace to a source, read the two major review aggregators and their complaint themes, checked the pricing page and the affiliate terms, and read the acquisition coverage. Where GPTZero's own numbers are the only ones available we label them as the vendor's.

Accuracy: lab benchmarks versus independent results

GPTZero's marketing cites a Chicago Booth benchmark at 99.5 percent and the RAID academic benchmark at 95.7 percent recall. Those are real results on those datasets. Scribbr's independent comparison, which used 30 texts across six categories and published its scoring sheet, put GPTZero at 52 percent overall. A 2,400-sample third-party test landed at 87. A reviewer at ToolsForHumans gave it 2.5 out of 5 and warned against standalone use after seeing it disagree with other detectors by up to 100 points on the same passage.

If the spread in that table bothers you, good. It bothered us. The 52 and the 95 are both real measurements of the same product on different text. That is the whole story of AI detection in one row.

Those numbers are not contradictory so much as they are measuring different things. Long, unedited output from a mainstream model is easy and GPTZero catches it. Short passages, paraphrased text, mixed human-and-AI documents and formal human prose are hard, and that is where the 52 percent comes from. If your use case is the first, GPTZero is excellent. If it is the second, no detector is reliable and GPTZero is not an exception.

SourceResultWho ran it
Chicago Booth benchmark99.5% accuracyCited by GPTZero
RAID (ACL 2024)95.7% recallAcademic benchmark
Third-party 2,400-sample test87% accuracyIndependent reviewer
Scribbr comparison52% overallIndependent (Scribbr sells a competing detector)
False-positive range1% to 12%Across independent tests

False positives and non-native English writers

This is the part of the review that matters most and the reason the Trustpilot page reads the way it does. Detectors flag writing that is statistically predictable, and predictable is what formal, careful, grammatically conservative prose looks like. A 2023 study in Patterns found that detectors flagged 61.3 percent of essays by non-native English speakers as AI-generated. GPTZero was one of several tools in that class of problem. The company has worked on it since and publishes its own false-positive figures, but independent tests still range from 1 to 12 percent, and 12 percent of a class of 200 students is 24 people wrongly accused.

An open bilingual dictionary on a desk with a pen
The 2023 Patterns study found 61 percent of non-native English essays wrongly flagged. Formal, careful prose looks machine-like to a classifier.

GPTZero's own guidance says its results should not be the sole basis for a decision. We would go further: if your institution uses GPTZero as the trigger for an academic misconduct process without a human reading the work, that is a process failure, not a software one.

Pricing, word limits and the billing complaints

The pricing page did not show dollar figures when we loaded it, only an "annual, save 45 percent" toggle and tier names, so the figures below are from a third-party tracker that last verified them in late September 2026. Confirm on the live page.

A stack of unopened envelopes on a desk
Trustpilot's recurring theme is not accuracy. It is charges after cancellation and support that does not write back.
PlanMonthlyAnnual (per month)Words per monthNotes
Free$0$010,000About 5,000 characters per scan
Essential$14.99$8.33150,000Plagiarism, extension, Docs
Premium$23.99$12.99300,000Adds Writing Replay, team features
Professional$45.99$24.99500,000Adds API
Enterprise / EducationCustomCustomLMS integration, SSO

The prices are fair against Originality.ai and Copyleaks. The complaints are about what happens around them. Trustpilot's 140 reviews are 49 percent one-star and the recurring themes are charges after cancellation, trouble reaching support, usage limits on paid plans feeling lower than advertised, and the app lagging or crashing. The company does not reply to negative reviews, which is unusual and, we think, a mistake. We found no published refund policy.

Features

Sentence highlighting is the headline feature and it is the reason to prefer GPTZero over detectors that return a single number. Writing Replay reconstructs the editing history of a Google Doc so an instructor can see whether an essay was typed over three evenings or pasted in at once; it is the most useful anti-misconduct tool here and it does not depend on the classifier being right. The Chrome extension and Google Docs integration work well. Canvas and LMS integrations are the reason schools buy it. A hallucination and citation checker arrived in 2025 and is being folded into Superhuman's products. The API is on the Professional tier and up.

The Superhuman acquisition

On June 23, 2026, Superhuman (the company formerly called Grammarly) announced it had acquired GPTZero for an undisclosed sum; reporting suggested at least a tenfold return on the $13.5 million raised. GPTZero continues as a standalone product, and its detection and hallucination-checking are being integrated into Superhuman Go. Two implications. First, the product is not going anywhere, which matters for institutions signing multi-year contracts. Second, Superhuman also sells an AI humanizer, so the same group now owns a tool that rewrites AI text to evade detection and the detector most likely to be used against it. We do not think that is sinister. We do think it tells you how much confidence the industry itself has in detection as a permanent solution.

What users say

Trustpilot, 2.2 out of 5: false positives on legitimately written academic work, charges after cancellation, difficulty reaching support, lag and crashes, limits that feel low. G2, 4.3 out of 5 from 101 reviews: praise for sentence-level transparency and integrations, complaints about false positives and the free character limit. The two audiences are different (students versus paying educators and businesses) and both are telling the truth about their experience.

Who should use GPTZero

Educators: yes, as one input alongside Writing Replay and a conversation with the student, never as the trigger on its own. Students: yes, on the free tier, to see how your own writing scores before someone else runs it, and to keep a record. Publishers and agencies: it works, but Originality.ai's paraphrase detection, site scans and team features fit that job better; see our Originality.ai review. Developers: the API exists on Professional, but Sapling's usage-based API is cheaper for small volumes.

GPTZero vs Originality.ai vs Copyleaks vs Turnitin

GPTZeroOriginality.aiCopyleaksTurnitin
Free tier10,000 words/mo60 credits/moNot shownInstitutional only
Paid from$14.99/mo$14.95/mo$16.99/moVia institution
Independent accuracy52 to 95%76 to 92%; 96.7% paraphrased65 to 66%About 77% AI, 93% human
False positives1 to 12%4.8 to 5.7%1 to 2%Conservative
Best forEducatorsPublishersLow false positivesInstitutions
Trustpilot2.24.6n/an/a

If you have been flagged by GPTZero

Do not argue about the tool in the abstract; produce evidence. Open your Google Docs or Word version history and export it. Gather notes, outlines and sources. Ask which threshold your institution uses and whether a human has read the work. Point to the published false-positive range on this page and to the Patterns study if English is not your first language. Request an oral discussion of the essay, which most instructors will grant and which resolves most of these cases. And quote GPTZero's own position that its output should not be the sole basis for a decision.

Questions readers ask about GPTZero

How accurate is it?

Anywhere from 52 to 95 percent depending on who tested what. Strong on long unedited AI text, weak on short, paraphrased or formal human prose.

Is it free?

Yes, 10,000 words a month. Paid from $14.99, or $8.33 on annual.

Who owns it?

Superhuman, formerly Grammarly, since June 2026. They also sell a humanizer.

Why the terrible Trustpilot score?

Flagged students and billing complaints. G2's paying educators give it 4.3. Both are real.

Does it catch humanized text?

Partly. Its own test of Undetectable AI output still read 91 percent AI. Independent testers find it struggles with light edits.

It flagged me. What now?

Version history, drafts, ask the threshold, cite the false-positive range, request a human read. GPTZero says its score should not decide alone.

Verdict

GPTZero is the best detector for a classroom because of the free tier, the sentence highlighting, Writing Replay and the integrations, and it is a poor instrument for a verdict because its accuracy depends enormously on what you feed it and its false positives land on real people. The Superhuman acquisition secures its future and complicates its story. If you teach, use it, with a human in the loop. If you publish, look at Originality.ai first. If you have been flagged, read the section above and gather your drafts.

Check the free tierWe may earn a commission.