Quillbot added its own AI detector alongside its long-standing paraphraser, and markets it with a specific number: a 99% detection rate, credited to RAID, an independent academic benchmark. That's a real benchmark and a real citation — but the headline figure leaves out some details that change how much weight it should carry.
Trusted by students at
Quillbot's AI detector product page states it can "detect AI text reliably with a 99% detection rate, according to independent evaluations from RAID," and describes RAID as covering output from ChatGPT, GPT-4, GPT-5, Claude, Gemini, Mistral, Llama, and others. RAID is a real, peer-reviewed benchmark (published at ACL 2024) built specifically to stress-test AI-text detectors, so citing it is a meaningfully stronger claim than an internal, self-reported number alone.
What the marketing page doesn't state is which operating point that 99% figure comes from. RAID reports detector accuracy at different false-positive rate settings — for instance a 5%-false-positive setting versus a stricter 1% setting — and those produce different headline numbers for the same detector. Without knowing which setting Quillbot is quoting, the 99% figure is hard to compare against any other detector's published numbers, including its own.
RAID's benchmark specifically includes adversarial cases — AI-generated text that's been paraphrased or lightly edited to evade detection — because that's a realistic thing detectors need to handle, not just clean, unedited AI output. Independent analysis of Quillbot's public RAID submission data found meaningfully lower accuracy against that paraphrased-text category than against plain AI-generated text, a gap the 99% headline figure doesn't surface.
That distinction matters in practice: text that started as AI output and was then run through a paraphraser (Quillbot's own or another one) is a different, harder detection case than freshly generated AI text, and it's exactly the case a lot of real submissions fall into.
To Quillbot's credit, its own product page pairs the 99% claim with real caveats: "no AI Detector can provide 100% accuracy," results are "probability estimates, not verdicts," and it explicitly advises against relying on AI detection alone "to make decisions that could impact someone's career or academic standing." It also flags that heavily paraphrased or lightly AI-edited content is harder for any checker to catch, and that formulaic human writing — academic definitions, legal boilerplate, structured templates — can score higher than expected.
Those disclaimers are the more useful part of the page, honestly, and they apply to essentially every detector in this category, not just Quillbot's.
Check your writing against Turnitin, GPTZero, Originality.ai, Copyleaks, and ZeroGPT methods at once, free, and see where they agree.
Compare Against 5 Detectors Free →