Pangram Labs is a newer entrant in AI detection, and it's drawn notably more positive independent coverage than most detectors in this category — including a mention in Nature. Here's what Pangram claims about its own accuracy, what outside coverage actually says, and where the two lines start to blur.
Trusted by students at
Pangram Labs publishes a specific false-positive figure on its own site: approximately 1 in 10,000 overall, and as low as 0.004% specifically on academic essays. That's a notably more precise, and more favorable, number than most detectors in this category publish — many competitors either don't publish a false-positive rate at all, or publish one in the range of several percent.
Like every self-reported figure in this category, it comes from Pangram's own testing methodology, not an outside audit. The company was co-founded by Max Spero and Bradley Emi, who according to published interviews trained the detection model on a broad range of writing rather than academic text alone, using text predating ChatGPT's release as a clean human-written baseline.
A 2026 Nature article covering AI-detection reliability in universities describes Pangram as claiming a near-zero false-positive rate and notes that "independent assessments have judged the product to be among the most accurate available" — but Nature's own reporting doesn't publish a separately verified percentage for Pangram specifically. It's citing the reputation Pangram has built, not running its own head-to-head test.
The same article cites a 2025 study putting GPTZero's false-positive rate at around 16% — useful context for how much variation exists across detectors in this category, though it isn't a direct Pangram-vs-GPTZero test, so it shouldn't be read as a scored comparison between the two.
You may see a widely-cited figure — that roughly 61% of pre-ChatGPT-era essays by Chinese students were wrongly flagged as AI-generated — attached to specific detectors online. That number actually comes from a 2023 Stanford study averaging results across seven different AI detectors on 91 essays, not a score for any single tool, and it predates Pangram's current model. It's evidence of a category-wide bias problem, not a review of one product, and shouldn't be pinned on Pangram or any one detector specifically.
The honest summary: Pangram's own published numbers are unusually strong, and its reputation among reviewers is unusually positive for this category, but there's no fully independent, apples-to-apples audit of Pangram we could verify — the same caveat that applies to nearly every detector on the market.
See how your writing scores against Turnitin, GPTZero, Originality.ai, Copyleaks, and ZeroGPT methods at once — free.
Compare Against 5 Detectors Free →