Issue #195

AI Detectors Flagged the Declaration of Independence as AI

A Stanford study found these tools misclassify non-native essays and even flag the 1776 Declaration of Independence as AI-written.

SocietyAI Detectors Flagged the Declaration of Independence as AI

The People Who Deliberately Made Their Essays “Less Perfect”

Lauren Jaeger, a chemistry major at Idaho State University, ran into something strange while applying to PhD programs. Some application portals warned that “if AI-generated text is detected in your personal statement, your entire application will be voided.” Jaeger had never used AI. But the warning made her nervous enough to run her own writing through a few online detectors — and every single one came back with “almost 100% AI.”

So she rewrote her personal statement. Not to make it better, but to make it deliberately worse — “less perfect” — so it wouldn’t trip the detectors. She kept revising until her AI score dropped to 30%, decided that was good enough, and submitted it. (She got into the University of Utah’s PhD program, for what it’s worth.)

Around the same time, something even more absurd was making the rounds online: someone ran the US Declaration of Independence, written in 1776, through a detector — and it came back “95-100% written by AI.” A 250-year-old document.

So how accurate are these tools that universities rely on to catch cheating? The bigger problem isn’t that detectors are occasionally wrong — it’s that a score from a tool that can be wrong is being treated as settled proof.

🔍 Why Universities Are Betting on Detectors

With generative AI tools like ChatGPT, the cost of having someone else write your text has effectively dropped to zero. Cass Ellis, who oversees academic integrity at Western Sydney University, calls this “a fundamental shift in scale.” Ghostwriting existed before, sure—but now, as she puts it, “the sheer volume of submissions that are at least substantially fabricated or manufactured has exploded.”

The problem is that the old weapons don’t work anymore. Plagiarism checkers like Turnitin work by cross-referencing text against a massive database of existing publications, hunting for traces of copying. But sentences generated by AI aren’t lifted wholesale from anywhere—they’re assembled fresh, on the spot, each time. With no original to compare against, there’s nothing for a plagiarism check to catch. AI builds its sentences by training on a vast body of past writing, but the output isn’t identical, sentence-for-sentence, to anything it learned from. That’s why plagiarism detection, which relies on matching against existing text, can’t catch it.

That’s the gap detectors claiming to identify “AI-written” text stepped into. Copyleaks, GPTZero, ZeroGPT, and tools built by Grammarly, QuillBot, and even Turnitin itself are all on the market now. Facing a flood of assignments and applications, universities and graduate schools have started leaning on these tools.

Why “well-written” text gets flagged as AI

Many detectors rely on a metric called perplexity1. Put simply, it measures “how predictable is the next word.” AI-written text tends to follow statistically smooth, predictable patterns. So the lower a text’s perplexity (the more predictable it is), the more it’s suspected of being “machine-written,” and the more jarring or unusual its phrasing, the more it’s judged to be “human-written.”

See the problem? By this logic, the more precisely someone follows grammar rules and writes cleanly without excess, the more likely they are to be suspected of using AI. This is exactly what Jaeger meant when he said, “I studied grammar books and wrote by strictly following the rules, which I think is why I was mistaken for AI.” You get caught for writing well.

The problem gets even more serious for people who aren’t native English speakers. In 2023, a Stanford research team ran 91 TOEFL essays—all written before ChatGPT existed, prior to 2020—through 7 detectors.

On average, 61.3% of these 91 essays were incorrectly classified as “AI-written.” Meanwhile, English writing by American 8th-graders (ages 13-14) was classified as “human-written” with near-perfect accuracy.

Non-native writers’ text tends to show less variety in expression, which produces lower perplexity scores—and detectors mistook this for being “AI-like.” A false-positive rate2 of 61% means more than half of this sample of essays, all genuinely written by humans, was wrongly judged to be AI-generated. That figure comes from a limited sample of 91 TOEFL essays.

The Declaration of Independence got flagged for the same reason. A text that’s been quoted and refined countless times over 250 years is, statistically speaking, extremely “predictable.” Nature ran a portion of the actual 1776 text through ZeroGPT multiple times, and every single time it came back 95-100% AI.

These misjudgments aren’t isolated flukes. A 2025 study tested GPTZero, reportedly the most widely used detector, and found that while it caught fully AI-generated text with high confidence, it mistook human-written text for AI roughly 16% of the time. The researchers’ conclusion: “reliability in distinguishing human-written text is limited.”

⚡ The Problem That Remains Even With Better Detectors

Not every detector is this bad, of course. Pangram Labs, based in New York, claims a false-positive rate of “close to zero.” Their approach is different: they train on massive datasets of human-written text alongside AI-rewritten versions of the same text, so the model learns each new chatbot’s style as it appears. They don’t lean on perplexity alone. Independent evaluations by the University of Chicago’s Booth School of Business and the University of Maryland found accuracy never dropped below 99.8%.

But detectors getting more accurate doesn’t make the problem go away.

First, there’s hybrid text. Mike Perkins (British University Vietnam), who studies AI’s impact on academia, says: “Detectors catch pure AI writing pretty well, but once you start touching that text up, detection falls apart.” A human doesn’t even need to do the editing. You can just ask another AI to “rewrite this,” or run it through a “humanizer3” tool built specifically to lower detection scores. Whenever detection companies try to catch these workaround tools, a new workaround shows up. In Perkins’s words, it’s “a massive arms race that helps nobody.”

Second, there’s a more fundamental problem. This year, Marzena Karpinska and her research team used Pangram to analyze 186,000 articles from 1,500 American newspapers published in the summer of 2025. About 9% were flagged as partially or entirely AI-written. But Karpinska herself was explicit that these results can be used to spot “large-scale trends” — not as grounds for pinning guilt on an individual writer.

“You can’t use this to fail people en masse.”

The same tool can be useful for reading overall trends while being unsuitable as grounds for judging an individual case of misconduct. A rate calculated at the group level doesn’t prove that one particular person’s writing was AI-generated.

Oswarld’s Lens

Honestly, I don’t read this story as a debate over “detector performance.” As someone who’s worked with data for years, what jumps out at me first is “threshold” and “base rate.”

There’s a trap I ran into constantly while doing data analysis. Even a model with 99.9% accuracy can wrongly implicate a lot of innocent people if the population is large enough and true positives are rare.

Say you have an excellent detector with a false-positive rate of just 0.1%. Apply it to 10,000 students who all wrote their essays honestly by hand, and about 10 of them get flagged as “AI cheaters” despite doing nothing wrong. For those 10 people, 99.9% accuracy is no comfort at all.

This is exactly the point Perkins raised. People see a score and just believe it. In the plagiarism-detection era, there was visible evidence — “this sentence here matches that paper word for word.” But AI-detection scores offer no such evidence. Neither the student nor the professor has any way to check what the number “87% AI” is actually based on. In my own work building GTM strategy, I’ve seen many times how easily “certainty dressed up as a number” can mislead an organization. A score is supposed to aid judgment, but it often ends up replacing it.

So here’s where I land. A detector can be used as an “investigative lead.” But the moment it’s used as “evidence of guilt,” the first people caught in the false-positive net are exactly the ones like Reijer who write with precise, correct grammar, and non-native speakers generally. This isn’t an argument against using the tool — it’s an argument for treating the detector’s score as a reference point, not a settled conclusion.

Closing

First, many AI detectors mistake grammatically precise writing and non-native prose for AI-generated text. In a Stanford study, an average of 61.3% of 91 TOEFL essays were wrongly flagged as AI-written. Second, even more accurate detectors still struggle to catch hybrid text that’s been touched up by a human, or writing that’s been run through a humanizer. And identifying trends across a group is a different problem from judging one individual’s misconduct. Third, the real danger isn’t that the tool can be wrong — it’s trusting a score that can be wrong as if it were proof.

If you’re in a position to evaluate or submit writing, just remember one thing: an AI detection score should prompt further checking, not serve as the basis for a conclusion on its own.

Jaeger deliberately rewrote his essay to be “less perfect” so he could get admitted. Think about that for a second — sabotaging your own writing on purpose just to dodge a detector. Have you ever actually experienced your own writing — or a student’s or colleague’s — being unfairly flagged as AI-generated? Tell me about it in the comments. I’ll consider it for a future issue.


💬 Tell me in the comments if you’ve ever been unfairly flagged as AI. I’ll use it for a future issue.

📨 If you think this could help a fellow writer, pass it along.

The English draft matches the Korean source accurately in meaning, structure, numbers, links, and glossary terms. No Hangul characters remain, all footnotes and headings align, and terminology follows the glossary. No corrections needed.

Your take shapes the next issue

What resonated most in this issue, or where has your experience been different?

Any registered reader can comment for free.

References & Further Reading

Primary sources

Background

  • Pangram Labs’ collection of third-party evaluations, “Pangram Evaluations”, “Chicago Booth Review”. : If you’re curious how the “near-zero false positive rate” claim was verified, check here.

Original article

Illustrated portrait of Kwangseob Ahn (Oswarld)

The author is Oswarld (Kwangseob Ahn). Current roles: Adjunct Professor at Sejong University, Strategy Consultant at INLEVEL9. Career, research, books, and recent work are kept current on the About page. Latest · July 2026: HEMA-2: A Consolidation-Aware Tri-Memory Architecture with Multi-Channel Scheduling for Lifelong Conversational AI.

Footnotes

  1. Perplexity: a measure of how predictable the next word in a sequence is. The lower it is, the more “obvious” the sentence seems — and the more AI-like it’s judged to be. The higher it is, the more it’s taken as human writing.

  2. False positive rate: the rate at which text actually written by a human is wrongly judged to be “AI-written.” Think of it as the rate at which innocent people get caught.

  3. Humanizer: a tool that rewrites AI-generated text to look human, lowering its detection score.