Substack shipped an AI detector this week. Every post, note, and comment over 100 words can now be scanned through Pangram to see how much of it reads as human or AI. Chris Best's launch post frames it as giving readers a choice: the platform isn't banning AI use, just surfacing it. Worth knowing going in: Pangram's own data already ranks Substack as the cleanest of the platforms it scans, a fraction of LinkedIn's AI-content rate. They're launching transparency tooling from the platform with the least to hide.
I read that and thought: I already know exactly how this goes. I lived it on a smaller scale, with a person instead of a platform integration.
A few months ago I got flagged twice in one day by "Sloan," DEV.to's moderation-warning system. Not a bot quietly scoring posts in the background. A specific community member, reading articles and running them through GPTZero, then sending the same message a blunt classifier would have sent.
The two pieces that got flagged were the ones that generated the most technical discussion I'd published all year. Short paragraphs. Named data points. Rhetorical questions doing real work. The features that make an argument land are the same features that read as "AI-shaped" to anyone calibrated to notice them, human or model.
Write worse, look more human. Write well, get flagged.
That thread also surfaced the part nobody had a clean answer for: the policy creates a dishonesty incentive. Two equally AI-assisted pieces, equally good — the one with a disclosure gets flagged, because now there's something to catch. The one without doesn't. The system was catching transparency, not AI use.
And then Marco showed up in the comments. Forty years in tech, writing in his second language, using AI to make sure his Italian didn't flatten into something stiffer than he meant. Same Sloan message. Same classifier verdict. Nothing to do with what the policy was built for.
Pangram is a real classifier with real engineering behind it: hard negative mining against its own false positives, training data deliberately mirrored so it can't just learn "formal writing = AI." That's more rigor than one guy running GPTZero between article reads. I'll give it that.
But it inherits the same structural problem Sloan had, because it's answering the same narrow question: does this text look AI-shaped. Not: did a human do the thinking. Chris Best's own post admits as much. Pangram can't tell you whether care went into something, only whether the sentences pattern-match to a machine's output.
That gap is where Marco lives. Detectors trained without deliberately mirrored data have a documented habit of flagging non-native English writing, since careful, formal phrasing correlates with both AI output and someone translating in their head before they type — enough of a problem that several major universities have stopped letting instructors use AI detectors at all. Pangram claims their mirror-prompt method fixes it. Maybe. Most of the numbers backing that claim trace back to Pangram or a study Pangram commissioned.
Someone with no stake in the answer already looked. The Atlantic's Matteo Wong traced a recent wave of AI-writing accusations back to Pangram itself, including a horror novel pulled from a major publisher days before its release. His argument wasn't that the tool is broken. It's that a detector that's mostly reliable can be more dangerous than one that's obviously unreliable, because people stop checking. A 99.98% accuracy rate sounds like certainty. Applied across millions of posts, the failures are still real people, still real reputations, just quieter about it.
That's Marco's risk, and mine, in one sentence: the false positive doesn't feel like a statistic when it's your byline.
I write from Port Harcourt, in English, the language I was taught in and think in, using AI as part of an actual workflow: not to generate opinions I don't have, but to get from a rough draft to a clean one without losing the argument along the way. Sloan already showed me what a false positive costs, close enough that I don't need to imagine it. Marco is the version of that risk I can't unsee.
It's also the whole reason I stopped using a generic humanizer and built one calibrated to my own published corpus instead. A tool trained to strip "AI-shaped" patterns from anyone's writing will also strip the parts of your writing that are just yours, an em dash you use structurally, a habit of compressing three examples into two. Voice-humanizer checks against what I actually sound like, not against a mirrored dataset of nobody in particular.
Sloan and Pangram are both answering "does this look like AI." I don't think that's the question that matters. The question is whether someone can be asked "did you know what you were writing about, and do you stand behind it," and answer yes.
I do.

Top comments (19)
I ran into the same issue: blog posts from 2017 that detectors claimed were AI-written with 97% certainty—based on sentence structure, phrasing, and likely other criteria—even though I had written them myself!
That said, I don't claim never to use AI; I use it primarily for formatting. French is my native language, so I do ask for rephrasing to make my English flow more smoothly—but this applies to articles I have already written, often spending anywhere from 6 to 20 hours on them before submission.
Pascal, the 2017 detail is the part that actually breaks the argument. A stylistic case, sentence rhythm, phrasing, can always be argued around. A publish date from before the technology existed can't. That's not a false positive with room for doubt, it's one with an alibi.
The 97% is doing more damage than a lower number would, too. Confident and wrong reads as authoritative in a way "uncertain and wrong" never does.
Did whichever tool flagged it change its read once you pointed out the year, or did the number just sit there unmoved?
Oh! The score was so absurd that I didn't even bother arguing with it. It mostly made me question how reliable these tools really are. I tried the same test with several other pieces I'd written, and every single one came back as "AI-generated." Then again... maybe I'm an AI after all. 😄
if every piece you've written comes back flagged, that's not really a false positive rate anymore. That's the tool deciding your writing style is what AI writing looks like. It's not making isolated mistakes on your work, it's using you as the reference class.🤣
That's both the funniest and the most disturbing explanation I've heard so far. 😄
If AI detectors have decided that my writing is what AI writing looks like, then we've reached a rather ironic point in history: I spent 35+ years learning to write clearly, and now clarity is apparently suspicious.
35 years is the detail that sticks. Clarity isn't incidental to what got flagged, it's the actual mechanism. A classifier trained on regularities reads consistency as your fingerprint and consistency as AI's fingerprint the same way, because to a statistical model those look identical.
That's a fascinating way to put it.
Maybe we've reached the point where "writing like an experienced human" and "writing like an AI" occupy the same statistical space. If that's true, then AI detectors aren't really detecting authorship anymore—they're detecting stylistic regularity.
"stylistic regularity" is the more honest name for what's actually being measured. Authorship implies the classifier knows something about origin. What it actually has is a distribution enough examples of both categories that it can tell you which cluster your sentence lands nearer to. That's a real signal. It's just not the signal the word "detector" implies it is.
Which means the tool works exactly as advertised until the moment a human writes with the same regularity a model does and then it stops being able to tell the difference in principle not just in this one case.
Exactly. The signal is real, but the interpretation is where things get complicated.
A detector can tell you "this looks statistically similar to AI output." It cannot tell you "an AI produced this." Those are two very different claims.
The irony is that experienced writers often train themselves toward the same qualities models are designed to reproduce: clarity, structure, and consistency.
This is the same failure mode from the devto AI-flag debate a few weeks back: the detector measures "does this look statistically consistent with model output," and consistency is exactly what careful writers optimize for. Marco's case is the cleanest proof, a non-native speaker using AI to clean up phrasing gets flagged the same as someone who wrote nothing themselves, because the tool has no way to tell polish apart from authorship.
The fix probably isn't a better classifier. It's moving the question from "does this look like AI" to "can the author defend what they wrote, line by line." Harder to automate, but it's the only thing that actually separates the two cases.
Mike, "can the author defend what they wrote" is the better question and it's also a heavier one to hand someone.
Right now the accused defends nothing until Pangram or Sloan, makes a claim. Move the standard to defend-it-line-by-line and defense becomes the default state every writer owes, whether or not anyone's accused them yet.
That's less an alternative to the detector than a description of what happens once cases like Marco's start piling up: authorship turns into something you have to be ready to prove, not something assumed until challenged.
Is there a version of this that doesn't just relocate the burden from the tool to the writer??
"Write worse, look more human. Write well, get flagged." — This paradox perfectly captures why probabilistic AI detectors are fundamentally broken.
As a Russian systems engineer with twenty years of experience, writing complex technical documentation in English, this hits incredibly close to home. When you are documenting hard engineering, your language must be highly structured, logical, and precise.
I deal with this exact demand for precision in my domain. I build strictly deterministic C++ state sync cores for Medium-Frequency Trading (MFT). I recently finalized the monolithic architecture for my own engine (TolmachЁv SDK v36.0.0). When your baseline requirement is hitting 41.5M TPS with ~24ns physical RTT and absolute zero CPU validation waste, your technical specs cannot afford a single ambiguous word.
I actively use AI to ensure my English explanations of this C++ architecture are as flawless and deterministic as the code itself. If a detector like Pangram or Sloan flags my documentation because it is "too perfectly structured" or lacks human "flaws," it is actively penalizing engineering rigor. They are treating precision as a crime.
You absolutely nailed the conclusion. The only metric that actually matters in technical writing is accountability: "Did you architect this system, and do you own the logic behind these words?" Excellent piece.
Puts me on abstract mode 🤔 forensic. I think a 2017 article claimed AI generated was based on the fact that a human knows a human voice to a good extent similarly an assistant was used to translate to French would be a case of grammar formating not idea. But how do you really know if an egg was actually fried or boiled by looking at the eggshell 🤔
Ekong, the eggshell metaphor is dope and it points at the real limit of it. You can't tell frying from boiling by the shell because the shell was never part of the cooking, it's packaging, not process.
Writing doesn't work that way. The words are the process. Every choice, which clause comes first, what gets cut, how a sentence closes, is baked directly into the same surface a detector is scanning.
That's what makes the grammar-versus-idea line so hard to hold in practice.
Pascal's case wasn't formatting laid over finished thinking, it was rephrasing woven into the sentence itself. A classifier reading the sentence has no way to separate the 2 because there's no shell to look at separately from the egg.
I don't think this gets solved by looking harder at the surface. You'd need to watch it get made.
The bit about disclosure creating a dishonesty incentive is the piece I keep coming back to, because that's a labeled dataset problem, not just a policy one. If the classifier was trained on posts marked AI-assisted, it's calibrated to the population that admits it, and the invisible population (equally assisted, undisclosed) becomes the false-negative tail nobody measures. Curious if Pangram publishes precision at the writers-who-know-their-craft slice, because that's where structural failure lives.
Kartik, that's a different failure mode than the one the piece is actually about, and it's the sharper one. I was pointing at false positives: real human writing that happens to read AI-shaped.
You're pointing at the mirror case: AI-assisted writing skilled enough to read human, undisclosed, sitting quietly in whatever "human" bucket the training set uses. That second tail is close to unlabelable in principle, since ground truth requires someone admitting to something they have no incentive to admit.
The mirror-prompt method Pangram describes doesn't obviously close that gap either. It generates synthetic AI counterparts matched to real human documents, that's a different distribution than a real writer's genuine, undisclosed, AI-assisted draft. One is manufactured contrast. The other is a population that opted itself out of the dataset by staying quiet...
I haven't seen Pangram publish precision sliced by disclosure status or writer skill, just an aggregate false-positive cap. That's exactly the number that would tell you whether the classifier is calibrated against real skilled writers or just against its own synthetic negatives, and it's the one number missing from everything they've published.
As someone who's starting to build with AI every day, I don't think AI use itself is the problem. The question should be whether the author understands what they're publishing and can defend it. AI should help people communicate better, not replace their thinking.
If it flags everything, then the tool isn't really doing its job.
Was it vibe coded 🤣