
My Client Said Originality.ai Proved I Used ChatGPT — So I Ran the Same Text Through 5 Detectors
Try WriteMask free
500 words/day. No credit card required. Paste AI text and see the difference.
Marcus had been writing content for the same SaaS marketing agency for eight months. Clean record. Good reviews. Then, on a Tuesday in March, the client emailed: "We ran your last three posts through Originality.ai. Average score: 84% AI. We're ending the contract."
That was $3,200 gone. And Marcus had written every word himself.
What followed was a two-week experiment that taught him something most content writers never figure out: AI detectors don't have a single accuracy rating. They disagree with each other — wildly, sometimes on the same sentence.
What "Checking AI Accuracy" Actually Means
AI detector accuracy is not one number. Each tool has its own training data, its own threshold for flagging text, and its own false positive rate. When someone asks how to check AI accuracy, they're usually asking the wrong question. The better question is: which detector should I trust, and under what conditions?
Researchers who benchmark these tools find false positive rates on human-written text that range from low single digits to over 30%, depending on the tool and the writing style being tested. Academic writing, non-native English, and heavily edited prose get flagged most often — not because they look like ChatGPT, but because they pattern-match to the same statistical signals.
Marcus Tests Five Detectors on the Same 800-Word Article
He took one of the flagged articles — a piece about product-led growth he'd written in a coffee shop over two hours — and ran it through five tools back to back. No edits. Same file every time.
- Originality.ai: 84% AI
- GPTZero: 41% AI
- Copyleaks: 18% AI
- Turnitin (via a university portal): Low AI likelihood
- WriteMask's free AI detector: 22% AI
Same text. Scores ranging from 18% to 84%. That's not a rounding error. That's the fundamental problem with AI detection: there is no shared ground truth between tools.
How to Actually Check Whether an AI Detector Is Accurate
The most reliable way to check AI detector accuracy is to run known-human text and known-AI text through the same tool and compare results. This is called a personal benchmark, and it takes about twenty minutes. Here's the exact method Marcus used:
- Take five pieces of text you wrote yourself — old emails, journal entries, any draft with no AI involvement whatsoever
- Take five pieces of raw ChatGPT output on similar topics, unedited
- Run all ten through the detector in question
- See how cleanly it separates the two groups
When Marcus did this with Originality.ai, it correctly flagged all five AI samples. But it also flagged three of his five human samples as AI. That's a 60% false positive rate on his specific writing style. Not a reliable signal.
Understanding how AI detectors work at a technical level explains exactly why this happens. Most tools learn statistical patterns from a snapshot of text. Your writing style might match those patterns even when it's 100% human — especially if you write with clarity and structure.
Why Some Writers Get Flagged More Than Others
Marcus noticed something in his own testing. His most structured, list-heavy posts scored highest for AI. His conversational, anecdote-driven pieces scored lowest. Every time.
This matches what researchers consistently observe: clean organization, predictable transitions, and uniform sentence rhythm are the signals these models key on. They were also signals of good writing long before ChatGPT existed — which is why AI detection false positives hit certain writers disproportionately hard. Technical writers. Non-native English speakers. Anyone who's spent years learning to write clearly and efficiently.
What Marcus Did Next
He couldn't undo the lost contract. But he could protect himself going forward. Three changes.
First, he now asks new clients upfront which detector they use, if any. Standard part of his onboarding process. Some clients don't use one at all — that's useful to know before you spend energy worrying.
Second, for clients who do check for AI, he runs his drafts through WriteMask before delivery. The tool restructures phrasing at the sentence level while keeping his meaning intact. It achieves a 93% pass rate across major detectors including Originality.ai and GPTZero — which covers the tools most of his clients actually use.
Third, he includes a brief note with each delivery showing detection results from three different tools, with the scores side by side. Most clients, when they see 18%, 41%, and 84% on the same article, quickly understand that the technology isn't an objective test. It's a probability estimate. Framing it that way has ended two potential disputes before they started.
If you're already past prevention and dealing with a formal accusation, there's a process for proving your writing is human that goes well beyond detector scores — timestamped drafts, version history, and direct appeals that put the burden of proof where it belongs.
The Practical Takeaway
If you want to check AI accuracy — meaning, if you want to know whether any given detector score actually means something — run the same text through at least three different tools and compare. Scores clustered together suggest real signal. Scores that span 60 percentage points are noise.
A single detector result is not evidence of AI use. It's a probabilistic guess made by a model that was not trained on your writing. Treat it accordingly, and make sure the people reviewing your work do too.
Marcus learned that the hard way. You don't have to.