
My Client Said Their AI Detector Flagged My Articles — Here's What Was Actually Happening
Try WriteMask free
500 words/day. No credit card required. Paste AI text and see the difference.
The email arrived on a Tuesday morning. The client — a SaaS startup I'd been writing for across six months — sent three lines: "Hi, we ran your last batch through our internal AI checker. It flagged four of the five articles as AI-generated. Can we jump on a call?"
The problem? I had written every word myself.
If you've landed here, something similar probably happened to you. Maybe a client, a professor, an editor, or your own anxiety made you wonder: how do you actually check if something is written by AI? And if you wrote it yourself — or used AI as a starting draft — what do those checks even measure?
What AI Detectors Actually Check For
AI detectors don't read writing the way humans do. They look for statistical patterns — not meaning, not ideas, not argument quality. Just math.
The two main signals are perplexity (how predictable each word choice is, compared to what a language model would expect) and burstiness (how much sentence length varies). AI-generated text tends to have low perplexity and low burstiness — it's smooth, consistently structured, and rarely surprises. Human writing is messier. Short punchy sentence. Then a longer one that wanders a bit before landing on the point.
Tools like GPTZero, Copyleaks, and Originality.ai all use variations of this approach. Understanding how AI detectors work changes how you react when one flags your writing — because you stop taking it personally and start treating it as a solvable technical problem.
Why Your Perfectly Human Writing Gets Flagged Anyway
Here's the uncomfortable part: false positives are common. Academic writing, formal business copy, and technical content all share surface features with AI output — precise word choices, consistent sentence structure, minimal slang. If your writing style is naturally polished and formal, you're at higher statistical risk of a flag.
That was my exact situation. B2B SaaS copy has a tonal profile that overlaps heavily with how GPT-4 writes. The detector wasn't wrong about the patterns. It was wrong about what caused them.
The research on AI detection false positives is sobering — some tools flag non-native English writers and highly structured writers at dramatically higher rates. The checker your client or professor used might be one of them.
How to Check If Something Is Written by AI (The Right Way)
Running text through a single AI detector and trusting the result is like getting one doctor's opinion on a complicated diagnosis. Here's a more reliable process:
- Run it through multiple tools. If one flags something at 85% AI and two others say 30%, you have conflicting data — and that's actually useful to share with whoever questioned you.
- Check section by section, not the whole document. Detectors often flag specific paragraphs. Knowing which section triggered it tells you what to actually address.
- Look at the confidence level, not just the label. A "likely AI" verdict at 62% confidence is very different from 96%. Most tools bury this.
- Use a tool that shows per-sentence breakdown. WriteMask's free AI detector highlights individual sentences so you can see exactly what tripped the algorithm — not just a blanket verdict on the whole piece.
What to Do After Your Writing Gets Flagged
Whether you used AI or not, the practical question is the same: what now?
If you didn't use AI at all, you have a documentation problem, not a writing problem. Timestamps from drafts, revision history in Google Docs, browser search records from your research — these can help establish authorship. There's a solid guide on how to prove your writing is human that covers exactly what to gather and how to present it credibly.
If you used AI as a starting point and then edited it, you're in grayer territory — and the text itself probably needs reworking regardless. This is where WriteMask earns its keep. It reworks the statistical fingerprint of AI-assisted text at the sentence level, adjusting perplexity and burstiness so the writing reads as human-authored — without gutting your meaning or making the content worse. It passes at a 93% rate across major detectors including GPTZero, Copyleaks, and Turnitin.
In my case: I reran the flagged articles through three separate detectors. Two gave clean results. One still flagged two pieces. I used WriteMask on those specific sections, shared the before-and-after scan results with my client, and the contract stayed intact.
The Real Problem with "Just Check It"
The assumption behind "we'll run it through an AI checker" is that these tools have a reliable answer. They don't — not yet. Detection accuracy varies wildly by tool, by writing domain, and by which AI model generated the text. A detector trained heavily on GPT-3.5 output may miss GPT-4o writing entirely. One tuned for student essays may consistently over-flag technical copy.
This matters whether you're the one being checked or the one doing the checking. A single detector score is not a verdict. It's a data point that needs context.
For writers, freelancers, and students — the practical answer is: check your own work first, understand what the score actually measures, and know how to address it before someone else makes it a problem.