
A Client Refused to Pay Me After GPTZero Flagged My Writing — Here's Exactly How AI Detection Works
Try WriteMask free
500 words/day. No credit card required. Paste AI text and see the difference.
The email arrived on a Thursday afternoon. Marcus, a freelance content writer based in Austin, had just delivered a 1,500-word case study — three hours of actual work, no AI involved. His client's reply was two sentences: "We ran this through GPTZero. It came back 87% AI. We're not paying."
Marcus had not used AI to write that piece. But he had no idea how to prove it. Or even what the detector had found in his words.
That dispute turned into a two-week deep dive into how AI detection actually works, why perfectly human writing sometimes triggers it, and what you can do about it — whether you're trying to catch AI writing or protect your own.
What Do AI Detectors Actually Look For?
AI detectors flag writing by measuring two main signals: perplexity and burstiness. Perplexity measures how predictable each word choice is — AI tends to pick the statistically safest word every time. Burstiness measures whether sentence lengths vary naturally — humans mix short punchy sentences with longer complex ones, while AI output has a more uniform rhythm.
Marcus pulled up several resources on how AI detectors work and started testing his own archived writing samples. Pieces written under deadline pressure — tighter sentences, less varied structure — scored 40–60% AI on multiple tools. His more relaxed, exploratory pieces scored under 10%. Same writer. Wildly different results.
The Manual Signs of AI Writing (That Detectors Also Catch)
If you want to spot AI writing without a tool, look for these patterns:
- Perfect topic coverage. AI rarely skips important points. If a piece covers every angle with equal emphasis, that's a flag.
- No uncomfortable opinions. Human writers hedge, contradict themselves, take unpopular stances. AI tends to be diplomatically vague.
- Missing specifics. No names, no dates, no dollar amounts, no messy real-world details. AI generalizes.
- Uniform sentence rhythm. Read it aloud. If every sentence takes roughly the same amount of time to say, that's suspicious.
- Filler transitions. Phrases like "it's worth noting," "in today's world," and "this is important because" appear constantly in unedited AI output.
Marcus recognized two of those patterns in his own flagged piece — not because he'd used AI, but because deadline pressure had pushed him into safe, predictable phrasing. The uncomfortable realization: the detector wasn't entirely wrong about the style. It was wrong about the cause. That's exactly the AI detection false positive problem — human writing that statistically resembles AI output, especially from ESL writers, technical writers, and anyone under pressure to be concise and clear.
Which Tools People Actually Use to Detect AI Writing
The main tools right now are GPTZero, Originality.ai, Copyleaks, and Turnitin's built-in AI detector. Each uses slightly different models, which is why the same paragraph can score 85% on one and 30% on another. There is no universal standard, no certification body, no appeals process.
WriteMask's free AI detector runs the same kind of analysis — paste any text and see what a detector flags before anyone else does. Marcus started running every piece through it before delivery. Not to hide anything. To catch writing patterns that might look suspicious even when the work was entirely his own.
What Marcus Actually Did to Fix the Problem Going Forward
He couldn't prove the disputed piece was human retroactively — no draft history, no notes saved. That cost him the invoice. But his workflow changed entirely after that.
First, he started saving timestamped Google Docs drafts to build a visible edit trail. Second, he ran every deliverable through a detector before submission. Third, when a piece scored above 20% on any tool, he used WriteMask to identify the specific sentences triggering the score and rewrote those passages. WriteMask carries a 93% pass rate across major detection tools — meaning the rewritten pieces consistently came back clean.
He also started having a conversation with new clients upfront: AI detectors are probabilistic, not forensic. They output a likelihood score, not a verdict. If you plan to use detector results in a payment dispute, that policy needs to be agreed on before work starts — not surfaced in a two-sentence rejection email.
If You're the One Trying to Catch AI Writing
Run the text through at least two different tools. If both flag the same passages, that's a meaningful signal. If they disagree sharply, the signal is weak. Use the manual checklist above alongside any tool score — patterns plus a high score together is far stronger evidence than either one alone.
And keep in mind: a high score does not prove AI authorship. It proves a style match. Non-native speakers, writers under pressure, and people writing in highly formal registers trip these tools regularly. If the stakes are real — a payment withheld, an academic sanction — understand what to do if you're accused of using AI before making or accepting any final decision.
Marcus eventually rebuilt the client relationship. He sent a thorough breakdown of how detectors work, shared his new draft-history workflow, and offered a second project with milestone check-ins. The client agreed. The next three pieces all scored under 5% on every tool — and he had the document history to back it up if anyone ever asked again.