
Our Content Agency Charged $4,000/Month. Then I Ran Their Work Through an AI Detector.
Try WriteMask free
500 words/day. No credit card required. Paste AI text and see the difference.
Jordan runs content for a mid-size direct-to-consumer brand. In early 2026, she was paying a content agency $4,000 a month for blog posts — 16 articles a month, delivered consistently, always on deadline. Good enough, she thought. Then a Slack message changed everything.
A peer at a competing brand sent a screenshot: "We caught our agency submitting raw ChatGPT output. Check yours." Jordan didn't want to believe it. But that night, she copied three of the agency's most recent deliverables and ran them through a detector. Two came back at 91% AI. One hit 97%.
How Do You Actually Check If Writing Is AI Generated?
Checking if writing is AI generated means running text through a detection tool that analyzes statistical patterns — things like sentence rhythm, predictability, and word choice distributions — that differ between human and AI output. Most detectors assign a percentage score indicating how likely the text is to have been written by a model like GPT-4 or Claude.
The fastest starting point is a free AI detector — no login, no paid plan. Paste the text, wait a few seconds, get a score. Simple. But the score alone doesn't tell the whole story, and that's exactly where Jordan's situation got complicated.
What Jordan Found — And Why the Score Alone Wasn't Enough
Three articles flagged high. But Jordan didn't immediately fire the agency. She'd read enough about AI detection false positives to know that a high score isn't always definitive. Technical writing, listicles, and content with predictable structure often scores high even when written by humans.
So she tested further. She pasted her own internal brief documents — definitely human-written — and they came back clean. She tested a recent article from a journalist she trusted. Also clean. She ran the agency's content again on a different detector. Still high. Consistent across tools. That's when she knew.
She also noticed something the score couldn't capture: the articles ignored every brand voice guideline she'd included in the brief. They were correct. Organized. But they didn't sound like anyone in particular. That's the qualitative tell that almost always accompanies a genuinely high AI score.
The Problem With Checking AI Writing in 2026
Detectors have gotten better. So has AI. Understanding how AI detectors work helps explain the gap: they're trained on patterns from earlier models, but carefully prompted outputs from newer systems can look statistically closer to human writing. A score of 60–75% is genuinely ambiguous. Above 85%, repeated across multiple tools, is harder to explain away.
There's also a business-specific wrinkle that Jordan hadn't considered. Her brand publishes content that Google crawls. She'd seen the data on Google and AI content SEO — the algorithm doesn't automatically penalize AI-written content, but it does reward demonstrable expertise and original perspective. Mass-generated AI articles often underperform on organic search because they're thin on genuine insight, regardless of whether they pass a detector.
What Jordan Did Next
She confronted the agency with the evidence. They admitted to "AI-assisted drafting" — a phrase that apparently covered running prompts and making minimal edits. The relationship ended.
But Jordan still had a backlog of published content sitting on her site with high AI scores. She needed to either rewrite it or humanize it. She ran the flagged posts through WriteMask, which rewrites AI-pattern text while preserving original meaning and structure. The platform's 93% pass rate on major detectors is why she chose it over basic paraphrasing tools. She worked through the entire backlog in a single afternoon.
A Practical Checklist for Checking Someone Else's Writing
Here's the process Jordan now uses before approving any outside deliverable:
- Run the text through at least two different detectors. Consistency across tools matters more than one high score from one platform.
- Test known-human writing from the same source — previous samples, emails, notes — and compare the scores. Baseline matters.
- Look for qualitative tells: missing brand voice, generic examples, no specific data, no real opinion, no hedging.
- Ask the contractor to share a rough draft or their drafting notes. Writers who write have them. AI runners rarely do.
- If flagged content is already published, humanize it rather than delete it — the SEO value in aged URLs is worth preserving.
The score is a signal, not a verdict. But several signals pointing the same direction? That's a pattern. Jordan learned that lesson the $4,000-a-month way. You don't have to.