
I Thought I Could Spot AI Writing by Eye. The Research Proved Me Wrong — Here's What Actually Works
Try WriteMask free
500 words/day. No credit card required. Paste AI text and see the difference.
Here is a number worth sitting with: in controlled studies, human readers correctly identify AI-generated text roughly 50 to 60 percent of the time. That is barely above a coin flip. If you are relying on gut feeling to tell whether a piece of writing came from ChatGPT or a person, the research says you will be wrong a significant portion of the time.
Sarah manages content for a mid-size e-commerce brand in Austin. Three months ago, she started suspecting two of her freelance writers were submitting AI-drafted copy. The writing was not bad — that was the problem. It was polished, technically correct, and hit every brief. Something felt off. She ran the files through three different AI detectors. One returned 92% AI. One said 11%. One said 68%. She had no idea which to trust.
That confusion is exactly where most people get stuck. Here is what the science actually says.
What Does Research Say About Spotting AI Writing?
You cannot reliably identify AI-written text through reading alone. Human judgment consistently hovers around 55 to 60 percent accuracy — close enough to chance that it is not a useful standalone signal. This is not a failure of intelligence. Modern language models are trained specifically to produce text that reads as human.
What researchers have identified are statistical patterns, not stylistic ones. The two most documented:
- Low perplexity: Language models select words that are predictable given the surrounding context. Human writers take risks, use unexpected word choices, and make idiosyncratic moves. AI text hugs the statistical mean. Most detectors are scoring exactly this.
- Low burstiness: Human writing varies dramatically in sentence length and complexity. A long, winding sentence followed by a short one. AI text is more metronomic — sentences cluster around similar lengths. Researchers quantify this as burstiness, and AI scores consistently lower than human writers on this metric.
Knowing this reframes the question. Instead of asking "does this sound like a robot?