AITextTools
Back to Blog
guide

How AI Detection Tools Work: Perplexity, Burstiness & More

AI detectors score probability, not proof. They measure perplexity (word predictability) and burstiness (sentence rhythm) — two statistical signals, not a fact-check.

AI Text Tools Team
Updated July 20, 2026
10 min read

I spent an afternoon last month running the same paragraph through five different AI detectors. Same text, word for word. One tool flagged it as 97% AI-generated. Another said 12%. A third couldn't decide and landed on "mixed."

That inconsistency isn't a glitch — it's the whole story of how AI detection actually works.

At AI Text Tools, we build software that writers, students, and content teams use every day, so we get asked constantly: how do these detectors even know? The honest answer is they don't "know" anything the way a human editor would. They're making statistical guesses based on patterns in your sentences — mainly two patterns called perplexity and burstiness. Once you understand both, the black box gets a lot less mysterious.

This guide covers how AI detectors work, why they sometimes flag human-written text as robotic (and miss actual AI text entirely), and what that means for writing that sounds genuinely human.

The Real Question About AI Detectors

Short answer: an AI detector is a classifier, not an interrogator. It doesn't understand what you wrote. It runs your text against statistical patterns associated with large language models and returns a probability — not a ruling.

Think of it less like a lie detector and more like a weather forecast. A forecast can't actually tell you whether it will rain tomorrow — it calculates a probability of rain based on patterns it recognizes. AI detectors work the same way. They're trained on huge amounts of human and LLM text, and they learn which numerical patterns show up more often on one side than the other.

That's probability estimation, not a fact-check — and it's wrong a lot, especially on shorter chunks of text. Every major detector — ZeroGPT, GPTZero, Originality.ai, Copyleaks, Turnitin's AI module — leans on two core metrics to build that probability: perplexity and burstiness. Everything else is refinement layered on top of those two ideas.

Perplexity Unpacked: The Predictability Score

Definition: perplexity measures how surprised a language model is by the next word in a sentence. Low perplexity means highly predictable language. High perplexity means less predictable, more varied word choice.

Easy way to picture it: for "The cat sat on the ___," almost anyone — human or AI — fills in "mat" or "floor." That's low perplexity, very predictable. For "The cat sat on the windowsill, watching pigeons argue over a dropped bagel," almost no one predicts "bagel." That's high perplexity.

Models like GPT-4 or Claude are trained to predict the statistically most likely next word, so their default output skews toward low perplexity — safer, less surprising language. Humans tend to score more predictable too when they write plainly, so a competent human writer using clear, direct language can end up with a low perplexity score and get flagged anyway.

Burstiness Explained: The Rhythm Score

Definition: burstiness measures how much sentence length and structure vary across a piece of text. High burstiness means a mix of short, punchy sentences and longer, complex ones. Low burstiness means uniform length and rhythm throughout.

Read this out loud: "AI detection is imperfect. It relies on statistical patterns. These patterns can be inconsistent. The results can vary between tools." Every sentence is roughly the same length, same rhythm — low burstiness, a classic tell of unedited AI output.

Now compare: "AI detection is imperfect. Genuinely, deeply imperfect — the kind of imperfect where two tools can look at identical text and land on opposite conclusions, because they were never measuring truth in the first place, just probability." One short sentence, then one long, winding one. That's burstiness — something humans do almost without thinking, because real speech and real thought don't move at a constant pace.

Language models, left alone, tend to smooth this out. They produce sentences of fairly consistent length and complexity because that's what "average, safe" writing looks like statistically. That evenness is exactly what burstiness scoring is built to catch.

How Detectors Combine These Signals

No detector relies on perplexity or burstiness alone. Most run your text through a scoring model that weighs both, then layers on additional checks:

  • Token-level probability mapping — checking how "expected" each word is given the words before it
  • Sentence embedding comparisons — measuring how closely your sentence structures match known AI-generation patterns
  • Repetition and pattern detection — flagging repeated sentence openers, transition words, or paragraph structures
  • Formatting fingerprints — some tools flag overuse of bullet lists, bolded "key takeaway" phrasing, or transition words that GPT-family models default to

The output is usually a single percentage: "78% likely AI-generated." That number feels precise. It isn't — it's an average of several imperfect signals, rounded to look more confident than the underlying math actually is.

Why AI Detectors Get It Wrong

Quick answer: AI detectors both misidentify human-written text as AI and fail to catch actual AI-written text. Independent studies have found error rates high enough to rule out relying on any single tool as definitive.

  • Non-native English writers get misidentified more often — a 2023 Stanford study found AI detectors disproportionately flagged essays by non-native English speakers as AI-written, due to their naturally lower perplexity
  • Editing an AI draft breaks the pattern detectors look for — once someone runs AI text through a human editing pass, adds specific examples, and varies sentence length, the statistical fingerprint shifts and detectors often miss it entirely
  • Short text is unreliable to score — perplexity and burstiness need enough sentences to establish a pattern; a 100-word passage doesn't give the model much to work with, so accuracy drops sharply below a few hundred words
  • Detectors are trained on specific model families — a detector tuned mostly on GPT output may perform worse on text from other models, and always lags behind whatever the newest model generation writes like

None of this makes detectors useless. It means they're one signal among several, not a courtroom verdict.

ZeroGPT, Turnitin, GPTZero, and Originality.ai Compared

ToolPrimary Use CaseKnown Tendency
ZeroGPTGeneral web text, free checksFrequently flags short or simple human text as AI
GPTZeroEducation, essaysBuilt specifically around perplexity/burstiness scoring
Turnitin AIAcademic institutionsConservative flagging to reduce false accusations, still imperfect
Originality.aiContent marketing, publishersStronger on longer-form content, weaker on edited AI drafts
  • Pros: fast, free or low-cost initial screening; useful as one data point among several; can catch obvious, unedited AI dumps
  • Cons: no published detector has independently verified 100% accuracy; false-positive rates disproportionately affect certain writing styles; scores shift between tools on identical text; newer AI models are partly trained to reduce these exact statistical signatures

Can AI-Generated Text Ever Pass as Human?

Yes — and this is what makes detection an ongoing arms race rather than a solved problem. When AI output goes through genuine human revision — restructuring sentences, adding specific personal detail, varying rhythm, cutting the "smooth but empty" phrasing — the statistical signature changes enough that detectors frequently miss it.

This is also why detection alone shouldn't be the only quality bar. Text can score "0% AI" and still be bland, inaccurate, or unhelpful. Text can score "high AI" while being carefully researched and genuinely useful. Perplexity and burstiness measure writing style, not writing quality or truthfulness.

How to Write Naturally (Without Gaming the System)

The goal isn't to trick a detector. It's to write the way people actually write, which happens to score better because it's the real thing, not an imitation of it.

  • Vary your sentence length on purpose — follow a long sentence with a short one, let some sentences run past what feels "clean"
  • Use specific, concrete details — names, numbers, small sensory details raise perplexity naturally because they're not the "safe average" word choice
  • Cut the transition words that show up in every AI draft — moreover, furthermore, in today's world, delve into, unlock, seamlessly — these are statistically overused by language models and read as filler to human readers too
  • Read it out loud — if it sounds like a press release, it'll likely score low on burstiness; real speech has bumps in it
  • Let opinions and small imperfections stay in — a slightly informal aside, a personal anecdote, a mild tangent are hard for models to fake convincingly and easy for humans to produce naturally

Common Mistakes Writers Make With AI Detectors

  • Trusting a single tool's score as final proof
  • Chasing a "0%" score instead of writing genuinely useful content
  • Assuming a high AI score means the writer is lying about authorship
  • Ignoring that heavily editing an AI draft still counts as AI-assisted, ethically, even if it scores as human
  • Using AI-humanizing tricks (random synonym swaps, forced complexity) that make writing worse, not more human

Key Takeaways

  • AI detectors score probability, not proof — they measure statistical patterns, not authorship
  • Perplexity measures how predictable your word choices are; higher usually reads as more human
  • Burstiness measures sentence rhythm variation; consistent sentence length is a common AI tell
  • Detectors have real, documented error rates, including flagging non-native English writers unfairly
  • No detector score should be treated as a final verdict on whether text was AI-generated
  • Writing with specific detail, varied rhythm, and a genuine point of view naturally improves both readability and detector scores

Frequently Asked Questions

How high should an AI detector score be before I worry?

There's no universal safe threshold — different tools use different scales for different text lengths. Focus on making your writing clear and specific rather than chasing a target number.

Can ZeroGPT be wrong?

Yes. Like every AI detector, ZeroGPT can misclassify human-written text as AI-generated and miss actual AI-generated text.

Does using Grammarly or a spell-checker trigger AI detectors?

Generally no. Basic grammar tools don't rewrite sentence structure or word choice enough to meaningfully change perplexity or burstiness patterns.

Why did my own writing get flagged as AI?

Simple, direct sentences with consistent length register as low perplexity and low burstiness — the same patterns detectors associate with AI output, even when every word was human-written.

Do AI detectors work on languages other than English?

Most existing AI detectors are trained primarily on English text and are less accurate on other languages, since perplexity and burstiness patterns differ by language.

Can AI-generated text be made undetectable?

Text with genuine specific detail, structure, and a real point of view tends to score lower on detectors, but there's no guarantee of a specific score on any given tool.

Can schools rely on AI detectors to catch cheating?

Given documented false-positive rates, most guidance recommends using a detector score alongside other evidence — drafts, writing history, in-person discussion — rather than as the sole basis for a decision.

Conclusion

AI detectors aren't magic, and they were never designed to be. Perplexity and burstiness are useful statistical indicators, but they evaluate the structure of your sentences, not the validity of your authorship. At AI Text Tools, we'd rather help you produce content that's genuinely coherent, clear, and worth reading — content that scores well in those tools because it's your natural writing, not a trick aimed at the algorithm.

If you produce content at volume and want another set of eyes on tone, structure, and readability before you publish, that's exactly what AI Text Tools is built for.

Ready to Try AI Text Tools?

Use AI Text Tools to detect AI-generated content or rewrite your text in seconds. No sign-up required.