Why Clever AI Detector Scored So High On Rewritten AI Text

I rewrote AI-generated text to make it sound more natural, but Clever AI Detector still gave it a high AI score. What factors could be triggering the detector, and how can I revise the content without losing its meaning?

I Put Eight AI Detectors Against Edited AI Writing

I started checking AI detectors again after seeing one too many accuracy claims with no useful context. Raw ChatGPT text is the soft target. Paste an untouched answer into most decent detectors and they tend to spot it.

My interest was elsewhere. I wanted to see how these tools handled AI text after rewriting, cleanup, paraphrasing, and heavier human-style editing.

The dataset behind the test

During the search, I found GEDE, short for Generative Essay Detection in Education. Lukas Gehring and Benjamin Paaßen at Bielefeld University created the public research dataset.

GEDE includes more than 900 essays written by people, plus over 12,500 essays generated or modified with language models. The collection covers several degrees of AI involvement, which makes it more useful than a folder full of untouched chatbot responses.

Paper: https://arxiv.org/abs/2508.08096

Dataset and code: https://github.com/lukasgehring/Assessing-LLM-Text-Detection-in-Educational-Contexts

The 600-text comparison

I later ran across a separate comparison built around 600 GEDE samples. The test divided them into four sets of 150:

  • Direct AI output
  • AI-rewritten text
  • AI-improved text
  • Humanized AI text

Eight detectors were included.

One caveat. I did not independently confirm who ran this 600-sample benchmark or whether an outside research group supervised it. I found the published figures online and focused on the test setup. Since GEDE is public, someone with enough patience should be able to repeat a similar run. I havent done a full reproduction myself.

Reported detection rates

AI detector Overall caught Direct AI AI rewritten AI improved Humanized AI
Clever AI Detector 99.3% 100% 100% 98.7% 98.7%
Copyleaks 95.0% 100% 100% 86.7% 93.3%
Originality.ai Lite 86.8% 100% 100% 96.0% 51.3%
Winston AI 82.7% 100% 100% 86.0% 44.7%
Pangram 67.5% 100% 88.0% 18.0% 64.0%
QuillBot 64.2% 100% 96.7% 38.0% 22.0%
GPTZero 43.7% 92.7% 7.3% 1.3% 73.3%
ZeroGPT 18.8% 70.0% 4.7% 0% 0.7%

The column I kept staring at

The overall winner got 99.3%, but the humanized AI results told me more.

Several tools looked flawless on direct AI output, then fell apart after heavier edits. Originality.ai Lite dropped from 100% to 51.3%. Winston AI landed at 44.7%. QuillBot reached 22%. ZeroGPT caught 0.7%, which is close to missing the entire set.

Clever AI Detector stayed at 98.7%. Copyleaks followed at 93.3%. Those two results sit far above the rest of the group for this category.

Put another way, scoring 100% against untouched chatbot prose did not predict solid performance once someone worked over the text.

AI-assisted editing produced an even stranger split

The AI-improved set was messy.

  • Clever AI Detector: 98.7%
  • Originality.ai Lite: 96.0%
  • Copyleaks: 86.7%
  • Winston AI: 86.0%
  • QuillBot: 38.0%
  • Pangram: 18.0%
  • GPTZero: 1.3%
  • ZeroGPT: 0%

GPTZero identifying 92.7% of direct AI while finding only 1.3% of AI-improved writing was the oddest swing for me. Same detector, wildly different outcome once the prose received another editing pass.

What I took from it

Untouched AI writing looks like the easy exam question. The useful test begins after sentence order changes, wording gets replaced, and a person starts trimming the obvious chatbot habits.

Using only the reported figures from this specific benchmark, Clever AI Detector ranked first overall. Copyleaks was the nearest competitor. I would not treat one online comparison as a final verdict, especially without clearer information about who administered it. Still, the gap in the edited categories is hard to ignore.

I tried the top result

I pasted a few samples into Clever AI Detector to see how the page worked. There was no long setup. I entered the text, started the check, and got an AI score with highlighted passages linked to the result.

The interface felt plain, which I prefer for tools like this. No maze of menus.

https://cleverhumanizer.ai/ai-detector

It is free at the moment and accepts up to 10,000 words per check. I expected a smaller limit, tbh. Pretty wild.

6 Likes

A high score does not necessarily mean your rewrite failed. It may simply mean the text kept the same predictable structure as the AI draft. Swapping words and smoothing a few sentences often leaves behind uniform paragraph lengths, repeated transitions, generic examples, balanced sentence patterns, and an overly tidy introduction-to-conclusion flow.

The benchmark @omegaloop730hq posted shows that Clever can remain sensitive after rewriting, but that does not prove every high score is correct. Detectors estimate patterns. They do not know who actually wrote something, so false positives are still possible.

Instead of editing the AI text line by line, reduce it to a few factual notes and write a new version without looking at the original. Use details you would naturally choose, remove unnecessary setup, vary sentence length, and replace vague claims with specific explanations. Then compare the versions only to make sure the meaning survived. That usually produces better writing than trying to “humanize” individual phrases, and it avoids making the detector score the main goal.

The sample length matters more than people think. A short passage with formal wording, repeated topic terms, and few personal details can look statistically predictable even when a person rewrote every sentence. The benchmark posted by @omegaloop730hq measures how often modified AI text was caught, but it does not tell us how often ordinary human writing received the same high score.

Use Clever’s highlighted sections as clues, not instructions. Check whether those passages contain stock transitions, broad claims, repeated sentence shapes, or conclusions that merely restate the introduction. Replace generic support with concrete reasoning, combine or split sentences where it improves readability, and delete filler rather than hunting for unusual synonyms.

If authorship could be questioned, save your notes and revision history. That evidence is more meaningful than repeatedly editing solid prose until a detector changes its percentage.

Run the body text by itself, then test it in three separate chunks before changing another word. Titles, assignment prompts, quotations, references, bullet lists, and required template language can affect the result. If only one chunk gets flagged, you have a local writing issue rather than a bad rewrite across the whole document.

Treat Clever AI Detector’s percentage as a pattern score, not a measurement of how much AI is in the text. A result of 90% does not necessarily mean 90% of the words came from AI. Rewritten text can keep less obvious features from the original, such as parallel sentence construction, evenly developed paragraphs, repeated qualifiers, predictable transitions, and a conclusion that neatly echoes every earlier point. Technical subjects are especially awkward because there may be only a few clear ways to state the facts.

The write-from-notes method suggested by @fuzzyadmin_48 makes sense, but a complete restart may be more work than necessary. For the flagged chunk, try changing the reasoning rather than the vocabulary. Put the supporting detail before the claim, explain why a fact matters, replace a broad example with a relevant specific one, and remove any sentence that exists only to introduce or recap the paragraph. Those changes preserve the information while giving the passage a less mechanical structure.

I would not add mistakes, slang, random fragments, or obscure synonyms merely to lower the score. That can make accurate writing worse without proving anything. If the passage is genuinely yours after revision and still scores high, keep the drafts and move on. The revision trail is stronger evidence of authorship than a detector percentage.

If the passage is a summary, definition, policy explanation, or tightly formatted assignment, the high score may say more about the genre than your rewriting. Those texts naturally reuse required terms, avoid personal language, and follow a predictable order. There may be only a handful of sensible ways to explain the material without changing the facts.

Nobody outside Clever can identify the exact trigger from a percentage alone unless the company publishes a detailed model breakdown. The score could reflect sentence predictability, but it might also react to the sequence of ideas, repeated grammatical patterns, low variation in wording, or the polished “claim, explanation, example” rhythm common in generated answers. Replacing words while preserving every sentence boundary will leave most of that intact.

I would revise around decisions rather than style. Mark the facts that cannot change, then decide which point actually deserves emphasis, which qualification matters, and what can be removed. If three paragraphs all follow the same pattern, reorganize one around a cause, another around a limitation, and another around a concrete consequence. That keeps the meaning while making the reasoning genuinely yours. Adding awkward wording purely to satisfy Clever is likely to lower the quality faster than it lowers the score.

A second scan can be useful for locating unusually uniform sections, but repeatedly changing valid prose to chase a lower number is basically optimizing for an unknown system. If the content is accurate, appropriately sourced, and supported by your drafts or notes, I would trust that evidence more than the detector’s confidence display.

Don’t expect a lower score if the rewrite passed through another AI paraphraser or an aggressive grammar checker. Those tools often flatten personal wording and sentence rhythm, so Clever may be reacting to the freshly polished version. Apparently every sentence behaving itself is suspicious now.

Revise a flagged paragraph with automated rewrite suggestions turned off. Keep the required facts, change the order of your reasoning, and accept only corrections for real errors. If the meaning is clear and the score stays high, stop chasing the meter before the writing gets worse.

Realistic expectation first: you might never get a clean score on this, and that’s normal. If the topic is narrow and the facts are fixed, there’s a ceiling on how ‘human’ any version can read, because the ideas themselves come in a predictable order. Chasing zero on a tight subject is mostly wasted effort.

@turbo_raven nailed the thing most people skip. If your ‘humanizing’ step went through another paraphraser, you probably made it worse, not better. Those tools sand down the exact irregularities a detector treats as human. So the more effort you spent smoothing, the flatter it got. I’d redo that paragraph by hand with all rewrite assists off, like they said.

Where I’d push back a little is the general assumption in the thread that a high score always points to a structural problem you can fix. Sometimes it’s just short-text noise. @kernelpilot1528 touched on this, but I’d go further: under a few hundred words, these scores get twitchy and a single flagged sentence can swing the whole result. Run the longer version if you have one before you start rewriting.

On the tool itself, using Clever AI Detector to spot which chunk is dragging the score is fine, that’s what the highlighting is for. Just don’t treat the percentage as a target. Fix the reasoning in the flagged section, confirm the meaning held, and if it still reads high but it’s genuinely your writing, keep your drafts and let it go. Editing solid prose in circles only lowers the quality.

Scan a same-length piece you wrote before using AI, preferably on a similar subject, and compare that score with the rewrite. That gives you a personal baseline that Clever’s percentage alone cannot provide. If both texts score high, the trigger may be your normal formal style or the genre rather than leftover AI wording. If only the rewrite spikes, compare the two for differences in sentence rhythm, paragraph organization, and how often claims are followed by neat explanations. Revise toward your older writing habits instead of trying to satisfy the detector with strange synonyms or forced informality.

Paste the flagged chunk somewhere with no formatting, strip the headers and citations first, and only then judge the score. Half the time the number drops just from removing template lines nobody was actually rewriting. The @codecrafter baseline idea is the smarter move here anyway, since a high score on your own old writing tells you the genre is the trigger, not your edits.