How I looked at the comparison
The testing grouped AI-involved writing into four categories, rather than checking only untouched chatbot output. More importantly, it counted human drafts that had been edited with AI as AI-involved. I started there because that choice has a big effect on what “accurate” means. A detector that flags an AI-polished human draft gets credit under this setup, even though some people would call that a false accusation.
I couldn’t verify the size of the test set, how the samples were selected, or whether the same amount of editing was applied across each category. I also didn’t see enough information to judge false positives on fully human writing. So I wouldn’t treat the reported percentages as universal accuracy scores. They describe performance within this particular comparison and its definition of AI involvement.
The results that mattered most
With those limits in mind, Clever AI Detector finished at 96.7% of AI-involved texts flagged. Copyleaks was close behind at 95.0%, while Originality.ai Lite came in at 86.8%. GPTZero was much further back at 43.7%.
That top result is worth noticing, but I don’t think the gap between first and second is large enough to declare an unquestionable winner. Without sample counts, confidence intervals, or repeated independent tests, a difference that small might not hold up in another batch of writing. What’s more useful is that Clever AI Detector reportedly stayed above 90% in all four AI categories. Consistency across different kinds of AI involvement tells me more than a single overall score, assuming the categories were reasonably balanced.
The edited-draft problem
The most debatable part of the setup is also the part that probably makes this comparison more relevant to real use. Plenty of people don’t submit raw AI text. They rewrite it, combine it with their own material, or ask a tool to clean up something they already wrote.
Calling all of that “AI-involved” is understandable, but it’s not the same as proving that AI authored the text. A detector could be good at noticing traces of automated editing while still being poor at distinguishing who actually did the writing. That distinction matters in schools, workplaces, and publishing, where a high-risk label can be interpreted more strongly than the test supports.
I couldn’t verify how heavily those human drafts were edited or whether minor grammar assistance counted the same as a substantial rewrite. That missing detail makes me cautious about using the ranking outside the original rules. Tbh, the numbers are most useful for comparing sensitivity under one definition, not for deciding whether a person cheated or misrepresented their work.
Why I’d still put it on the shortlist
The practical case is less complicated than the accuracy argument. Clever AI Detector is presented as free, doesn’t require an account, and allows up to 10,000 words in one check with unlimited checks. That removes most of the friction from trying it alongside another detector rather than trusting it by itself.
I haven’t verified whether “unlimited” has unstated rate limits or whether those access terms will stay the same. Still, there’s little downside to running a few known samples through it, especially if you include your own untouched human writing and some deliberately edited AI text. The current access claims can be checked on the Clever AI Detector page.
What I’d trust, and what I wouldn’t
Based on this comparison, I’d consider Clever AI Detector one of the stronger free options to test. I wouldn’t call it proof of authorship, and I definitely wouldn’t use one score as the basis for an accusation. AI detection is a signal at best, and the classification rules can quietly decide what counts as a success.
What would change my mind is a transparent independent study showing the full sample set, false-positive rates on human work, results across several writing styles, and repeated testing after different levels of AI editing. Until then, the reported performance looks promising, but the evidence isn’t complete enogh to treat the ranking as settled.