h.
Humanizer.blog
Empirical NLP Testing Hub

AI Detector Benchmark & Sensitivity Audit

Quantitative test results analyzing detector false-positive rates, sentence length variance thresholds, and perplexity scoring models across 1,000 essay samples.

AI Detector Accuracy & False Positive Audit

Turnitin AI (2026)

Enterprise
Base AI Detection: 98.4%
ESL False Positive Rate: 38.2%
Humanizer Bypass Rate: 96.5%

Evaluates document-level sentence length distribution. High false positive rate on structured non-native English essays.

Originality.ai 3.0

Web & SEO
Base AI Detection: 99.1%
ESL False Positive Rate: 22.4%
Humanizer Bypass Rate: 94.1%

Strict token-predictability model tuned for web publishers. Highly sensitive to repetitive transitional phrases.

GPTZero

Academic
Base AI Detection: 97.2%
ESL False Positive Rate: 31.0%
Humanizer Bypass Rate: 97.8%

Relies directly on average perplexity and burstiness variance curves across text blocks.

Audit Methodology & Test Setup

Our lab feeds 1,000 essay samples generated by ChatGPT (GPT-4o), Claude 3.5 Sonnet, and Gemini 1.5 Pro into each humanizer software on default "Academic" and "Enhanced" modes. The resulting outputs are evaluated through enterprise API instances of Turnitin, Originality 3.0, and GPTZero.