Evaluating AI Detectors on AI-Generated and Human-Written Academic Texts
The growing use of AI in academic writing is becoming a challenge for higher education institutions. This study evaluates the reliability of four AI detection tools in distinguishing human from AI-generated texts. It shows why detector scores should support academic judgement, not replace it.
Alfred, Piracintha, 2026
Type of Thesis Bachelor Thesis
Client FHNW
Supervisor Karg, Jona
Views: 3
Higher education institutions increasingly deploy AI detection tools to verify student authorship. Previous research shows that the performance of AI detectors varies significantly depending on the tool, model, text type, and test conditions. False positives, false negatives, and inconsistent classifications are widespread. Many studies rely on short texts or controlled datasets, while less is known about detector behaviour on complete academic papers and how lecturers interpret conflicting results in academic integrity decisions.
The study examined the reliability and limitations of AI detection tools and how meaningful their results are across different text versions. It combined a literature review, a case-based experiment, and a lecturer survey. One original ToBIT paper from 2024 was compared with two AI-generated versions using the same topic and references. Claude Sonnet 5 generated one freely, while ChatGPT-5.5 was guided by the student’s writing style. All three were tested with Turnitin, GPTZero, Originality.ai, and QuillBot. Five lecturers were surveyed about tool use and authorship verification.
The results show that AI-detection scores can provide useful signals, but they are not reliable enough to establish authorship on their own. Turnitin scored the human-written paper at 0% AI, compared with 88% for ChatGPT-5.5 and 72% for Claude Sonnet 5 versions. GPTZero classified the human paper as 96% human and both AI versions as 100% AI. Originality.ai assigned the human-written paper an estimated 62% AI score, indicating a possible false positive, while QuillBot under-detected the AI versions at 16.3% and 39%. The five surveyed lecturers also did not treat detector scores as definitive proof. They relied on oral discussion, exams, drafts and contextual evidence. For higher education institutions, the study therefore supports using AI detectors only as screening aids within a transparent, documented review process that combines tool output with policy, source checks, writing history and student explanation. AI detection can support academic judgement, but it cannot replace it.
Studyprogram: Business Information Technology (Bachelor)
Keywords Public Management Summary
Confidentiality: öffentlich