← Back to articles
news4 min read

The Turnitin Score Isn't Evidence: What JCQ Guidance Actually Says

A federal lawsuit over a false AI-cheating flag is a warning UK schools should heed: JCQ guidance says a detection score alone is not proof, but many teachers still treat it as one.

Q
Quill

A sophomore at Palo Alto High School submitted an essay on Arthur Miller's The Crucible in October 2025. Two weeks later, when it went through Turnitin, the tool flagged 76 per cent of it as AI-generated. His teacher, following the school's "non-punitive" policy, had him retake the assignment in class. He scored a D, and his overall grade dropped from a B to a C.

His father, Takashi Kato, submitted nearly 1,200 pages of evidence — drafts, notes, and Google Docs revision history — to show the essay was his son's own work. In May 2026 he sued Palo Alto Unified in federal court, alleging discrimination and denial of due process. The district is now fighting a $150 million claim, and in late June it filed its response denying wrongdoing.

This is not an isolated dispute over one essay. It is a preview of what happens when a probabilistic tool gets treated as a verdict.

The score was never meant to be a verdict

Turnitin's own documentation states a margin of error of roughly plus or minus 15 percentage points on its AI-writing indicator. A 76 per cent score, by the company's own admission, could reflect anywhere between 61 and 91 per cent — or, given how these models are trained on population-level patterns rather than individual students, it could simply be wrong. Turnitin says its own testing puts the false positive rate at under 4 per cent for full documents, but independent samples have found much higher rates, and the company itself is explicit that a score "should never be the sole basis for a decision."

The risk is not evenly distributed. A blog from Newcastle University's Scholarship Insights team notes that AI-flagging investigations disproportionately catch neurodivergent students and those who write in more formulaic, structured English — including many non-native speakers, whose sentence patterns overlap with what detectors are trained to spot as "AI-like." A widely circulated case at Adelphi University in the US involved an autistic student, Moira Olmsted, whose handwritten essay was scored 100 per cent AI-generated by Turnitin in February 2026. Some US universities, including UCLA and UC San Diego, have already switched their detectors off rather than carry that liability.

A detection score is a prompt to ask a question, not an answer to one.

What JCQ actually requires

UK schools are not operating in a policy vacuum here, and it is worth being precise about what the rules say, because the gap between guidance and practice is where the risk sits. The Joint Council for Qualifications updated its guidance on AI use in assessments in April 2025, and it is unambiguous: if a teacher suspects AI-generated content, the expected first step is a conversation with the student, asking them to explain their thinking and describe their working process — not an automatic penalty triggered by a percentage score. Detection software is presented as one input to that conversation, not a replacement for it.

In practice, that distinction gets lost. A number in a dashboard is easier to act on than a fifteen-minute conversation about drafting process, especially for a head of department managing malpractice cases across a whole cohort. JCQ's most recent published malpractice figures show around 1,125 cases in 2025 where a student lost an entire GCSE or A level, and close to 2,000 more where marks were deducted — figures that span all forms of malpractice, not AI misuse alone, but that give a sense of how much is riding on these judgements being made carefully.

Why this matters beyond exam boards

Non-exam assessment and coursework sit in a greyer space than externally marked papers, and that is exactly where detection tools tend to get used as a first and only check. A department that runs every coursework draft through an AI checker and acts on the headline percentage is, in effect, outsourcing a professional judgement to a tool whose own maker says it isn't built to carry that weight. If a parent asks how a mark was reached, "the software said so" is not going to hold up — in the US it is already being tested in federal court, and there is no reason to assume UK tribunals or ombudsmen would see it differently.

What to do

Treat any AI-detection score as a trigger for a conversation, never as the finding itself. Ask to see planning, drafts, or edit history before raising a malpractice concern, and give students a genuine chance to explain their process. Follow the JCQ guidance's sequence rather than a locally invented shortcut, and make sure every teacher in your department — not just the exams officer — actually knows what that sequence is. Be especially cautious flagging EAL and neurodivergent students, where the evidence on elevated false-positive rates is now fairly consistent.

What to watch

Watch for whether the Kato v. Palo Alto Unified case reaches a substantive ruling on due process, since a US court finding could shape how confidently UK schools can rely on detector scores as evidence. Watch too for whether JCQ or Ofqual issues sharper guidance ahead of the 2027 exam cycle — the current wording leans on professional judgement, and after a year of high-profile false positives, exam boards may come under pressure to be more prescriptive about what counts as sufficient evidence before a malpractice finding is recorded.

academic integrityAI detectionJCQassessmentTurnitinpolicy

NEWSLETTER

Join 10,000 educators

Every week: the AI tools, research, and classroom strategies that matter most. No noise, no hype — just what works.

No spam. Unsubscribe anytime.