AI Grading Tools Miss Half the Marks Without a Real Rubric
A University of Georgia study found an AI grader was accurate just 33.5% of the time on its own, rising to only 50% with a detailed rubric — evidence against unsupervised AI marking.
Grading is the task most teachers hope AI will take off their plate. Gallup's 2026 research puts the average teacher's marking load at 5.9 hours a week, and vendors selling AI grading tools now claim roughly 72 percent of schools use some form of automated marking. The pitch is simple: upload student work, get a scored draft back in seconds.
A study out of the University of Georgia, published in the journal Technology, Knowledge and Learning, gives that pitch a hard number to reckon with. Researchers fed middle school science responses — students explaining what happens to particles when heat energy transfers between them — to the large language model Mixtral. Left to grade on its own, without a rubric written by a teacher, the model scored the answers accurately just 33.5 percent of the time. Given a detailed human-made rubric, accuracy rose, but only to about 50 percent.
That is the headline finding, and it should temper any assumption that AI grading is a solved problem in 2026, however polished the marketing around it looks.
Why the model kept getting it wrong
The UGA team, whose findings were reported by UGA Today and the university's Mary Frances Early College of Education, found the model wasn't reasoning through student answers the way a teacher does. Instead, it leaned on shortcuts: spotting a keyword like "particles move faster" and assuming the student understood the underlying concept, even when the reasoning around that phrase was muddled or wrong.
This matters because science and humanities answers rarely hinge on a single correct phrase. A student can use the right vocabulary while misunderstanding the mechanism, or explain the mechanism correctly using imprecise language. Human markers weigh both. The study suggests Mixtral, even with a rubric in hand, struggled to replicate that judgement consistently — it needed rubrics that encoded the specific analytical steps a human grader takes, not just a list of criteria.
There is a genuine upside in the same research: teachers who used the AI-assisted process reported it freed up time for what the study's lead author described as more meaningful work, and grading did move faster. The tool is not useless. It is unreliable as a sole grader, and reliable mainly as a first-pass drafting aid that a teacher must check.
The rubric a teacher writes for a human marker to interpret is not automatically a rubric an AI model can execute accurately — the two need to be written differently, and even then the ceiling on accuracy in this study stopped at one in two.
What this means for schools buying grading tools
Marketing copy for AI grading products routinely cites "time saved" without addressing accuracy against a defensible rubric. The UGA figures — 33.5 percent unsupervised, 50 percent with a rubric — are from one model on one type of open-ended science question, not a verdict on every product on the market. But they are a rare instance of independently published, peer-reviewed accuracy data in a field mostly driven by vendor claims, and they land in the range other researchers have found with similarly structured, open-ended responses.
The gap between "AI can grade quickly" and "AI can grade correctly" is exactly where schools tend to get burned: a tool that returns a confident-looking score is more dangerous than one that visibly struggles, because the confident wrong answer doesn't get double-checked.
What to do
- Treat any AI grading output as a first draft, not a mark, for anything beyond simple right/wrong or multiple-choice items. The UGA study's own recommendation is human review of every AI-generated score before it goes on record.
- If you use an AI grading tool, write the rubric for the machine separately from the rubric you'd give a human marker. Spell out the reasoning steps explicitly rather than assuming the model will infer them from criteria alone.
- Ask any vendor for accuracy data against a known human-marked dataset, ideally for question types similar to what your department actually sets, not aggregate "time saved" figures.
- Reserve AI grading for high-volume, low-stakes formative work — homework checks, practice quizzes — and keep summative and exam-style marking under full human control until accuracy evidence improves.
What to watch
Whether accuracy climbs meaningfully as newer models and purpose-built education tools are tested against the same kind of open-ended, reasoning-heavy questions the UGA study used, rather than the multiple-choice or short-factual-answer formats where AI grading already performs well.