Generic essay rubrics reward exactly what large language models do best. The fix isn't a detector — it's reweighting the rubric toward the things AI can't fake.
Ask a large language model to write a five-paragraph essay on almost any assigned topic and it will hand back prose that is fluent, correctly punctuated, logically ordered, and free of spelling errors. Now look at a typical essay rubric: it awards points for a clear thesis, coherent structure, grammatical accuracy, and smooth transitions. The uncomfortable overlap is the whole problem. A generic rubric is, in effect, a checklist of the exact things current models do well — which means an AI-written essay can score highly against it without the student having thought about the subject at all.
The instinct is to reach for an "AI detector" and police submissions after the fact. That instinct is a trap, for reasons covered below. The durable answer is to redesign the rubric so that a high score requires evidence of a process a model can't reproduce: engagement with specific sources, personal or local knowledge, drafts made in the room, and reasoning tied to what actually happened in the course. This article walks through why generic rubrics fail and how to reweight them.
Why generic rubrics reward AI
Language models are, at their core, fluency machines. They were trained to produce text that reads well. So the criteria most rubrics lean on hardest are the criteria they satisfy most reliably:
- Grammar and mechanics. Near-perfect by default. Grading heavily on this rewards the tool, not the learner.
- Structure and organisation. The five-paragraph form, topic sentences, signposted transitions — models produce textbook structure on request.
- Clarity and fluency. Readable, confident prose is the model's native output.
- Coverage of the "expected" points. For any common assignment, the standard talking points are well represented in training data.
Meanwhile, the things models genuinely struggle with tend to be under-weighted or absent from the rubric:
- Accurate, specific citation. Models routinely invent plausible-looking sources, page numbers, and quotations. Real, verifiable references to assigned readings are hard to fake.
- Genuinely personal or local detail. A model doesn't know what happened in your classroom, in the student's town, or in last week's lab.
- Reasoning tied to course-specific material. A response that engages a framework introduced in week seven, or contradicts a claim made in a specific lecture, signals real participation.
- Original synthesis under constraint. Connecting two ideas the course never explicitly linked is harder to outsource than reciting a standard argument.
The takeaway is not that grammar and structure don't matter. It's that when they dominate the point allocation, the rubric is measuring surface polish — the one dimension where a student and a model are indistinguishable.
Skip the detectors — they don't work well enough to grade on
Before reweighting, it's worth being blunt about the shortcut. AI-text detectors are not reliable enough to make grading or misconduct decisions. They produce false positives — flagging human-written work as AI — at rates high enough to harm real students, and false negatives are trivial to induce with light paraphrasing.
The problem is well documented. OpenAI withdrew its own "AI Text Classifier" in 2023, citing a low rate of accuracy. Multiple university teaching centres now advise faculty against relying on detector scores in academic-integrity cases, and Turnitin's own materials acknowledge a false-positive rate on its AI-writing indicator. Research has repeatedly found that these tools disproportionately misflag text from non-native English writers, whose more formulaic phrasing reads as "machine-like" to a classifier.
Put together, that means a detector score is not evidence you can defend in a grade dispute or a conduct hearing. Treat it, at most, as a private prompt to look more closely — never as a verdict, and never as something you show the student as proof. The rest of this article assumes you are not grading on detection at all.
Design assessments the model can't easily fake
The most effective move is upstream of the rubric: build the assignment so that AI-generated text simply can't satisfy it. A rubric can only reward process evidence if the assignment demands process evidence. Practical levers:
- Require citations to specific assigned sources. Not "cite three sources" but "engage the argument in this chapter and quote the passage you're responding to, with page number." Fabricated citations become obvious when they must map to a fixed, known reading list.
- Ask for personal or local examples. An example drawn from the student's own experience, town, workplace, or fieldwork is something a model can't supply and a marker can sanity-check.
- Grade in-class drafts and revision. Collect a handwritten or in-room draft, then have the polished version show visible development from it. The trajectory is the evidence.
- Tie prompts to course-specific material. Reference a specific lecture, a class discussion, a dataset the class collected, or a guest speaker. Generic prompts invite generic (and outsourceable) answers.
- Add a short oral defence. A two-minute conversation about the argument — "walk me through why you chose this source" — reveals authorship faster and more fairly than any classifier. Weight it in the rubric so it counts.
None of this is about surveillance. It's about assessing the thinking, not the artifact. A student who wrote the essay can defend it; the assessment simply makes that defence part of the grade.
A sample reweighted rubric
The table below takes common criteria, names what AI does well against each, and suggests how to reweight or redefine so the points land on the human contribution. Treat it as a starting template, not a fixed scheme — the weights depend on your subject and level.
| Criterion | What AI does well | How to reweight |
|---|---|---|
| Grammar & mechanics | Near-perfect by default | Cut to a small pass/fail band; stop making it a scoring lever |
| Structure & organisation | Textbook five-paragraph form on request | Reduce weight; reward structure that serves an original argument, not the template itself |
| Fluency & clarity | Confident, readable prose | Keep modest; pair with a "voice consistent with in-class writing" check |
| Use of sources | Plausible-looking but often fabricated citations | Raise weight; require verifiable quotes from the assigned reading list with page numbers |
| Specific / local evidence | Cannot supply genuine personal or local detail | Add as a high-weight criterion; demand concrete examples a marker can check |
| Course-specific reasoning | Generic argument, weak on class context | Add; reward engagement with a named lecture, dataset, or discussion |
| Process & defence | No draft trail; can't answer follow-ups | Add as a scored component: in-class draft + short oral walkthrough |
Notice the shift in centre of gravity. In a generic rubric, the first three rows carry most of the points. In the reweighted version, they shrink to a competency floor and the bottom four rows — the ones a model can't satisfy — carry the grade.
Tool walkthrough
Turning that reweighting into a usable, shareable rubric is where Toolhub's rubric generator helps. It builds a structured criteria-and-levels table you can adapt: drop the mechanics weighting down to a pass/fail band, promote source verification and local-evidence criteria to the top rows, and add explicit process and oral-defence components so students see from the outset that the draft trail and the follow-up conversation count. Because the output is a plain criteria grid, it's easy to hand to students up front — which is itself a deterrent, since it signals that polish alone won't score.
The other tool worth pairing in is the readability checker. It is emphatically not an AI detector and should never be treated as one — a readability score cannot tell you who wrote something. What it can do is give you a neutral baseline for your own class: run a few known in-class writing samples through it to see the typical readability range for your students, so that a submission wildly outside that band becomes a private cue to have a conversation, not an accusation. Used that way — as a prompt to look closer, never as proof — it supports the human judgement the rubric is built around.
Where to read further
- Vanderbilt University: guidance on AI detection, and why we disabled Turnitin's AI detector — a university teaching centre's account of detector unreliability and false positives, and what to do instead.
- Wikipedia: Rubric (academic) — background on rubric design, criteria, and performance-level scales.
- Wikipedia: Authentic assessment — the broader case for assessing process and applied reasoning rather than a single polished artifact.
You will not out-detect a language model, and you shouldn't try — the tools that promise it are wrong often enough to hurt honest students. What you can do is grade the parts of the work a model can't produce: verifiable engagement with real sources, evidence that is specific to the student and the course, and a process the student can stand behind and explain. Reweight the rubric toward those, and the question of "did a model write this?" mostly stops mattering — because a model can't earn the points that count.
← All articles