AI grading that instructors trust: drafts, not decisions
“The model drafts and the instructor reviews” decays into “the model grades” unless the software stops it. Five conditions that keep the loop real.
Ten wpis nie został jeszcze przetłumaczony — wyświetlamy wersję angielską.
Every AI grading rollout begins as “the model drafts, the instructor reviews”. A term later, most have quietly become “the model grades”. Nobody decided that. The review step stayed on the screen; it stopped meaning anything. Whether a person is genuinely in the loop is a property of how the software is built rather than of what the rollout email promised — it comes down to whether a draft can reach a learner without a confirmation, what the model was shown before it wrote, and whether anyone can reconstruct afterwards what was suggested versus what was awarded.
How “the instructor reviews” becomes “the model grades”
The drift is not laziness. It is arithmetic. A hundred and eighty scripts, a marking window that was always two weeks and is now a weekend, and a queue that has already written a score into every box. The first twenty drafts look about right. That forms a prior. By script forty the reading has changed from marking to checking, and by ninety checking has become scanning for anything obviously wrong.
Anchoring does the rest. A number on the screen before you have formed your own is not a neutral starting point. Show a marker a suggested 68 and their independent 74 rarely survives intact; they land at 70, and it feels like their own reasoning. Human-factors research calls this automation bias: people accept a machine's recommendation more readily than a colleague's, and stop checking sooner.
Interface friction finishes the job. If confirming takes one click and disagreeing takes four, the ratio between those numbers is your actual grading policy, whatever the handbook says. An “accept all” button on a queue of 180 is the switch that turns human-in-the-loop grading into automated grading, placed where a tired person will find it at eleven at night.
How do you stop teachers rubber-stamping AI-suggested grades?
Make agreement deliberate and disagreement cheap. Confirm per criterion rather than per submission, so accepting five scores takes five decisions. Remove any bulk-accept control. Do not show a computed total until the criteria are confirmed, because the total is what people anchor on. Then measure the override rate per marker and per cohort, and treat near-total agreement as a warning rather than a success metric — a marker who changes nothing across 180 scripts is not agreeing with the model, they are not reading it.
What has to be true for the loop to hold
Human-in-the-loop grading
A grading design in which model output is a proposal with no release path of its own. The draft goes to a marking queue, never to the learner. A named person confirms each grade before it exists as a grade, the record keeps the suggestion and the award as separate fields, and no control turns a pile of proposals into results in one action. If any of those is missing, the model is grading and a person is watching it grade.
- The draft is never released. There is no code path from model output to a learner-visible grade. Not a setting that ships disabled — no path at all. A toggle marked auto-release for objective criteria will be switched on during results week by somebody who intends to switch it back.
- The instructor confirms every grade, one submission at a time, criterion by criterion, with the work open beside the suggestion.
- The model sees an anonymised submission. No name, no photograph, no class, no previous marks, no earlier comments from this teacher about this learner.
- Suggestions are per criterion against a stated rubric, with the evidence each score rests on — not a single number the marker can only accept or reject.
- The record keeps what was suggested and what was awarded, separately and permanently, along with who confirmed it and when.
Check the first one first. It is the only one that cannot be recovered by good practice: everything else on the list can be undermined by a rushed marker, but a release path can be used by a script.
A draft sits in a queue as five per-criterion suggestions with the text they refer to. An instructor changes two, confirms five, and the result becomes a grade at that moment — attributed to them.
A model writes a score into the gradebook, a notification asks somebody to review it before Friday, and on Friday nobody does.
The model should not know whose work it is
Marking anonymously is an old idea in education and a well-tested one. The AI version has a sharper edge: a model has no memory of a learner unless the system hands it one. It does not know that this pupil has been struggling since October. That ignorance is worth protecting, because the moment you paste in “previous grade: 42” you have built a machine that regresses every learner towards their own history.
Stripping the obvious identifiers is the easy half: name, learner id, photograph, group, prior results, the teacher's earlier feedback, and the file metadata on the upload, which carries an author name more often than people expect. The hard half is that names live inside the work — the reflective essay opening “my name is”, the code file with a header comment. An anonymisation claim that only covers database fields is worth testing with a real submission.
There is a second reason, and it is about the rubric rather than the data. A language model reads surface fluency with great confidence and is far less certain about whether an argument is any good. Asked for one overall judgement, it lets fluency colour everything, which penalises the same people it always penalises: second-language writers, dyslexic learners, anyone whose thinking is better than their prose. Scoring language in the language criterion keeps that contamination in one box.
Is AI grading unfair to students who write in a second language?
The risk is real, and it is mostly a rubric problem rather than a model problem. Models score surface fluency confidently, so a single holistic score lets weak phrasing drag down criteria that were never about phrasing. Score every criterion separately against a rubric that keeps language apart from content, and review override patterns by cohort. If markers consistently raise the model's scores for one group, the model has a bias and the rubric is where you fix it.
Per criterion, with the evidence attached
A single score is unreviewable. “72” makes no claim you can disagree with; you can only feel that it is roughly right, which is exactly the feeling that produces a click. Reviewing is only possible when the output is specific enough to be wrong.
“Three out of four on use of evidence: paragraphs two and four cite the source directly, paragraph five makes the strongest claim in the essay and cites nothing” is a different object. It names a criterion, states a level and points at a location. A marker checks paragraph five in fifteen seconds and either agrees or does not. That is a review. The disagreement is worth something too — it usually means the rubric line is ambiguous.
Rubric shape matters more than people expect. An analytic grid gives the model a described level to match against, which is the easiest thing for it to do well. A single-point rubric — one column describing proficient work, with notes on what fell short and what went beyond — suits drafting better still, because it asks for the deviation rather than a number, and a deviation is a sentence with evidence in it. A holistic rubric is the worst fit: it asks for the opaque overall judgement nobody can review.
One more thing to require: let the model decline. “I cannot assess this criterion from the submission” is more useful than a confident 3, and it lets the queue sort itself — the easy half fast, the hard half slowly.
The audit trail is what survives the appeal
Six months later somebody appeals. The question in the room is rarely “was the AI right”. It is “who decided this”, and a system that stored only the final number cannot answer it — by then the drafts have been overwritten by the grades they became.
So the record has to hold both. Suggested and awarded as separate fields, per criterion. The named person who confirmed, with a timestamp. The rubric version in force at the time, because rubrics get edited. The model and version that produced the draft, recorded at draft time rather than looked up later. And whether the confirmation was per criterion or a bulk action — the field that tells you the truth about your own process.
The data subject shall have the right not to be subject to a decision based solely on automated processing, including profiling, which produces legal effects concerning him or her or similarly significantly affects him or her.
A grade that decides progression, a classification or a professional licence sits comfortably inside “similarly significantly affects”. If an instructor confirmed each criterion with the work in front of them, the decision was not solely automated. If they cleared 180 with one action, a regulator may read it differently — and your logs, not your policy document, decide which story you can tell.
One result, after confirmation
submission anonymised · no name, photo or prior marks sent
rubric "Extended essay" v3 · analytic · 5 criteria
model provider + version, recorded when the draft was written
criterion suggested awarded changed
Thesis and argument 4 4 —
Use of evidence 3 4 yes
Structure 4 4 —
Language 2 3 yes
Referencing 3 3 —
suggested total 16 / 20 never shown to the learner
awarded total 18 / 20 the grade
confirmed by instructor 118 · per criterion · 5 of 5
released by the same person, in the same action
Two totals, two columns, kept forever.
A system that stores only the second cannot answer an appeal.This is the shape Lurno is built to. AI drafts rubric scores and written feedback over anonymised context and puts them in the unified grading queue beside everything else waiting to be marked; an instructor confirms and releases every grade, and there is no mode in which a draft becomes a result on its own. Grading policies are ordered passes — automatic, manual, peer, AI-assisted — with an acceptance rule choosing the final grade, so AI assistance is one pass among several rather than the pipeline. Moderation thresholds hold anything above a set stake for a second marker.
The constraint that matters is the one you cannot configure away: the model has no path to release. How it is fenced in is on AI in Lurno; the hash-chained audit log behind the trail, which cannot be quietly edited afterwards, is on security.
What AI grading is good at, and what it is not
The genuine strength is consistency across volume. A marker at script three and the same marker at script 140 are not the same marker — standards drift, criteria get skipped, the fifth identical mistake gets a shorter comment than the first. A model does not tire that way, and used as a draft it catches what fatigue costs you.
Reliably useful:
- Applying the same rubric to the hundred and fortieth script as to the third.
- Catching a criterion a tired marker left blank — the most common marking error and the least interesting one.
- Flagging the outlier, the submission unlike every other in the batch. That is worth a person's eyes for several reasons, only one of which is the grade.
- Drafting the feedback paragraph. Rewriting a draft is far quicker than composing from nothing, and this is where most of the real time saving comes from.
Not to be trusted, and no amount of prompt work changes it:
- Anything where the judgement is the point — originality, an argument that breaks the rubric because it is better than the rubric, work that is technically weak and intellectually interesting.
- The first answer of its kind in a cohort. Models are good at typical and poor at unprecedented, and a genuinely new approach reads to them as an error.
- Academic integrity. A suspicion of misconduct is an accusation about a person, and it needs an investigator and a process, not a confidence score.
- The pastoral read — the essay that is technically fine and quietly signals that the learner is not.
The honest summary: AI grading does not remove marking time, it moves it. The mechanical part shrinks. The part that needed a person still needs a person. A rollout that banks the saved hours as a headcount saving, rather than redirecting them to the hard pile, will produce worse marking than the process it replaced — and it will take a term for anyone to notice.
Is AI grading accurate enough to replace human markers?
No, and average accuracy is the wrong test. A model can track a rubric's mechanical criteria closely on typical work, which makes headline agreement figures look good. It fails in the tail — unusual, original or borderline work — and the tail is where consequential grades live. Ask a vendor not how often the model agrees with a human, but what happens on the scripts where it does not, because those decide classifications and produce appeals.
What to ask before you switch it on
- Show me the code path from a model draft to a released grade. If one exists behind a setting, it exists.
- What exactly is sent to the model? Ask for a real payload from a real submission, not a diagram.
- Can a marker clear a queue in one action? Try it in the demo tenant yourself.
- Can I report override rate by marker, by cohort and by criterion? That report is the only early warning that review has stopped happening.
The short version: a model that drafts is a good marking assistant and a bad marker, and the difference is structural rather than cultural. No release path, per-criterion confirmation, an anonymised submission, rubric-anchored suggestions with the evidence attached, and a record that keeps the suggestion beside the award. Get those five right and the review step still means something in March, when everyone is tired and the queue is 180 deep.