
GCSE grades are often a coin flip
We are 96% confident your 5 is a 4, a 5 or a 6. Imagine a pilot announcing you will be landing at Gatwick, give or take a mile.
Deep in gov.uk, the exam regulator has quietly published children's odds.
Not in a press release. In a 2018 PDF, written for an audience of roughly eleven people, at least four of whom I suspect were the authors. It has been sitting there ever since, doing nothing, like a smoke alarm nobody wired in.
Here is the bit that should have made the news.
"The probability of receiving the 'definitive' qualification grade varies by qualification and subject, from 0.96 (a mathematics qualification) to 0.52 (an English language and literature qualification)."
If you are the sort of person who reads a decimal and immediately feels a jolt of excitement, you have already put your tea down and reached for the biscuits. If not, let me translate.
What that means looking at a group of hopeful children
Exam boards have a concept called the definitive grade. It is the grade a senior examiner, taking their time, says is correct. The gold standard.
Ofqual measured how often students actually get it. In Maths, 96% of the time. In the essay subjects, only between 52% and 60%.
So picture your Year 11 English class. Thirty of them. Blazers on, one radiator, optimistically some shirts tucked in.
Roughly 12 of those 30 will receive a grade the exam board's own senior examiner would not have given them. Not because anyone was lazy or cheated. Because a different qualified adult read the same essay and saw it differently.
Now the Maths class is next door. Same 30 kids, same school, same summer.
About 1.
Twelve in one room. One in the next. The difference is which subject a 14 year old picked at an options evening while mainly thinking about who else was picking it.
Two notes so nobody accuses me of fiddling the numbers. The 0.52 is combined English Language and Literature, an A level, so GCSE English teachers do not get to claim that one. Yours sits nearer 0.6, which I appreciate is not the reassurance you were hoping for. And to be clear, this is not 40 people in a hall in Wolverhampton. It is 453 components, 16.4 million marking events, all 4 exam boards, collected during live marking.
The reassurance is worse than the accusation
Raise this with anybody official and you get the same reply, which is that grades are accurate to within one grade. And they are. The probability of landing on the definitive grade or one either side is above 0.95 everywhere, and for the essay subjects it runs between 0.96 and 0.99.
Sit with the shape of that sentence, since it is genuinely offered up as a comfort.
We are 96% confident your 5 is a 4, a 5 or perhaps even a 6.
Your blood test was accurate to within one diagnosis. Your mortgage rate is right, give or take 2%. The pilot has landed at approximately Gatwick, give or take a mile. The butcher smiles and says that is 1kg of beef, give or take a bag of sugar. You have got to be joking.
No joke.
I am being unfair, and I will explain why shortly. But the instinct that something has gone slightly mad here is a sound instinct, and every English teacher has had it while handing back results and saying something encouraging about resilience.
What this looks like from a chair in Year 11
We will call a student Sarah, and she needs a 6 to do English at sixth form. She gets a 5.
Her teacher, who has read every essay Sarah has written for two years, is surprised. The department pays to send the script back. It returns unchanged, and quickly, with the swift confidence of a system that has not changed its mind because it was never asked to.
Here is what is wrong with that story. There is a decent chance nothing went wrong.
The 5 was a legitimate mark. Another examiner, equally trained, equally conscientious, equally tired, would have given a 6, and that would have been legitimate too. Both defensible. One of them on her certificate. She is not told any of this. Sarah is told she got a 5.
The thing nobody says out loud about judging
Why does Olympic gymnastics use a panel of judges, throw away the highest score, throw away the lowest, and average the rest?
Because everyone involved accepts, openly and without embarrassment, that one human watching one performance once will sometimes be off. Not corrupt. Not incompetent. Human, on that day, in that seat. So the sport builds the wobble into the design and cancels it out.
Diving does it. Dressage does it. Figure skating does it.
Your child's English Literature essay gets one judge. One person, one reading, one Sunday evening, and whatever mark they land on follows that student into a sixth form application and much of their future.
We have known the fix forever, and it is also impossible. Use more than one marker and compare. The reason we do not is not scientific. A second marker for every script in the country costs a fortune nobody has, so we quietly agreed to pretend one look is enough.
That is the real scandal, and it has nothing to do with technology.
Three admissions in the same document
None of this is hidden from you. It is written in a dialect nobody speaks.
On whether the right answer is right:
"in instances where the most frequent mark awarded by examiners differs from the definitive mark there is a possibility that the definitive mark is wrong"
The regulator conceding, in writing, that the gold standard might be plated.
On whether there is one right answer at all:
"although it is possible that there is more than one legitimate mark for some responses, the system does not capture these"
Several defensible marks usually exist. The system picks one, bins the others, then prints the survivor on a certificate as though it had won something.
And then the one that will properly irritate you. Since 2016, boards may only change a mark on review where there has been a clear marking error. A difference of professional judgement does not count. Ofqual defended this by explaining that thinking in terms of right and wrong marks is a misunderstanding, because more than one mark can be fair for the same script.
Which is true. It is also the reason your appeal fails.
The single biggest source of variation, two experts reading the same essay and landing in different places, is the one thing specifically excluded from appeal. You may appeal. You may not appeal about that.
At question level it is no better. On a 6 mark question, two examiners agree on the exact same mark somewhere between 46% and 75% of the time. Six marks. Two professionals. Sometimes barely better than tossing a coin.
Right. So about AI
Every conversation I have about AI marking reaches the same question inside ninety seconds. Is it accurate enough?
Fair question. Half a question.
Accurate compared to what, exactly?
The benchmark is one senior examiner's judgement. Which the regulator concedes may be wrong. Which another equally qualified examiner regularly disagrees with. And which you are not allowed to challenge on the grounds that someone disagrees with it.
We have spent three years holding a new technology to a standard the old one never claimed to meet.
Now the careful bit, because there is a cheap version of this argument and I have watched people in my industry make it.
This is not examiners being bad at their jobs. Marking extended writing is structurally hard. The descriptors say things like perceptive and developed, then invite you to decide, at eleven at night, on script 94, with a cold cup of tea and a dog that needs letting out, whether this paragraph is perceptive or merely quite good. I have marked mocks. I have disagreed with myself on a Tuesday about a script I marked on the Sunday.
The task is hard. That is the finding. Not that the people are careless.
Machines get it wrong, in a completely different way
Human marking error scatters. Ofqual's data shows average differences from the definitive mark close to zero, because generous readings and harsh readings cancel each other out across a big enough pile. Noisy, but not leaning in any particular direction.
Machines do the opposite, and they do it with total confidence.
A Cambridge led study published in May tested 3 frontier models on 761 undergraduate psychology essays from Cambridge, Nottingham and Manchester Metropolitan. The models were extraordinarily steady, giving near identical marks when handed the same essay again, and agreeing more closely with each other than any of them managed with the human examiners. They were also dragged toward the middle. Strong essays marked down, weak essays marked up, least reliable at exactly the top and bottom of the class, which is where everything interesting happens. Asked simply to place each essay in a broad degree band, a first, a 2:1, a 2:2, they matched the human examiners between 35% and 65% of the time depending on the university.
Then the detail that should embarrass the industry. The frontier, tip top, out of the box AI models gave higher marks for longer answers, wider vocabulary and more complicated sentences, whether or not any of it was any good.
They mark the way a Year 11 thinks marking works. Write more. Use bigger words. If in doubt, deploy a semicolon.
That is precisely what levels of response mark schemes were invented to stop a tired human doing at eleven at night. It is not marking. It is a very confident impression of marking, in the way a parrot is a very confident impression of a conversation.
Two things about that study are worth holding onto, though.
The essays were undergraduate, marked against degree classification bands, which are about as broad as descriptors get. And the models were handed an essay and asked for a grade with no mark scheme doing any of the work at all.
That is not marking. That is asking a stranger what they reckon.
School level mark schemes are a different animal entirely. Levels of response, assessment objectives, worked examples of what each band actually looks like. That whole apparatus exists precisely because "is this perceptive" is too vague a question to ask anybody, human or machine, and the scheme narrows it until it becomes answerable.
Which is the argument for building a marking system around the mark scheme rather than around the essay. It is not, to be clear, a claim that we have solved what Cambridge found. Our own accuracy study will answer that, and it is running now. But a study that removed the mark scheme was never testing the thing that makes marking consistent in the first place.
And those two failures are not equally hopeless.
Random scatter cannot be corrected, because nobody knows which script went astray or by how much. A consistent lean can be measured, and anything that can be measured can be adjusted for and checked again. That is not a fudge. It is calibration, and it is how every measuring instrument on earth is made trustworthy.
Consistent and wrong is a fixable problem. Unpredictable is not.
Why AI critically deserves a seat at the table
Nobody wanted Video Assisted Referees in football.
The argument against it was that football is a human game, that referees have judgement, that a machine has no business in it. All reasonable, and all beside the point, because VAR does not referee the match. The referee referees the match. VAR exists because we finally admitted that one person, one angle, one moment, at full speed, will sometimes miss something.
Nobody calls that an insult to referees. It is an acknowledgement that the job is hard and one pair of eyes is not many.
Marking has the same shape and none of the same support.
So put each where they are strong. Machines do not get tired on script 94. They read the same descriptor the same way at eleven on a Sunday night as at nine on Monday morning, and consistency at volume is the one thing humans structurally cannot deliver and the one thing a machine gets for free.
Judgement at the boundary is what machines are worst at, and exactly what an experienced teacher is irreplaceably good at, particularly with their own students.
Consistent first pass. Teacher reads it, disagrees where they disagree, confirms every mark before it exists as a mark at all.
That is not AI marking. It is the first draft, and the second marker every department is supposed to have and nobody can afford.
And the same trick works inside your own department. Standardisation as it really happens is a room in March, 4 teachers, two exemplars chosen by whoever booked the room, forty minutes, and a plate of biscuits that arrived at 3.45 and was gone by 3.47. Everyone leaves reassured, and nobody has the faintest idea whether that agreement survived the other 118 scripts. If the department marks against the same scheme in the same place, you can see where colleagues actually disagree and standardise on those scripts, instead of on two exemplars that happened to be near the photocopier.
What we are and are not claiming
We publish our repeatability figure because it is the thing we can measure honestly. The same 40 mark essay, marked 5 times at production settings, lands within plus or minus 0.94 marks. In plain terms, hand it the same essay five times and the marks land inside a band about one mark wide. Not the right mark, necessarily. The same mark, and as Ofqual states, there is often more than one right mark. Our accuracy study, DeepMark against a trained teacher working from the same scheme, is running now, and we will publish it whether or not we enjoy the result.
We are not claiming DeepMark marks better than your best examiner. The evidence does not support it and I will not pretend otherwise. We are claiming it marks more consistently than any human manages across 30 scripts on a Sunday evening, and that a teacher stays in charge of every mark that counts.
None of this is a scandal about examiners, who are doing a hard job under conditions no sane person would design. It is a scandal about a system that has known for a decade how much variance sits inside its own grades, published it in a PDF written for nobody, and left teachers and students to absorb the difference in silence.
Twelve kids in one classroom. One in the next.
So. Do you feel lucky?
Your students did not get asked.
20 papers free trial, no strings. You will know inside a free period.
Sources
- Ofqual, Marking consistency metrics: an update, November 2018 (Ofqual/18/6449/2).
- Ofqual, Marking mistakes must be corrected, legitimate marks should stand, 2016.
- National Association for the Teaching of English, Making marks: the quality of GCSE and A Level English exam marking, August 2025.
- University of Cambridge, AI in University Assessment: Evaluating the Opportunities and Risks of Automated Marking, May 2026.
See it live
This is DeepMark marking a real script
Scroll through as it marks. Feedback and annotations appear as they land. Click anything, highlight, edit a mark to get a feel for how it works.