From governess to grammar school: what is next for British education
educationAI markingpersonalised learningGCSEedtechteaching

From governess to grammar school: what is next for British education

By Geoff Waugh13 min read

Somewhere in England in 1850, two children of similar intelligence woke up on the same morning and received entirely different educations. The first was the son of a Wiltshire landowner. By nine o'clock he was seated with a governess who had spent six months learning the precise shape of his mind: where he was quick, where he was not, which approach would land and which would not. She caught misconceptions before they became permanent gaps. When he was ready to move on, they moved on. He went to Oxford and did not find it particularly difficult.

The second child walked two miles to a church schoolroom. Fifty children, one teacher, one piece of chalk. The lesson was the same for all fifty. The child who kept up, kept up. The child who did not fell behind quietly and stayed there. This second child was, by most accounts, the more naturally curious of the two. It made very little difference.

This is not primarily a story about inequality, though it is that too. It is a story about a structural failure that two centuries of educational reform have not resolved. We have always known what works. We have never been able to scale it.

The research no one acted on

In 1984, Benjamin Bloom published a finding that should have changed everything. Students receiving one to one tutoring with mastery learning performed two standard deviations above those in conventional classrooms. The average tutored student outperformed 98% of students taught in the standard way. A student who would have scored a C scored an A.

Bloom called this the two sigma problem, not because the result was troubling but because it raised a question nobody could answer: how do you deliver this at scale? For forty years, nobody could. The technology had not arrived. It has now.

The design choice we forgot we made

The Elementary Education Act of 1870 brought two million children into schooling who had none. That was a genuine achievement. It also embedded a model of education built entirely on industrial logic: one supervisor, many workers, standardised process, uniform output. This was not cynical. It was the only model available for anything at scale in 1870.

What is remarkable is that we are still running it a hundred and fifty years later. The batch production classroom, thirty students, one teacher, one curriculum sequence, one terminal examination, was built to solve the access problem. It was never designed to solve the learning problem. Those are different things.

What every other industry figured out

Every large-scale industry has faced the same dilemma. Toyota resolved it in manufacturing by embedding quality at every step of production rather than inspecting failures out at the end. Dell resolved the personalisation problem by building computers to individual specification at industrial volume. Mass customisation: not a compromise between scale and quality, but a third option that made the choice unnecessary. Education has not made this transition. It is about to.

The data problem

Personalisation has always failed in education for one practical reason. It requires data, real time, high quality, individual level data about where each student actually is at any given moment. A good teacher has always had some of this intuitively. They read the room, notice who has gone quiet, catch the look that means something has not landed. But no teacher can hold thirty adaptive learning pathways simultaneously in their head while managing behaviour, delivering content, and completing the administrative demands of the role. This is the specific problem AI addresses: not by replacing the teacher, but by generating the data infrastructure that makes personalisation operationally possible.

The loop works like this. A student is taught. The work is assessed consistently, with annotation that reveals not just whether an answer is right but why it is wrong and exactly where the gap lies. That data feeds back into the tutoring. The next session is calibrated to the actual gap. The loop tightens continuously around each student's real position relative to the learning destination. The governess knew her student completely because she had time to observe, adapt, and respond continuously. The AI loop does the same thing across thirty students simultaneously, without requiring the teacher to carry it all as a manual process.

The false tension

The apparent conflict between personalised learning and standardised outcomes is false. Consistent, high-information marking does not prevent personalisation. It enables it. The GCSE specification is the shared destination. The mark scheme is the map. Consistent annotated marking tells you with precision where each student stands on their individual route to that shared endpoint. Without that data, personalisation is guesswork. You are adjusting without knowing what to adjust toward. The students sit the same exam. The mark schemes are the same. What changes is everything that happens between now and results day.

The honest part

The technology is further ahead than the institutions, and that is where the real work lies. Data governance, exam board validity, parental consent, staff training, procurement cycles, union agreements, and the question of who owns the data generated when a student's work enters an AI system: these are not technical problems. They are political and structural ones, and they will not resolve themselves.

The exam boards face their own uncomfortable version of this. Students already use AI to reverse-engineer question patterns, generate practice papers, and predict likely topics. The traditional unseen examination faces a validity crisis it has not yet publicly acknowledged. What does standardised assessment mean when preparation for it can be systematically automated? That question is not going away.

The teacher

The question that generates the most anxiety in this conversation is also the easiest to answer. The teacher does not become redundant. The role becomes more distinctly human.

The teaching role currently bundles things that require genuine human judgment with things that do not. Marking thirty scripts consistently requires accuracy and subject knowledge. It does not require the capacity to read a room, recognise distress, or make the relational judgment that determines whether a student feels seen or invisible. When the marking infrastructure is handled systematically, the teacher's time and attention concentrate on what only a human being can do: the mentor, the interpreter of data, the person who knows that this student's drop in performance is not a gap in knowledge but a situation at home.

The governess was not a delivery mechanism. She was a sustained human relationship. Mass education replaced her with something that could scale. What comes next preserves the scale and recovers something of the relationship.

The problem that is not yet solved

Before the vision runs ahead of the evidence, there is a reckoning worth making. A Cambridge-led study published in May 2026 tested three of the most capable AI systems on 761 long-form undergraduate essays across three UK universities. The models did not fail because they were inconsistent with themselves. They were highly consistent, and the three systems converged closely with each other. The failure was systematic bias: a persistent central tendency, assigning middling marks to most submissions, least accurate precisely at the grade boundaries that matter most. They rewarded essay length and vocabulary over the quality of argument. They were reliably wrong in the same direction.

This matters because it distinguishes two problems that are often conflated. Repeatability, marking the same script consistently on different occasions, is a necessary precondition for trustworthy assessment. But it does not fix central tendency bias. A perfectly consistent system can still systematically misgrade at grade boundaries. The Cambridge study also used broad degree-classification descriptors rather than the granular, point by point level of response schemes that GCSE marking is anchored to, which is a substantially harder calibration task.

Teacher-in-the-loop tools address this at both levels: granular mark schemes that anchor assessment to precise criteria, and teacher confirmation of every mark before it is recorded. Ofqual's marking consistency research finds that roughly 75% of GCSE grades would be confirmed if scripts were independently remarked. At the mark level, DeepMark's published methodology shows that when the same 40 mark essay is assessed twice, the result falls within plus or minus 0.94 marks. Consistency is the foundation. The mark scheme anchoring and teacher confirmation loop are what convert it into trustworthy assessment.

The toolkit

On the tutoring side of the loop, a new generation of adaptive AI platforms now offers students something previously available only to those whose families could afford private tuition: a conversational tutor calibrated to a specific exam board, generating practice questions adjusted in real time to the student's demonstrated ability, tracking understanding at specification point level, and feeding back without waiting for a teacher to be available. A teacher with access to that data before the lesson begins knows, with specificity, where each of their thirty students stands. Not a show of hands. A map.

The most immediate entry point is homework, where students consolidate and explore with an adaptive tutor before returning to the classroom with a clearer picture of what they know and what they do not. AI is effective at helping students explore ideas, test understanding, and identify gaps. It is considerably less effective at the deeper work: retention, consolidation, the kind of understanding that stays. That distinction shapes how these tools should be used.

Hugh Grant is not wrong

He is answering the wrong question. Close Screens, Open Minds, the campaign group he backs alongside Sophie Winkleman and Jonathan Haidt, has built a movement on a legitimate grievance: 350,000 members, a celebrity platform, and a genuine argument that educational technology was sold to schools as a revolution and delivered, in many cases, distraction. The concerns about screen dependency, fractured attention, and the commercialisation of children's learning time are real. The instinct to protect children from a technology industry that has prioritised engagement over wellbeing is not hysterical. It is reasonable.

What does not hold up is the conclusion. The classroom these campaigners want to return to was the system that produced the problem in the first place: one teacher, thirty students, one pace, one outcome, a structural inability to adapt to any individual child. The chalk and the textbook did not solve the two sigma problem. They created it. Returning to that model is not protecting children from technology. It is returning them to a known inadequacy and calling it safety.

The question has never been screens or no screens. It is what the technology is doing and who is directing it. A child passively absorbing content without a teacher in sight is one thing. A student receiving targeted practice on the precise concept they misunderstood in last week's assessment, while their teacher uses the resulting data to reshape tomorrow's lesson, is another thing entirely. The evidence suggests handwriting aids retention in ways typing does not, and that kinesthetic engagement, discussion, and the social dimensions of the classroom are not inefficiencies. They are what learning is.

The governess was not a delivery mechanism. She was a relationship. No software in 2026 reliably replicates that. But the alternative to bad AI in education is not no AI. It is better AI, under teacher direction, with purpose. AI can do things for student learning that no teacher standing in front of thirty students has ever been able to do: track every student at specification point level, provide consistent feedback at scale, and return the hours currently consumed by manual processing to the teacher's actual work. That is not a threat. It is an amplifier. Teachers are the interface. The glue. The judgment in the room that reads thirty different states and makes the call no algorithm replicates. AI without the teacher is a screen. AI with the teacher is, for the first time, the governess at scale.

The box is open. The technology is in students' pockets, in their homework, in their revision, whether or not any school has a policy covering it. The only question is whether school leaders engage deliberately, with governance and purpose, or inherit whatever arrives without their input.

Where this goes

The governess had one student. The classroom has thirty. For the first time in the history of mass education, the tools exist to begin closing that gap: not by reducing the teacher's role, but by extending the reach of their judgment, generating the data that makes adaptive learning possible, and freeing the hours currently consumed by manual infrastructure.

The research challenges are real, the institutional questions substantial, and the technology is running ahead of the governance frameworks designed to manage it. The schools that move now, that build the data infrastructure and train their teachers to interpret information rather than just generate it, will be in a different position from those that wait for consensus. The two children who woke up in 1850 had different lives because of the circumstances they were born into. The technology now exists to make that particular gap smaller. That is worth acting on.

DeepMark. Built by teachers, for the marking problem.

support@getdeepmark.com

See it live

This is DeepMark marking a real script

Scroll through as it marks. Feedback and annotations appear as they land. Click anything, highlight, edit a mark to get a feel for how it works.

Loading the marking editor…