
We built an AI detector and binned it
An AI detector is the most tempting thing on our entire roadmap, because from the outside it looks like a solved problem. Competitors ship them, the detection methods are all published and free to implement, and you can have something producing confident looking output before lunch. So I built one, and then I spent days testing it properly against real student work, and the data told me to throw it away.
What we ended up with was not an integrity check at all. It was a filter that selected almost perfectly for inexperience, catching the students who had not thought to cover their tracks while waving through everybody who had. Here is everything we tried, what it actually caught, and why I think anyone selling you a number on this is overstating what their tool can do.
449 real essays, and 61 fakes
Two piles.
The first was 449 real GCSE and A Level answers pulled from work teachers had already put through DeepMark, covering Business, English Literature, English Language, History, Geography, Religious Studies and Food Preparation. Real writing, with real spelling in it. "Manufacter". "Spechial occasions". "Persentage". One pupil had written out a full essay plan before their answer, with arrows and shorthand, exactly as they had been taught to do.
The second was 61 GCSE answers written by AI, and we did not sit there prompting until we got something that looked suspicious. We had three different AI assistants write them, in three different ways, and we deliberately did not tell any of them what we were testing for. That last detail matters more than anything else in this post, and I will come back to it.
The three ways were the three ways a student actually uses ChatGPT, or anything like it, on a piece of homework or coursework:
- Paste the question in and ask for an answer.
- Ask for a grade 9 model answer with PEE paragraphs and embedded quotations.
- Ask for an answer that sounds like a real Year 11 wrote it.
The four AI detection tricks everyone recommends
If you search for how to spot AI writing you will find the same four ideas everywhere, and they are appealing because they cost nothing and need no AI at all. We built all four.
Perplexity and burstiness. The theory is that humans write unpredictably, mixing a four word sentence with a rambling thirty word one, while AI supposedly holds closer to an average. Burstiness measures exactly that, the variation in sentence length across a piece.
Phrase spotting. The "delve into" theory. There is a well known list of words that AI reaches for constantly, including tapestry, multifaceted, underscores the importance, and it is important to note that. You search the text and count the hits.
Connective density. AI loves a linking word, so you count however, moreover, furthermore and consequently per hundred words.
Invisible characters. This one is genuinely clever, because text copied out of a chat window carries fingerprints a keyboard cannot produce, including non breaking spaces and zero width characters. You cannot see them, but they are sitting in the file.
Across 61 AI-written GCSE answers, all four methods combined caught one.
That is not a typo. One of them, out of 61, which is a catch rate of 1.6%.
0.429 against 0.392
This is the one I want to show you properly, because it is the method most often recommended and it is the one I was most confident about going in.
Across 381 real student answers, being the ones long enough for sentence variation to mean anything at all, the median burstiness score was 0.429. Across the AI-written answers, it was 0.392.
There is a difference there, and it is nowhere near big enough to use. Student scores ranged from 0.166 all the way up to 0.896, while the AI scores sat between 0.256 and 0.831, almost entirely inside that same span. There is no line you can draw through that which catches machine writing without also catching children, so the theory is not wrong exactly, it is just far too weak to accuse anyone with.
The phrase spotting failed for a more embarrassing reason, which is that my list was a caricature of how AI writes. An AI asked to write about the Treaty of Versailles writes about the Treaty of Versailles, talking about reparations and the war guilt clause, because the subject drives the vocabulary rather than some house style. It does not go delving into rich tapestries. We got one hit across all 61 answers, and that single hit was the entire haul of the whole deterministic approach.
The invisible character trick, for what it is worth, is the only one still standing on paper, and it cannot help here anyway, because the moment a pupil copies an AI essay out by hand into their exercise book every one of those fingerprints is gone.
So we asked an AI instead
No word lists, no arithmetic. Just read the essay, tell us who wrote it, and explain the reasoning.
This worked far better, correctly identifying most of the AI-written answers, and across 40 real student essays sampled from the larger pile it wrongly accused nobody. I was ready to be pleased with that.
Then we split the results by how the essay had been prompted.
| How the student asked | Caught |
|---|---|
| "Answer this question" | 20 out of 20 |
| "Write me a grade 9 model answer" | 18 out of 18 |
| "Make it sound like a Year 11 wrote it" | 5 out of 23 |
One hundred percent. One hundred percent. Twenty two percent.
Pool those three rows together and you get 70.5%, which is exactly the sort of figure a vendor would put on a slide, and it is close to meaningless. It only looks respectable because two thirds of our AI essays were not disguised in any way. Weight it by how a student who is genuinely trying to cheat would actually behave and you are looking at something nearer 22%.
We watched this happen live, because an earlier run of the same test gave 57%, and all we did was add more undisguised essays to the pile before it climbed to 70%. The number moves with whatever mix you feed it, which is a very good reason not to trust anyone quoting one at you, including us.
One sentence, and the whole thing falls over
That third row is the entire story, because the difference between being caught every single time and getting away with it 78% of the time is one extra sentence typed into a chat box:
Write it like a Year 11 student would.
Ten seconds, no technical skill, no special tool, and nothing that a fifteen year old would not work out for themselves within about a week of trying.
Plagiarism was easy to catch. This is not.
It is worth being clear about why this is a fundamentally different problem to the one schools solved twenty years ago. Plagiarism detection was never really detection at all, it was matching, and Turnitin worked because a copied paragraph exists somewhere else and can be pointed at directly. There was a source document, a percentage overlap, and a screen you could turn round and show to a parent.
AI writing has no source. Every machine written essay is genuinely original, in the narrow and unhelpful sense that nothing identical to it has ever existed before, so there is nothing to match it against and all anyone can do is guess at the odds from the shape of the prose.
Teachers are not wrong that they can often feel it. A language model picks each word by working out which one is statistically most likely to come next, and it does that several thousand times in a row, which is why so much AI writing has that flat, evenly weighted, faintly committee-written quality that experienced markers pick up within a paragraph. The trouble is that this is a tendency rather than a fingerprint, and it is the first thing to disappear the moment somebody asks the model to write a bit more scruffily.
Which is precisely what that third row of the table is. The same instruction that defeats our detector is the one that defeats the sense in the back of your head, and neither of you gets any warning that it has happened.
It catches the wrong children
Here is the part that actually settled it for me.
A tool with these numbers does not catch cheating, it catches students who did not think to hide it. The pupil who added one sentence sails straight through, while the pupil who did not know to add it gets flagged, so what you have built is not an integrity check but a filter that selects almost perfectly for inexperience. The students it catches are disproportionately the ones with less exposure to these tools, less confidence with them, and fewer people around them explaining how they work.
Because it produces a confident looking output it would also be trusted, and a teacher would quite reasonably assume that no flag means no problem. It does not mean that. It means nothing at all.
I would rather ship nothing than ship that.
The number we are not going to print
We got no false positives at all across those 40 real student essays, with not one child wrongly accused and no confident wrong answers anywhere in the set.
That was not luck, it was the one thing we designed for from the very beginning. We decided early that this tool would never output a percentage, never give a verdict on its own, and always show the evidence pointing the other way alongside anything incriminating, because a miss costs nothing while a false accusation costs a child.
I am not going to dress that up as a safety guarantee though, because 40 essays is not enough to make one. Zero out of 40 is statistically consistent with a real false positive rate as high as one essay in eleven, and if I put that number on a marketing page I would be doing the exact thing this post is warning you about. The instinct held up. The sample was simply too small to prove it, and the thing it was protecting turned out not to work anyway.
This is not a lonely finding, incidentally. OpenAI quietly withdrew its own AI Text Classifier in 2023 because of low accuracy, several universities have since switched AI detection off in their marking systems over false positives, and Stanford researchers found detectors flagging over half of the essays written by non native English speakers as machine written while barely flagging native speakers at all. If you teach EAL students, that last finding should stop you cold, because it is the same inexperience filter showing up in somebody else's data.
What to do on Monday morning
I do not have a neat answer for you, but I do have some honest ones.
Ask for the process, not just the product. Plans, drafts, notes, a photo of the whiteboard. The essay itself is trivially easy to generate, whereas three weeks of visible thinking is not.
A conversation beats a percentage every time. Two minutes spent asking a student to explain their third paragraph will tell you more than any tool currently on the market, and it is considerably fairer, because they get to answer.
Write in class sometimes. Low tech, thoroughly unglamorous, and still the only genuinely reliable method anyone has come up with.
Watch this space on provenance. The one avenue we have not written off has nothing to do with the words at all. A Word document quietly records how long it was open and edited, and across how many separate sittings, so an essay built up over ninety minutes and four days looks nothing whatsoever like one that existed for ninety seconds. That is evidence about how the document was made rather than a guess about how it reads, which makes it a far more honest thing to put in front of a parent. We are looking at it, and we will tell you if it works, just as we will tell you if it does not.
So we binned it
There is real money in telling teachers that a machine can spot cheating, and that is exactly why the claim deserves testing rather than repeating. We tried to build it with real student work and a real test, and what we ended up with was a method that catches the careless while waving everybody else straight through, which is not an integrity tool by any definition worth the name.
You deserve to know that before you trust anything with a percentage attached to it, and that includes anything we might have shipped.
See it live
This is DeepMark marking a real script
Scroll through as it marks. Feedback and annotations appear as they land. Click anything, highlight, edit a mark to get a feel for how it works.