What AI is actually good at in a classroom — and what it isn't

James Connolly · · 6 min read

A teacher leaning over a pupil's desk in a secondary classroom. A Classfolio live lesson is on the board showing a multiple-choice starter question, and pupils are answering on laptops.
Illustration generated with AI. The Classfolio screens are real; the classroom and the people in it are not.

I expected AI marking to be inconsistent. That was the whole reason I checked it properly instead of trusting it — I went through 350 pupil attempts against it, one at a time — a single assessment, run with my classes across a week. The marking was consistent, the feedback was consistent, and both were about what I would have written if I had marked all 350 myself. Not better than me. On par with me, three hundred and fifty times in a row, which is the part I did not expect.

That is one teacher checking one subject's worth of work, not a study, and I would not claim otherwise. But 350 is enough attempts that I stopped opening each one braced for something stupid.

And it changed nothing about what happens next, which is the point. When a pupil submits, multiple-choice and true/false answers are marked and approved on the spot, because those have a right answer and the machine either matched it or did not. Every AI-marked written answer lands in a different state: needs_review. It waits for a teacher. I could have taken those 350 as licence to let the marks stand on their own. I did not, and I would not now, because a mark on a child's work is a judgement with consequences and the only person who should sign it is the one who knows the child. Being right almost every time is not the same as being the one who should decide.

Most writing about AI in schools argues about whether it should be there at all. That argument has been overtaken by events: it is already in your classroom, in your pupils' pockets, and in half the tools your school pays for. The more useful question is narrower and more practical — which parts of teaching is it genuinely good at, and which parts should it be kept away from?

What it is genuinely good at

Starting things

The gap between a blank screen and a rough first version is where a lot of planning time disappears. A model will produce a serviceable first draft of a lesson, a set of questions or a mark scheme in seconds. It will not be the lesson you teach. It will be something to react to, and reacting is far faster than starting.

Volume without fatigue

Marking thirty short answers is not hard. Marking thirty short answers as carefully at 9pm as you did at 4pm is hard, and that is a human limit, not a professional failing. A model gives the thirtieth answer exactly as much attention as the first. Where the task is repetitive and the standard is consistency, that is a real advantage.

There is a ceiling, though, and it is fairer to say so than to let you find it. AI marking is capped at two hundred answers a day per teacher, counted per written question rather than per pupil — so a class of thirty with four written questions is a hundred and twenty of them. Past that, marking simply stops and the remaining answers sit waiting for you; nothing is lost and nothing is marked badly in a hurry. The cap exists because every one of those answers costs real money to mark, and a product that quietly runs up that bill is one that either raises its price or disappears.

Patience

A pupil who needs the same thing explained a fourth way, at 11pm, before an exam, is asking for something no teacher can always give. Not because we do not want to — because there is one of us and thirty of them. Machines are unbothered by the fourth explanation.

Access

Translation and read-aloud have quietly become very good. For a pupil who arrived in the country in March, or one who reads two years below their chronological age, that difference is not a productivity gain — it is whether they can take part in the lesson at all.

What it is bad at

Knowing when it does not know

This is the one that matters most, and it is not really fixable. A model produces a confident, fluent, well-structured answer whether or not it has any basis for it. There is no wobble in its voice when it is guessing. Every other weakness on this list is manageable; this one has to be designed around.

Judgement about a particular child

A model can tell you an answer is worth three marks out of five. It cannot tell you that this pupil has written more here than in the whole of last term, and that the mark matters less today than the fact they tried. That is the actual job, and it is not a data problem.

Anything requiring exactness it was not given

Ask an image model for a labelled diagram and you will get confident, misspelled labels. We ask ours for the fewest words it can manage, precisely because the pictures are good and the writing inside them is not. Ask a model to write for Year 3 without telling it the year group and it will write for a fifteen-year-old, every time.

The principle: put AI where being wrong is cheap

That is the whole of it. AI is safe where a mistake is visible, reversible and caught by someone who would have done the work anyway. It is dangerous where a mistake is invisible, lands on a child, and nobody is looking.

A useful test before letting AI touch anything in school: if this is wrong, who notices, how quickly, and what does it cost? If the answer is "nobody, eventually, quite a lot" — do not automate it.

How that shapes Classfolio

Every AI feature in the product is built on the same rule, and you can check each of these:

  • AI marks are never final. A multiple-choice answer is marked and approved instantly, because it has a right answer. An AI-marked written answer arrives as "needs review" and waits for a teacher to agree, adjust or overrule it.
  • AI-generated lessons are flagged, not blocked. When generated slides look thin, the lesson is saved and marked for review rather than thrown away — you decide, not the quality checker.
  • Anything the assistant creates for your planner is marked as needing review before it counts as planned.
  • AI never sees more than it needs. Pupil work is not reachable through our AI integrations at all, and there is a test in the codebase that fails if that ever changes.

None of this is caution for its own sake. It is that the useful version of AI in teaching is the one that does the first eighty per cent of a job and then gets out of the way — not the one that pretends to do the last twenty as well.

So: is it worth it?

Yes, with a boundary. Used for first drafts, for volume, for patience and for access, AI gives back time that goes straight into the parts of teaching only a teacher can do. Used as a decision-maker about individual children, it is a liability dressed up as a saving.

The tools worth adopting are the ones that know the difference and are honest about which side of the line each feature sits on.


Classfolio is a teaching platform built by a teacher. See what it costs, or read why it exists.