Bandly
Criteria

How accurate are AI IELTS writing scores compared with a real examiner?

AI IELTS writing scores are not equally accurate across the four criteria. Where AI marking tracks an examiner, where it drifts, and how to use it.

The Bandly team11 min read

AI IELTS writing scores are not equally accurate across the board: they tend to track a real examiner most closely on grammar and vocabulary, and drift most on Task Response, the criterion that asks whether you actually answered the question. So "how accurate is AI marking?" has no single honest answer — the useful question is how accurate it is on each of the four criteria, and what you should check yourself.

Below: why one accuracy figure misleads, what automated marking can see well, where it is weakest, why "the examiner's score" is not a fixed target either, and how to use an AI estimate without being misled by it.

A student at a desk holding an essay sheet of grey placeholder bars and looking towards an examiner, with an open laptop on the left showing four score bars and the examiner on the right holding a clipboard with four matching bars, dotted lines joining each pair, and the one pair of different heights ringed in thin terracotta

Compare criterion by criterion, not total against total.

Are AI IELTS writing scores accurate? The wrong way to ask

An examiner does not give your essay one mark. For each task they award a whole band on four criteria — Task Response (Task Achievement in Task 1), Coherence and Cohesion, Lexical Resource, and Grammatical Range and Accuracy — and those four are averaged into the task band. Your reported Writing band then combines the two tasks, with Task 2 counting double: (Task 1 + Task 2 × 2) ÷ 3.

That structure is why a single accuracy figure hides the thing you need to know. Suppose an automated marker gets the overall band "right" on an essay. It may have done so by rating grammar half a band too low and Task Response half a band too high. The total matches; the diagnosis is wrong in two places. If you then spend a month on grammar, you have practised the wrong thing on the strength of a score that looked accurate.

The opposite also happens: an estimate can miss the overall band by half a band and still identify the lowest criterion correctly, which is the more valuable result.

So when you see a claim that an AI marker "matches examiners", three questions decide whether it means anything: matches on which criteria, on what kind of scripts, and measured against whose marks? A headline percentage without those answers is not something you can check. This is also why this page contains no accuracy figure for Bandly's own checker: a number we could not show you how to verify would be exactly the kind of claim this article warns about.

A more useful way to think about it is to ask, criterion by criterion, how much of the evidence is on the surface of the text and how much depends on reading for meaning. The more a criterion depends on surface evidence, the more consistent automated marking can be. The more it depends on meaning, the more room there is to drift.

What AI marking gets right: grammar and vocabulary

Grammatical Range and Accuracy and Lexical Resource are the two criteria where most of the evidence sits inside individual sentences.

Grammar is largely countable. The band 7 grammar descriptor reads: "A variety of complex structures is used with some flexibility and accuracy." It adds that error-free sentences are frequent. The band 6 line says "a mix of simple and complex sentence forms is used but flexibility is limited." Whether a sentence contains an error, whether it is simple or complex, and whether the errors cluster in the complex ones are all questions that can be answered one sentence at a time. A machine applies the same standard to the last paragraph as to the first. On this kind of evidence, automated marking can be at least as consistent as a human reader, and sometimes more so.

Vocabulary is mostly visible. The band 7 Lexical Resource wording is: "The resource is sufficient to allow some flexibility and precision. There is some ability to use less common and/or idiomatic items. An awareness of style and collocation is evident, though inappropriacies occur." Repetition, word-choice errors, non-standard collocations and shifts into an informal register can all be found in the text. A phrase like make a crime instead of commit a crime is a pattern, and patterns are what automated marking is good at spotting.

There are limits: precision sometimes depends on the argument around a word, and a marker that rewards "less common" words without checking fit can be fooled by ambitious vocabulary used slightly wrongly. Still, these are the two criteria where an AI estimate is most likely to land where an examiner would.

Coherence and Cohesion sits in the middle. Cohesion — reference, substitution, linking words used or overused — is visible in the text; the band 6 line that cohesion "may be faulty or mechanical due to misuse, overuse or omission" describes something you can point at. Coherence — whether the ideas progress logically across the whole essay — requires following the argument, and that is closer to the judgement discussed next. The linking-word side of this is set out in linking words are keeping you at band 6.

Where AI marking is weakest: Task Response

A student holding an essay sheet of three grey placeholder blocks while an examiner beside them traces lines with a pen from each block to the left half of a propped-up prompt card, the empty right half with no lines reaching it ringed in thin terracotta, a closed laptop pushed to the edge of the table

Task Response is a judgement about meaning.

Task Response asks a different kind of question. Band 7 reads: "The main parts of the prompt are appropriately addressed. A clear and developed position is presented." Band 6 reads: "The main parts of the prompt are addressed (though some may be more fully covered than others)." The difference between those two lines is not in any sentence. It is in the relationship between the whole essay and the question that was set.

To mark it, a reader has to do three things that are hard to reduce to patterns:

  1. Break the prompt into its parts. "Discuss both views and give your own opinion" has three. "Why is this happening, and is it a positive or negative development?" has two, and the second is not the same as "what are the advantages and disadvantages?"
  2. Decide whether each part is answered, or only mentioned. A paragraph can use every keyword from the prompt and still answer a neighbouring question — the causes of a problem instead of its solutions, or the topic in general instead of the specific claim.
  3. Decide whether the position is clear and developed. That means following the argument from the introduction to the conclusion and checking that it holds together.

This is where automated marking is most likely to disagree with an examiner, and the typical error is generosity. A fluent, well-organised, nearly error-free essay produces strong surface signals in three criteria, and those signals can pull the Task Response estimate up with them. But the descriptors mark Task Response on its own terms. An essay that answers a slightly different question can have 7-level grammar and a 5 or 6 in Task Response, and the examiner will not average the problem away. How examiners check relevance is covered in off-topic IELTS essays.

The same gap explains a common complaint: that a general-purpose chatbot rates essays higher than a teacher does. We look at that pattern in a separate article on whether ChatGPT overrates IELTS writing, but the short version is that a model asked "what band is this?" tends to respond to how well the essay reads rather than to whether it does the task.

In Task 1, the equivalent criterion is Task Achievement, and the same logic applies: whether the overview picks out the right key features is a judgement about the data, not about the sentences.

Real examiners vary too: compare the spread, not the "right answer"

The question "is the AI accurate?" assumes there is one correct band to compare against. For many scripts there effectively is. For scripts close to a boundary, there is not.

Examiners mark by best fit: for each criterion they decide which band's description fits the script as a whole. Real essays are uneven — two developed body paragraphs and one thin one, precise vocabulary with a couple of misfires — and where the evidence points to two adjacent bands, placing the script is a trained judgement. Two trained examiners can make that judgement differently on the same criterion. That is not a flaw in the system so much as the nature of marking writing against descriptive criteria, and it is one reason IELTS offers candidates the option of an Enquiry on Results — a formal request to have the test marked again.

We are not aware of a published official figure for how often examiners disagree, and we will not guess one. What matters for your purposes is the principle: the fair comparison is not "does the AI hit the examiner's number?" but "does the AI's spread look like the spread between examiners?" On a clear band 7 essay, a good estimate should land on 7 in every criterion. On a boundary essay, an estimate of 6 or 7 in one criterion may be no less defensible than a human reader's — and a useful marker should show you that the essay is on the boundary, rather than hiding it behind one confident number.

The same boundary logic explains why two teachers can mark one essay differently without either being careless — a case we cover in a separate article. For what a checker can and cannot judge on a boundary script, see IELTS writing band score checkers: what they can and cannot tell you.

How to use an AI score: find your weakest criterion

Given all of this, the right use of an AI estimate is diagnostic, not predictive. You are not asking it what you will score. You are asking it which of the four criteria is holding the essay down, and then checking its reasoning.

Before you paste anything, write the essay the way you would in the test: to an unseen prompt, in 40 minutes, without polishing. An estimate of your best draft is not an estimate of your test-day script.

Then read the result in this order:

  1. Find the lowest criterion. Ignore the average for now. The lowest criterion is where your next hour of practice should go.
  2. Trust the grammar and vocabulary evidence, but verify it. Open each cited sentence and check it against the descriptor line it is tied to. If a flagged collocation looks fine to you, look it up in a learner's dictionary before dismissing it.
  3. Check Task Response yourself. Write the prompt's parts in a list and put the paragraph that answers each one next to it. If a part has no paragraph, or only a sentence, your Task Response is weaker than any estimate that rated it highly — this is the one criterion where your own check should override a generous number.
  4. Look for boundary signals. If the reasoning for a criterion mixes 6-level and 7-level evidence, treat the band as unsettled and work on the 6-level evidence first.

If the same criterion comes out lowest on two different essays, it is a habit, not an accident. How to get band 7 in IELTS Writing sets out what has to change in each criterion, and the full descriptor wording is in IELTS Writing band descriptors explained.

FAQ

Are AI IELTS writing scores accurate compared to a real examiner?

Not equally across the four criteria. AI marking tends to be steadiest on what is visible sentence by sentence — grammatical errors, sentence structures, collocations — and least steady on Task Response, which depends on judging whether the essay actually answers the prompt. Ask how accurate it is on each criterion, not overall.

Can an AI score replace an official IELTS writing band?

No. An AI score is an estimate of one essay against the public band descriptors. An official band comes only from a certified examiner marking your test scripts, and it combines Task 1 and Task 2 with Task 2 counting double.

Why does an AI give my essay a higher band than my teacher?

The most common reason is Task Response. An essay can be fluent and nearly error-free while answering a slightly different question from the one set, or developing one part of the prompt much less than the other. That is the judgement automated marking is most likely to get wrong, and it usually errs on the generous side.

Do human IELTS examiners always agree with each other?

Not always on scripts that sit near a boundary between two bands. Examiners are trained to place a script by best fit, and where the evidence is split, two trained readers can land on adjacent bands for a criterion. IELTS offers a remark service for results a candidate wants checked.

What is the best way to use an AI IELTS writing score?

Use it to find your lowest criterion, not to predict your result. Check the sentences it points to against the descriptor wording, and give extra scrutiny to its Task Response estimate by checking each part of the prompt against your paragraphs yourself.

Are the band estimates on this site official?

No. They are estimates made against the public band descriptors. Only a certified examiner produces an official band.

All descriptor wording quoted here is from the public Writing band descriptors updated in May 2023, published by IELTS and the British Council.

Frequently asked questions

Are AI IELTS writing scores accurate compared to a real examiner?
Not equally across the four criteria. AI marking tends to be steadiest on what is visible sentence by sentence — grammatical errors, sentence structures, collocations — and least steady on Task Response, which depends on judging whether the essay actually answers the prompt. Ask how accurate it is on each criterion, not overall.
Can an AI score replace an official IELTS writing band?
No. An AI score is an estimate of one essay against the public band descriptors. An official band comes only from a certified examiner marking your test scripts, and it combines Task 1 and Task 2 with Task 2 counting double.
Why does an AI give my essay a higher band than my teacher?
The most common reason is Task Response. An essay can be fluent and nearly error-free while answering a slightly different question from the one set, or developing one part of the prompt much less than the other. That is the judgement automated marking is most likely to get wrong, and it usually errs on the generous side.
Do human IELTS examiners always agree with each other?
Not always on scripts that sit near a boundary between two bands. Examiners are trained to place a script by best fit, and where the evidence is split, two trained readers can land on adjacent bands for a criterion. IELTS offers a remark service for results a candidate wants checked.
What is the best way to use an AI IELTS writing score?
Use it to find your lowest criterion, not to predict your result. Check the sentences it points to against the descriptor wording, and give extra scrutiny to its Task Response estimate by checking each part of the prompt against your paragraphs yourself.
Are the band estimates on this site official?
No. They are estimates made against the public band descriptors. Only a certified examiner produces an official band.

Find out where your own writing sits

Bandly grades your essay or letter against all four criteria and shows the gap to your target band, sentence by sentence.

Keep reading

All articles