IELTS writing band score checkers: what they can and cannot tell you
A band score checker can estimate your four IELTS writing criteria but not predict your test score. What it can judge, what it can't, and how to pick one.
AI IELTS writing scores are not equally accurate across the four criteria. Where AI marking tracks an examiner, where it drifts, and how to use it.
AI IELTS writing scores are not equally accurate across the board: they tend to track a real examiner most closely on grammar and vocabulary, and drift most on Task Response, the criterion that asks whether you actually answered the question. So "how accurate is AI marking?" has no single honest answer — the useful question is how accurate it is on each of the four criteria, and what you should check yourself.
Below: why one accuracy figure misleads, what automated marking can see well, where it is weakest, why "the examiner's score" is not a fixed target either, and how to use an AI estimate without being misled by it.

Compare criterion by criterion, not total against total.
An examiner does not give your essay one mark. For each task they award a whole band on four criteria — Task Response (Task Achievement in Task 1), Coherence and Cohesion, Lexical Resource, and Grammatical Range and Accuracy — and those four are averaged into the task band. Your reported Writing band then combines the two tasks, with Task 2 counting double: (Task 1 + Task 2 × 2) ÷ 3.
That structure is why a single accuracy figure hides the thing you need to know. Suppose an automated marker gets the overall band "right" on an essay. It may have done so by rating grammar half a band too low and Task Response half a band too high. The total matches; the diagnosis is wrong in two places. If you then spend a month on grammar, you have practised the wrong thing on the strength of a score that looked accurate.
The opposite also happens: an estimate can miss the overall band by half a band and still identify the lowest criterion correctly, which is the more valuable result.
So when you see a claim that an AI marker "matches examiners", three questions decide whether it means anything: matches on which criteria, on what kind of scripts, and measured against whose marks? A headline percentage without those answers is not something you can check. This is also why this page contains no accuracy figure for Bandly's own checker: a number we could not show you how to verify would be exactly the kind of claim this article warns about.
A more useful way to think about it is to ask, criterion by criterion, how much of the evidence is on the surface of the text and how much depends on reading for meaning. The more a criterion depends on surface evidence, the more consistent automated marking can be. The more it depends on meaning, the more room there is to drift.
Grammatical Range and Accuracy and Lexical Resource are the two criteria where most of the evidence sits inside individual sentences.
Grammar is largely countable. The band 7 grammar descriptor reads: "A variety of complex structures is used with some flexibility and accuracy." It adds that error-free sentences are frequent. The band 6 line says "a mix of simple and complex sentence forms is used but flexibility is limited." Whether a sentence contains an error, whether it is simple or complex, and whether the errors cluster in the complex ones are all questions that can be answered one sentence at a time. A machine applies the same standard to the last paragraph as to the first. On this kind of evidence, automated marking can be at least as consistent as a human reader, and sometimes more so.
Vocabulary is mostly visible. The band 7 Lexical Resource wording is: "The resource is sufficient to allow some flexibility and precision. There is some ability to use less common and/or idiomatic items. An awareness of style and collocation is evident, though inappropriacies occur." Repetition, word-choice errors, non-standard collocations and shifts into an informal register can all be found in the text. A phrase like make a crime instead of commit a crime is a pattern, and patterns are what automated marking is good at spotting.
There are limits: precision sometimes depends on the argument around a word, and a marker that rewards "less common" words without checking fit can be fooled by ambitious vocabulary used slightly wrongly. Still, these are the two criteria where an AI estimate is most likely to land where an examiner would.
Coherence and Cohesion sits in the middle. Cohesion — reference, substitution, linking words used or overused — is visible in the text; the band 6 line that cohesion "may be faulty or mechanical due to misuse, overuse or omission" describes something you can point at. Coherence — whether the ideas progress logically across the whole essay — requires following the argument, and that is closer to the judgement discussed next. The linking-word side of this is set out in linking words are keeping you at band 6.

Task Response is a judgement about meaning.
Task Response asks a different kind of question. Band 7 reads: "The main parts of the prompt are appropriately addressed. A clear and developed position is presented." Band 6 reads: "The main parts of the prompt are addressed (though some may be more fully covered than others)." The difference between those two lines is not in any sentence. It is in the relationship between the whole essay and the question that was set.
To mark it, a reader has to do three things that are hard to reduce to patterns:
This is where automated marking is most likely to disagree with an examiner, and the typical error is generosity. A fluent, well-organised, nearly error-free essay produces strong surface signals in three criteria, and those signals can pull the Task Response estimate up with them. But the descriptors mark Task Response on its own terms. An essay that answers a slightly different question can have 7-level grammar and a 5 or 6 in Task Response, and the examiner will not average the problem away. How examiners check relevance is covered in off-topic IELTS essays.
The same gap explains a common complaint: that a general-purpose chatbot rates essays higher than a teacher does. We look at that pattern in a separate article on whether ChatGPT overrates IELTS writing, but the short version is that a model asked "what band is this?" tends to respond to how well the essay reads rather than to whether it does the task.
In Task 1, the equivalent criterion is Task Achievement, and the same logic applies: whether the overview picks out the right key features is a judgement about the data, not about the sentences.
The question "is the AI accurate?" assumes there is one correct band to compare against. For many scripts there effectively is. For scripts close to a boundary, there is not.
Examiners mark by best fit: for each criterion they decide which band's description fits the script as a whole. Real essays are uneven — two developed body paragraphs and one thin one, precise vocabulary with a couple of misfires — and where the evidence points to two adjacent bands, placing the script is a trained judgement. Two trained examiners can make that judgement differently on the same criterion. That is not a flaw in the system so much as the nature of marking writing against descriptive criteria, and it is one reason IELTS offers candidates the option of an Enquiry on Results — a formal request to have the test marked again.
We are not aware of a published official figure for how often examiners disagree, and we will not guess one. What matters for your purposes is the principle: the fair comparison is not "does the AI hit the examiner's number?" but "does the AI's spread look like the spread between examiners?" On a clear band 7 essay, a good estimate should land on 7 in every criterion. On a boundary essay, an estimate of 6 or 7 in one criterion may be no less defensible than a human reader's — and a useful marker should show you that the essay is on the boundary, rather than hiding it behind one confident number.
The same boundary logic explains why two teachers can mark one essay differently without either being careless — a case we cover in a separate article. For what a checker can and cannot judge on a boundary script, see IELTS writing band score checkers: what they can and cannot tell you.
Given all of this, the right use of an AI estimate is diagnostic, not predictive. You are not asking it what you will score. You are asking it which of the four criteria is holding the essay down, and then checking its reasoning.
Before you paste anything, write the essay the way you would in the test: to an unseen prompt, in 40 minutes, without polishing. An estimate of your best draft is not an estimate of your test-day script.
Then read the result in this order:
If the same criterion comes out lowest on two different essays, it is a habit, not an accident. How to get band 7 in IELTS Writing sets out what has to change in each criterion, and the full descriptor wording is in IELTS Writing band descriptors explained.
Are AI IELTS writing scores accurate compared to a real examiner?
Not equally across the four criteria. AI marking tends to be steadiest on what is visible sentence by sentence — grammatical errors, sentence structures, collocations — and least steady on Task Response, which depends on judging whether the essay actually answers the prompt. Ask how accurate it is on each criterion, not overall.
Can an AI score replace an official IELTS writing band?
No. An AI score is an estimate of one essay against the public band descriptors. An official band comes only from a certified examiner marking your test scripts, and it combines Task 1 and Task 2 with Task 2 counting double.
Why does an AI give my essay a higher band than my teacher?
The most common reason is Task Response. An essay can be fluent and nearly error-free while answering a slightly different question from the one set, or developing one part of the prompt much less than the other. That is the judgement automated marking is most likely to get wrong, and it usually errs on the generous side.
Do human IELTS examiners always agree with each other?
Not always on scripts that sit near a boundary between two bands. Examiners are trained to place a script by best fit, and where the evidence is split, two trained readers can land on adjacent bands for a criterion. IELTS offers a remark service for results a candidate wants checked.
What is the best way to use an AI IELTS writing score?
Use it to find your lowest criterion, not to predict your result. Check the sentences it points to against the descriptor wording, and give extra scrutiny to its Task Response estimate by checking each part of the prompt against your paragraphs yourself.
Are the band estimates on this site official?
No. They are estimates made against the public band descriptors. Only a certified examiner produces an official band.
All descriptor wording quoted here is from the public Writing band descriptors updated in May 2023, published by IELTS and the British Council.
Bandly grades your essay or letter against all four criteria and shows the gap to your target band, sentence by sentence.
A band score checker can estimate your four IELTS writing criteria but not predict your test score. What it can judge, what it can't, and how to pick one.
You can check if your IELTS essay is band 7 before the exam, but only criterion by criterion. A four-pass self-check, its two blind spots and a timeline.