In short
- The CEFR says what a reader can do at each level; it gives no word lists or grammar lists for texts.
- A level is an estimate built from three checks: the share of known words, the structures used, and features of the text such as length, topic and how much is left unsaid.
- Public reference tools help, but they were built from learners' writing and from word frequency, so they guide the decision and do not make it.
- Two tools can give one text two levels, which is why a teacher's reading and a trial with pupils still matter.
A label such as 'B1' on a story looks like a measurement. It is closer to a judgement. Someone has decided that a reader at B1 will follow this text without too much struggle. This article explains how that decision is made and how far you can trust it. It works through one short passage at B1 and at A2, describes the public reference tools any school can use, and explains why two tools can disagree.
The short answer
A story is levelled to the CEFR by estimating which readers can follow it. The CEFR describes what learners can do at each level, not what a text must contain. So the person levelling a story checks three things. The words: what share of them a reader at that level is likely to know. The grammar: which structures appear. The text itself: its length, its sentence length, how familiar the topic is, and how much the reader must infer.
The result is an estimate, not a measurement. Good practice adds a fourth step. A person reads the text with a particular class in mind, and a few pupils at the target level try it.
Why does the CEFR not level texts for you?
The CEFR Companion Volume (Council of Europe, 2020) is written as can-do statements about people. At A2, a reader "can understand short, simple texts containing the highest frequency vocabulary". The scale never says how long 'short' is, or which words are 'high frequency'.
It could not. O'Keeffe and Mark (2017) point out that the CEFR's level descriptors were intuitively derived and were not designed for one specific language. A framework written for every language cannot list English words.
So every levelled story rests on a translation, from a statement about readers to a decision about a text. Claridge (2012) interviewed editors at four major publishers of graded readers. All graded their series by the number of headwords and the structures expected at each level.
Check one: the words
The first check asks what share of the running words a reader at the target level will know. Hu and Nation (2000) tested this with 66 adults on a pre-university English course. They read a 673-word story in which 80%, 90%, 95% or 100% of the words were familiar. The other words had been replaced by nonsense words.
With 80% familiar, no reader reached adequate comprehension. With 90% or 95%, some did but most did not. The authors concluded that around 98% is needed to read fiction without help. That is about two unknown words in a hundred. The figure comes from adults reading one story, so treat it as a guide for school pupils.
To run the check, you need a list of words by level. One public reference is the English Profile Wordlists, later published as the English Vocabulary Profile. Capel (2010) describes how the lists for A1 to B2 were compiled from learners' writing, word frequency and exam word lists. They record what learners at each level do know, not what they should know.
The lists give levels to meanings, not just to words. Capel (2010) takes the verb 'keep'. The meaning in 'May I keep this?' is placed at A2, 'keep doing something' at B1, and 'This product will keep for three days' at B2. A checker that only matches spellings cannot tell these apart. For estimates of how many words learners know, see How Many Words Do You Need at Each CEFR Level?.
Check two: the grammar
The second check lists the structures in the text and asks whether a reader at the level has met them. The public reference is the English Grammar Profile. O'Keeffe and Mark (2017) explain that it was built from the Cambridge Learner Corpus, a large collection of learners' writing. It contains over 1,200 statements of grammatical competence, each tied to a CEFR level.
It records what learners write, not what they can read. A pupil may follow a structure in a story before using it in an essay. Use the profile as a cautious guide. Grammar by CEFR Level: What Learners Meet, A1 to B2 sets out the structures.
Check three: the text itself
Two texts with the same words and grammar can still differ in difficulty. The Companion Volume descriptors suggest what else to look at.
- Length and sentence length. The A1 and A2 descriptors speak of very short and short texts.
- Topic. The A2 descriptors name familiar, concrete, everyday matters.
- What is left unsaid. The B1 descriptors mention "explicitly expressed feelings". Implicit meaning is first named at C1.
- Support. The A1 descriptors mention illustrated stories in which the images help the reader to guess.
Automatic readability formulas cover part of this check. Crossley, Greenfield and McNamara (2008) note that many rest narrowly on surface features. They tested a formula built on word frequency, similarity between sentences and words shared between sentences. Using data from Japanese EFL students, it predicted reading difficulty more accurately than traditional formulas. The authors call the study exploratory.
Natova (2021) offers tools made for the CEFR: a qualitative checklist and figures from online text analysis, adjusted to CEFR levels. She tested them at B1 and B2 only, by asking students to translate texts at their level. Most understood at least 90% of the vocabulary. She presents the tools as a help with a preliminary estimate.
A worked example: one passage at B1 and at A2
The two passages below were written for this article. The levels are our judgement against the descriptors. The passages have not been tried with pupils, which is the step a school would add.
B1 version (86 words)
Although the rain had stopped, the forest path was still wet, so Lina walked slowly. She had promised her grandmother that she would be home before dark. Then she noticed a small dog that was hiding under a bush. It was shaking, and one of its legs looked injured. If she carried it, she would be late. She wrapped the dog in her jacket and kept walking. By the time she reached the village, the street lights were on, but she did not regret her choice.
What makes this B1 and not A2?
- Two events happened before the story starts ('had stopped', 'had promised'), so the reader must hold two points in time.
- Several words are less common: 'noticed', 'injured', 'wrapped', 'regret'. The B1 descriptors expect a reader to work out occasional unknown words from context.
- Lina's choice is never stated. The reader must link the dog, the promise and the last sentence.
A2 version (74 words)
The rain stopped, but the path in the forest was wet. Lina walked slowly. She wanted to be home before dark. Then she saw a small dog under a tree. It was cold and afraid. One of its legs was hurt. 'I can't leave it here,' Lina said. She put the dog in her jacket and walked on. She got home late, but her grandmother was not angry. 'Can we keep it?' Lina asked.
| Feature | B1 version | A2 version |
|---|---|---|
| Average sentence length | About 12 words | About 7 words |
| Verb forms | Past simple, past continuous, past perfect, 'would' | Past simple only |
| Joining words | although, so, that, if, by the time | but, then, and |
| Word choice | noticed, injured, bush, wrapped, regret | saw, hurt, tree, put; 'regret' is removed |
| Lina's choice | Left for the reader to work out | Spoken aloud: 'I can't leave it here' |
| The verb 'keep' | 'kept walking' (keep doing something) | 'Can we keep it?' (have and not give back) |
Look at the row for 'keep'. Both versions use the word, and a checker that matches spellings would treat them alike. The word swaps are our own judgement of which words are more common. A school would check each against a vocabulary reference. For one story at A1, A2 and B1, see What Can an A2 Reader Actually Read?.
Why do two tools give one text two levels?
Put the same story into two checkers and you may get B1 from one and A2 from the other. Neither has to be broken. They can differ in what they count.
- Different word lists. One may be built from learners' writing, another from word frequency in general English.
- Words or meanings. A tool that matches spellings grades 'keep' once. A reference that grades meanings grades it several times.
- What is weighed. A readability formula leans on the length of words and sentences. A vocabulary checker ignores sentence length.
- Where the line is drawn. One tool may label a text by its hardest words, another by most of its words.
Publishers disagree in the same way. Claridge (2012) compared the levels on publishers' websites. One classed books of 400 headwords as A1/A2. Another put books of 1,100 headwords in its A2 list. How to Match Students to the Right Graded Readers explains how to handle this when you choose books.
Why a human check still matters
Tools count what can be counted. They do not know that your pupils have never seen snow, or that a joke depends on a second meaning. A person has to read the text with one class in mind.
- Run the word check. For each word above the level, decide: replace it, teach it first, or leave it because the context explains it.
- List the structures. Look twice at any your pupils have not been taught.
- Read the text aloud. Mark each place where the reader must work something out.
- Ask whether the topic and setting are familiar to this class.
- Give the text to three or four pupils at the level. Ask them to underline unknown words and retell the story.
Classroom example (an illustration, not a case study)
A teacher wants a story for an A2 class. A checker marks it A2. She gives a 200-word extract to four pupils. Each underlines nine or ten words, about 5% of the text. At that density, most readers in Hu and Nation (2000) did not reach adequate comprehension. She keeps the story for guided reading and picks an easier one for reading at home.
A trial is the only check that involves real readers. The guide to CEFR reading levels sets levelling beside placing pupils and reporting progress.
How iRead handles this
In iRead, every story, quiz and game is levelled to the CEFR, from A1 to C2. A pack is a CEFR-levelled set of stories, taken from the iRead library or built by the school from its own books. Vocabulary and quizzes are mapped to CEFR bands, so a B1 story means the same thing whichever pack it comes from. Teachers see retention per word, per game and per class.
The core principle: A CEFR level on a story is an estimate of who can read it. Check the words, the grammar and the text, then let real readers confirm it.
See it in iRead: Every story, quiz and game in iRead is levelled to the CEFR, from A1 to C2. A B1 story means the same thing whichever pack it comes from.
See how iRead worksKey takeaway
Treat a level label as a careful estimate. Know what was checked, know what the tools cannot see, and try the story with a few pupils before you give it to the class.
Frequently asked questions
How do I level a text to the CEFR?
Check three things. What share of the words would a reader at that level know? Which grammar structures appear? How long, how familiar and how explicit is the text? Then read it with your class in mind, and try it with a few pupils at the target level.
Is there a CEFR text level checker I can trust?
An automatic checker gives a useful first estimate. It compares the words with a graded list or measures sentence length. It cannot judge how familiar the topic is or how much is implied. A checker that matches spellings cannot separate the meanings of a word. Confirm the result with readers.
How are graded readers levelled?
Claridge (2012) interviewed editors at four major publishers. All graded their series by the number of headwords and the structures expected at each level. The lists were guides for writers, not strict rules. The headword counts attached to a CEFR level differed widely from one publisher to another.
Why does the same story get different levels from different tools?
Tools use different word lists and weigh different features. One may grade by vocabulary only, another by sentence length. Some match spellings, while a reference such as the English Vocabulary Profile grades separate meanings. A difference of one level does not mean a tool is broken.
Sources
- Council of Europe (2020). Common European Framework of Reference for Languages: Learning, teaching, assessment. Companion volume. Council of Europe Publishing, Strasbourg. rm.coe.int/common-european-framework-of-reference-for-languages-learning-teaching/16809ea0d4
- Capel, A. (2010). A1-B2 vocabulary: Insights and issues arising from the English Profile Wordlists project. English Profile Journal, 1(1), e3. doi.org/10.1017/S2041536210000048
- O'Keeffe, A., & Mark, G. (2017). The English Grammar Profile of learner competence: Methodology and key findings. International Journal of Corpus Linguistics, 22(4), 457-489. doi.org/10.1075/ijcl.14086.oke
- Hu, M. H.-C., & Nation, P. (2000). Unknown vocabulary density and reading comprehension. Reading in a Foreign Language, 13(1), 403-430. doi.org/10.64152/10125/66973
- Claridge, G. (2012). Graded readers: How the publishers make the grade. Reading in a Foreign Language, 24(1), 106-119. doi.org/10.64152/10125/66668
- Natova, I. (2021). Estimating CEFR reading comprehension text complexity. The Language Learning Journal, 49(6), 699-710. doi.org/10.1080/09571736.2019.1665088
- Crossley, S. A., Greenfield, J., & McNamara, D. S. (2008). Assessing text readability using cognitively based indices. TESOL Quarterly, 42(3), 475-493. doi.org/10.1002/j.1545-7249.2008.tb00142.x