In short
- No research paper gives one correct number of questions; the purpose of the quiz decides it.
- Longer tests give more reliable scores, but each extra question adds less than the one before.
- For a daily check after one story, five to eight questions is a sensible working range.
- Several short quizzes teach more than one long test, as long as pupils get feedback.
Teachers ask this question because they want a rule, and there is no rule in the research. What the research does give is a way to think about it. A quiz has a purpose: to check that pupils understood today's story, to practise, or to produce a mark that counts. Each purpose needs a different number of questions. This article explains why, in plain words, and ends with a table of working ranges that are judgement, not findings.
The short answer
For a daily check on one short story, five to eight questions is enough, and ten is a ceiling. For a weekly practice quiz across two or three stories, eight to fifteen. For an end-of-term test that decides a grade, twenty to forty items across several texts. These are working ranges, not research findings. No study has compared reading quizzes of different lengths with school-age learners of English and reported a best number. What the research does offer is reasons, and they follow.
Why the purpose decides the number
Hughes (2003) separates two jobs for assessment. Formative assessment checks progress so that the teacher can adjust teaching and give feedback; informal tests and quizzes have a part to play here. Summative assessment measures what has been achieved at the end of a term or year, and here formal tests are usually called for. The two jobs need different numbers of questions.
A daily quiz is formative. Its job is to tell you, and the pupil, whether the story was understood. Two wrong answers out of six is useful information the same day, and nobody's grade depends on it, so a rough score is fine. An end-of-term test is summative. A pupil's report may rest on it, so it needs enough questions to make the score dependable.
Hughes (2003) also notes that the results of a summative test should not be read in isolation, and that a complete view draws on as many sources as possible. Many small quizzes over the term build that record; see formative assessment for reading in EFL.
Why longer tests give more reliable scores
Hughes (2003) puts it simply: a test is reliable if it measures consistently, so a pupil gets much the same score on one day as on the next. Two things spoil reliability. One is the pupil: mood, sleep, luck. The other is the test: unclear instructions, ambiguous questions, and items that invite guessing. Hughes (2003) adds that small-scale testing tends to be less reliable than it should be.
Think of each question as one small sample of what the pupil can do. With three questions, one lucky guess or one careless slip moves the score by a third. With thirty, the same slip moves it by a thirtieth, and luck tends to cancel out. This is why the Council of Europe (2011) manual for test developers lists the number of items as a design decision. Enough items are needed to cover the content and give reliable information, within practical limits on test length.
The same manual gives a rule of thumb for formal tests: a reliability estimate in the top third of the 0 to 1 range, from 0.6 up, is often considered acceptable. It also says something teachers should hear. When the number of items or test takers is low, reliability usually cannot be estimated at all. The manual's advice is then to treat the test as only part of the evidence, alongside other work over time (Council of Europe, 2011).
Why the gain shrinks as the test grows
Adding questions helps, but each new question helps less than the last. The Spearman-Brown formula, published separately by Spearman (1910) and Brown (1910), predicts how reliability changes when a test is lengthened with similar items.
Here is the shape, with an assumed starting point for illustration only. Suppose a five-question quiz had a reliability of 0.50. The formula predicts about 0.67 for ten questions, 0.80 for twenty, 0.89 for forty and 0.94 for eighty. The first doubling adds 0.17; the last adds 0.05 and costs forty more questions. This is arithmetic, not a finding about your class, and it assumes the new questions are as good as the old.
The practical lesson is simple. Going from three questions to eight buys a lot. Going from twenty to forty buys a little, and costs half a lesson.
Age, attention and the length of the text
No study fixes how many questions a ten-year-old can face before attention fades, so treat what follows as classroom judgement. Younger and lower-level pupils read slowly, reread often, and tire sooner, so fewer questions on a shorter text is the safe side. The CEFR reading levels describe an A1 reader as handling very short texts one phrase at a time, and an A2 reader as handling short, simple texts on familiar topics.
Match the number of questions to the length of the text. Day and Park (2005) recommend no more than ten questions for a text of about 600 words, and warn that even a keen pupil gets bored answering twenty questions on a three-paragraph text. A rough working ratio is one question per 80 to 100 words of story for a daily check, and fewer at A1.
Time per question matters more than the count. A multiple-choice question on a fact takes seconds; a short written answer about a moral takes a minute or more. A quick way to set the length:
- Time yourself on the quiz, then allow pupils three to four times as long.
- If a daily check runs past ten minutes, cut questions rather than rushing.
- Cut again if you see the signs below.
- Scores drop on the last questions, whatever their difficulty.
- Pupils stop looking back at the text and start guessing.
- The quiz eats the time you wanted for talking about the story.
- Pupils remember the quiz, not the story.
How to spread question types
Once you know the number, spread it across types. A daily check of six questions might carry three literal questions, one that joins two parts of the story, one inference, and one on the moral. A graded test needs the same spread at a larger scale, across several texts. See reading comprehension question types, with examples.
Whatever the count, keep the text in front of pupils. Day and Park (2005) are clear that comprehension questions test reading, not memory. And avoid tricky wording: a statement with one true half and one false half measures reading of the question, not of the story.
Classroom example: a six-question daily check on a 400-word A2 story
1. Where did the story happen? (literal, multiple choice)
2. What did the girl lose? (literal, short answer)
3. Who helped her, and when? (reorganisation, joins two paragraphs)
4. Why did the shopkeeper smile at the end? (inference, yes or no plus why)
5. True or false: the girl went home before she found the bag. (reorganisation)
6. What does the story teach? (evaluation, one sentence)
Sensible ranges by purpose
| Purpose | Questions | Texts | Why |
|---|---|---|---|
| Daily check after one story | 5 to 8 | One story | Feedback the same day; a rough score is fine |
| Practice between sessions | 6 to 12 | The current story | Low stakes; repetition matters more |
| Weekly or unit quiz | 8 to 15 | Two or three stories | Some weight on the score; spread the types |
| End-of-term graded test | 20 to 40 | Several texts | The score must be dependable |
| Placement or level check | 30 or more | Texts at several levels | A wrong placement is costly; take the time |
Read the table as a floor and a ceiling for each purpose, not a target.
Several short quizzes beat one long test
Black and Wiliam (1998) report a meta-analysis by Bangert-Drowns, Kulik and Kulik (1991) of 40 studies of frequent classroom testing. Performance improved with more frequent testing, up to a point somewhere beyond one or two tests a week, and several short tests were more effective than fewer longer ones. Bangert-Drowns, Kulik and Kulik (1991) themselves report that the gain from each added test grew smaller. Pupils who met the same items over many short tests also did better than those who met them in a few long ones.
Two cautions. The studies were mostly college courses in mathematics, science and social science, not English reading in schools. And Black and Wiliam (1998) note that such reviews ignore the quality of the questions; a quiz never used for feedback is just frequent testing. The lesson still holds: prefer a short quiz after every story to a long test once a term. See why short daily quizzes beat termly tests.
How iRead handles this
In iRead, a quiz runs after every session, and practice questions are available between sessions. Questions are generated from the story the class has just read: its real sentences, the grammar that appears in it, comprehension of the plot, and the moral. Every quiz is levelled to the CEFR. Teachers see retention per word, per game and per class.
The core principle: Match the number of questions to the purpose. A daily check needs a few good questions and fast feedback; a graded test needs enough questions to make the score dependable.
See it in iRead: iRead runs a quiz after every session and offers practice questions between sessions, all generated from the story the class is reading and levelled to the CEFR.
See how iRead worksKey takeaway
Five to eight questions after a story, twenty or more when a grade is at stake, and a short quiz often rather than a long one rarely.
Frequently asked questions
Is a five-question reading quiz too short?
Not for a daily check. Five questions on one story tell you whether the class understood it, and the feedback arrives the same day. Five questions are too few for a mark that counts, because one guess moves the score by 20 per cent.
How many questions should an end-of-term reading test have?
There is no fixed number in the research. As working guidance, twenty to forty items across several texts is a reasonable range. The Council of Europe (2011) manual gives a rule of thumb for acceptable reliability, and more items help you reach it. Pilot the test and check the timing.
How long should a reading quiz take?
For a daily check after one story, aim for under ten minutes, including time to look back at the text. Time yourself and allow pupils three to four times as long. If the quiz crowds out discussion of the story, cut questions rather than hurrying the class.
Does a quiz with more questions teach more?
Not by itself. Black and Wiliam (1998) report evidence that several short tests beat fewer long ones, and that the gain depends on feedback being used. More questions on one day mostly buy a more dependable score. More quizzes across the term, each with feedback, is what helps learning.
Sources
- Hughes, A. (2003). Testing for Language Teachers (2nd ed.). Cambridge University Press. doi.org/10.1017/CBO9780511732980
- Council of Europe (produced by ALTE for the Language Policy Division) (2011). Manual for Language Test Development and Examining: For use with the CEFR. Council of Europe, Strasbourg. rm.coe.int/manual-for-language-test-development-and-examining-for-use-with-the-ce/1680667a2b
- Black, P., & Wiliam, D. (1998). Assessment and classroom learning. Assessment in Education: Principles, Policy & Practice, 5(1), 7-74. doi.org/10.1080/0969595980050102
- Bangert-Drowns, R. L., Kulik, J. A., & Kulik, C.-L. C. (1991). Effects of frequent classroom testing. The Journal of Educational Research, 85(2), 89-99. doi.org/10.1080/00220671.1991.10702818
- Day, R. R., & Park, J.-s. (2005). Developing reading comprehension questions. Reading in a Foreign Language, 17(1), 60-73. doi.org/10.64152/10125/66599
- Spearman, C. (1910). Correlation calculated from faulty data. British Journal of Psychology, 3(3), 271-295. doi.org/10.1111/j.2044-8295.1910.tb00206.x
- Brown, W. (1910). Some experimental results in the correlation of mental abilities. British Journal of Psychology, 3(3), 296-322. doi.org/10.1111/j.2044-8295.1910.tb00207.x