How to Run a Pilot of a Reading Platform

A pilot is a small, time-limited trial with one question decided in advance. Write the pass mark before you start, and build the pilot so that a poor product would fail it.

In short

  • A pilot is a small trial with one question, a pass mark written in advance and a fixed end date.
  • Use one typical class, take a baseline, run four to six weeks and keep a comparison class if you can.
  • Early results flatter a new tool: novelty, small samples and low starting scores can all push the numbers up.
  • A short pilot can show whether pupils use the tool and teachers can run it, but not a change of CEFR level.

A pilot that cannot fail is not a test. If a school tries a product for a few weeks, the pupils enjoy it and the purchase goes ahead, nothing has been tested. Morrison, Ross and Cheung (2019) studied how 54 school districts chose education technology. Decisions were often made on small-scale pilot tryouts and peer references, and less often on rigorous evaluation evidence. A pilot is still worth running, if a poor product would fail it. This article gives the steps, three reasons early results look better than they are, a table of measures and a one-page plan.

What is a pilot, and how do you run one?

A pilot of a reading platform is a small trial with a fixed end date and one question decided in advance. You run it with one typical class for four to six weeks. You write the pass mark before you start, take a baseline, keep a comparison if you can, and decide on the evidence.

  1. Write the question and the pass mark. One question and one number, dated before any pupil logs in.
  2. Pick one typical class. Not your best.
  3. Take a baseline. Use the same test you will use at the end.
  4. Run four to six weeks. Fix the end date.
  5. Keep a comparison if you can. A similar class that reads in the usual way.
  6. Decide on the evidence. Judge the result against the pass mark.

The pilot comes after the shortlist, not before it. Use a buyer's checklist to choose what to pilot. Put the student data questions to the vendor before any pupil gets an account.

Write the question and the pass mark before you start

A pilot answers one question. 'Is this platform good?' is not a question you can answer in six weeks. 'Will a typical class use it three times a week without being chased?' is.

The Education Endowment Foundation's guide to implementation (Sharples, Albers, Fraser and Kime, 2019) recommends that schools define the problem they want to solve before they choose a program. It also says that the outcomes you will monitor need to be agreed, and understood by the staff involved, before monitoring begins.

Write the pass mark as a number with a date. The figures below are examples, so set your own.

  • Use. Will pupils read and practise without being chased? Pass: 18 of 26 pupils are active three times a week in weeks 4 to 6.
  • Workload. Can the teacher run it in normal planning time? Pass: setup takes 20 minutes a week or less by week 3.
  • Retention. Do pupils keep the words from the stories? Pass: the pilot class beats the comparison class on the same word test.

Give the written pass mark to someone outside the pilot, such as the head. A target that someone else holds is harder to move.

Choose a typical class, take a baseline and keep a comparison

If you pilot with your keenest teacher and your strongest class, you learn what the tool does in the best case. Most of your school is not the best case. Pick an ordinary class, with a teacher who is willing but not an enthusiast.

Take the baseline in the week before the pilot starts. Keep it short: a test of words from the stories the class will read, and a check of reading level. Use the same test, in the same conditions, at the end.

A comparison class makes the result easier to read. Choose a class of the same grade and a similar level that reads the same stories in the usual way. Without it, you cannot tell a gain from the tool apart from a gain from six more weeks of English lessons.

One class against one class is still not an experiment, because the classes differ in ways you cannot see. Treat a difference between them as a hint.

Why early results flatter a new tool

Three effects can push pilot numbers up. None of them is dishonest, and none of them is the product's doing.

New things get more effort

Clark (1983) reviewed research on learning from media. He described how pupils give more effort and attention to a medium that is new to them, and how those gains tend to shrink as it becomes familiar. He reported one review of computer-assisted instruction in grades 6 to 12. Effects were 0.56 standard deviations in studies of four weeks or less, about 0.3 at five to eight weeks, and about 0.2 after eight weeks.

Clark also noted that a review of college courses found no such pattern. His evidence is more than forty years old, and computers were newer to pupils then. Still, compare the first two weeks of your pilot with the last two. If use falls steadily, novelty was doing much of the work.

Small trials show bigger effects

Kraft (2020) collected 1,942 effect sizes from 747 randomised controlled trials of education interventions with standardised achievement outcomes. The median effect was 0.10 standard deviations. In studies with 100 students or fewer, the median was 0.24. In studies with more than 2,000 students, it was 0.03. Kraft suggests that publication bias explains part of the gap. Cost explains part too, because intensive programs are more likely to be tested on small samples.

Your pilot is one class. Expect a smaller effect when thirty classes use the tool with less attention from you.

The lowest scorers seem to improve most

Barnett, van der Pols and Dobson (2005) explain regression to the mean. Unusually high or low measurements tend to be followed by measurements closer to the average. This can make natural variation look like real change. The effect is more noticeable when measurement is imprecise, and when you follow up only a group selected for its baseline score. Their paper comes from epidemiology, not education.

So if your ten weakest pupils gain the most, do not conclude that the tool works best for weak readers. Check whether the weakest pupils in the comparison class rose as well.

What size of effect is realistic?

Cheung and Slavin (2013) reviewed 20 studies of educational technology for struggling readers, with about 7,000 students in grades 1 to 6. They found a positive but small effect on reading skills (effect size 0.14). Those pupils may not match your classes, so use the figure as a guide to scale only. Escueta, Nickow, Oreopoulos and Quan (2020) reviewed randomised trials and regression discontinuity studies of education technology in developed countries. They note that technology can help learning, or in some cases hinder it.

What to measure in a pilot, and what result would make you stop

The EEF guide (Sharples et al., 2019) lists common implementation outcomes: fidelity, acceptability, reach, feasibility and costs. You can see each one within six weeks. Add one narrow measure of learning, and decide the stop signal for each measure before you start.

What to measure in a pilot, how to collect it, and an example stop signal (example figures, set your own)
What to measureHow to collect itExample stop signal
Reach: how many pupils use itLogins or sessions attended, per pupil, per weekFewer than two-thirds of the class use it weekly by week 4
Fidelity: is it used as plannedSessions held against sessions plannedFewer than half the planned sessions took place
Feasibility: can the teacher run itThe teacher's log of setup minutes each weekSetup still takes over 30 minutes a week in week 4
Acceptability: what pupils thinkThe same five-question survey in week 1 and the final weekMost pupils say they would not continue
Learning on a narrow measureA test of words from the stories, before and after, in both classesNo gain, or no difference from the comparison class
CostHours of staff time, plus the fee quotedCost per pupil is above the agreed budget

Collect the first three every week. The EEF guide notes that measures which are simple and quick to collect are more likely to be accepted by staff. See also the guide to measuring reading progress.

What a short pilot can and cannot show

Four to six weeks is long enough for some questions and far too short for others. A short pilot can show:

  • Use. How many pupils log in, how often, and whether that holds after the second week.
  • Whether teachers can run it. Setup time, and sessions held against sessions planned.
  • What pupils and teachers think. A short survey at the start and the end.
  • A hint of learning on a narrow measure. For example, words from the stories the class read.

It cannot show a change of CEFR level. A CEFR level is a broad band of ability, and you should not expect a class to cross one in six weeks. If a pilot report shows levels rising in that time, ask how the level was measured. For longer periods, see measuring CEFR progress by cohort.

A narrow measure needs care too. Kraft (2020) found larger effects on narrow measures than on broad achievement tests, with medians of 0.17 and 0.10. A gain on a test of the pilot's own words does not prove a gain in reading. Nor can a pilot show that a result will last.

A one-page pilot plan

Write the plan on one page and have the head sign it before the pilot starts.

Example pilot plan (invented school)

School and class: Greenfield School (invented), class 7B, 26 pupils reading at A2.

Question: Will pupils read and practise at least three times a week without being chased?

Pass mark, written on 1 September: 18 of 26 pupils active three times a week in weeks 4 to 6.

Baseline, week 0: a 20-word test on words from the first three stories.

Comparison: class 7C reads the same stories on paper and takes the same test on the same days.

Length: six weeks.

Stop signals: fewer than half the planned sessions held, or setup over 30 minutes a week in week 4.

Decision: week 7, owned by the English coordinator. Go on, extend once, or stop.

Decide on the evidence

Hold the decision meeting in the week after the pilot ends. Bring the signed plan, the weekly figures and both sets of test results. There are three honest outcomes.

  • Pass. The pass mark was met and no stop signal was triggered. Plan the next stage.
  • Fail. The pass mark was missed. Stop, and record why.
  • Unclear. The result is close, or the data are incomplete. Extend once with the same pass mark, or stop.

A pass is not a licence to roll out to the whole school at once. The EEF guide advises schools to treat scale-up as a new implementation process. If the pilot fails, stopping is a good result. One class has spent six weeks, and the school has kept a year of budget and staff time.

If the board will see the result, read how to write a term reading report for your board.

How iRead handles this

A school can run a free iRead pilot with one class for 4 to 6 weeks, before any school-wide decision. The pilot comes with full dashboards, including attendance and retention. Attendance at sessions is recorded automatically. Teachers see retention per word, per game and per class.

The core principle: A pilot that cannot fail tells you nothing. Write the pass mark first, then run the trial.

See it in iRead: A school can run a free pilot with one class for 4 to 6 weeks, with full dashboards for attendance and retention, before any school-wide decision.

See iRead for school leaders

Key takeaway

Set one question and one pass mark, use a typical class for four to six weeks, and compare with a baseline and a second class. Then decide on what the figures show, even if the answer is no.

Frequently asked questions

How do you pilot an edtech product in a school?

Write one question and a pass mark before you start. Choose one typical class, take a baseline and run the tool for four to six weeks. Keep a comparison class if you can. Then judge the result against the pass mark.

How long should a pilot of a reading platform last?

Four to six weeks is enough to see whether pupils use the tool and whether teachers can run it. It is too short to show a change of CEFR level. Clark (1983) reported larger effects in studies of four weeks or less, so watch the later weeks.

What should an edtech pilot checklist include?

One question and a dated pass mark. The class, its teacher and a comparison class. A baseline test and a fixed end date. A stop signal for each measure. A named owner and a date for the decision meeting. Keep it to one page.

What should you do if a pilot fails?

Stop, and record why. A failed pilot has done its job, because it cost one class six weeks and not the whole school a year. The Education Endowment Foundation's guide (Sharples et al., 2019) suggests that schools make fewer, more strategic choices.

Sources

  1. Sharples, J., Albers, B., Fraser, S., & Kime, S. (2019). Putting Evidence to Work: A School's Guide to Implementation. Guidance Report (2nd ed.). Education Endowment Foundation, London. d2tic4wvo1iusb.cloudfront.net/eef-guidance-reports/implementation/EEF_Implementation_Guidance_Report_2019.pdf
  2. Morrison, J. R., Ross, S. M., & Cheung, A. C. K. (2019). From the market to the classroom: How ed-tech products are procured by school districts interacting with vendors. Educational Technology Research and Development, 67(2), 389-421. doi.org/10.1007/s11423-019-09649-4
  3. Clark, R. E. (1983). Reconsidering research on learning from media. Review of Educational Research, 53(4), 445-459. doi.org/10.3102/00346543053004445
  4. Kraft, M. A. (2020). Interpreting effect sizes of education interventions. Educational Researcher, 49(4), 241-253. doi.org/10.3102/0013189X20912798
  5. Barnett, A. G., van der Pols, J. C., & Dobson, A. J. (2005). Regression to the mean: What it is and how to deal with it. International Journal of Epidemiology, 34(1), 215-220. doi.org/10.1093/ije/dyh299
  6. Cheung, A. C. K., & Slavin, R. E. (2013). Effects of educational technology applications on reading outcomes for struggling readers: A best-evidence synthesis. Reading Research Quarterly, 48(3), 277-299. doi.org/10.1002/rrq.50
  7. Escueta, M., Nickow, A. J., Oreopoulos, P., & Quan, V. (2020). Upgrading education with technology: Insights from experimental research. Journal of Economic Literature, 58(4), 897-996. doi.org/10.1257/jel.20191507
Ready to see it in practice?

We'll show your team how iRead puts this into practice.