Curriculum Hub → Maths Lessons → Statistics → Collecting Data → Teaching the Large Data Set
If you are new to teaching A-Level Statistics, the large data set can look like an extra spreadsheet to mention in September and then forget. That is not how it is examined.
Ofqual wanted statistics taught with real data, not invented tables. The exam cannot put Excel in front of a candidate, so papers test familiarity. Examiners will not define tr, oktas or Beaufort. If a student has never opened the sheet, they are meant to be at a disadvantage.
This post is for teachers who want the subject knowledge to go beyond a one-off “here is the Information tab” lesson. I will walk through four exam-style problems from the collecting-data lessons, and I will keep coming back to the same examiner comments: deleting tr, treating May to October as the year, and writing a model comment that is too vague to score.
Three for Key Stage 3, one for GCSE and two for A-Level. Presentation, worksheet and answers for each, plus one free resource every Monday.
Unsubscribe any time.
A two-sided A4 you can print for a department meeting or photocopy for a class. One side is the key facts from the large data set. The other has the three classroom questions below, plus a QR code that opens the worked solutions.
They have to know the sheet well enough to answer questions with no extract in front of them. That means the five UK stations — Camborne, Heathrow, Hurn, Leeming and Leuchars — and the three overseas ones: Beijing, Jacksonville and Perth. It also means knowing what is missing.
The data covers May to October only, in 1987 and 2015. Students cannot talk about “the year”, and they cannot talk about winter. Not every variable exists at every station. There is no sunshine overseas. Beijing’s temperature is a daily average; Heathrow’s is a daily maximum. That is why “compare Heathrow with Beijing” or “sunnier in the UK than Perth” falls over.
They also have to clean the data. A reading of tr is a trace of rain, less than 0.05 mm. Replace it with 0, or with 0.025. Do not delete the day. A dash or n/a is the missing value. Visibility uses a dash, and a systematic sample of 20 can easily come back short.
Cloud cover is discrete, in oktas from 0 to 8. Wind direction is the direction the wind blows from. Humidity above 95% goes with mist and fog. Typical pressure is about 1013 hPa. Wind speeds come in knots and need converting if the question is in mph. Beaufort force is not a sensible partner for a linear relationship with temperature.
I start Year 12 with ten minutes and no computer: which station is further north, what tr means, why you cannot compare sunshine. Then the same ideas go into sampling, location and spread, scatter graphs and hypothesis tests. When the textbook offers two questions, pick the large data set one — and give them the whole sheet, not just the extract.
The video below is the teaching sequence I use for that first lesson.
Examiner reports make the same complaints again and again. Candidates delete tr and call it cleaning. They treat May to October as a representative year. They forget that cloud cover includes 0. On the June 2018 paper, a question that was meant to be a gentle starter left a lot of scripts blank, because students did not know that cloud cover is measured in oktas. Circle the Information tab before anyone opens Excel. Get them to write “tr is a value, n/a is a gap” at the top of the page.
This is a good first stretch question. The sampling method is routine, but the interval is not, and the second half of the question is really about whether they have ever cleaned the rainfall column.

a The large data set for Heathrow in 2015 is May to October, which is 184 days. The interval is 184 ÷ 15 ≈ 12. Generate a random starting number between 1 and 12, then take every 12th date.
b ‘tr’ means trace (rainfall less than 0.05 mm). Ravi’s method is not suitable. By removing these values entirely rather than replacing them, he reduces his sample size below 15 and biases the sample towards higher rainfall days, since the ‘tr’ days are still valid (very low but real) rainfall readings.
c Any data value given as ‘tr’ should be recorded as 0 for the purpose of carrying out any calculations. Recording it as 0.025 is also accepted.
The trap is 365 ÷ 15. Students who have never opened the sheet treat “Heathrow in 2015” as a calendar year and write an interval of 24. The large data set does not contain winter, so the population they are sampling from is 184 days, and the interval is 12.
The June 2023 A-Level paper had a script that said “get rid of days with no rain recorded”. That scored 0 out of 2. Trace is a value. Missing is the gap. If they delete tr they also shrink the sample, so a systematic sample of size 15 is no longer size 15.
Exam tip
Write the size of the data set before you write the interval. May to October is 31 + 30 + 31 + 31 + 30 + 31 = 184. Then 184 ÷ 15 gives 12. A random start from 1 to 12, then every 12th row. If they skip the 184, they are guessing.
Where marks are lost
This is the one I use to show a class that familiarity is not the same as remembering a number from the spreadsheet. The table is already in the question. They cannot look January 2015 up on the large data set, because the large data set has no January. They have to know the stations.

a Location X → Perth (hot, dry — Southern Hemisphere summer)
b Location Y → Camborne (cool, very wet — exposed UK coastal winter)
c Location Z → Jacksonville (mild, moderate rainfall — subtropical winter)
January is summer in Perth, so 21.2 °C and almost no rain is the Southern Hemisphere giveaway. Camborne is the UK coastal station: cool and very wet in winter. Jacksonville sits in between, a mild subtropical winter rather than a British one.
The teaching point is the map, not the January figures. Once they can place Perth, Jacksonville and Camborne without a prompt, questions about “which location is this?” stop being a lottery.
Exam tip
Start with the hemisphere, not the rainfall. Perth is the only Southern Hemisphere station in the large data set, so a hot January is Perth before you look at the second row. Then use rainfall to separate the two Northern Hemisphere stations.
Where marks are lost
The large data set sits next to sampling in the scheme of work, and papers like to test both in the same hour. This question is not about weather. It is about whether they can design a quota sample that an examiner will actually mark.

a) It is the morning rush hour on a weekday, so more people are likely to be commuting past.
The station is a location most people pass through, giving a good spread of the local population.
b) For example, male commuters under 40, female commuters under 40, male commuters 40 and over, female commuters 40 and over.
c) For example, use census data to determine the size of each group as a proportion of the whole population, then assign the quotas as the same proportion of the whole sample.
Part (a) wants two reasons that belong to this time and this place, not a textbook line about quota sampling. Rush hour gives volume. A station gives a mix of the town that a single workplace would not.
Part (b) needs named groups, and she can only ask adults. Gender by two age bands is enough. Students who write “different ages and genders” without listing four groups usually lose the mark. Part (c) is the proportional step: census figures to get the share of each group, then the same share of the sample. That is what stops quota sampling collapsing into “ask anyone”.
Exam tip
Quota sampling is not random. The quotas can be representative; the people she stops are still her choice. If a later part asks for a limitation, “only commuters at 8 am” scores. “It might be biased” without saying who is missing does not.
Where marks are lost
This is the June 2018 A-Level Statistics opener. It looks gentle until you realise the first two marks are only there if they already know that cloud cover is measured in oktas. There is no definition in the question.
Helen believes that the random variable C, representing cloud cover from the large data set, can be modelled by a discrete uniform distribution.
(a) Write down the probability distribution for C.
(b) Using this model, find the probability that cloud cover is less than 50%
Helen used all the data from the large data set for Hurn in 2015 and found that the proportion of days with cloud cover of less than 50% was 0.315
(c) Comment on the suitability of Helen’s model in the light of this information.
(d) Suggest an appropriate refinement to Helen’s model.
a Cloud cover is measured in oktas, so C takes the values 0, 1, 2, 3, 4, 5, 6, 7, 8. Under a discrete uniform model,
P(C=c)=\dfrac{1}{9},\quad c=0,1,2,\ldots,8b Four oktas is 50% cloud cover, so less than 50% is C = 0, 1, 2 or 3.
P(C<4)=\dfrac{4}{9}c 0.315 is substantially lower than 4/9 ≈ 0.444, so the discrete uniform model is not a good fit for Hurn in 2015.
d Use a discrete non-uniform distribution, with probabilities taken from the observed frequencies, and allow the model to vary by month or location. “Not uniform” on its own does not score.
The examiner report on this paper is blunt. Plenty of candidates left the question blank. Those who knew about oktas often forgot 0 and wrote a sample space from 1 to 8. Part (d) was poor because “use a different distribution” is not a refinement. The mark needs a discrete model that is not uniform, or probabilities based on the actual frequencies, ideally with a mention of month or place.
Including 4 in part (b) is the other common slip. Four oktas is exactly half the sky, so less than 50% stops at 3. The mark scheme wants 4/9, not 5/9.
Here is the fully worked solution from the June 2018 paper.
Exam tip
Write the nine values of C before you write any probabilities. If 0 is on the page, the 1/9 follows. For part (d), name the new model: discrete and based on frequencies, or higher probabilities for larger oktas. A continuous model such as the normal scores 0.
Where marks are lost
Once students can name the stations, the last marks on a real paper are often not calculation at all. Examiner reports are consistent on this.
If you want to write harder questions for your own class, those are the levers. Ask for a systematic sample from a column that contains n/a. Put tr in the raw data and ask them to clean it. Give three unnamed stations and force the hemisphere. Then add one sentence that forces a judgement about the model.
Use this as a Starter for 10, or as a five-minute recap before a sampling lesson. Level 1 is the Information tab. Level 3 is the material they still guess in Year 13.
Open full screen ↗. Use this link when projecting in class for the best experience.
Print it for the next faculty meeting. Key facts on one side, the three classroom questions on the other.
It is a Met Office weather data set used in Edexcel AS and A-Level Mathematics Statistics. It covers five UK stations and three overseas stations for May to October in 1987 and 2015. Papers test familiarity with the sheet, not Excel skills on the day.
tr means a trace of rainfall: less than 0.05 mm. Replace it with 0, or with 0.025, before calculating. Do not delete the day. Deleting tr shrinks the sample and biases it towards wetter days.
May to October is 184 days (June and September have 30 days). A systematic sample of size 15 from Heathrow 2015 therefore uses an interval of about 12, not 24. There is no winter data and no full calendar year.
Cloud cover is measured in oktas: eighths of the sky covered, taking the discrete values 0, 1, 2, 3, 4, 5, 6, 7 and 8. Examiners will not define this in the question. Four oktas is 50% cloud cover, so less than 50% is 0 to 3.
UK coastal: Camborne, Hurn and Leuchars. UK inland: Heathrow and Leeming. Overseas: Beijing (China), Jacksonville (USA) and Perth (Australia). Perth is in the Southern Hemisphere, so its seasons are reversed.
No. Daily total sunshine is not recorded at the overseas stations. Beijing’s temperature is also a daily average, while Heathrow’s is a daily maximum, so temperature comparisons need care too.
Download the free two-sided faculty handout from this page: https://mr-mathematics.com/free_files/Teaching-the-Large-Data-Set-Faculty-Handout.pdf. One side has the key facts. The other has three classroom questions and a QR code linking to worked solutions. No membership is required.
The three AS Statistics lessons take this in order: the large data set, populations and samples, then random and non-random sampling. Each pack has the presentation, the worksheet and the worked solutions.
A research-backed case exploring why over-reliance on automated math homework platforms and repetitive worksheets lowers student expectations, and how departments can build genuine mathematical thinking.
How to teach converting between fractions, decimals and percentages.
Four problem solving lessons to develop student’s mathematical reasoning and communication skills.