Teaching the Large Data Set

Curriculum Hub → Maths Lessons → Statistics → Collecting Data → Teaching the Large Data Set

If you are new to teaching A-Level Statistics, the large data set can look like an extra spreadsheet to mention in September and then forget. That is not how it is examined.

Ofqual wanted statistics taught with real data, not invented tables. The exam cannot put Excel in front of a candidate, so papers test familiarity. Examiners will not define tr, oktas or Beaufort. If a student has never opened the sheet, they are meant to be at a disadvantage.

This post is for teachers who want the subject knowledge to go beyond a one-off “here is the Information tab” lesson. I will walk through four exam-style problems from the collecting-data lessons, and I will keep coming back to the same examiner comments: deleting tr, treating May to October as the year, and writing a model comment that is too vague to score.

Six extra free lessons

Three for Key Stage 3, one for GCSE and two for A-Level. Presentation, worksheet and answers for each, plus one free resource every Monday.

Get six extra lessons

Unsubscribe any time.

Free worksheet (PDF)

A two-sided A4 you can print for a department meeting or photocopy for a class. One side is the key facts from the large data set. The other has the three classroom questions below, plus a QR code that opens the worked solutions.

What students actually have to be able to do

They have to know the sheet well enough to answer questions with no extract in front of them. That means the five UK stations — Camborne, Heathrow, Hurn, Leeming and Leuchars — and the three overseas ones: Beijing, Jacksonville and Perth. It also means knowing what is missing.

The data covers May to October only, in 1987 and 2015. Students cannot talk about “the year”, and they cannot talk about winter. Not every variable exists at every station. There is no sunshine overseas. Beijing’s temperature is a daily average; Heathrow’s is a daily maximum. That is why “compare Heathrow with Beijing” or “sunnier in the UK than Perth” falls over.

They also have to clean the data. A reading of tr is a trace of rain, less than 0.05 mm. Replace it with 0, or with 0.025. Do not delete the day. A dash or n/a is the missing value. Visibility uses a dash, and a systematic sample of 20 can easily come back short.

Cloud cover is discrete, in oktas from 0 to 8. Wind direction is the direction the wind blows from. Humidity above 95% goes with mist and fog. Typical pressure is about 1013 hPa. Wind speeds come in knots and need converting if the question is in mph. Beaufort force is not a sensible partner for a linear relationship with temperature.

Start with familiarity, not a spreadsheet tour

I start Year 12 with ten minutes and no computer: which station is further north, what tr means, why you cannot compare sunshine. Then the same ideas go into sampling, location and spread, scatter graphs and hypothesis tests. When the textbook offers two questions, pick the large data set one — and give them the whole sheet, not just the extract.

The video below is the teaching sequence I use for that first lesson.

Examiner reports make the same complaints again and again. Candidates delete tr and call it cleaning. They treat May to October as a representative year. They forget that cloud cover includes 0. On the June 2018 paper, a question that was meant to be a gentle starter left a lot of scripts blank, because students did not know that cloud cover is measured in oktas. Circle the Information tab before anyone opens Excel. Get them to write “tr is a value, n/a is a gap” at the top of the page.

Question 1: systematic sampling and tr

This is a good first stretch question. The sampling method is routine, but the interval is not, and the second half of the question is really about whether they have ever cleaned the rainfall column.

a The large data set for Heathrow in 2015 is May to October, which is 184 days. The interval is 184 ÷ 15 ≈ 12. Generate a random starting number between 1 and 12, then take every 12th date.

b ‘tr’ means trace (rainfall less than 0.05 mm). Ravi’s method is not suitable. By removing these values entirely rather than replacing them, he reduces his sample size below 15 and biases the sample towards higher rainfall days, since the ‘tr’ days are still valid (very low but real) rainfall readings.

c Any data value given as ‘tr’ should be recorded as 0 for the purpose of carrying out any calculations. Recording it as 0.025 is also accepted.

The trap is 365 ÷ 15. Students who have never opened the sheet treat “Heathrow in 2015” as a calendar year and write an interval of 24. The large data set does not contain winter, so the population they are sampling from is 184 days, and the interval is 12.

The June 2023 A-Level paper had a script that said “get rid of days with no rain recorded”. That scored 0 out of 2. Trace is a value. Missing is the gap. If they delete tr they also shrink the sample, so a systematic sample of size 15 is no longer size 15.

Exam tip

Write the size of the data set before you write the interval. May to October is 31 + 30 + 31 + 31 + 30 + 31 = 184. Then 184 ÷ 15 gives 12. A random start from 1 to 12, then every 12th row. If they skip the 184, they are guessing.

Where marks are lost

  • Using 365 days and an interval of 24.
  • Omitting the random start and just counting every 12th day from the first row.
  • Saying tr means “no rain”, or treating it as missing.
  • Deleting the tr days and then still calling it a sample of size 15.

Question 2: which weather station?

This is the one I use to show a class that familiarity is not the same as remembering a number from the spreadsheet. The table is already in the question. They cannot look January 2015 up on the large data set, because the large data set has no January. They have to know the stations.

a Location X → Perth (hot, dry — Southern Hemisphere summer)
b Location Y → Camborne (cool, very wet — exposed UK coastal winter)
c Location Z → Jacksonville (mild, moderate rainfall — subtropical winter)

January is summer in Perth, so 21.2 °C and almost no rain is the Southern Hemisphere giveaway. Camborne is the UK coastal station: cool and very wet in winter. Jacksonville sits in between, a mild subtropical winter rather than a British one.

The teaching point is the map, not the January figures. Once they can place Perth, Jacksonville and Camborne without a prompt, questions about “which location is this?” stop being a lottery.

Exam tip

Start with the hemisphere, not the rainfall. Perth is the only Southern Hemisphere station in the large data set, so a hot January is Perth before you look at the second row. Then use rainfall to separate the two Northern Hemisphere stations.

Where marks are lost

  • Putting Camborne as the hot station because they remember it is in Cornwall.
  • Swapping Jacksonville and Perth: both can be “not the UK”, but only Perth is in summer in January.
  • Trying to recall a January sheet that does not exist, then leaving the question blank.
  • Naming Beijing, which is not in the list.

Question 3: quota sampling

The large data set sits next to sampling in the scheme of work, and papers like to test both in the same hour. This question is not about weather. It is about whether they can design a quota sample that an examiner will actually mark.

a) It is the morning rush hour on a weekday, so more people are likely to be commuting past.

The station is a location most people pass through, giving a good spread of the local population.

b) For example, male commuters under 40, female commuters under 40, male commuters 40 and over, female commuters 40 and over.

c) For example, use census data to determine the size of each group as a proportion of the whole population, then assign the quotas as the same proportion of the whole sample.

Part (a) wants two reasons that belong to this time and this place, not a textbook line about quota sampling. Rush hour gives volume. A station gives a mix of the town that a single workplace would not.

Part (b) needs named groups, and she can only ask adults. Gender by two age bands is enough. Students who write “different ages and genders” without listing four groups usually lose the mark. Part (c) is the proportional step: census figures to get the share of each group, then the same share of the sample. That is what stops quota sampling collapsing into “ask anyone”.

Exam tip

Quota sampling is not random. The quotas can be representative; the people she stops are still her choice. If a later part asks for a limitation, “only commuters at 8 am” scores. “It might be biased” without saying who is missing does not.

Where marks are lost

  • Giving advantages of sampling in general, rather than reasons for 8 am at the station.
  • Listing fewer than four groups, or groups that include children.
  • Calling the method stratified, or saying the people in each quota are chosen at random.
  • In part (c), writing “make the groups equal” instead of matching the population proportions.

Question 4: the challenge

This is the June 2018 A-Level Statistics opener. It looks gentle until you realise the first two marks are only there if they already know that cloud cover is measured in oktas. There is no definition in the question.

Helen believes that the random variable C, representing cloud cover from the large data set, can be modelled by a discrete uniform distribution.

(a) Write down the probability distribution for C.

(b) Using this model, find the probability that cloud cover is less than 50%
Helen used all the data from the large data set for Hurn in 2015 and found that the proportion of days with cloud cover of less than 50% was 0.315

(c) Comment on the suitability of Helen’s model in the light of this information.

(d) Suggest an appropriate refinement to Helen’s model.

a Cloud cover is measured in oktas, so C takes the values 0, 1, 2, 3, 4, 5, 6, 7, 8. Under a discrete uniform model,

P(C=c)=\dfrac{1}{9},\quad c=0,1,2,\ldots,8

b Four oktas is 50% cloud cover, so less than 50% is C = 0, 1, 2 or 3.

P(C<4)=\dfrac{4}{9}

c 0.315 is substantially lower than 4/9 ≈ 0.444, so the discrete uniform model is not a good fit for Hurn in 2015.

d Use a discrete non-uniform distribution, with probabilities taken from the observed frequencies, and allow the model to vary by month or location. “Not uniform” on its own does not score.

The examiner report on this paper is blunt. Plenty of candidates left the question blank. Those who knew about oktas often forgot 0 and wrote a sample space from 1 to 8. Part (d) was poor because “use a different distribution” is not a refinement. The mark needs a discrete model that is not uniform, or probabilities based on the actual frequencies, ideally with a mention of month or place.

Including 4 in part (b) is the other common slip. Four oktas is exactly half the sky, so less than 50% stops at 3. The mark scheme wants 4/9, not 5/9.

Here is the fully worked solution from the June 2018 paper.

Exam tip

Write the nine values of C before you write any probabilities. If 0 is on the page, the 1/9 follows. For part (d), name the new model: discrete and based on frequencies, or higher probabilities for larger oktas. A continuous model such as the normal scores 0.

Where marks are lost

  • Leaving the question blank because oktas were never taught.
  • Using 1 to 8 and writing 1/8.
  • Taking less than 50% as 0 to 4 and writing 5/9.
  • Part (d): “not uniform”, or suggesting a normal distribution.

The familiarity marks at the end of the paper

Once students can name the stations, the last marks on a real paper are often not calculation at all. Examiner reports are consistent on this.

  • tr is a value, n/a is a gap. Replacing tr with 0 or 0.025 scores. Deleting the day does not, and it can wreck a sample size.
  • May to October is not the year. “Only 50% of the days” is not enough. They have to say it is not a representative sample of the year, and that winter rainfall is missing, so an annual estimate will be too low.
  • A limitation has to name who or what is missing. “It might be biased” does not score. “Only commuters at 8 am” does.
  • Describe the correlation in context. “Negative correlation” without pressure and temperature is thin. A small sample can look like no correlation when the full set is clearly negative.
  • Interpolation needs a value that actually appears. They guess 50 hPa and it is wrong. They have to have looked at the range.
  • Read the accuracy demand. A mean given as 2.1 when they asked for 3 significant figures loses a mark even when the method is fine.

If you want to write harder questions for your own class, those are the levers. Ask for a systematic sample from a column that contains n/a. Put tr in the raw data and ask them to clean it. Give three unnamed stations and force the hemisphere. Then add one sentence that forces a judgement about the model.

Practise recalling the large data set

Use this as a Starter for 10, or as a five-minute recap before a sampling lesson. Level 1 is the Information tab. Level 3 is the material they still guess in Year 13.

Open full screen ↗. Use this link when projecting in class for the best experience.

Print it for the next faculty meeting. Key facts on one side, the three classroom questions on the other.

Frequently Asked Questions

What is the Edexcel large data set?

It is a Met Office weather data set used in Edexcel AS and A-Level Mathematics Statistics. It covers five UK stations and three overseas stations for May to October in 1987 and 2015. Papers test familiarity with the sheet, not Excel skills on the day.

What does tr mean in the large data set?

tr means a trace of rainfall: less than 0.05 mm. Replace it with 0, or with 0.025, before calculating. Do not delete the day. Deleting tr shrinks the sample and biases it towards wetter days.

How many days are in the large data set for one year?

May to October is 184 days (June and September have 30 days). A systematic sample of size 15 from Heathrow 2015 therefore uses an interval of about 12, not 24. There is no winter data and no full calendar year.

What are oktas in the large data set?

Cloud cover is measured in oktas: eighths of the sky covered, taking the discrete values 0, 1, 2, 3, 4, 5, 6, 7 and 8. Examiners will not define this in the question. Four oktas is 50% cloud cover, so less than 50% is 0 to 3.

Which weather stations are in the large data set?

UK coastal: Camborne, Hurn and Leuchars. UK inland: Heathrow and Leeming. Overseas: Beijing (China), Jacksonville (USA) and Perth (Australia). Perth is in the Southern Hemisphere, so its seasons are reversed.

Can students compare sunshine at Heathrow and Perth?

No. Daily total sunshine is not recorded at the overseas stations. Beijing’s temperature is also a daily average, while Heathrow’s is a daily maximum, so temperature comparisons need care too.

Where can I download the free large data set worksheet?

Download the free two-sided faculty handout from this page: https://mr-mathematics.com/free_files/Teaching-the-Large-Data-Set-Faculty-Handout.pdf. One side has the key facts. The other has three classroom questions and a QR code linking to worked solutions. No membership is required.

Lesson packs

The three AS Statistics lessons take this in order: the large data set, populations and samples, then random and non-random sampling. Each pack has the presentation, the worksheet and the worked solutions.

Mr Mathematics Blog

Developing Mathematical Thinking Beyond Procedural Fluency

A research-backed case exploring why over-reliance on automated math homework platforms and repetitive worksheets lowers student expectations, and how departments can build genuine mathematical thinking.

Converting Between Fractions, Decimals and Percentages

How to teach converting between fractions, decimals and percentages.

From Key Skills to Deep Connections: Problem Solving in Secondary Maths

Four problem solving lessons to develop student’s mathematical reasoning and communication skills.