Large Data Set

Includes hypothesis testingIncludes sampling

Edexcel

AQA

OCR A

OCR MEI

June 2025 Paper 3 Q4

EdexcelCurrent spec9 marksIncludes hypothesis testingIncludes samplingCorrelation & RegressionData ProcessingLarge Data Set

4. Kay is studying the variables Daily Total Sunshine (\(x\)) and Daily Total Rainfall (\(y\)) from the large data set for Leeming in 2015

Kay starts with 5th May and then selects every 10th day thereafter.

(a) State the name of the sampling technique Kay uses. (1)

Kay wants to find the regression line of \(y\) on \(x\) for these data.

(b) Using your knowledge of the large data set, explain how Kay might need to clean these data before finding the equation of the regression line. (1)

The equation of the regression line Kay finds is \(y = 0.741 + 0.199x\)

(c) Using your knowledge of the large data set,
(i) state the units of the gradient of the regression line,
(ii) give an interpretation of the \(y\)-intercept of the regression line. (2)

Kay’s teacher claimed that the greater the amount of sunshine in a day the lower the amount of rain there should be.

(d) State, giving a reason, whether the teacher’s claim is true for Kay’s data. (1)

The teacher used all the data for these variables from the large data set for Leeming in 2015, as a sample.
The teacher calculated the product moment correlation coefficient for \(x\) and \(y\) to be \(-0.160\)

In a suitable test to determine whether there is evidence to support the teacher’s claim, the \(p\)-value was 0.015

(e) Using a 5% level of significance, state the hypotheses and conclusion for this test. (2)
(f) For the test in part (e) describe
(i) the sample,
(ii) a possible population. (2)

June 2024 Paper 3 Q3

EdexcelCurrent spec6 marksData ProcessingLarge Data Set

3. Ming is studying the large data set for Perth in 2015

He intended to use all the data available to find summary statistics for the Daily Mean Air Temperature, \(x\) °C.
Unfortunately, Ming selected an incorrect variable on the spreadsheet.
This incorrect variable gave a mean of 5.3 and a standard deviation of 12.4

(a) Using your knowledge of the large data set, suggest which variable Ming selected. (1)

The correct values for the Daily Mean Air Temperature are summarised as

\[n = 184 \qquad \sum x = 2801.2 \qquad \sum x^2 = 44\,695.4\]
(b) Calculate the mean and standard deviation for these data. (3)

One of the months from the large data set for Perth in 2015 has

  • mean \(\bar{x} = 19.4\)
  • standard deviation \(\sigma_x = 2.83\)

for Daily Mean Air Temperature.

(c) Suggest, giving a reason, a month these data may have come from. (2)

June 2025 Paper 3 Q20

AQACurrent spec9 marksLarge Data SetNormal Distribution

20 The oxides of nitrogen emissions, \(F\) g/km, of a car registered in 2000 can be modelled by a normal distribution with mean 0.41 and standard deviation 0.07

(a) Find \(\mathrm{P}(F \lt 0.39)\) [1 mark]
(b) Find \(\mathrm{P}(0.3 \lt F \lt 0.5)\) [1 mark]
(c)
(i) Find \(\mathrm{P}(F \gt 0.6)\) [1 mark]
(ii) Explain why \(\mathrm{P}(F \geqslant 0.6) = \mathrm{P}(F \gt 0.6)\) in this model. [2 marks]
(d) The oxides of nitrogen emissions, \(F\) g/km, of a car in 2010 can be modelled by a normal distribution with mean 0.36 and standard deviation 0.09

Compare the oxides of nitrogen emissions in 2000 with those in 2010.

[2 marks]
(e) A researcher collected data on the oxides of nitrogen emissions from the Large Data Set for cars registered in 2002 and 2016.

The researcher cleaned the data by removing the details for some cars.

Using your knowledge of the Large Data Set, give one reason why the researcher cleaned the data.

[1 mark]
(f) The researcher wished to make a comparison between the oxides of nitrogen emissions for a sample of Nissan cars in their local town from 2016 with the data from all cars registered in the same year in the Large Data Set.

Using your knowledge of the Large Data Set, give one reason why a meaningful comparison could not be made.

[1 mark]

June 2024 Paper 3 Q19

AQACurrent spec9 marksIncludes hypothesis testingBinomial DistributionLarge Data Set

19 It is known that 80% of all diesel cars registered in 2017 had carbon monoxide (CO) emissions less than 0.3 g/km.

Talat decides to investigate whether the proportion of diesel cars registered in 2022 with CO emissions less than 0.3 g/km has changed.

Talat will carry out a hypothesis test at the 10% significance level on a random sample of 25 diesel cars registered in 2022.

(a)
(i) State suitable null and alternative hypotheses for Talat’s test. [1 mark]
(ii) Using a 10% level of significance, find the critical region for Talat’s test. [5 marks]
(iii) In his random sample, Talat finds 18 cars with CO emissions less than 0.3 g/km.

State Talat’s conclusion in context. [1 mark]

(b) Talat now wants to use his random sample of 25 diesel cars, registered in 2022, to investigate whether the proportion of diesel cars in England with CO emissions more than 0.5 g/km has changed from the proportion given by the Large Data Set.

Using your knowledge of the Large Data Set, give two reasons why it is not possible for Talat to do this. [2 marks]

June 2023 Paper 3 Q15

AQACurrent spec11 marksIncludes samplingData ProcessingDRVsLarge Data Set

15

(a) A random sample of eight cars was selected from the Large Data Set.

The masses of these cars, in kilograms, were as follows.

950989124714151506168018332040

It is given that, for the population of cars in the Large Data Set:

\[\begin{aligned}\text{lower quartile} &= 1167 \\ \text{median} &= 1393 \\ \text{upper quartile} &= 1570\end{aligned}\]
(i) It was decided to remove any of the masses which fall outside the following interval.\[\text{median} - 1.5 \times \text{interquartile range} \leqslant \text{mass} \leqslant \text{median} + 1.5 \times \text{interquartile range}\]

Show that only one of the eight masses in the sample should be removed. [3 marks]

(ii) Write down the statistical name for the mass that should be removed in part (a)(i). [1 mark]
(b) The table shows the probability distribution of the number of previous owners, \(N\), for a sample of cars taken from the Large Data Set.
\(n\)0123456 or more
\(\mathrm{P}(N = n)\)0.140.37\(0.9k\)0.25\(0.4k\)\(1.7k\)0

Find the value of \(\mathrm{P}(1 \leqslant N \lt 5)\) [4 marks]

(c) An expert team is investigating whether there have been any changes in CO2 emissions from all cars taken from the Large Data Set.

The team decided to collect a quota sample of 200 cars to reflect the different years and the different makes of cars in the Large Data Set.

(i) Using your knowledge of the Large Data Set, explain how the team can collect this sample. [2 marks]
(ii) Describe one disadvantage of quota sampling. [1 mark]

June 2022 Paper 3 Q13

AQACurrent spec2 marksData ProcessingLarge Data Set

13 A reporter is writing an article on the CO2 emissions from vehicles using the Large Data Set.

The reporter claims that the Large Data Set shows that the CO2 emissions from all vehicles in the UK have declined every year from 2002 to 2016.

Using your knowledge of the Large Data Set, give two reasons why this claim is invalid. [2 marks]

June 2025 Paper 2 Q13

OCR ACurrent spec6 marksData ProcessingLarge Data Set

13 The table shows excerpts from the census data for Age Structure in 2001 and 2011 for four Local Authorities (LAs) in South West England.

Age
LAYear0 to 45 to 78 to 910 to 141516 to 1718 to 1920 to 2425 to 29
Bournemouth2001
2011
8 171
10 275
5 103
4 862
3 579
2 999
8 752
8 399
1 681
1 732
3 255
3 517
4 289
6 141
12 901
17 130
11 785
14 935
City of Bristol2001
2011
23 453
29 633
12 887
14 371
8 794
8 466
22 871
21 703
4 791
4 408
8 690
8 922
11 812
13 711
34 798
44 371
32 001
40 752
Plymouth2001
2011
13 213
15 336
8 535
7 956
6 091
4 991
16 078
13 645
3 106
2 954
6 103
6 011
7 454
8 839
17 245
24 343
14 885
18 888
Swindon2001
2011
11 392
14 083
7 194
7 551
4 862
4 722
12 121
12 433
2 178
2 593
4 337
5 141
3 798
4 690
10 212
12 859
13 816
15 075

Researchers want to investigate whether there is evidence that those people who were residents in these LAs in 2001 were still resident in the same LAs in 2011.

(a) One researcher suggests that, for each of the LAs, they should compare the sum of the data for 2001 for Age 15 to 19 with the data for 2011 for Age 25 to 29.

Explain why this comparison is relevant. [1]
(b) The census data also includes data for the following age ranges:
Age
30 to 4445 to 5960 to 6465 to 7475 to 8485 to 8990 +
Explain why none of these data ranges are helpful in this context. [1]
(c) The data for one of these four LAs show that some of the 0 to 4 year olds living in that LA in 2001 are definitely no longer living there in 2011.

Explain which LA this is. [1]
(d) One of the researchers says that there has been little movement in or out of Bournemouth between 2001 and 2011 for those who were aged 5 to 7 in 2001.
(i) Use values from the table to show that this statement is not contradicted by the data. [2]
(ii) Explain why the statement is not necessarily correct. [1]

June 2023 Paper 2 Q13

OCR ACurrent spec10 marksData ProcessingLarge Data Set

13 The scatter diagram uses information about all the Local Authorities (LAs) in the UK, taken from the 2011 census.

For each LA it shows the percentage (\(x\)) of employees who used public transport to travel to work and the percentage (\(y\)) who used motorised private transport.

“Public transport” includes train, bus, minibus, coach, underground, metro and light rail.
“Motorised private transport” includes car, van, motorcycle, scooter, moped, taxi and passenger in a car or van.

Scatter diagram of percentage using motorised private transport (y, 0 to 90) against percentage using public transport (x, 0 to 80): a dense cluster at x about 2 to 15, y about 55 to 80, with points trailing down to x about 50 to 63, y about 11 to 23; outlier A at about (1, 23) and outlier B at about (29, 4)
(a) Most of the points in the diagram lie on or near the line with equation \(x + y = k\), where \(k\) is a constant.
(i) Give a possible value for \(k\). [1]
(ii) Hence give an approximate value for the percentage of employees who either worked from home or walked or cycled to work. [1]
(b) The average amount of fuel used per person per day for travelling to work in any LA is denoted by F.
Consider the two groups of LAs where the percentages using motorised private transport are highest and lowest.
(i) Using only the information in the diagram, suggest, with a reason, which of these two groups will have greater values of F than the other group. [1]

A student says that it is not possible to give a reliable answer to part (b)(i) without some further information.

(ii) Suggest two kinds of further information which would enable a more reliable answer to be given. [2]
(c) Points \(A\) and \(B\) in the diagram are the most extreme outliers. Use their positions on the diagram to answer the following questions about the two LAs represented by these two points.
(i) The two LAs share a certain characteristic.
Describe, with a justification, this characteristic. [2]
(ii) The environments in these two LAs are very different.
Describe, with a justification, this difference. [2]
(d) A student says that it is difficult to extract detailed information from the scatter diagram.
Explain whether you agree with this criticism. [1]

June 2022 Paper 2 Q10

OCR ACurrent spec10 marksData ProcessingLarge Data Set

10 The table shows the age structure of usual residents of 18 Local Authorities (LAs) in the North West region of the UK in 2011.

Local AuthorityAge 0 to 17Age 18 to 24Age 25 to 64Age 65 and over
A26.20%9.06%51.81%12.92%
B23.32%8.99%52.32%15.37%
C22.24%8.96%52.56%16.23%
D22.67%8.10%53.27%15.96%
E20.70%7.77%54.77%16.76%
F18.14%6.51%51.13%24.21%
G18.96%14.20%48.51%18.33%
H19.06%14.79%52.12%14.04%
I25.15%9.04%51.16%14.65%
J22.93%8.81%52.22%16.04%
K21.48%13.98%50.82%13.73%
L23.98%9.20%52.26%14.56%
M21.67%11.19%52.94%14.19%
N17.82%6.01%51.93%24.23%
O22.83%7.30%53.86%16.01%
P21.76%8.28%54.03%15.93%
Q21.42%8.43%53.90%16.25%
R18.61%7.33%49.35%24.71%

Percentage of residents

(a) Without reference to any other columns, explain how you would use only the columns for the age ranges 0 to 17 and 18 to 24 to decide whether an LA might be one of the following.
(i) An LA that includes a university [1]
(ii) An LA that attracts young couples to live [1]
(iii) An LA that attracts retired people to live [1]
(b) Using your answers to part (a), identify the following.
(i) Four LAs that might include a university [1]
(ii) Three LAs that might be attractive to retired people [1]
(c) Explain why your answer to part (b)(ii), based only on the columns for the age ranges 0 to 17 and 18 to 24, may not be reliable. [1]
(d) The lower quartile, median and upper quartile of the percentages in the column “Age 65 and over” are 14.56%, 15.99% and 16.76% respectively.
Use this information to comment on your answers to part (b)(ii) and part (c). [2]

In a magazine article, a councillor plans to describe a typical LA in the North West region. He wants to quote the average percentage of residents aged 65 or over.

(e) The mean of the percentages in the column “Age 65 and over” is 16.90%.
Use this information, and the information given in part (d), to explain whether the median or the mean better represents the data in the column “Age 65 and over”. [2]

October 2021 Paper 2 Q13

OCR ACurrent spec7 marksData ProcessingLarge Data Set

13 The four pie charts illustrate the numbers of employees using different methods of travel in four Local Authorities in 2011.

Four pie charts for Local Authorities A, B, C and D with key: public transport (dark), private motorised transport (hatched), bicycle (dotted), all other methods of travel (light grey). A: mostly private, small bicycle and very small public. B: mostly private, large other, very small bicycle and public. C: mostly private, about a fifth public, tiny bicycle. D: half public, rest other, private and a small bicycle
(a) State, with reasons, which of the four Local Authorities is most likely to be a rural area with many hills. [2]
(b) Explain why pie charts are more suitable for answering part (a) than bar charts showing the same data. [1]
(c) Two of the Local Authorities represent urban areas.
(i) State with a reason which two Local Authorities are likely to be urban. [2]
(ii) One urban Local Authority introduced a Park-and-Ride service in 2006. Users of this service drive to the edge of the urban area and then use buses to take them into the centre of the area. A student claims that a comparison of the corresponding pie charts for 2001 (not shown) and 2011 would enable them to identify which Local Authority this was.
State with a reason whether you agree with the student. [2]

June 2025 Paper 2 Q10

OCR MEICurrent spec5 marksData ProcessingLarge Data Set

10 The pre-release material contains information about life expectancy at birth for countries of the world at 10-year intervals from 1960 until 2020.

The life expectancy at birth of the population of Sudan for this period, together with a line of best fit, is shown in Fig. 10.1.

Fig. 10.1: life expectancy at birth in Sudan plotted against year from 1960 to 2020, rising from about 48 to about 65, with a dotted line of best fit
Fig. 10.1

The equation of the line of best fit is \(L = 0.275X - 490\), where \(X\) is the year and \(L\) is the life expectancy at birth, measured in years.

(a) Use the equation of the line of best fit to estimate the life expectancy at birth in Sudan in 1975. [1]
(b) Explain whether this estimate is likely to be close to the true value of life expectancy at birth in Sudan in 1975. [1]
(c) Use your knowledge of the pre-release material to explain whether your answer to part (a) is likely to be a close approximation to the life expectancy at birth in the United Kingdom in 1975. [1]

The pre-release material also gives the median age of the population in countries of the world.

The table shows the median age of the population and the life expectancy at birth in 2020 for some countries in Africa.

CountryMedian ageLife expectancy at birth 2020
Nigeria18.655.02
Rwanda19.769.33
Saint Helena, Ascension and Tristan da Cunha43.2#N/A
Sao Tome and Principe19.370.58
Senegal19.468.21
(d) Explain how the data in the table should be cleaned before these data can be included in a scatter diagram for life expectancy at birth against median age for the countries in Africa. [1]

Fig. 10.2 shows a scatter diagram for the cleaned data of life expectancy at birth against median age in 2020 for the countries in Africa. It also shows a line of best fit.

The product moment correlation coefficient for these data is 0.680.

Fig. 10.2: scatter diagram of life expectancy at birth 2020 (L) against median age in 2020 (X) for African countries, median ages about 15 to 37, with a dotted line of best fit
Fig. 10.2
(e) A student decides to use the line of best fit to estimate the life expectancy at birth in 2020 for Saint Helena, Ascension and Tristan da Cunha.

Explain whether this estimate is likely to be reliable. [1]

June 2024 Paper 2 Q14

OCR MEICurrent spec8 marksData ProcessingLarge Data Set

14 The pre-release material contains medical data for 103 women and 97 men.

The boxplot represents the weights in kg of 101 of the women from the pre-release material.

Boxplot of weight in kg on a scale from 0 to 140: minimum 41.4, lower quartile 57.7, median 69.5, upper quartile 82.05, maximum 132.2
(a) Use your knowledge of the pre-release material to give a reason why the weights of all 103 women were not included in the diagram. [1]
(b) Determine the range of values in which any outliers lie. [3]
(c) Use your knowledge of the pre-release material to explain whether these outliers should be removed from any further analysis of the data. [1]
(d) The median weight of men in the sample was found to be 79.9 kg.
Explain what may be inferred by comparing the median weight of men with the median weight of women. [1]

Further analysis of the weights of both men and women is carried out. The table shows some of the results.

meanstandard deviation
men82.69 kg19.98 kg
women72.5 kg19.95 kg
(e) Use the information in the table to make two inferences about the distribution of the weights of men compared with the distribution of the weights of women. [2]

June 2023 Paper 2 Q14

OCR MEICurrent spec8 marksCorrelation & RegressionLarge Data Set

14 The pre-release material contains information concerning the median income of taxpayers in £ and the percentage of all pupils at the end of KS4 achieving 5 or more GCSEs at grade A*–C, including English and Maths, for different areas of London.

Some of the data for 2014/15 is shown in Fig. 14.1.

Fig. 14.1

Median Income of Taxpayers in £Percentage of Pupils Achieving 5 or more A*–C, including English and Maths
City of London61 100#N/A
Barking and Dagenham21 80054.0
Barnet27 10070.1
Bexley24 40055.0
Brent22 70060.0
Bromley28 10068.0

A student investigated whether there is any relationship between median income of taxpayers and percentage of pupils achieving 5 or more GCSEs at grade A*–C, including English and Maths.

(a) With reference to Fig. 14.1, explain how the data should be cleaned before any analysis can take place. [1]

After the data was cleaned, the student used software to draw the scatter diagram shown in Fig. 14.2.

Fig. 14.2: scatter diagram of percentage of students (40 to 75) against median income in pounds (15 000 to 40 000); most points cluster between 20 000 and 30 000 with percentages 52 to 73, and a few points between 31 000 and 39 000
Fig. 14.2

The student calculated that the product moment correlation coefficient for these data is 0.3743.

(b) Give two reasons why it may not be appropriate to use a linear model for the relationship between median income of taxpayers in £ and the percentage of all pupils at the end of KS4 achieving 5 or more GCSEs at grade A*–C. [2]

The student carried out some further analysis. The results are shown in Fig. 14.3.

Fig. 14.3

median income of taxpayers in £percentage of pupils achieving 5+ A*–C
mean27 21661.0
standard deviation4177.55.32

The student identified three outliers in total.

(c)
  • Use the information in Fig. 14.3 to determine the range of values of the median income of taxpayers in £ which are outliers.
  • Use the information in Fig. 14.3 to determine the range of values of the percentage of all pupils at the end of KS4 achieving 5 or more GCSEs at grade A*–C which are outliers.
  • On the copy of Fig. 14.2 in the Printed Answer Booklet, circle the three outliers identified by the student.
[4]

The student decided to remove these outliers and recalculate the product moment correlation coefficient.

(d) Explain whether the new value of the product moment correlation coefficient would be between 0.3743 and 1 or between 0 and 0.3743. [1]

June 2023 Paper 2 Q9

OCR MEICurrent spec5 marksIncludes samplingData ProcessingLarge Data Set

9 The pre-release material contains information concerning the median income of taxpayers in different areas of London. Some of the data for Camden is shown in the table below. The years quoted in this question refer to the end of the financial years used in the pre-release material. For example, the year 2004 in the table refers to the year 2003/04 in the pre-release material.

Year20042005200620072008200920102011
Median Income in £21 30023 20024 20025 90026 900#N/A28 40029 400
(a) Explain whether these data are a sample or a population of Camden taxpayers. [1]

A time series for the data is shown below.

Time series titled Median income of taxpayers in Camden 2004–2011: median income in pounds (0 to 35 000) against year (2003 to 2012), points for 2004 to 2008, 2010 and 2011 rising from about 21 300 to 29 400, with no point for 2009

The LINEST function on a spreadsheet is used to formulate the following model for the data:

\(I = 1115Y - 2\,212\,950\), where \(I =\) median income of taxpayers in £ and \(Y =\) year.

(b) Use this model to find an estimate of the median income of taxpayers in Camden in 2009. [1]
(c) Give two reasons why this estimate is likely to be close to the true value. [2]

The median income of taxpayers in Croydon in 2009 is also not available.

(d) Use your knowledge of the pre-release material to explain whether the model used in part (b) would give a reasonable estimate of the missing value for Croydon. [1]

June 2022 Paper 2 Q15

OCR MEICurrent spec9 marksCorrelation & RegressionLarge Data Set

15 The pre-release material includes information on life expectancy at birth in countries of the world. Fig. 15.1 shows the data for Liberia, which is in Africa, together with a time series graph.

Time series graph “Life expectancy at birth – Liberia”: life expectancy at birth against year from 1960 to 2010, increasing from about 35 to about 60
196019701980199020002010
34.6739.2546.0047.1852.4259.63

Fig. 15.1

Sundip uses the LINEST function on a spreadsheet to model life expectancy as a function of calendar year by a straight line.

The equation of this line is \(L = 0.473y - 892\), where \(L\) is life expectancy at birth and \(y\) is calendar year.

(a) Use this model to find an estimate of the life expectancy at birth in Liberia in 1995. [1]

According to the model, the life expectancy at birth in Liberia in 2025 is estimated to be 65.83 years.

(b) Explain whether each of these two estimates is likely to be reliable. [2]
(c) Use your knowledge of the pre-release material to explain whether this model could be used to obtain a reliable estimate of the life expectancy at birth in other countries in 1995. [1]

Fig. 15.2 shows the life expectancy at birth between 1960 and 2010 for Italy and South Africa.

Fig. 15.2: “Life expectancy at birth – Italy and South Africa”, 1960 to 2010; Series 1 (solid line) rises from about 53 to about 62 in 1990 then falls to about 56; Series 2 (dashed line) rises steadily from about 69 to about 82
Fig. 15.2
(d) Use your knowledge of the pre-release material to
  • Explain whether series 1 or series 2 represents the data for Italy.
  • Explain how the data for South Africa differs from the data for most developed countries.
[2]

Sundip is investigating whether there is an association between the wealth of a country and life expectancy at birth in that country. As part of her analysis she draws a scatter diagram of GDP per capita in US$ and life expectancy at birth in 2010 for all the countries in Europe for which data is available. She accidentally includes the data for the Central African Republic. The diagram is shown in Fig. 15.3.

Fig. 15.3: scatter diagram of life expectancy at birth in 2010 (40 to 85) against GDP per capita in US dollars (0 to 160 000); most points between 70 and 82 for GDP up to about 75 000, two points near 80 at about 105 000 and 140 000, and one isolated point at about 47.5 near GDP 0
Fig. 15.3
(e) On the copy of Fig. 15.3 in the Printed Answer Booklet, use your knowledge of the pre-release material to circle the point representing the data for the Central African Republic. [1]

Sundip states that as GDP per capita increases, life expectancy at birth increases.

(f) Explain to what extent the information in Fig. 15.3 supports Sundip’s statement. [2]

October 2021 Paper 2 Q12

OCR MEICurrent spec7 marksData ProcessingLarge Data Set

12 Fig. 12.1 shows an excerpt from the pre-release material.

ABCDEFGH
1SexAgeMaritalWeightHeightBMIWaistPulse
2Female34Married60.3173.420.0582.574
3Female85Widowed64.7161.224.9#N/A#N/A
4Female48Divorced100.6171.434.24105.692
5Male61Married70.9169.524.6892.270
6Male68Divorced96.8181.629.35112.968

Fig. 12.1

There was no data available for cell H3.

(a) Explain why #N/A is used when no data is available. [1]

Fig. 12.2 shows a scatter diagram of pulse rate against BMI (Body Mass Index) for females.
All the available data was used.

Fig. 12.2: scatter diagram titled Pulse rate against BMI for females; BMI 0 to 50, pulse rate 0 to 140; most points have BMI 17 to 45 and pulse 44 to 106, with one point at about (1.3, 88) and one at about (30.6, 128)
Fig. 12.2

There are two outliers on the diagram.

(b) On the copy of Fig. 12.2 in the Printed Answer Booklet, ring these outliers. [1]
(c) Use your knowledge of the pre-release material to explain whether either of these outliers should be removed. [2]
(d) State whether the diagram suggests there is any correlation between pulse rate and BMI. [1]

The product moment correlation coefficient between waist measurement, \(w\), in cm and BMI, \(b\), for females was found to be 0.912. All the available data was used.

(e) Explain why a model of the form \(w = mb + c\) for the relationship between waist measurement and BMI is likely to be appropriate. [1]

The LINEST function on a spreadsheet gives \(m = 2.16\) and \(c = 33.0\).

(f) Calculate an estimate of the value for cell G3 in Fig. 12.1. [1]

October 2020 Paper 2 Q13

OCR MEICurrent spec7 marksIncludes hypothesis testingCorrelation & RegressionLarge Data Set

13 The pre-release material contains information concerning median house prices, recycling rates and employment rates. Fig. 13.1 shows a scatter diagram of recycling rate against employment rate for a random sample of 33 regions.

Fig. 13.1: scatter diagram of recycling rate (10 to 55) against employment rate (60 to 85) for 33 regions, showing weak positive correlation
Fig. 13.1

The product moment correlation coefficient for this sample is 0.37154 and the associated \(p\)-value is 0.033.

Lee conducts a hypothesis test at the 5% level to test whether there is any evidence to suggest there is positive correlation between recycling rate and employment rate. He concludes that there is no evidence to suggest positive correlation because \(0.033 \approx 0\) and \(0.37154 > 0.05\).

(a) Explain whether Lee’s reasoning is correct. [2]

Fig. 13.2 shows a scatter diagram of recycling rate against median house price for a random sample of 33 regions.

Fig. 13.2: scatter diagram of recycling rate against median house price (120 000 to 1 320 000) for 33 regions; most points lie between 250 000 and 550 000, with isolated points near 920 000 and 1 200 000
Fig. 13.2

The product moment correlation coefficient for this sample is \(-0.33278\) and the associated \(p\)-value is 0.058.

Fig. 13.3 shows summary statistics for the median house prices for the data in this sample.

Statistics
n33
Mean465467.9697
σ201236.1345
s204356.2606
Σx15360443
Σx28486161617387
Min243500
Q1342500
Median410000
Q3521000
Max1200000

Fig. 13.3

(b) Use the information in Fig. 13.3 and Fig. 13.2 to show that there are at least two outliers. [2]
(c) Describe the effect of removing the outliers on
  • the product moment correlation coefficient between recycling rate and median house price,
  • the \(p\)-value associated with this correlation coefficient,
in each case explaining your answer. [2]

All 33 items in the sample are areas in London. A student suggests that it is very unlikely that only areas in London would be selected in a random sample.

(d) Use your knowledge of the pre-release material to explain whether you think the student’s suggestion is reasonable. [1]

October 2020 Paper 2 Q11

OCR MEICurrent spec10 marksCorrelation & RegressionLarge Data Set

11 The pre-release material contains information concerning median house prices over the period 2004 – 2015. A spreadsheet has been used to generate a time series graph for two areas: the London borough of “Barking and Dagenham” and “North West”. This is shown together with the raw data in Fig. 11.1.

Fig. 11.1: time series graph of median house price from 2004 to 2015; Barking and Dagenham (dashed) is always above North West (solid)
YearBarking and DagenhamNorth West
2004160 000107 000
2005163 000118 000
2006168 000127 000
2007185 000134 750
2008190 000129 950
2009160 000130 000
2010171 000130 000
2011170 000127 000
2012174 995130 000
2013180 995131 000
2014215 000138 500
2015243 500140 000

Fig. 11.1

Dr Procter suggests that it is unusual for median house prices in a London borough to be consistently higher than those in other parts of the country.

(a) Use your knowledge of the large data set to comment on Dr Procter’s suggestion. [1]

Dr Procter wishes to predict the median house price in Barking and Dagenham in 2016. She uses the spreadsheet function LINEST to find the equation of the line of best fit for the given data. She obtains the equation

\(P = 4897Y - 9\,657\,847\), where \(P\) is the median house price in pounds and \(Y\) is the calendar year, for example 2015.

(b) Use Dr Procter’s equation to predict the median house price in Barking and Dagenham in
  • 2016
  • 2017.
[2]

Professor Jackson uses a simpler model by using the data from 2014 and 2015 only to form a straight-line model.

(c) Find the equation Professor Jackson uses in her model. [2]
(d) Use Professor Jackson’s equation to predict the median house price in Barking and Dagenham in
  • 2016
  • 2017.
[2]

Professor Jackson carries out some research online. She finds some information about median house prices in Barking and Dagenham, which is shown in Fig. 11.2.

20162017
£290 000£300 000
Fig. 11.2

(e) Comment on how well
  • Dr Procter’s model fits the data,
  • Professor Jackson’s model fits the data.
[2]
(f) Explain which, if any, of the models is likely to be more reliable for predicting median house prices in Barking and Dagenham in 2020. [1]