Correlation & Regression

Includes hypothesis testingIncludes sampling

Edexcel

AQA

OCR A

OCR MEI

June 2025 Paper 3 Q4

EdexcelCurrent spec9 marksIncludes hypothesis testingIncludes samplingCorrelation & RegressionData ProcessingLarge Data Set

4. Kay is studying the variables Daily Total Sunshine (\(x\)) and Daily Total Rainfall (\(y\)) from the large data set for Leeming in 2015

Kay starts with 5th May and then selects every 10th day thereafter.

(a) State the name of the sampling technique Kay uses. (1)

Kay wants to find the regression line of \(y\) on \(x\) for these data.

(b) Using your knowledge of the large data set, explain how Kay might need to clean these data before finding the equation of the regression line. (1)

The equation of the regression line Kay finds is \(y = 0.741 + 0.199x\)

(c) Using your knowledge of the large data set,
(i) state the units of the gradient of the regression line,
(ii) give an interpretation of the \(y\)-intercept of the regression line. (2)

Kay’s teacher claimed that the greater the amount of sunshine in a day the lower the amount of rain there should be.

(d) State, giving a reason, whether the teacher’s claim is true for Kay’s data. (1)

The teacher used all the data for these variables from the large data set for Leeming in 2015, as a sample.
The teacher calculated the product moment correlation coefficient for \(x\) and \(y\) to be \(-0.160\)

In a suitable test to determine whether there is evidence to support the teacher’s claim, the \(p\)-value was 0.015

(e) Using a 5% level of significance, state the hypotheses and conclusion for this test. (2)
(f) For the test in part (e) describe
(i) the sample,
(ii) a possible population. (2)

June 2024 Paper 3 Q2

EdexcelCurrent spec6 marksIncludes hypothesis testingCorrelation & Regression

2. Amar is studying the flight of a bird from its nest.

He measures the bird’s height above the ground, \(h\) metres, at time \(t\) seconds for 10 values of \(t\)
Amar finds the equation of the regression line for the data to be \(h = 38.6 - 1.28t\)

(a) Interpret the gradient of this line. (1)

The product moment correlation coefficient between \(h\) and \(t\) is \(-0.510\)

(b) Test whether or not there is evidence of a negative correlation between the height above the ground and the time during the flight.
You should
  • state your hypotheses clearly
  • use a 5% level of significance
  • state the critical value used
(3)

Jane draws the following scatter diagram for Amar’s data.

scatter diagram of h against t for 10 points at about (1, 28), (1.5, 34), (2, 36), (2.5, 38), (3.5, 40), (4.5, 39), (5, 37), (6.5, 33), (7, 27), (8.5, 20), rising then falling
(c) With reference to the scatter diagram, state, giving a reason, whether or not the regression line \(h = 38.6 - 1.28t\) is an appropriate model for these data. (1)

Jane suggests an improved model using the variable \(u = (t - k)^2\) where \(k\) is a constant.

She obtains the equation \(h = 38.1 - 0.78u\)

(d) Choose a suitable value for \(k\) to write Jane’s improved model for \(h\) in terms of \(t\) only. (1)

June 2024 Paper 3 Q16

AQACurrent spec4 marksIncludes hypothesis testingCorrelation & Regression

16 A medical student believes that, in adults, there is a negative correlation between the amount of nicotine in their blood stream and their energy level.

The student collected data from a random sample of 50 adults.

The correlation coefficient between the amount of nicotine in their blood stream and their energy level was −0.45

Carry out a hypothesis test at the 2.5% significance level to determine if this sample provides evidence to support the student’s belief.

For \(n = 50\), the critical value for a one-tailed test at the 2.5% level for the population correlation coefficient is 0.2787 [4 marks]

June 2024 Paper 2 Q10

OCR ACurrent spec11 marksIncludes hypothesis testingCorrelation & Regression

10 Each month, the manager of a large store records the number, \(c\), of customers who visit the store, and the amount, £\(h\), spent on heating during that month. The manager wants to test whether there is linear correlation between \(c\) and \(h\).

For a randomly chosen year the value of Pearson’s product-moment correlation coefficient, \(r\), between \(c\) and \(h\) was \(-0.798\), correct to 3 significant figures.

(a) Using the table below, carry out the test at the 1% significance level. [5]
(b) Describe briefly two main features of a scatter diagram that could be drawn to illustrate the values of \(c\) and \(h\) for this year. There is no need to draw a diagram. [2]
(c) The manager makes the following statement.
“The value of \(r\) shows that when we spend more on heating, fewer customers visit the store. So we should spend less on heating.”
Comment briefly on this statement, making reference to the context. [2]
(d) Give a statement about a probability to explain the meaning of the value 0.7155 in the table below. [2]

Critical values of Pearson’s product-moment correlation coefficient

1-tail test5%2.5%1%0.5%
2-tail test10%5%2%1%
\(n\)100.54940.63190.71550.7646
110.52140.60210.68510.7348
120.49730.57600.65810.7079
130.47620.55290.63390.6835

October 2021 Paper 2 Q10

OCR ACurrent spec6 marksIncludes hypothesis testingCorrelation & Regression

10 A researcher plans to carry out a statistical investigation to test whether there is linear correlation between the time (\(T\) weeks) from conception to birth, and the birth weight (\(W\) grams) of new-born babies.

(a) Explain why a 1-tail test is appropriate in this context. [1]

The researcher records the values of \(T\) and \(W\) for a random sample of 11 babies. They calculate Pearson’s product-moment correlation coefficient for the sample and find that the value is 0.722.

(b) Use the table below to carry out the test at the 1% significance level. [5]

Critical values of Pearson’s product-moment correlation coefficient.

1-tail test5%2.5%1%0.5%
2-tail test10%5%2.5%1%
\(n\)100.54940.63190.71550.7646
110.52140.60210.68510.7348
120.49730.57600.65810.7079
130.47620.55290.63390.6835

June 2025 Paper 2 Q15

OCR MEICurrent spec10 marksIncludes samplingCorrelation & RegressionData Processing

15 A personal trainer is investigating whether, in the general population, there is any association between resting pulse rate in beats per minute and mean hours per week spent running.

One week he collects a sample by asking 25 of his clients for the relevant data.

(a)
(i) State the name of the sampling technique used by the personal trainer. [1]
(ii) Explain why this sampling technique might introduce bias. [1]

A biologist researching the same topic collects a random sample of size 47. She represents the data using the scatter diagram shown below.

Scatter diagram of resting pulse rate against mean hours per week spent running, 0 to 11 hours: higher, more spread pulse rates below about 2.5 hours and lower pulse rates (about 40 to 60) above
(b) Describe the association between resting pulse rate and mean hours per week spent running. [1]
(c) The biologist identifies two distinct regions on the scatter diagram. Identify these regions on the copy of the scatter diagram in the Printed Answer Booklet by drawing an appropriate vertical line between them. [1]

According to medical research, the normal resting pulse rate for an adult is between 60 and 110 beats per minute.

The biologist separates the data into a group to the left of the vertical line on the scatter diagram and a group to the right of the vertical line on the scatter diagram. She calculates Spearman’s rank correlation coefficient, \(r_s\), and the associated \(p\)-value for each group.

The results are shown in the table.

Group\(r_s\)\(p\)-value
To the left of the vertical line0.209 360.419 98
To the right of the vertical line–0.613 490.000 31
(d) With reference to the two groups identified by the biologist and to the values in the table, explain what may be inferred about the association between resting pulse rate and mean number of hours per week spent running. [6]

June 2024 Paper 2 Q11

OCR MEICurrent spec5 marksIncludes hypothesis testingCorrelation & Regression

11 A householder is investigating whether there is any relationship between his monthly cost of gas and his monthly cost of electricity, both measured in pounds (£). The householder collects a random sample of monthly costs and presents them in the scatter diagram below.

Scatter diagram of gas cost in pounds (0 to 100) against electricity cost in pounds (0 to 140); most points lie between electricity 50 and 135 with gas 45 to 92, and one point at about (46, 15)

One of the points on the diagram represents the energy costs in a month when the householder was away on holiday for three weeks. The other points represent the energy costs in months when the householder did not go away on holiday.

(a) On the copy of the diagram in the Printed Answer Booklet, circle the point which represents the month when the householder was most likely to have been away on holiday for three weeks. [1]
(b) With reference to the diagram, describe the relationship between the cost of gas and the cost of electricity. [1]

The householder decides to test whether there is evidence to suggest that there is any association between the monthly cost of gas and the monthly cost of electricity. The value of Spearman’s rank correlation coefficient for this sample is 0.4359 and the associated \(p\)-value is 0.091 95.

(c) Determine whether there is any evidence to suggest, at the 5% level, that there is any association between the monthly cost of gas and the monthly cost of electricity. [3]

June 2023 Paper 2 Q14

OCR MEICurrent spec8 marksCorrelation & RegressionLarge Data Set

14 The pre-release material contains information concerning the median income of taxpayers in £ and the percentage of all pupils at the end of KS4 achieving 5 or more GCSEs at grade A*–C, including English and Maths, for different areas of London.

Some of the data for 2014/15 is shown in Fig. 14.1.

Fig. 14.1

Median Income of Taxpayers in £Percentage of Pupils Achieving 5 or more A*–C, including English and Maths
City of London61 100#N/A
Barking and Dagenham21 80054.0
Barnet27 10070.1
Bexley24 40055.0
Brent22 70060.0
Bromley28 10068.0

A student investigated whether there is any relationship between median income of taxpayers and percentage of pupils achieving 5 or more GCSEs at grade A*–C, including English and Maths.

(a) With reference to Fig. 14.1, explain how the data should be cleaned before any analysis can take place. [1]

After the data was cleaned, the student used software to draw the scatter diagram shown in Fig. 14.2.

Fig. 14.2: scatter diagram of percentage of students (40 to 75) against median income in pounds (15 000 to 40 000); most points cluster between 20 000 and 30 000 with percentages 52 to 73, and a few points between 31 000 and 39 000
Fig. 14.2

The student calculated that the product moment correlation coefficient for these data is 0.3743.

(b) Give two reasons why it may not be appropriate to use a linear model for the relationship between median income of taxpayers in £ and the percentage of all pupils at the end of KS4 achieving 5 or more GCSEs at grade A*–C. [2]

The student carried out some further analysis. The results are shown in Fig. 14.3.

Fig. 14.3

median income of taxpayers in £percentage of pupils achieving 5+ A*–C
mean27 21661.0
standard deviation4177.55.32

The student identified three outliers in total.

(c)
  • Use the information in Fig. 14.3 to determine the range of values of the median income of taxpayers in £ which are outliers.
  • Use the information in Fig. 14.3 to determine the range of values of the percentage of all pupils at the end of KS4 achieving 5 or more GCSEs at grade A*–C which are outliers.
  • On the copy of Fig. 14.2 in the Printed Answer Booklet, circle the three outliers identified by the student.
[4]

The student decided to remove these outliers and recalculate the product moment correlation coefficient.

(d) Explain whether the new value of the product moment correlation coefficient would be between 0.3743 and 1 or between 0 and 0.3743. [1]

June 2022 Paper 2 Q15

OCR MEICurrent spec9 marksCorrelation & RegressionLarge Data Set

15 The pre-release material includes information on life expectancy at birth in countries of the world. Fig. 15.1 shows the data for Liberia, which is in Africa, together with a time series graph.

Time series graph “Life expectancy at birth – Liberia”: life expectancy at birth against year from 1960 to 2010, increasing from about 35 to about 60
196019701980199020002010
34.6739.2546.0047.1852.4259.63

Fig. 15.1

Sundip uses the LINEST function on a spreadsheet to model life expectancy as a function of calendar year by a straight line.

The equation of this line is \(L = 0.473y - 892\), where \(L\) is life expectancy at birth and \(y\) is calendar year.

(a) Use this model to find an estimate of the life expectancy at birth in Liberia in 1995. [1]

According to the model, the life expectancy at birth in Liberia in 2025 is estimated to be 65.83 years.

(b) Explain whether each of these two estimates is likely to be reliable. [2]
(c) Use your knowledge of the pre-release material to explain whether this model could be used to obtain a reliable estimate of the life expectancy at birth in other countries in 1995. [1]

Fig. 15.2 shows the life expectancy at birth between 1960 and 2010 for Italy and South Africa.

Fig. 15.2: “Life expectancy at birth – Italy and South Africa”, 1960 to 2010; Series 1 (solid line) rises from about 53 to about 62 in 1990 then falls to about 56; Series 2 (dashed line) rises steadily from about 69 to about 82
Fig. 15.2
(d) Use your knowledge of the pre-release material to
  • Explain whether series 1 or series 2 represents the data for Italy.
  • Explain how the data for South Africa differs from the data for most developed countries.
[2]

Sundip is investigating whether there is an association between the wealth of a country and life expectancy at birth in that country. As part of her analysis she draws a scatter diagram of GDP per capita in US$ and life expectancy at birth in 2010 for all the countries in Europe for which data is available. She accidentally includes the data for the Central African Republic. The diagram is shown in Fig. 15.3.

Fig. 15.3: scatter diagram of life expectancy at birth in 2010 (40 to 85) against GDP per capita in US dollars (0 to 160 000); most points between 70 and 82 for GDP up to about 75 000, two points near 80 at about 105 000 and 140 000, and one isolated point at about 47.5 near GDP 0
Fig. 15.3
(e) On the copy of Fig. 15.3 in the Printed Answer Booklet, use your knowledge of the pre-release material to circle the point representing the data for the Central African Republic. [1]

Sundip states that as GDP per capita increases, life expectancy at birth increases.

(f) Explain to what extent the information in Fig. 15.3 supports Sundip’s statement. [2]

October 2020 Paper 2 Q13

OCR MEICurrent spec7 marksIncludes hypothesis testingCorrelation & RegressionLarge Data Set

13 The pre-release material contains information concerning median house prices, recycling rates and employment rates. Fig. 13.1 shows a scatter diagram of recycling rate against employment rate for a random sample of 33 regions.

Fig. 13.1: scatter diagram of recycling rate (10 to 55) against employment rate (60 to 85) for 33 regions, showing weak positive correlation
Fig. 13.1

The product moment correlation coefficient for this sample is 0.37154 and the associated \(p\)-value is 0.033.

Lee conducts a hypothesis test at the 5% level to test whether there is any evidence to suggest there is positive correlation between recycling rate and employment rate. He concludes that there is no evidence to suggest positive correlation because \(0.033 \approx 0\) and \(0.37154 > 0.05\).

(a) Explain whether Lee’s reasoning is correct. [2]

Fig. 13.2 shows a scatter diagram of recycling rate against median house price for a random sample of 33 regions.

Fig. 13.2: scatter diagram of recycling rate against median house price (120 000 to 1 320 000) for 33 regions; most points lie between 250 000 and 550 000, with isolated points near 920 000 and 1 200 000
Fig. 13.2

The product moment correlation coefficient for this sample is \(-0.33278\) and the associated \(p\)-value is 0.058.

Fig. 13.3 shows summary statistics for the median house prices for the data in this sample.

Statistics
n33
Mean465467.9697
σ201236.1345
s204356.2606
Σx15360443
Σx28486161617387
Min243500
Q1342500
Median410000
Q3521000
Max1200000

Fig. 13.3

(b) Use the information in Fig. 13.3 and Fig. 13.2 to show that there are at least two outliers. [2]
(c) Describe the effect of removing the outliers on
  • the product moment correlation coefficient between recycling rate and median house price,
  • the \(p\)-value associated with this correlation coefficient,
in each case explaining your answer. [2]

All 33 items in the sample are areas in London. A student suggests that it is very unlikely that only areas in London would be selected in a random sample.

(d) Use your knowledge of the pre-release material to explain whether you think the student’s suggestion is reasonable. [1]

October 2020 Paper 2 Q11

OCR MEICurrent spec10 marksCorrelation & RegressionLarge Data Set

11 The pre-release material contains information concerning median house prices over the period 2004 – 2015. A spreadsheet has been used to generate a time series graph for two areas: the London borough of “Barking and Dagenham” and “North West”. This is shown together with the raw data in Fig. 11.1.

Fig. 11.1: time series graph of median house price from 2004 to 2015; Barking and Dagenham (dashed) is always above North West (solid)
YearBarking and DagenhamNorth West
2004160 000107 000
2005163 000118 000
2006168 000127 000
2007185 000134 750
2008190 000129 950
2009160 000130 000
2010171 000130 000
2011170 000127 000
2012174 995130 000
2013180 995131 000
2014215 000138 500
2015243 500140 000

Fig. 11.1

Dr Procter suggests that it is unusual for median house prices in a London borough to be consistently higher than those in other parts of the country.

(a) Use your knowledge of the large data set to comment on Dr Procter’s suggestion. [1]

Dr Procter wishes to predict the median house price in Barking and Dagenham in 2016. She uses the spreadsheet function LINEST to find the equation of the line of best fit for the given data. She obtains the equation

\(P = 4897Y - 9\,657\,847\), where \(P\) is the median house price in pounds and \(Y\) is the calendar year, for example 2015.

(b) Use Dr Procter’s equation to predict the median house price in Barking and Dagenham in
  • 2016
  • 2017.
[2]

Professor Jackson uses a simpler model by using the data from 2014 and 2015 only to form a straight-line model.

(c) Find the equation Professor Jackson uses in her model. [2]
(d) Use Professor Jackson’s equation to predict the median house price in Barking and Dagenham in
  • 2016
  • 2017.
[2]

Professor Jackson carries out some research online. She finds some information about median house prices in Barking and Dagenham, which is shown in Fig. 11.2.

20162017
£290 000£300 000
Fig. 11.2

(e) Comment on how well
  • Dr Procter’s model fits the data,
  • Professor Jackson’s model fits the data.
[2]
(f) Explain which, if any, of the models is likely to be more reliable for predicting median house prices in Barking and Dagenham in 2020. [1]