Browse Data presentation

116 questions at your level

Measures of location — mean, median, mode and quartiles

33 questions

LessonNot started
A ferry company records the number of passengers on each of 30 morning crossings and each of 30 evening crossings of a river.

The results are summarised in the back-to-back stem and leaf diagram below.
TotalsMorningStemEveningTotals
(0)18(1)
(2)8 624 7 9(3)
(6)9 7 5 5 3 130 2 6 6 8(5)
(10)9 8 6 4 4 4 4 3 2 041 3 5 5 7 7 9(7)
(5)8 6 5 2 150 2 3 4 6 8(6)
(3)7 3 061 4 6(3)
(2)5 270 3(2)
(2)9 385(1)
(0)92 6(2)
Key: 3∣4∣53 \mid 4 \mid 5 means 43 passengers on a morning crossing and 45 passengers on an evening crossing
(a) Write down the modal number of passengers for these morning crossings.

(1 mark)

Some of the quartiles for these two distributions are shown in the table below.
MorningEvening
Lower quartileaa36
Medianbb48
Upper quartile58cc
(b) Find the value of aa, the value of bb and the value of cc

(3 marks)
(c) For these morning crossings find, to one decimal place,

(i) the mean number of passengers,

(ii) the standard deviation of the number of passengers.

(You may use ∑x=1489\sum x = 1489 and ∑x2=81135\sum x^2 = 81135 where xx is the number of passengers on a morning crossing.)

(3 marks)

One measure of skewness is found using

3(mean−median)standard deviation\frac{3(\text{mean} - \text{median})}{\text{standard deviation}}
(d) Evaluate this measure and describe the skewness of the numbers of passengers on these morning crossings.

(2 marks)
(e) Comment on one difference between the distribution of the numbers of passengers on these morning crossings and the distribution of the numbers of passengers on these evening crossings. State the values of any statistics you have used to support your comment.

(1 mark)
●●●●●Level 410 marksStart

Measures of spread, standard deviation and coding

38 questions

LessonNot started
A ferry company records the number of passengers on each of 30 morning crossings and each of 30 evening crossings of a river.

The results are summarised in the back-to-back stem and leaf diagram below.
TotalsMorningStemEveningTotals
(0)18(1)
(2)8 624 7 9(3)
(6)9 7 5 5 3 130 2 6 6 8(5)
(10)9 8 6 4 4 4 4 3 2 041 3 5 5 7 7 9(7)
(5)8 6 5 2 150 2 3 4 6 8(6)
(3)7 3 061 4 6(3)
(2)5 270 3(2)
(2)9 385(1)
(0)92 6(2)
Key: 3∣4∣53 \mid 4 \mid 5 means 43 passengers on a morning crossing and 45 passengers on an evening crossing
(a) Write down the modal number of passengers for these morning crossings.

(1 mark)

Some of the quartiles for these two distributions are shown in the table below.
MorningEvening
Lower quartileaa36
Medianbb48
Upper quartile58cc
(b) Find the value of aa, the value of bb and the value of cc

(3 marks)
(c) For these morning crossings find, to one decimal place,

(i) the mean number of passengers,

(ii) the standard deviation of the number of passengers.

(You may use ∑x=1489\sum x = 1489 and ∑x2=81135\sum x^2 = 81135 where xx is the number of passengers on a morning crossing.)

(3 marks)

One measure of skewness is found using

3(mean−median)standard deviation\frac{3(\text{mean} - \text{median})}{\text{standard deviation}}
(d) Evaluate this measure and describe the skewness of the numbers of passengers on these morning crossings.

(2 marks)
(e) Comment on one difference between the distribution of the numbers of passengers on these morning crossings and the distribution of the numbers of passengers on these evening crossings. State the values of any statistics you have used to support your comment.

(1 mark)
●●●●●Level 410 marksStart
A company asked 1111 of its employees for the distance, dd km, from home to the office and the time, tt minutes, of their journey to work on one morning. The results are shown in the table below.
EmployeeABCDEFGHIJK
Distance (dd km)35689111214171922
Time (tt minutes)1215212095262831334043
On that morning, employee E was delayed by a cancelled train.

An outlier is defined as a value that is greater than Q3+1.5×(Q3−Q1)Q_3 + 1.5 \times (Q_3 - Q_1) or smaller than Q1−1.5×(Q3−Q1)Q_1 - 1.5 \times (Q_3 - Q_1)
(a) Show that 9595 is an outlier for the journey times.

(3 marks)

Leaving out employee E, the company calculated the following summary statistics for the other 1010 employees.

∑d=117∑t=269Sdd=360.1Sdt=572.7\sum d = 117 \qquad \sum t = 269 \qquad S_{dd} = 360.1 \qquad S_{dt} = 572.7
(b) Use these summary statistics to show that the equation of the least squares regression line of tt on dd for these 1010 employees is

t=8.29+1.59dt = 8.29 + 1.59d

where the values of the intercept and gradient are given to 3 significant figures. You must show your working.

(3 marks)
(c) Give an interpretation of the gradient of the regression line.

(1 mark)

Two new employees live 1616 km and 3030 km from the office.
(d) Using the equation given in part (b), estimate the journey time for

(i) the employee who lives 1616 km from the office,

(ii) the employee who lives 3030 km from the office.

(3 marks)
(e) State, giving a reason, which of the two estimates found in part (d) would be the more reliable estimate.

(2 marks)
●●●●●Level 412 marksStart

Cumulative frequency, box plots and outliers

18 questions

LessonNot started
A ferry company records the number of passengers on each of 30 morning crossings and each of 30 evening crossings of a river.

The results are summarised in the back-to-back stem and leaf diagram below.
TotalsMorningStemEveningTotals
(0)18(1)
(2)8 624 7 9(3)
(6)9 7 5 5 3 130 2 6 6 8(5)
(10)9 8 6 4 4 4 4 3 2 041 3 5 5 7 7 9(7)
(5)8 6 5 2 150 2 3 4 6 8(6)
(3)7 3 061 4 6(3)
(2)5 270 3(2)
(2)9 385(1)
(0)92 6(2)
Key: 3∣4∣53 \mid 4 \mid 5 means 43 passengers on a morning crossing and 45 passengers on an evening crossing
(a) Write down the modal number of passengers for these morning crossings.

(1 mark)

Some of the quartiles for these two distributions are shown in the table below.
MorningEvening
Lower quartileaa36
Medianbb48
Upper quartile58cc
(b) Find the value of aa, the value of bb and the value of cc

(3 marks)
(c) For these morning crossings find, to one decimal place,

(i) the mean number of passengers,

(ii) the standard deviation of the number of passengers.

(You may use ∑x=1489\sum x = 1489 and ∑x2=81135\sum x^2 = 81135 where xx is the number of passengers on a morning crossing.)

(3 marks)

One measure of skewness is found using

3(mean−median)standard deviation\frac{3(\text{mean} - \text{median})}{\text{standard deviation}}
(d) Evaluate this measure and describe the skewness of the numbers of passengers on these morning crossings.

(2 marks)
(e) Comment on one difference between the distribution of the numbers of passengers on these morning crossings and the distribution of the numbers of passengers on these evening crossings. State the values of any statistics you have used to support your comment.

(1 mark)
●●●●●Level 410 marksStart
A music streaming service takes a random sample of 160 songs from one of its playlists and records the length, in seconds, of each song.

The lengths of the songs in the sample are summarised in the following box plot.150200216248310Length (seconds)(a) Use linear interpolation to estimate the probability that a randomly chosen song from this sample is shorter than 232 seconds.

(2 marks)

The service takes the quartiles from this sample and decides to label any song whose length is at least Q3+1.5×(Q3−Q1)Q_3 + 1.5 \times (Q_3 - Q_1) as "extended".
(b) Find the shortest length of a song that the service would label as extended.

(1 mark)

An analyst suggests that the lengths of the songs on the playlist may be modelled by a normal distribution.
(c) Explain whether or not the box plot supports this suggestion.

(1 mark)

The analyst records the length of every song on the playlist and classifies any song whose length is more than 2.5 standard deviations from the mean as unusual.

Assuming that the lengths of the songs on the playlist may be modelled by a normal distribution,
(d) find the probability that a randomly selected song from the playlist would be classified as unusual.

(2 marks)

The mean length of the songs on the playlist is 226 seconds.

Given that a song of length 350 seconds is classified as unusual,
(e) find the maximum possible value of the standard deviation of the lengths of the songs on the playlist.

(2 marks)
●●●●●Level 48 marksStart

Histograms and frequency polygons

15 questions

LessonNot started

Correlation and regression

18 questions

LessonNot started
Callum is studying how the depth of water in a small reservoir changed during one summer.

He measures the depth of the water, yy metres, at time tt days after 1 June, for 1212 values of tt between t=0t = 0 and t=88t = 88

Callum finds the equation of the regression line of yy on tt for his data to be

y=8.13−0.0241ty = 8.13 - 0.0241t
(a) Interpret the gradient of this line.

(1 mark)

The product moment correlation coefficient between yy and tt is −0.6602-0.6602

The critical value for a sample of size 1212 at the 5%5\% level of significance is 0.49730.4973
(b) Test whether or not there is evidence of a negative correlation between the depth of the water and the time.

You should
  • state your hypotheses clearly
  • use a 5%5\% level of significance
  • state the critical value used
(3 marks)

Nina plots Callum's data on a scatter diagram. The points lie close to a curve: the depth falls steeply at first, reaches its lowest value about 5555 days after 1 June, and then rises again.
(c) With reference to Nina's scatter diagram, state, giving a reason, whether or not the regression line y=8.13−0.0241ty = 8.13 - 0.0241t is an appropriate model for these data.

(1 mark)

Nina suggests an improved model using the variable u=(t−k)2u = (t - k)^2, where kk is a constant.

She obtains the equation

y=6.10+0.00111uy = 6.10 + 0.00111u
(d) Choose a suitable value for kk to write Nina's improved model for yy in terms of tt only.

(1 mark)
●●●●●Level 46 marksStart
A meteorologist believes that windier days tend to be colder. She records the daily mean windspeed, ww knots, and the daily mean air temperature, T ∘T\,^\circC, on 1010 randomly chosen days at a weather station. The results are shown in the table.

w4679101214151820T16.213.814.917.013.114.411.915.510.812.6\begin{array}{|c|c|c|c|c|c|c|c|c|c|c|}\hline w & 4 & 6 & 7 & 9 & 10 & 12 & 14 & 15 & 18 & 20 \\ \hline T & 16.2 & 13.8 & 14.9 & 17.0 & 13.1 & 14.4 & 11.9 & 15.5 & 10.8 & 12.6 \\ \hline \end{array}

For these data the product moment correlation coefficient is r=−0.627r = -0.627 and the equation of the regression line of TT on ww is T=16.7−0.234wT = 16.7 - 0.234w.
(a) Test, at the 5%5\% level of significance, whether these data provide evidence to support the meteorologist's belief. State your hypotheses clearly. (4 marks)
(b) Give an interpretation of the gradient of the regression line in this context. (1 mark)
(c) Use the regression line to estimate the daily mean air temperature on a day when the daily mean windspeed is 12.512.5 knots, giving your answer to 11 decimal place. (2 marks)
(d) Explain why it would be unreliable to use this regression line to estimate (i) the daily mean air temperature on a day when the daily mean windspeed is 3535 knots, (ii) the daily mean windspeed on a day when the daily mean air temperature is 12 ∘12\,^\circC. (2 marks)
●●●●●Level 49 marksStart
An agricultural researcher is investigating the effect of a fertiliser on tomato plants. She grows tomato plants in 1010 plots. Each plot is given a different amount of fertiliser, ff grams per square metre, and the mean yield, yy kg per plant, is recorded for each plot. The values of ff used were between 2020 and 8080

The researcher summarises the data as follows

∑f=465∑y=45.6∑y2=217.44∑fy=2295Sff=3552.5\sum f = 465 \qquad \sum y = 45.6 \qquad \sum y^2 = 217.44 \qquad \sum fy = 2295 \qquad S_{ff} = 3552.5
(a) Calculate the exact value of SfyS_{fy} and the exact value of SyyS_{yy}

(3 marks)
(b) Calculate the value of the product moment correlation coefficient between ff and yy

(2 marks)
(c) Give an interpretation, in context, of your product moment correlation coefficient.

(1 mark)
(d) Show that the equation of the regression line of yy on ff can be written as

y=2.27+0.0491fy = 2.27 + 0.0491f

where the values of the intercept and gradient are given to 3 significant figures.

(3 marks)
(e) Give an interpretation, in context, of the gradient of the regression line.

(1 mark)

Using the equation of the regression line given in part (d)
(f) (i) estimate the mean yield per plant for a plot given 4545 grams of fertiliser per square metre,

(ii) explain why an estimate of the mean yield per plant for a plot given 150150 grams of fertiliser per square metre is not reliable.

(2 marks)
●●●●●Level 412 marksStart
A company asked 1111 of its employees for the distance, dd km, from home to the office and the time, tt minutes, of their journey to work on one morning. The results are shown in the table below.
EmployeeABCDEFGHIJK
Distance (dd km)35689111214171922
Time (tt minutes)1215212095262831334043
On that morning, employee E was delayed by a cancelled train.

An outlier is defined as a value that is greater than Q3+1.5×(Q3−Q1)Q_3 + 1.5 \times (Q_3 - Q_1) or smaller than Q1−1.5×(Q3−Q1)Q_1 - 1.5 \times (Q_3 - Q_1)
(a) Show that 9595 is an outlier for the journey times.

(3 marks)

Leaving out employee E, the company calculated the following summary statistics for the other 1010 employees.

∑d=117∑t=269Sdd=360.1Sdt=572.7\sum d = 117 \qquad \sum t = 269 \qquad S_{dd} = 360.1 \qquad S_{dt} = 572.7
(b) Use these summary statistics to show that the equation of the least squares regression line of tt on dd for these 1010 employees is

t=8.29+1.59dt = 8.29 + 1.59d

where the values of the intercept and gradient are given to 3 significant figures. You must show your working.

(3 marks)
(c) Give an interpretation of the gradient of the regression line.

(1 mark)

Two new employees live 1616 km and 3030 km from the office.
(d) Using the equation given in part (b), estimate the journey time for

(i) the employee who lives 1616 km from the office,

(ii) the employee who lives 3030 km from the office.

(3 marks)
(e) State, giving a reason, which of the two estimates found in part (d) would be the more reliable estimate.

(2 marks)
●●●●●Level 412 marksStart

Exponential models and regression

17 questions

LessonNot started

Measuring correlation (the PMCC)

29 questions

LessonNot started
A meteorologist believes that windier days tend to be colder. She records the daily mean windspeed, ww knots, and the daily mean air temperature, T ∘T\,^\circC, on 1010 randomly chosen days at a weather station. The results are shown in the table.

w4679101214151820T16.213.814.917.013.114.411.915.510.812.6\begin{array}{|c|c|c|c|c|c|c|c|c|c|c|}\hline w & 4 & 6 & 7 & 9 & 10 & 12 & 14 & 15 & 18 & 20 \\ \hline T & 16.2 & 13.8 & 14.9 & 17.0 & 13.1 & 14.4 & 11.9 & 15.5 & 10.8 & 12.6 \\ \hline \end{array}

For these data the product moment correlation coefficient is r=−0.627r = -0.627 and the equation of the regression line of TT on ww is T=16.7−0.234wT = 16.7 - 0.234w.
(a) Test, at the 5%5\% level of significance, whether these data provide evidence to support the meteorologist's belief. State your hypotheses clearly. (4 marks)
(b) Give an interpretation of the gradient of the regression line in this context. (1 mark)
(c) Use the regression line to estimate the daily mean air temperature on a day when the daily mean windspeed is 12.512.5 knots, giving your answer to 11 decimal place. (2 marks)
(d) Explain why it would be unreliable to use this regression line to estimate (i) the daily mean air temperature on a day when the daily mean windspeed is 3535 knots, (ii) the daily mean windspeed on a day when the daily mean air temperature is 12 ∘12\,^\circC. (2 marks)
●●●●●Level 49 marksStart
An agricultural researcher is investigating the effect of a fertiliser on tomato plants. She grows tomato plants in 1010 plots. Each plot is given a different amount of fertiliser, ff grams per square metre, and the mean yield, yy kg per plant, is recorded for each plot. The values of ff used were between 2020 and 8080

The researcher summarises the data as follows

∑f=465∑y=45.6∑y2=217.44∑fy=2295Sff=3552.5\sum f = 465 \qquad \sum y = 45.6 \qquad \sum y^2 = 217.44 \qquad \sum fy = 2295 \qquad S_{ff} = 3552.5
(a) Calculate the exact value of SfyS_{fy} and the exact value of SyyS_{yy}

(3 marks)
(b) Calculate the value of the product moment correlation coefficient between ff and yy

(2 marks)
(c) Give an interpretation, in context, of your product moment correlation coefficient.

(1 mark)
(d) Show that the equation of the regression line of yy on ff can be written as

y=2.27+0.0491fy = 2.27 + 0.0491f

where the values of the intercept and gradient are given to 3 significant figures.

(3 marks)
(e) Give an interpretation, in context, of the gradient of the regression line.

(1 mark)

Using the equation of the regression line given in part (d)
(f) (i) estimate the mean yield per plant for a plot given 4545 grams of fertiliser per square metre,

(ii) explain why an estimate of the mean yield per plant for a plot given 150150 grams of fertiliser per square metre is not reliable.

(2 marks)
●●●●●Level 412 marksStart
A company asked 1111 of its employees for the distance, dd km, from home to the office and the time, tt minutes, of their journey to work on one morning. The results are shown in the table below.
EmployeeABCDEFGHIJK
Distance (dd km)35689111214171922
Time (tt minutes)1215212095262831334043
On that morning, employee E was delayed by a cancelled train.

An outlier is defined as a value that is greater than Q3+1.5×(Q3−Q1)Q_3 + 1.5 \times (Q_3 - Q_1) or smaller than Q1−1.5×(Q3−Q1)Q_1 - 1.5 \times (Q_3 - Q_1)
(a) Show that 9595 is an outlier for the journey times.

(3 marks)

Leaving out employee E, the company calculated the following summary statistics for the other 1010 employees.

∑d=117∑t=269Sdd=360.1Sdt=572.7\sum d = 117 \qquad \sum t = 269 \qquad S_{dd} = 360.1 \qquad S_{dt} = 572.7
(b) Use these summary statistics to show that the equation of the least squares regression line of tt on dd for these 1010 employees is

t=8.29+1.59dt = 8.29 + 1.59d

where the values of the intercept and gradient are given to 3 significant figures. You must show your working.

(3 marks)
(c) Give an interpretation of the gradient of the regression line.

(1 mark)

Two new employees live 1616 km and 3030 km from the office.
(d) Using the equation given in part (b), estimate the journey time for

(i) the employee who lives 1616 km from the office,

(ii) the employee who lives 3030 km from the office.

(3 marks)
(e) State, giving a reason, which of the two estimates found in part (d) would be the more reliable estimate.

(2 marks)
●●●●●Level 412 marksStart