Statistics · Cumulative frequency & box plots
1 / 16
Quartiles & the interquartile range
Why the quartiles are a quarter and three quarters of the way *through the data* rather than along the number line, how to find them from an ordered list and from a cumulative frequency curve, and why the interquartile range beats the range as a measure of spread.
Statistics · Cumulative frequency & box plots
Quartiles & the interquartile range
Why the quartiles are a quarter and three quarters of the way *through the data* rather than along the number line, how to find them from an ordered list and from a cumulative frequency curve, and why the interquartile range beats the range as a measure of spread.
Why it works
The median splits ordered data in half. The quartiles do the same job at the quarter marks: the lower quartile (LQ) is the value a quarter of the way through the data, and the upper quartile (UQ) is the value three quarters of the way through. With the median between them they cut the data into four groups of equal size — a quarter of the values in each.That phrase — a quarter of the way through the data — is the whole topic, and it is where most of the marks are lost. It does not mean a quarter of the way along the number line.
The trap, made concrete. Here are eleven numbers, already in order:
The smallest is and the largest is , so the range is . A quarter of the way along that number line is . Is the lower quartile? Count the values below it: — six out of eleven, more than half the data. So is nowhere near a quarter of the way through the data; it only looks like a quarter because one enormous value, , has stretched the number line. Counting is what matters, not measuring.
Positions in an ordered list. For values written in order, the convention used at GCSE for a raw list is
The is there so the quartiles agree with the median rule you already use: with , the median is the th value, and the quartiles fall at positions and — the same distance in from each end. So LQ , median , UQ . Two warnings. First, these formulas give a position, not a value: means "the 3rd number", not "3". Second, the list must be in order first — reading the 3rd number off an unsorted list is meaningless.
If a position lands on a half — gives — take the number halfway between the two values it falls between.
On a cumulative frequency curve the convention changes. There you have no list to count along, just a smooth curve summarising or values, so you work on the cumulative frequency axis at
— not . Why is that allowed? Because the difference is a quarter of one unit. With it is against : a quarter of a frequency, which on the axis is thinner than the pencil line you draw with. A reading off a curve is an estimate anyway, so the correction is invisible. On a short list of eleven numbers the same correction moves you from the th value to the rd — a different number — so there it genuinely matters. Throughout this concept: on a raw ordered list, on a cumulative frequency curve.
Reading the curve has its own trap. Cumulative frequency (a how many) is on the vertical axis; the quantity (a how much) is on the horizontal. So start on the vertical axis at , go across to the curve, then drop down to read the value. Start on the horizontal axis instead and you come away with a frequency where a quartile should be.
The interquartile range. . It is the width of the interval holding the middle half of the data — a quarter of the values have been trimmed off the bottom and a quarter off the top before the measuring starts.
Why that beats the range. The range uses exactly two numbers, and they are the two most extreme ones — precisely the two most likely to be a freak, a mistake or a one-off. For the eleven numbers above, the range is and the IQR is . Now suppose that was really (a slip of the pen, or one genuinely unusual case). The range leaps to , more than four times bigger — but the LQ and UQ have not moved at all, so the IQR is still . One value out of eleven changed the range by and the IQR by nothing. That is the whole argument: the IQR is resistant to extreme values, because it throws the extremes away before it measures anything.
A value sitting well away from the rest, like that , is an outlier — at GCSE you spot one informally, by ordering the data or looking at the diagram and seeing whether a value sits far from the bulk. When there is one, quote the IQR rather than the range, and say why.
Comparing two sets of data — where the marks actually are. A comparison question wants two separate things: a comparison of an average (usually the median) and a comparison of a spread (usually the IQR), both written in context. "A's median is and B's is " scores nothing — that restates two numbers without comparing them and without saying what they mean. "A's median time is lower, so group A were generally quicker, and A's IQR is smaller, so their times were more consistent" scores both marks: two comparisons, each interpreted in the language of the question.