Statistics · Averages & range
1 / 16
Estimating averages from grouped data
Why a mean from a grouped frequency table can only ever be an estimate, how to use midpoints in $\frac{\sum fx}{\sum f}$, and why the mode and median come out as a class rather than a number.
Statistics · Averages & range
Estimating averages from grouped data
Why a mean from a grouped frequency table can only ever be an estimate, how to use midpoints in $\frac{\sum fx}{\sum f}$, and why the mode and median come out as a class rather than a number.
Why it works
A grouped frequency table is a summary, and summarising costs you something. Once 40 waiting times have been recorded as "13 of them were between 10 and 20 minutes", those 13 individual times are gone. Nobody — not you, not the examiner — can get them back out of the table. That one fact is the whole of this topic: grouping destroys information, so any average you calculate afterwards is a reconstruction rather than a measurement. This is why the exam always says "work out an estimate for the mean". The word is not padding, and an answer that never admits it is an estimate has missed the point of the question.So what can you do? You know how many values sit in each class, and you know the two ends of the class. What you don't know is where inside the class each value sits. The honest move is to pick one number to stand in for every value in that class — and the fairest stand-in is the midpoint, because it is the number that is least wrong. Every value in is within of , and no other choice keeps all of them that close.
Using the midpoint is exactly the same as assuming the values are spread evenly through the class. If they are, the ones above the midpoint and the ones below it cancel out and the estimate is spot on. If they aren't, the estimate drifts — and predicting which way it drifts is itself an exam skill (see the end of this section).
Finding the midpoint. Add the two class boundaries and halve them. For that is . Three wrong answers are tempting here, and each is worth naming:
- 10 — the lower boundary, "because the class starts there". That pretends
- 20 — the upper boundary. The same mistake in the other direction.
- 15.5 — treating the class as the whole numbers and
The estimated mean. An ordinary mean is total how many. Here the total has to be estimated too: a class with frequency and midpoint contributes about to the total, so the estimated total is and
That denominator is — the number of values, not the number of classes. A table with four rows describing 40 people is still 40 people, so you divide by 40, never by 4.
The modal class. The mode is "the most common value", but you no longer have values — you have classes. The most you can honestly say is which class holds the most values, and that answer is a class, not a number. For the table above the modal class is . It is not (that is the frequency — how many, not how long) and it is not (that is a midpoint you invented). One warning: a modal class only means something when the classes are the same width, because a wider class collects more values for free.
The class containing the median. The median is the middle value once the data is in order — and grouping does keep the order, because the classes are already listed from smallest to largest. So you can still say which class the middle value fell into, even though you can't say what it was. Find the position first: with values the middle is the th value, so for it is the th — that is, halfway between the 20th and the 21st. Then run a cumulative (running) total down the frequency column and find the first row whose running total reaches that position. Once more, the answer is an interval. That is precisely why the exam asks for "the class interval containing the median" rather than "the median": asking for a number would be asking you to invent one.
Too high or too low? Because the midpoint assumes an even spread, you can often predict the direction of the error when you know the spread isn't even. A wide class at the top of a table — say — usually mops up a handful of stragglers who are all bunched near its lower end. Its midpoint of then overstates every one of them, comes out too big, and the estimate is too high. Bunched near the top of their classes instead, and the estimate comes out too low. Saying which way, and why, is often the final mark on the question — and it is a mark you can only get if you have understood that the midpoint was an assumption all along.