Leave lesson

Statistics · Averages & range

1 / 16

Estimating averages from grouped data

Why a mean from a grouped frequency table can only ever be an estimate, how to use midpoints in $\frac{\sum fx}{\sum f}$, and why the mode and median come out as a class rather than a number.

Statistics · Averages & range

Estimating averages from grouped data

Why a mean from a grouped frequency table can only ever be an estimate, how to use midpoints in $\frac{\sum fx}{\sum f}$, and why the mode and median come out as a class rather than a number.

Why it works

A grouped frequency table is a summary, and summarising costs you something. Once 40 waiting times have been recorded as "13 of them were between 10 and 20 minutes", those 13 individual times are gone. Nobody — not you, not the examiner — can get them back out of the table. That one fact is the whole of this topic: grouping destroys information, so any average you calculate afterwards is a reconstruction rather than a measurement. This is why the exam always says "work out an estimate for the mean". The word is not padding, and an answer that never admits it is an estimate has missed the point of the question.

So what can you do? You know how many values sit in each class, and you know the two ends of the class. What you don't know is where inside the class each value sits. The honest move is to pick one number to stand in for every value in that class — and the fairest stand-in is the midpoint, because it is the number that is least wrong. Every value in 10<t2010 < t \le 20 is within 55 of 1515, and no other choice keeps all of them that close.

Using the midpoint is exactly the same as assuming the values are spread evenly through the class. If they are, the ones above the midpoint and the ones below it cancel out and the estimate is spot on. If they aren't, the estimate drifts — and predicting which way it drifts is itself an exam skill (see the end of this section).

Finding the midpoint. Add the two class boundaries and halve them. For 10<t2010 < t \le 20 that is 10+202=15\frac{10 + 20}{2} = 15. Three wrong answers are tempting here, and each is worth naming:
  • 10 — the lower boundary, "because the class starts there". That pretends
every one of the 13 values is the smallest it could possibly be, which drags the whole mean down.
  • 20 — the upper boundary. The same mistake in the other direction.
  • 15.5 — treating the class as the whole numbers 11,12,,2011, 12, \ldots, 20 and
averaging those. Time is continuous: 10<t10 < t does not mean "from 11", it means "anything bigger than 10, however slightly". The boundaries of the class are 1010 and 2020 whatever the inequality signs look like, so the midpoint is 1515. The << and \le signs are there to tell you which class a value of exactly 1010 or exactly 2020 belongs to — nothing more.

The estimated mean. An ordinary mean is total ÷\div how many. Here the total has to be estimated too: a class with frequency ff and midpoint xx contributes about f×xf \times x to the total, so the estimated total is fx\sum fx and

estimated mean=fxf.\text{estimated mean} = \frac{\sum fx}{\sum f}.

That denominator is f\sum f — the number of values, not the number of classes. A table with four rows describing 40 people is still 40 people, so you divide by 40, never by 4.

The modal class. The mode is "the most common value", but you no longer have values — you have classes. The most you can honestly say is which class holds the most values, and that answer is a class, not a number. For the table above the modal class is 10<t2010 < t \le 20. It is not 1313 (that is the frequency — how many, not how long) and it is not 1515 (that is a midpoint you invented). One warning: a modal class only means something when the classes are the same width, because a wider class collects more values for free.

The class containing the median. The median is the middle value once the data is in order — and grouping does keep the order, because the classes are already listed from smallest to largest. So you can still say which class the middle value fell into, even though you can't say what it was. Find the position first: with nn values the middle is the n+12\frac{n+1}{2}th value, so for n=40n = 40 it is the 20.520.5th — that is, halfway between the 20th and the 21st. Then run a cumulative (running) total down the frequency column and find the first row whose running total reaches that position. Once more, the answer is an interval. That is precisely why the exam asks for "the class interval containing the median" rather than "the median": asking for a number would be asking you to invent one.

Too high or too low? Because the midpoint assumes an even spread, you can often predict the direction of the error when you know the spread isn't even. A wide class at the top of a table — say 30<m6030 < m \le 60 — usually mops up a handful of stragglers who are all bunched near its lower end. Its midpoint of 4545 then overstates every one of them, fx\sum fx comes out too big, and the estimate is too high. Bunched near the top of their classes instead, and the estimate comes out too low. Saying which way, and why, is often the final mark on the question — and it is a mark you can only get if you have understood that the midpoint was an assumption all along.