Leave lesson

Statistics · Cumulative frequency & box plots

1 / 17

Cumulative frequency diagrams

How to build a running total from a grouped frequency table, why every point is plotted at the upper class boundary rather than the midpoint, and how to read estimates off the curve in both directions.

Statistics · Cumulative frequency & box plots

Cumulative frequency diagrams

How to build a running total from a grouped frequency table, why every point is plotted at the upper class boundary rather than the midpoint, and how to read estimates off the curve in both directions.

Why it works

An ordinary grouped frequency table tells you how many values fall inside each class. A cumulative frequency table answers a different question: how many values are at or below a given point? You build it by keeping a running total, adding the classes on from the bottom.

Here are the ages, in years, of the 60 members of a swimming club.
Age, aa (years)FrequencyCumulative frequency
0<a100 < a \le 1066
10<a2010 < a \le 201420
20<a3020 < a \le 302040
30<a4030 < a \le 401353
40<a5040 < a \le 50760
Each cumulative entry is the one above it plus the new frequency: 66, then 6+14=206 + 14 = 20, then 20+20=4020 + 20 = 40, then 40+13=5340 + 13 = 53, then 53+7=6053 + 7 = 60. The last cumulative frequency must come out as the total number of values — that is your free check that you have added correctly.

The one thing that really matters: what goes on the horizontal axis.

Say the 2020 in the third row out loud, in full. It does not mean "20 members are in the class 10<a2010 < a \le 20". It means "20 members are 20 years old or younger". Notice that the sentence contains an age — 20 years — and that age is the top of the class. Every cumulative frequency is a statement about a boundary, and it is always the upper boundary, because a running total sweeps up everything up to and including that point. So the point you plot is (20, 20)(20,\ 20): at a=20a = 20, the count so far is 2020.

The tempting alternative is the midpoint, 1515 — partly because midpoints are exactly what you use for an estimated mean from a grouped table, and this looks like the same sort of table. It isn't. A midpoint is a stand-in for the *values inside a class; a cumulative frequency is a count up to a boundary*. Different jobs, different xx.

The trap, made concrete. Plot at the midpoints and the five points become (5, 6)(5,\ 6), (15, 20)(15,\ 20), (25, 40)(25,\ 40), (35, 53)(35,\ 53), (45, 60)(45,\ 60):1020304050102030405060Age (years)Cumulative frequencyRead that last point aloud: "all 60 members are 45 or younger." Nothing in the table says that. The table says the oldest member is somewhere in 40<a5040 < a \le 50, so the only honest claim is that all 60 are 50 or younger. The midpoint graph has quietly made everybody five years younger. The first point is wrong in the same way: (5, 6)(5,\ 6) claims 6 members are 5 or younger, when the table only tells you that 6 of them are 10 or younger — for all you know every one of those 6 is nine years old. Each point has slid half a class to the left, so every reading taken off that curve comes out too small. This isn't untidiness; it is the wrong graph.

Where the curve starts. Below the first class there is nothing at all: no member is 0 years old or younger, so the curve starts at (0, 0)(0,\ 0) — the lower boundary of the first class, paired with a count of zero. Five classes, six points.

Joining them up. Join the points in order with a smooth increasing curve (ruled straight segments are accepted too, and give almost the same readings). A cumulative frequency graph can only ever go up or stay flat — never down — because a running total cannot shrink when you let more values in. If your curve dips anywhere, you have plotted a plain frequency instead of the running total, or slipped in the addition, so go back and re-add.1020304050102030405060Age (years)Cumulative frequencyReading it forwards. "How many members are 25 or younger?" Go up from 2525 on the horizontal axis to the curve, then straight across to the cumulative frequency axis: about 3030 members.

Reading it backwards. "The youngest 45 members are all below what age?" Now you start on the vertical axis: across from 4545 to the curve, then straight down: about 3434 years. The two readings run in opposite directions, and going up from 4545 on the age axis instead is one of the most common slips in this topic. Ask yourself which axis the number in the question lives on — a count belongs on the vertical axis, a measurement on the horizontal one.

Between two values. The curve only ever answers "how many so far?", so to get "how many members are between 15 and 35 years old" you read both and subtract: about 4646 are 35 or younger, about 1313 are 15 or younger, so about 4613=3346 - 13 = 33 members lie in between. Adding the two readings answers no question at all.

Why every answer is an estimate. Two things stack up. First, grouping has already destroyed the individual ages — the table cannot say where inside a class anyone sits, and drawing a curve between the boundary points quietly assumes the values are spread evenly through each class. Second, you are then reading a number off a drawn line by eye. That is why the question says "estimate" and why the mark scheme accepts a range rather than one number. The plotted points themselves are exact; everything read from between them is an estimate.