Statistics · Histograms
1 / 18
Interpreting histograms
Getting information back out of a histogram — a frequency is an area, so part of a class is the matching proportion of a bar's area. Why that rests on values being spread evenly within a class (and is therefore an estimate), and how to read off a total, a mean, the median class, a range that cuts across class boundaries, and a comparison of two groups.
Statistics · Histograms
Interpreting histograms
Getting information back out of a histogram — a frequency is an area, so part of a class is the matching proportion of a bar's area. Why that rests on values being spread evenly within a class (and is therefore an estimate), and how to read off a total, a mean, the median class, a range that cuts across class boundaries, and a comparison of two groups.
Why it works
Building a histogram turns frequencies into areas. Reading one turns areas back into frequencies. That is the whole of this page. The height of a bar is the frequency density, soand every question below is secretly the same question: which area am I being asked for?
Here is the histogram we will read all the way through. It shows the times, in minutes, that people took to finish a task.A whole bar is the easy case: is wide and tall, so it holds people. Do that for all five bars and you get , , , , ; adding those gives . The total frequency is the total area — there is no other way to get the total off a histogram, because the heights on their own () count nobody.
Part of a bar is where the thinking is. How many people took between and minutes? The histogram does not record that. It was drawn from a grouped table, and grouping threw away where inside each person actually landed. All that survived is " people, somewhere in there".
So you make the one assumption the picture itself suggests. The bar is a rectangle with a flat top: the same frequency density all the way across the class. Constant density means the values are spread evenly through the class. If they are, half the class width holds half the people, a quarter holds a quarter, and in general
For to that is people — which is just the area of that slice of the bar, . The same calculation said two ways.
Why the answer is only an estimate, made concrete. Suppose the truth was that all of those people finished between and minutes. The grouped table would still say ": people", the histogram would look identical, and yet the true number between and would be , not . Nothing in the diagram can tell those two worlds apart. That is why every answer here is an estimate, and why "assuming the times are spread evenly throughout each class" is the sentence an examiner wants. It is not a polite hedge — it is the assumption doing all the work.
When the question's range doesn't line up with the classes. "How many took between and minutes?" That range cuts through and through . Chop it at every class boundary inside it and do each piece on its own:
- to is of the wide class: .
- to is of the wide class: .
An estimate for the mean. You never see the individual times, so each class is represented by its midpoint — and that is the even-spread assumption again, because evenly spread values average out at the middle. Recover each frequency from its area, multiply by the midpoint, add, and divide by the total:
Both of those must be frequencies, not densities: what you multiply the midpoints by, and what you divide by. Dividing by because there are five bars is the classic wreck — a mean is per person, not per class.
The median. The median is the middle value once everything is in order, so build up running totals from the left until you pass half the total. Half of is ; the running totals are , , , so the th person sits in — that is the median class. Notice it is not the tallest bar: is taller, but height alone says nothing about where the middle of the data lies. If you want a value rather than a class, run the even-spread idea backwards: you need more people out of the in that class, so travel of the way across a class of width , giving minutes.
Comparing two groups. Two histograms drawn for groups of different sizes cannot be compared by raw counts: out of is a bigger share () than out of (), even though . Turn each count into a proportion of its own total first. And before you compare bar heights by eye, check the two frequency-density scales actually match — two histograms can look alike and mean very different things.