Statistics · Data presentation
1 / 11
Types of data and grouped frequency
Qualitative vs quantitative data, and within quantitative the split between discrete and continuous; how continuous data is grouped into classes, and how to read off the class boundaries, class width and midpoint you need before you can draw a histogram or estimate an average.
Statistics · Data presentation
Types of data and grouped frequency
Qualitative vs quantitative data, and within quantitative the split between discrete and continuous; how continuous data is grouped into classes, and how to read off the class boundaries, class width and midpoint you need before you can draw a histogram or estimate an average.
Why it works
Before you can summarise or draw data you have to know what kind it is, because that decides which tools are allowed.Qualitative vs quantitative. Quantitative data is numerical — heights, times, number of goals. Qualitative (or categorical) data describes a quality with no numerical value — eye colour, type of car, cardinal wind direction (N, SE, …). You can count how many fall in each category, but you can't take a mean of "north".
Discrete vs continuous. Quantitative data divides again:
- Discrete data can only take particular separate values — usually whole-number
- Continuous data can take any value in a range, limited only by the accuracy
Grouping data into classes. With a lot of continuous data we collect values into classes (intervals). Three numbers matter for each class, and getting them right is what makes histograms and grouped averages work:
- Class boundaries — the values where one class actually stops and the next
- Class width upper boundary lower boundary.
- Midpoint — used
When the class is written with inequalities, the boundaries are obvious. For the boundaries are and , the width is and the midpoint is .
The hidden-gap case. When continuous data is recorded to the nearest unit, the printed limits hide the real boundaries. Heights recorded to the nearest cm and grouped –, – look as if they have a cm gap, but a height of cm rounds to . So the class – really covers up to :
- boundaries and ,
- width ,
- midpoint .
Discrete grouped data is treated more simply — for –, – goals the "classes" are just blocks of values; for averages you use the midpoint of the stated limits (e.g. ). Histograms, though, are for continuous data, so when a histogram is asked of discrete-looking groups you use continuous boundaries.