Statistics · Scatter graphs & correlation
1 / 16
Scatter graphs & correlation
Why one point on a scatter graph is one individual measured twice, how to describe correlation by type and strength in the words of the question, why a correlation never proves a cause, and how to place a line of best fit — through the bulk of the data, balanced either side, ignoring a genuine outlier.
Statistics · Scatter graphs & correlation
Scatter graphs & correlation
Why one point on a scatter graph is one individual measured twice, how to describe correlation by type and strength in the words of the question, why a correlation never proves a cause, and how to place a line of best fit — through the bulk of the data, balanced either side, ignoring a genuine outlier.
Why it works
Almost every graph you have met so far plots one quantity. A scatter graph plots two measurements of the same individual at once, and that is the whole idea. Take ten students and record two numbers for each of them: the hours they revised and the mark they scored. Student four revised for hours and scored , so student four is the single point . Ten students, ten points, and nothing else on the graph.That already tells you what the graph can and cannot say. It says nothing about either quantity on its own — only about how the two travel together.Correlation is the shape of the cloud. Describing it takes two independent words:
- Type — the direction. Points rising from left to right (one goes up, so
- Strength — how tightly the points hug a straight line. A narrow band is
Read the direction off the graph, never off the story. Sprint times plotted against long-jump distances slope downwards, because a better sprinter has a smaller time — that is negative correlation, even though "better at one, better at the other" sounds positive.
Then write the sentence that actually earns the marks: say it in the variables of the question. "Positive correlation" is worth about half of "the longer a student revised, the higher the mark they tended to get". The word tended is doing real work — correlation is a claim about the group as a whole, not a promise about any one person in it.
Correlation is not causation. Go into a primary school, measure every child's shoe size and give them all the same reading test. You will get a strong positive correlation: bigger feet, better reading. This is not a fluke and it is not a badly collected sample — it would happen again in any primary school you tried. But buying a child bigger shoes will not teach them to read. The link is there because of a third factor sitting behind both: age. Eleven-year-olds have bigger feet than five-year-olds, and eleven-year-olds also read better than five-year-olds. Age moves both quantities, so the two move together, and neither one causes the other. Whenever a scatter graph is used to argue "A causes B", the question to ask is: *what else could be pushing both of these up at the same time?*
The line of best fit is a single ruled straight line, drawn by eye, that summarises the trend. It should run through the bulk of the points, with roughly as many points above it as below and the ones above no further off than the ones below, and it should stretch across the data. Three things go wrong.
Joining the dots. Nine straight segments zig-zagging from point to point is not one line, and it adds nothing — it just redraws the data you already have.
Forcing it through the origin. On the revision data above, an honest line passes close to and : gradient , crossing the vertical axis around . Now force a ruler through instead. To reach the right-hand end of the cloud it needs gradient , so it sits at when , at when and at when — while the real points there are at , and . It misses the first point by nearly marks, and it lies below every single point in the left-hand half of the graph:The origin is not a data point. Unless the question actually gives you a reading there, the line has no obligation to go anywhere near it. Placed properly, the same line looks like this — points scattered on both sides of it all the way along:Letting one odd point drag it. An outlier is a point that does not fit the pattern the other points make — which is not the same as the biggest or the smallest reading. The largest value in a data set usually sits happily at the far end of the trend; an outlier sits off the trend, often in the middle of the range, where nothing else is anywhere near it. Include it and the ruler gets pulled towards one observation, so the line ends up describing none of the data.
So the routine is: spot it, suggest what could have caused it (equipment failed, someone was ill, the shop shut early), draw the line through the other points and say that you have ignored it. What you must not do is rub it out. It is a real observation, and something real happened.