Probability · Combined events & tree diagrams
1 / 17
Conditional probability
Why "given that" shrinks the sample space to the group you were told about, so the denominator becomes that group's total — and why P(A given B) and P(B given A) are two different questions with two different answers.
Probability · Combined events & tree diagrams
Conditional probability
Why "given that" shrinks the sample space to the group you were told about, so the denominator becomes that group's total — and why P(A given B) and P(B given A) are two different questions with two different answers.
Why it works
Every probability you have met so far has the same shape:That bottom number is the sample space — the list of everything that could have happened. A "given that" question does not change how you count the top of the fraction. It changes the bottom, because the words given that tell you that a whole chunk of the sample space did not happen and is no longer in the running.
The idea in one sentence: "given that" shrinks the sample space to the group you were told about.
Here is a two-way table for pupils, recording their year group and whether they walk to school.
| Walks | Does not walk | Total | |
|---|---|---|---|
| Year 10 | |||
| Year 11 | |||
| Total |
Now change the question: given that the pupil is in Year 10, what is the probability they walk? The condition has done something physical to the table. The Year 11 row is gone. You are no longer choosing from pupils, you are choosing from the pupils in the Year 10 row, and of those walk:
Nothing clever happened. It is still favourable total. It is just that "total" now means the total of the row you were told about, not the grand total. Writing instead would be answering a different question — "what is the probability the pupil is in Year 10 and walks?" — because that fraction is still measuring against all pupils.
Why the formula is the same thing. You will sometimes see conditional probability written as
That is not a new rule; it is the counting above with everything divided by . Take and divide top and bottom by the grand total:
Same number, . So if a question hands you probabilities instead of frequencies, divide; if it hands you a table of counts, just read the row.
The reversal trap, with numbers that prove it. and are different questions and usually have different answers, because they are measured against different groups. From the same table:
- — out of the Year 10 row.
- — out of the walkers column.
A blunter example fixes the idea for good. Almost everyone who has just been struck by lightning is outdoors, so is close to . Almost nobody outdoors has just been struck by lightning, so is close to . Same two events, same overlap, wildly different answers — because the denominators are wildly different sizes.
Running it backwards. Because is (the count in the overlap) (the size of group ), knowing the conditional probability and the size of the group hands you the overlap by multiplying:
If people ordered a dessert and , then people had ice cream. Note the multiplier lands on — the group named after the word "given" — and never on the grand total. That single missing frequency is often all you need to finish a table.
Testing independence. Two events are independent when knowing one happened tells you nothing about the other. Written in this language, that is
The condition shrank the sample space, but the proportion inside the smaller group came out the same as in the whole group. In the table above, but — those are not equal, so walking to school and year group are not independent here: hearing "Year 10" really does change the odds. The comparison that decides independence is always the conditional against the plain probability. Comparing with tests nothing at all.