to the true (theoretical) probability
A coin is tossed and the relative frequency of heads recorded at various stages.
| Tosses | $10$ | $50$ | $100$ | $1000$ | $10\,000$ |
|---|---|---|---|---|---|
| Heads | $7$ | $27$ | $46$ | $509$ | $5013$ |
| Relative frequency | $0.7$ | $0.54$ | $0.46$ | $0.509$ | $0.5013$ |
The coin appears to be fair, and the best estimate always comes from the largest sample.
Three students test the same biased spinner.
| Student | Spins | Reds |
|---|---|---|
| Ali | $50$ | $18$ |
| Beth | $120$ | $39$ |
| Cara | $230$ | $81$ |
Find the best estimate of $P(\text{red})$.
Two people estimate the probability that a drawing pin lands point up. Sam drops it $30$ times and gets $0.6$. Priya drops it $500$ times and gets $0.54$. Which estimate should be used, and why?
Use Priya's estimate, $0.54$. It is based on a much larger sample and is therefore more reliable.
- Work out what the theoretical probability would be if the object were fair.
- Multiply by the number of trials to find the expected frequency.
- Compare with the actual frequency.
- Judge whether the gap is large relative to the number of trials.
- State a conclusion, mentioning the sample size.
A die is rolled $900$ times. A one comes up $205$ times. Is the die biased?
Yes, the die appears biased towards one. The difference is large and the sample of $900$ is big enough to make chance an unlikely explanation.
A die is rolled $30$ times and a one comes up $8$ times. Is the die biased?
There is not enough evidence to say the die is biased. Rolling it several hundred more times would give a much better test.
| Feature | Why it matters |
|---|---|
| Many trials | Reduces the effect of random variation. |
| Identical conditions | Drop the pin from the same height onto the same surface each time. |
| No selection by the experimenter | Stopping when you get the result you like biases the outcome. |
| Record every trial | Discarding "odd" results distorts the frequency. |
| Fix the number of trials in advance | Prevents stopping at a convenient moment. |
Tom wants to estimate the probability that a bent coin lands heads. He tosses it until he gets $10$ heads, which takes $16$ tosses, and concludes $P(\text{heads}) = \dfrac{10}{16} = 0.625$. Give two criticisms.
Better method: decide in advance to toss the coin, say, $500$ times, record every result, and then calculate the relative frequency.
A common exam application: use a sample to estimate how many of something there are.
A bag contains $80$ counters, some red and some blue. A counter is drawn, its colour noted, and it is replaced. This is done $200$ times, and red comes up $130$ times. Estimate how many red counters are in the bag.
About $52$ red counters (and therefore about $28$ blue).
A large box contains many identical beads, some white. A sample of $250$ beads is taken and $45$ are white. The box is known to hold $4000$ beads. Estimate how many are white, and say how the estimate could be improved.
The main idea
More trials β relative frequency converges on the true probability.
Best estimate
The one from the largest number of trials.
Combining experiments
Add all successes and all trials; never average the frequencies.
Detecting bias
Compare actual with expected, and consider the sample size.
Small samples
Cannot prove bias β differences are easily due to chance.
Careful wording
"Not enough evidence of bias", not "the die is fair".
Good experiment
Many trials, identical conditions, number fixed in advance.
Estimating contents
Relative frequency $\times$ total number in the population.
Replacement
Replacing keeps each trial identical, which the method requires.
A spinner is spun $600$ times and lands on red $148$ times. Estimate $P(\text{red})$.
βΆ Show solution
$P(\text{red}) \approx \dfrac{148}{600} = 0.247$ (3 d.p.)
Two students estimate the same probability. Ravi uses $40$ trials and gets $0.4$; Nina uses $600$ trials and gets $0.47$. Which estimate is better, and why?
βΆ Show solution
Nina's, $0.47$.
She used far more trials, so random variation has had more opportunity to average out. Larger samples give more reliable estimates.
Three people test a coin: $40$ tosses with $24$ heads, $60$ tosses with $33$ heads, and $100$ tosses with $53$ heads. Find the best combined estimate of $P(\text{heads})$.
βΆ Show solution
Total tosses $= 40 + 60 + 100 = 200$
Total heads $= 24 + 33 + 53 = 110$
$P(\text{heads}) \approx \dfrac{110}{200} = 0.55$
A fair die should give a six about $\tfrac{1}{6}$ of the time. In $1200$ rolls, a six comes up $284$ times. Is the die biased?
βΆ Show solution
Expected sixes $= \dfrac{1}{6} \times 1200 = 200$.
Actual $= 284$, which is $84$ more than expected.
Relative frequency $= \dfrac{284}{1200} = 0.237$, well above $0.167$.
Yes β the die appears biased towards six, and with $1200$ rolls the evidence is strong.
A coin is tossed $12$ times and lands tails $9$ times. Can you conclude that the coin is biased? Explain.
βΆ Show solution
Expected tails $= 6$; actual $= 9$, a difference of only $3$.
Twelve tosses is a very small sample, and this kind of result happens quite often with a fair coin.
No β there is not enough evidence. Many more tosses would be needed.
A bag holds $60$ counters. A counter is drawn and replaced $300$ times; green appears $105$ times. Estimate how many green counters are in the bag.
βΆ Show solution
$P(\text{green}) \approx \dfrac{105}{300} = 0.35$
Estimated green counters $= 0.35 \times 60 = 21$
Explain why a counter must be replaced after each draw when estimating a probability in this way.
βΆ Show solution
Replacing the counter keeps the contents of the bag the same for every draw, so the probability is identical each time.
Without replacement the probability changes after every draw, so the relative frequency would no longer estimate a single fixed probability.
Jess wants to estimate the probability that a piece of buttered toast lands butter-side down. She drops it until it lands butter-side down for the fifth time, taking $9$ drops. Give two criticisms of her method.
βΆ Show solution
1. Too few trials. Nine drops is far too small a sample for a reliable estimate.
2. She stopped on a butter-side-down result. Because the stopping rule depends on the outcome, the final drop is guaranteed to be butter-side down, which biases the estimate upwards.
She should fix the number of drops in advance (say $200$) and record every one.
The relative frequency of a spinner landing yellow is recorded as the number of spins increases.
| Spins | $20$ | $60$ | $150$ | $400$ | $1000$ |
|---|---|---|---|---|---|
| Rel. freq. | $0.45$ | $0.32$ | $0.29$ | $0.272$ | $0.268$ |
(a) Estimate $P(\text{yellow})$. (b) Describe the pattern. (c) How many yellows would you expect in $2500$ spins?
βΆ Show solution
(a) Use the largest sample: $P(\text{yellow}) \approx 0.268$.
(b) The early values vary a lot ($0.45$ down to $0.32$), but as the number of spins increases they settle down and change less and less, converging on roughly $0.27$.
(c) $0.268 \times 2500 = 670$ yellows
A factory claims that $2\%$ of its products are faulty. A shop receives a delivery of $1500$ items and tests a random sample of $250$, finding $9$ faulty.
(a) Estimate the probability that an item is faulty. (b) Estimate the number of faulty items in the whole delivery. (c) Does the sample support the factory's claim? (d) How could the shop improve the reliability of its check?
βΆ Show solution
(a) $P(\text{faulty}) \approx \dfrac{9}{250} = 0.036$, that is $3.6\%$.
(b) $0.036 \times 1500 = 54$ faulty items.
(c) The factory claims $2\%$, which would mean $0.02 \times 250 = 5$ faulty in the sample. The shop found $9$ β nearly double.
However, the sample is only $250$ items and the difference is just $4$ items, which could plausibly happen by chance. So the sample raises doubt about the claim but does not disprove it.
(d) Test a larger sample, chosen at random from throughout the delivery rather than from one box, so that any variation between batches is captured.