πŸ”¬ Experimental Probability and Sample Size

GCSE Maths Β· Probability (P5)

Ages 15–16 Β· Foundation & Higher

← Back to topic overview
1 The Law of Large Numbers
The central idea of this page
The more trials you carry out, the closer the relative frequency gets
to the true (theoretical) probability
A small experiment can give a wildly misleading answer purely by chance. A large one cannot β€” random ups and downs cancel each other out as the number of trials grows.
0.5 101005002000 Number of trials Relative frequency settles down towards the true value
Worked Example 1 β€” Watching it converge

A coin is tossed and the relative frequency of heads recorded at various stages.

Tosses$10$$50$$100$$1000$$10\,000$
Heads$7$$27$$46$$509$$5013$
Relative frequency$0.7$$0.54$$0.46$$0.509$$0.5013$
β‘ After $10$ tosses the estimate is $0.7$ β€” a long way from $0.5$.
β‘‘By $1000$ tosses it is $0.509$, and by $10\,000$ it is $0.5013$.
β‘’The values are homing in on $0.5$, the theoretical probability for a fair coin.

The coin appears to be fair, and the best estimate always comes from the largest sample.

2 Choosing the Best Estimate
The rule
Always use the estimate based on the largest number of trials
Do not average the relative frequencies. If three students get $0.4$, $0.6$ and $0.55$ from different numbers of trials, the answer is not the mean of those three. Combine the raw data instead: add all the successes and divide by all the trials.
Worked Example 2 β€” Combining several experiments

Three students test the same biased spinner.

StudentSpinsReds
Ali$50$$18$
Beth$120$$39$
Cara$230$$81$

Find the best estimate of $P(\text{red})$.

β‘ Combine the raw data: total spins $= 50 + 120 + 230 = 400$.
β‘‘Total reds $= 18 + 39 + 81 = 138$.
β‘’$P(\text{red}) \approx \dfrac{138}{400} = 0.345$
Using all $400$ spins gives a far better estimate than any student's individual result β€” and better than averaging $0.36$, $0.325$ and $0.352$.
Worked Example 3 β€” Choosing between estimates

Two people estimate the probability that a drawing pin lands point up. Sam drops it $30$ times and gets $0.6$. Priya drops it $500$ times and gets $0.54$. Which estimate should be used, and why?

β‘ Both are valid relative frequencies.
β‘‘Priya used far more trials, so random variation has had more chance to even out.

Use Priya's estimate, $0.54$. It is based on a much larger sample and is therefore more reliable.

3 Detecting Bias
Two things matter: how big the difference is, and how many trials there were. A difference of $10$ out of $30$ trials proves nothing; the same difference out of $3000$ trials is very strong evidence.
Worked Example 4 β€” Strong evidence of bias

A die is rolled $900$ times. A one comes up $205$ times. Is the die biased?

β‘ If fair, $P(1) = \dfrac{1}{6}$.
β‘‘Expected number $= \dfrac{1}{6} \times 900 = 150$.
β‘’Actual $= 205$ β€” that is $55$ more than expected.
β‘£Relative frequency $= \dfrac{205}{900} = 0.228$, compared with $0.167$.

Yes, the die appears biased towards one. The difference is large and the sample of $900$ is big enough to make chance an unlikely explanation.

Worked Example 5 β€” Weak evidence

A die is rolled $30$ times and a one comes up $8$ times. Is the die biased?

β‘ Expected $= \dfrac{1}{6} \times 30 = 5$.
β‘‘Actual $= 8$ β€” only $3$ more than expected.
β‘’With just $30$ rolls, a difference of $3$ is easily explained by chance.

There is not enough evidence to say the die is biased. Rolling it several hundred more times would give a much better test.

Note the careful wording. You can say "there is not enough evidence" β€” you cannot say "the die is fair". Absence of evidence is not proof of fairness.
4 Designing a Good Experiment
FeatureWhy it matters
Many trialsReduces the effect of random variation.
Identical conditionsDrop the pin from the same height onto the same surface each time.
No selection by the experimenterStopping when you get the result you like biases the outcome.
Record every trialDiscarding "odd" results distorts the frequency.
Fix the number of trials in advancePrevents stopping at a convenient moment.
Worked Example 6 β€” Criticising a method

Tom wants to estimate the probability that a bent coin lands heads. He tosses it until he gets $10$ heads, which takes $16$ tosses, and concludes $P(\text{heads}) = \dfrac{10}{16} = 0.625$. Give two criticisms.

β‘ Too few trials. Sixteen tosses is a very small sample, so the estimate is unreliable.
β‘‘He stopped on a head. Deciding to stop as soon as the tenth head appears guarantees the last toss is a head, which pushes the estimate upwards.

Better method: decide in advance to toss the coin, say, $500$ times, record every result, and then calculate the relative frequency.

5 Using Samples to Estimate Contents

A common exam application: use a sample to estimate how many of something there are.

Worked Example 7 β€” Estimating how many counters

A bag contains $80$ counters, some red and some blue. A counter is drawn, its colour noted, and it is replaced. This is done $200$ times, and red comes up $130$ times. Estimate how many red counters are in the bag.

β‘ Estimated $P(\text{red}) = \dfrac{130}{200} = 0.65$
β‘‘Estimated number of red counters $= 0.65 \times 80 = 52$

About $52$ red counters (and therefore about $28$ blue).

Note the counter is replaced each time. This keeps the probability the same for every draw, which is what makes the estimate valid.
Worked Example 8 β€” Capture–recapture style reasoning

A large box contains many identical beads, some white. A sample of $250$ beads is taken and $45$ are white. The box is known to hold $4000$ beads. Estimate how many are white, and say how the estimate could be improved.

β‘ Estimated $P(\text{white}) = \dfrac{45}{250} = 0.18$
β‘‘Estimated number white $= 0.18 \times 4000 = 720$ beads
β‘’To improve it: take a larger sample, and make sure the box is well mixed so the sample is genuinely random.
6 Quick Reference

The main idea

More trials β†’ relative frequency converges on the true probability.

Best estimate

The one from the largest number of trials.

Combining experiments

Add all successes and all trials; never average the frequencies.

Detecting bias

Compare actual with expected, and consider the sample size.

Small samples

Cannot prove bias β€” differences are easily due to chance.

Careful wording

"Not enough evidence of bias", not "the die is fair".

Good experiment

Many trials, identical conditions, number fixed in advance.

Estimating contents

Relative frequency $\times$ total number in the population.

Replacement

Replacing keeps each trial identical, which the method requires.

7 Practice Questions
Question 1

A spinner is spun $600$ times and lands on red $148$ times. Estimate $P(\text{red})$.

β–Ά Show solution

$P(\text{red}) \approx \dfrac{148}{600} = 0.247$ (3 d.p.)

Question 2

Two students estimate the same probability. Ravi uses $40$ trials and gets $0.4$; Nina uses $600$ trials and gets $0.47$. Which estimate is better, and why?

β–Ά Show solution

Nina's, $0.47$.

She used far more trials, so random variation has had more opportunity to average out. Larger samples give more reliable estimates.

Question 3

Three people test a coin: $40$ tosses with $24$ heads, $60$ tosses with $33$ heads, and $100$ tosses with $53$ heads. Find the best combined estimate of $P(\text{heads})$.

β–Ά Show solution

Total tosses $= 40 + 60 + 100 = 200$

Total heads $= 24 + 33 + 53 = 110$

$P(\text{heads}) \approx \dfrac{110}{200} = 0.55$

Question 4

A fair die should give a six about $\tfrac{1}{6}$ of the time. In $1200$ rolls, a six comes up $284$ times. Is the die biased?

β–Ά Show solution

Expected sixes $= \dfrac{1}{6} \times 1200 = 200$.

Actual $= 284$, which is $84$ more than expected.

Relative frequency $= \dfrac{284}{1200} = 0.237$, well above $0.167$.

Yes β€” the die appears biased towards six, and with $1200$ rolls the evidence is strong.

Question 5

A coin is tossed $12$ times and lands tails $9$ times. Can you conclude that the coin is biased? Explain.

β–Ά Show solution

Expected tails $= 6$; actual $= 9$, a difference of only $3$.

Twelve tosses is a very small sample, and this kind of result happens quite often with a fair coin.

No β€” there is not enough evidence. Many more tosses would be needed.

Question 6

A bag holds $60$ counters. A counter is drawn and replaced $300$ times; green appears $105$ times. Estimate how many green counters are in the bag.

β–Ά Show solution

$P(\text{green}) \approx \dfrac{105}{300} = 0.35$

Estimated green counters $= 0.35 \times 60 = 21$

Question 7

Explain why a counter must be replaced after each draw when estimating a probability in this way.

β–Ά Show solution

Replacing the counter keeps the contents of the bag the same for every draw, so the probability is identical each time.

Without replacement the probability changes after every draw, so the relative frequency would no longer estimate a single fixed probability.

Question 8

Jess wants to estimate the probability that a piece of buttered toast lands butter-side down. She drops it until it lands butter-side down for the fifth time, taking $9$ drops. Give two criticisms of her method.

β–Ά Show solution

1. Too few trials. Nine drops is far too small a sample for a reliable estimate.

2. She stopped on a butter-side-down result. Because the stopping rule depends on the outcome, the final drop is guaranteed to be butter-side down, which biases the estimate upwards.

She should fix the number of drops in advance (say $200$) and record every one.

Question 9

The relative frequency of a spinner landing yellow is recorded as the number of spins increases.

Spins$20$$60$$150$$400$$1000$
Rel. freq.$0.45$$0.32$$0.29$$0.272$$0.268$

(a) Estimate $P(\text{yellow})$.   (b) Describe the pattern.   (c) How many yellows would you expect in $2500$ spins?

β–Ά Show solution

(a) Use the largest sample: $P(\text{yellow}) \approx 0.268$.

(b) The early values vary a lot ($0.45$ down to $0.32$), but as the number of spins increases they settle down and change less and less, converging on roughly $0.27$.

(c) $0.268 \times 2500 = 670$ yellows

Question 10

A factory claims that $2\%$ of its products are faulty. A shop receives a delivery of $1500$ items and tests a random sample of $250$, finding $9$ faulty.

(a) Estimate the probability that an item is faulty.   (b) Estimate the number of faulty items in the whole delivery.   (c) Does the sample support the factory's claim?   (d) How could the shop improve the reliability of its check?

β–Ά Show solution

(a) $P(\text{faulty}) \approx \dfrac{9}{250} = 0.036$, that is $3.6\%$.

(b) $0.036 \times 1500 = 54$ faulty items.

(c) The factory claims $2\%$, which would mean $0.02 \times 250 = 5$ faulty in the sample. The shop found $9$ β€” nearly double.

However, the sample is only $250$ items and the difference is just $4$ items, which could plausibly happen by chance. So the sample raises doubt about the claim but does not disprove it.

(d) Test a larger sample, chosen at random from throughout the delivery rather than from one box, so that any variation between batches is captured.

Experimental Probability & Sample Size (P5) Β· GCSE Maths Revision Β· Created with MathJax