3.6 Probability Topics
SLO 3
Describe and apply probability concepts and distributions.
This lab has you apply the chapter's rules twice to the same seven events — once from the bag's known contents, once from 24 recorded draws — so you can say what a probability claim rests on and how far a 12-trial estimate can drift from it.
Learning Objectives
By the end of this section, you will be able to:
- compute a theoretical probability for a two-pick experiment, both with replacement and without replacement;
- run the experiment yourself and compute the matching empirical probability from your own record of results;
- compare the two estimates and explain, in terms of the number of trials, why they differ;
- predict how the gap between an empirical estimate and its theoretical target behaves as the number of trials grows.
This is a lab, so it is the one section of the chapter you cannot read your way through. You need a bag of about 40 mixed-color M&Ms, a cup, a pencil, and a partner or two. Record your class time and the names of everyone in your group at the top of your worksheet — the numbers you write down are your group's data and nobody else's.
Sections 3.1 through 3.5 gave you the machinery: sample spaces, the multiplication and addition rules, conditional probability, and the difference between sampling with and without replacement. Every one of those was a calculation done on paper. This lab does the same calculations twice — once on paper, once with actual candy — and then asks you to explain the gap between the two answers.
There are two honest ways to answer the question "how likely is this?", and this lab makes you use both: you can count what is in the bag, or you can run the experiment and keep score. The two definitions that open the next subsection name them.
3.6.1 Theoretical and Empirical Probability
The theoretical probability of an event is the ratio of the number of outcomes in that event to the total number of outcomes in the sample space, computed from the known makeup of the experiment:
$$P(\text{event}) = \frac{\text{number of outcomes in the event}}{\text{total number of outcomes}}$$It requires no trials, because it comes from the structure of the experiment rather than from doing it. If you know exactly what is in the bag, you can work out what fraction of all the possible two-pick outcomes are the ones you care about without ever picking anything. It is what the experiment should do in the long run.
Definition 3.6.1 — Theoretical probability is counted off the bag: 5 of the 40 candies are red, so P(red) = 5/40 without a single pick.
The empirical probability of an event is the ratio of the number of trials in which the event actually occurred to the total number of trials performed:
$$P(\text{event}) = \frac{\text{number of trials in which the event occurred}}{\text{total number of trials}}$$You can say heads is one-half because a coin has two sides, or you can flip it 12 times, get 7 heads, and say seven-twelfths. Neither of you is lying. The first person described the coin; the second described an afternoon.
It requires no knowledge of the bag's contents. You get it by running the thing and keeping score: pick two M&Ms twelve times, count how many of those picks came out the way you cared about, and divide. It is what the experiment did, and it is also called a relative frequency.
Definition 3.6.2 — Empirical probability is scored off the trials: 4 of the 12 recorded picks were doubles, so P(doubles) = 4/12.
Both numbers are legitimate answers to the same question, and they will almost never agree exactly. The whole point of the lab is that the disagreement is not a mistake — it is information about how many trials you ran.
Erin rolls a standard six-sided die 20 times. Four of her rolls come up 3. Write the theoretical probability of rolling a 3 and the empirical probability of rolling a 3 from this run, and say which one would change if you rolled another 20 times.
Solution
Step 1 — the theoretical probability. A fair die has six equally likely faces and exactly one of them is a 3, so
$$P(\text{roll a } 3) = \frac{1}{6}$$This comes from the structure of the die. Rolling it more times does not change it.
Step 2 — the empirical probability. Four of the 20 rolls came up 3, so
$$P(\text{roll a } 3) = \frac{4}{20}$$Step 3 — which one moves. Only the empirical one. A second run of 20 rolls will give her some other count — three 3s, or five, or two — and the empirical probability moves with it. The theoretical value stays at \(\frac{1}{6}\) forever, because it was never about any particular run.
Answer: theoretical \(\frac{1}{6}\), empirical \(\frac{4}{20}\); the empirical one changes on a re-run.
3.6.2 Counting Your Population
Count out 40 mixed-color M&Ms — about one small bag's worth — and record how many of each color you have in Table 3.6.1. These counts are your population. Everything you compute on paper for the rest of the lab depends on them, so count twice.
| Color | Quantity |
|---|---|
| Yellow (Y) | |
| Green (G) | |
| Blue (BL) | |
| Brown (B) | |
| Orange (O) | |
| Red (R) |
Every worked calculation from here on uses one specific demonstration bag so you can see each step done all the way through. Those numbers are not your numbers. Read the demonstration to learn the move, then redo it with the counts you wrote in Table 3.6.1.
| Color | Quantity |
|---|---|
| Yellow (Y) | 8 |
| Green (G) | 6 |
| Blue (BL) | 7 |
| Brown (B) | 5 |
| Orange (O) | 9 |
| Red (R) | 5 |
Throughout this lab, \(R_1\) means red on the first pick and \(G_2\) means green on the second pick. The subscript is the pick number, not a second color. A double means both picks came out the same color.
The subscripts matter more than they look. \(P(R_1 \text{ AND } G_2)\) and \(P(G_1 \text{ AND } R_2)\) describe different events, and in the without-replacement case they can even have different denominators depending on which color left the cup first. Reading the subscript wrong is the single most common way to lose points on this lab, so when you set up a calculation, say the event out loud in words before you write a fraction.
Using the demonstration bag in Table 3.6.2, how many M&Ms are not yellow, and what fraction of the bag is that?
Solution
Step 1 — total the bag. \(8 + 6 + 7 + 5 + 9 + 5 = 40\).
Step 2 — subtract the yellows. There are 8 yellow, so the non-yellow count is
$$40 - 8 = 32$$Step 3 — write it as a fraction of the bag.
$$\frac{32}{40}$$Answer: 32 of the 40 candies are not yellow, which is \(\frac{32}{40}\) of the bag.
3.6.3 Theoretical Probabilities With Replacement
Now do the paper half. For each of the seven events in Table 3.6.3 you will compute the probability twice — once assuming you put the first M&M back before picking the second, and once assuming you do not. This subsection does the first column; the next does the second.
Leave every answer in unreduced fractional form. Do not multiply the fractions out and do not cancel. The unreduced form keeps the denominator visible, and the denominator is the part that tells you which of the two designs you were working in.
| Event | With Replacement | Without Replacement |
|---|---|---|
| \(P(\text{2 reds})\) | ||
| \(P(R_1 B_2 \text{ OR } B_1 R_2)\) | ||
| \(P(R_1 \text{ AND } G_2)\) | ||
| \(P(G_2 \mid R_1)\) | ||
| \(P(\text{no yellows})\) | ||
| \(P(\text{doubles})\) | ||
| \(P(\text{no doubles})\) |
With replacement, the cup is full again before the second pick, so both picks face the same 40 candies. That gives \(40 \cdot 40 = 1600\) ordered pairs, and it makes the two picks independent — the multiplication rule from section 3.3 collapses to plain multiplication of unconditional probabilities.
Using the demonstration bag with replacement, compute \(P(B_1 \text{ AND } O_2)\) — brown on the first pick and orange on the second. Leave the answer unreduced.
Solution
Step 1 — read the counts. The bag has 5 brown and 9 orange out of 40.
Step 2 — multiply. The brown goes back in, so the second pick still faces all 40 candies:
$$\frac{5}{40} \cdot \frac{9}{40} = \frac{45}{1600}$$Answer: \(\frac{45}{1600}\).
Tayen's demonstration bag has 8 yellow, 6 green, 7 blue, 5 brown, 9 orange, and 5 red, for a total of 40. Help them compute all seven theoretical probabilities with replacement.
Solution
Step 1 — \(P(\text{2 reds})\). There are 5 reds, and the red goes back in before the second pick.
$$\frac{5}{40} \cdot \frac{5}{40} = \frac{25}{1600}$$Step 2 — \(P(R_1 B_2 \text{ OR } B_1 R_2)\). Red-then-brown and brown-then-red are mutually exclusive, so add them. There are 5 red and 5 brown.
$$\frac{5}{40} \cdot \frac{5}{40} + \frac{5}{40} \cdot \frac{5}{40} = \frac{50}{1600}$$Step 3 — \(P(R_1 \text{ AND } G_2)\). Red first, then green: 5 reds and 6 greens.
$$\frac{5}{40} \cdot \frac{6}{40} = \frac{30}{1600}$$Step 4 — \(P(G_2 \mid R_1)\). Red on the first pick is given, so you are only being asked about the second pick — and the red went back in, so the cup still holds all 40 candies including all 6 greens.
$$\frac{6}{40}$$Step 5 — \(P(\text{no yellows})\). From Try It Now 3.6.2, 32 of the 40 are not yellow, and that is true for both picks.
$$\frac{32}{40} \cdot \frac{32}{40} = \frac{1024}{1600}$$Step 6 — \(P(\text{doubles})\). A double is two of the same color, so add up the six single-color cases.
$$\frac{8 \cdot 8 + 6 \cdot 6 + 7 \cdot 7 + 5 \cdot 5 + 9 \cdot 9 + 5 \cdot 5}{1600} = \frac{64 + 36 + 49 + 25 + 81 + 25}{1600} = \frac{280}{1600}$$Step 7 — \(P(\text{no doubles})\). The complement of Step 6, over the same 1600 pairs.
$$\frac{1600 - 280}{1600} = \frac{1320}{1600}$$Answer: the with-replacement column reads \(\frac{25}{1600}\), \(\frac{50}{1600}\), \(\frac{30}{1600}\), \(\frac{6}{40}\), \(\frac{1024}{1600}\), \(\frac{280}{1600}\), \(\frac{1320}{1600}\).
3.6.4 Theoretical Probabilities Without Replacement
Now redo the same seven events with the first candy left out of the cup. The first pick is unchanged — the cup was full when you reached in — but every second pick now faces 39 candies instead of 40, and the color you removed is one short.
That gives \(40 \cdot 39 = 1560\) ordered pairs, and it makes the two picks dependent. You need the full multiplication rule, \(P(A \text{ AND } B) = P(A) \, P(B \mid A)\), because the second factor genuinely depends on what came out first.
Redo Try It Now 3.6.3 without replacement: compute \(P(B_1 \text{ AND } O_2)\) for the demonstration bag, and say why the numerator did not change even though the design did.
Solution
Step 1 — set up the two factors. Brown first is \(\frac{5}{40}\). Given a brown is gone, the cup holds 39 candies and all 9 oranges are still there, so orange second is \(\frac{9}{39}\).
Step 2 — multiply.
$$\frac{5}{40} \cdot \frac{9}{39} = \frac{45}{1560}$$Step 3 — explain the numerator. Removing a brown candy reduces the brown count, not the orange count. Since the event asks for orange on the second pick, the count of favorable second picks is untouched — only the pool it is drawn from shrank.
Answer: \(\frac{45}{1560}\). Same numerator as with replacement, smaller denominator, so the without-replacement probability is the larger of the two.
Wes is redoing the same demonstration bag without replacement. Compute all seven theoretical probabilities the way he would.
Solution
Step 1 — \(P(\text{2 reds})\). After a red comes out, only 4 reds remain among 39 candies.
$$\frac{5}{40} \cdot \frac{4}{39} = \frac{20}{1560}$$Step 2 — \(P(R_1 B_2 \text{ OR } B_1 R_2)\). Taking a red out does not remove any browns, so the numerator of the second factor stays at 5 — only the denominator drops.
$$\frac{5}{40} \cdot \frac{5}{39} + \frac{5}{40} \cdot \frac{5}{39} = \frac{50}{1560}$$Step 3 — \(P(R_1 \text{ AND } G_2)\). Same idea: all 6 greens are still in the cup after a red leaves.
$$\frac{5}{40} \cdot \frac{6}{39} = \frac{30}{1560}$$Step 4 — \(P(G_2 \mid R_1)\). Red is given and stays out, so the cup holds 39 candies, 6 of them green.
$$\frac{6}{39}$$Step 5 — \(P(\text{no yellows})\). The first pick avoids yellow with probability \(\frac{32}{40}\). Given that, one non-yellow candy is gone, so 31 non-yellows remain among 39.
$$\frac{32}{40} \cdot \frac{31}{39} = \frac{992}{1560}$$Step 6 — \(P(\text{doubles})\). For each color, the second candy must be one of the remaining candies of that color.
$$\frac{8 \cdot 7 + 6 \cdot 5 + 7 \cdot 6 + 5 \cdot 4 + 9 \cdot 8 + 5 \cdot 4}{1560} = \frac{56 + 30 + 42 + 20 + 72 + 20}{1560} = \frac{240}{1560}$$Step 7 — \(P(\text{no doubles})\). The complement, over the same 1560 pairs.
$$\frac{1560 - 240}{1560} = \frac{1320}{1560}$$Answer: the without-replacement column reads \(\frac{20}{1560}\), \(\frac{50}{1560}\), \(\frac{30}{1560}\), \(\frac{6}{39}\), \(\frac{992}{1560}\), \(\frac{240}{1560}\), \(\frac{1320}{1560}\).
3.6.5 Where the Two Designs Diverge
Put your two columns side by side. Four of the seven rows changed only in the denominator, one changed in both, and one changed in a way worth stopping over.
Look at the last row. The two numerators are both 1320, which is not a coincidence. Going from with-replacement to without-replacement removes exactly 40 ordered pairs from the sample space — the 40 pairs where you would have drawn the very same candy twice — and every one of those 40 pairs was a double. So the doubles count falls by exactly 40 and the total falls by exactly 40, leaving the number of two-different-colors pairs untouched at 1320. The probability still changes, because the denominator moved, but the count of favorable outcomes did not.
Banning replacement removes exactly the outcomes where you pick the same physical candy twice. Those were all doubles, so the non-doubles shelf is untouched — 1320 either way.
The general pattern is worth naming. Removing the first candy can only affect the second pick's count for that candy's own color. If the event you are computing asks for a different color on the second pick, only the denominator moves. If it asks for the same color, both move.
Alexis wants to settle this without computing anything. Help them predict whether each of these is larger with replacement or without, for the demonstration bag: (a) \(P(\text{2 reds})\), and (b) \(P(R_1 \text{ AND } G_2)\).
Solution
Step 1 — part (a), same color twice. Removing the first red leaves one fewer red for the second pick, so the numerator shrinks and the denominator shrinks. The numerator falls from \(5 \cdot 5 = 25\) to \(5 \cdot 4 = 20\), a 20% drop, while the denominator falls from 1600 to 1560, only a 2.5% drop. The numerator loses more, so the without-replacement value is smaller.
Step 2 — part (b), two different colors. Removing a red takes nothing away from the greens, so the numerator holds at \(5 \cdot 6 = 30\) while the denominator drops from 1600 to 1560. Same top, smaller bottom, so the without-replacement value is larger.
Answer: (a) larger with replacement; (b) larger without replacement. Checking against Examples 3.6.1 and 3.6.2 confirms both: \(\frac{25}{1600} > \frac{20}{1560}\) and \(\frac{30}{1600} < \frac{30}{1560}\).
3.6.6 Running the Experiment
Now the candy half. Put all 40 M&Ms in the cup. The experiment is to pick two M&Ms, one at a time, without looking at them as you pick.
First pass — with replacement. Pick one M&M, record its color, put it back, then pick the second and record that color. Put the second one back too. Do this 12 times and record each pair in the "With Replacement" column of Table 3.6.4.
Second pass — without replacement. Pick one M&M and record its color, and this time do not put it back before picking the second. Record the second color, then return both to the cup before starting the next pair. Do this 12 times as well and record each pair in the "Without Replacement" column.
The "return both before the next pair" instruction is what makes the 12 pairs comparable to each other. Each pair starts from a full cup of 40; the only thing that changes inside a pair is whether the first candy is available for the second pick.
| Pick | With Replacement | Without Replacement |
|---|---|---|
| 1 | ||
| 2 | ||
| 3 | ||
| 4 | ||
| 5 | ||
| 6 | ||
| 7 | ||
| 8 | ||
| 9 | ||
| 10 | ||
| 11 | ||
| 12 |
Suppose you run the without-replacement pass with a bag that holds only one red candy. Could one of your pairs come out (R, R)? Explain what that tells you about reading a double.
Solution
Step 1 — check what a without-replacement pair requires. The first candy is removed before the second is drawn, so the two candies in a pair are always two different physical objects.
Step 2 — apply that to a one-red bag. With only one red candy in the bag, once it is out there is no red left to draw, so (R, R) is impossible.
Step 3 — say what a double actually means. A double means the two candies matched in color, not that the same candy was drawn twice. In the with-replacement pass a double can be either. In the without-replacement pass it is always two distinct candies of the same color, which is why it takes at least two candies of that color to be possible at all.
Answer: no — without replacement, (R, R) needs at least two reds in the bag, because a double there always means two different candies that share a color.
Malik and his partner Devon ran both passes with the demonstration bag and wrote down the 24 pairs below. His record is the input for every empirical calculation in the next two subsections.
Solution
Step 1 — the with-replacement pass. Twelve pairs, first pick then second pick:
| Pick | With Replacement | Pick | With Replacement |
|---|---|---|---|
| 1 | (R, G) | 7 | (Y, O) |
| 2 | (O, O) | 8 | (BL, BL) |
| 3 | (Y, B) | 9 | (O, G) |
| 4 | (BL, R) | 10 | (B, Y) |
| 5 | (G, O) | 11 | (R, B) |
| 6 | (R, R) | 12 | (G, G) |
Notice pick 6: with replacement he could have drawn the same red candy twice, and a repeat of that kind still counts as a double.
Step 2 — the without-replacement pass. Twelve more pairs, this time with no return between the two picks:
| Pick | Without Replacement | Pick | Without Replacement |
|---|---|---|---|
| 1 | (G, R) | 7 | (BL, Y) |
| 2 | (O, O) | 8 | (G, O) |
| 3 | (B, O) | 9 | (R, B) |
| 4 | (R, G) | 10 | (Y, Y) |
| 5 | (Y, G) | 11 | (B, G) |
| 6 | (R, R) | 12 | (O, BL) |
Picks 2, 6, and 10 are doubles even though nothing was replaced — the bag holds several candies of each color, so two different orange candies still read as (O, O).
Answer: 12 pairs recorded under each design, ready to be counted.
3.6.7 Counting Your Empirical Probabilities
Go through your Table 3.6.4 record and count. For six of the seven events, count how many of your 12 pairs satisfied it and write that count over 12. Leave it unreduced — the 12 in the denominator is the number you are going to argue about in the discussion questions, so keep it in view.
The seventh event, \(P(G_2 \mid R_1)\), works differently, and it gets its own subsection after this one.
| Event | With Replacement | Without Replacement |
|---|---|---|
| \(P(\text{2 reds})\) | ||
| \(P(R_1 B_2 \text{ OR } B_1 R_2)\) | ||
| \(P(R_1 \text{ AND } G_2)\) | ||
| \(P(G_2 \mid R_1)\) | ||
| \(P(\text{no yellows})\) | ||
| \(P(\text{doubles})\) | ||
| \(P(\text{no doubles})\) |
Rosalia is reading Malik's record. His empirical \(P(\text{doubles})\) came out higher with replacement (\(\frac{4}{12}\)) than without (\(\frac{3}{12}\)). Does that match what the theoretical values predicted?
Solution
Step 1 — pull the theoretical values. From Examples 3.6.1 and 3.6.2, \(\frac{280}{1600} = 0.1750\) with replacement and \(\frac{240}{1560} \approx 0.1538\) without.
Step 2 — compare the directions. Theory says with-replacement doubles are more likely, and his data also came out higher with replacement. The direction matches.
Step 3 — do not overread it. The theoretical gap is about 0.021, while one trial out of 12 is worth about 0.083 — four times the gap. A single 12-trial run cannot reliably detect a difference that small, so the agreement here is encouraging rather than confirming.
Answer: yes, the direction matches, but 12 trials is far too coarse to treat that as evidence — the two theoretical values differ by less than a quarter of one trial's worth of the empirical estimate.
Fill in the six out-of-12 rows of the empirical table from the 24 pairs in his record in Example 3.6.3.
Solution
Step 1 — \(P(\text{2 reds})\). With replacement, only pick 6 is (R, R), so \(\frac{1}{12}\). Without replacement, only pick 6 is (R, R), so \(\frac{1}{12}\).
Step 2 — \(P(R_1 B_2 \text{ OR } B_1 R_2)\). Look for red-then-brown or brown-then-red, remembering that BL is blue and B is brown. With replacement, pick 11 is (R, B), so \(\frac{1}{12}\). Without replacement, pick 9 is (R, B), so \(\frac{1}{12}\).
Step 3 — \(P(R_1 \text{ AND } G_2)\). Red first, green second. With replacement, pick 1 is (R, G), so \(\frac{1}{12}\). Without replacement, pick 4 is (R, G), so \(\frac{1}{12}\).
Step 4 — \(P(\text{no yellows})\). Count the pairs containing no Y at all. With replacement, picks 3, 7, and 10 contain a yellow, leaving 9, so \(\frac{9}{12}\). Without replacement, picks 5, 7, and 10 contain a yellow, leaving 9, so \(\frac{9}{12}\).
Step 5 — \(P(\text{doubles})\). With replacement, picks 2, 6, 8, and 12 match in color, so \(\frac{4}{12}\). Without replacement, picks 2, 6, and 10 match, so \(\frac{3}{12}\).
Step 6 — \(P(\text{no doubles})\). The complement of Step 5 out of the same 12 trials: \(\frac{8}{12}\) with replacement and \(\frac{9}{12}\) without.
Answer: the six out-of-12 rows, with replacement then without: \(\frac{1}{12}\) and \(\frac{1}{12}\); \(\frac{1}{12}\) and \(\frac{1}{12}\); \(\frac{1}{12}\) and \(\frac{1}{12}\); \(\frac{9}{12}\) and \(\frac{9}{12}\); \(\frac{4}{12}\) and \(\frac{3}{12}\); \(\frac{8}{12}\) and \(\frac{9}{12}\).
Let rāSHio do the tallying
Open rāSHio, enter your results as a column, and choose Graph → Frequency Table. It counts each distinct value for you and reports the matching relative frequency alongside it, which is the pair of columns you are building by hand here. Do the counting yourself first — that is the skill this subsection teaches — then use the table to check that your counts still add back to the number of trials you ran.
3.6.8 Reading a Conditional Probability from Your Record
The row \(P(G_2 \mid R_1)\) is not out of 12, and this is the step most groups get wrong. A conditional probability only concerns the trials where the given condition actually happened. The pairs that did not start with red never met the condition, so they are not evidence for or against it — they are simply not part of the question.
So count in two stages. First, how many of your 12 pairs started with red? That count is your denominator. Second, how many of those had green on the second pick? That count is your numerator.
Jordan is checking Malik's conditional row. Compute the empirical \(P(G_2 \mid R_1)\) for both of his passes in Example 3.6.3, and explain the denominator you used.
Solution
Step 1 — the with-replacement pass. Pairs starting with red: pick 1 (R, G), pick 6 (R, R), pick 11 (R, B). That is 3 pairs, and exactly one of them has green second.
$$P(G_2 \mid R_1) = \frac{1}{3}$$Step 2 — the without-replacement pass. Pairs starting with red: pick 4 (R, G), pick 6 (R, R), pick 9 (R, B). Again 3 pairs, again exactly one with green second.
$$P(G_2 \mid R_1) = \frac{1}{3}$$Step 3 — justify the denominator. The other nine pairs in each pass started with some other color, so they never satisfied the condition \(R_1\). Including them would answer a different question — "how often did red-then-green happen out of all trials?" — which is \(P(R_1 \text{ AND } G_2)\), not \(P(G_2 \mid R_1)\).
Answer: \(\frac{1}{3}\) in both passes; the denominator is 3 because only three trials met the given condition.
3.6.9 Comparing the Two Estimates
You now have two full tables of numbers for the same seven events. Put them side by side and look at where they agree and where they do not.
Mai plans to repeat the with-replacement pass 240 times instead of 12. What values could her empirical \(P(\text{doubles})\) take, and would you expect it to land closer to the theoretical 0.1750?
Solution
Step 1 — find the new step size. With 240 trials the empirical probability is a count out of 240, so it can be any of \(\frac{0}{240}, \frac{1}{240}, \frac{2}{240}, \dots, \frac{240}{240}\) — steps of about 0.0042 instead of 0.0833.
Step 2 — check whether the target is now reachable. The theoretical value 0.1750 times 240 is 42, a whole number, so \(\frac{42}{240}\) hits it exactly. With 12 trials there was no count that could.
Step 3 — say what the Law of Large Numbers adds. More trials do not just make the target reachable, they make it likely: as the number of trials grows, the relative frequency of an event settles toward its theoretical probability.
Answer: the estimate could take any multiple of \(\frac{1}{240}\), and yes — you would expect it much closer to 0.1750, both because \(\frac{42}{240}\) is now a possible landing spot and because the Law of Large Numbers pulls long-run relative frequency toward the theoretical value.
For Malik's demonstration bag and his record, compare the with-replacement estimates of \(P(\text{no yellows})\) and \(P(\text{doubles})\) in decimal form.
Solution
Step 1 — convert \(P(\text{no yellows})\).
$$\text{theoretical: } \frac{1024}{1600} = 0.6400 \qquad \text{empirical: } \frac{9}{12} = 0.7500$$The empirical estimate is high by 0.1100.
Step 2 — convert \(P(\text{doubles})\).
$$\text{theoretical: } \frac{280}{1600} = 0.1750 \qquad \text{empirical: } \frac{4}{12} \approx 0.3333$$The empirical estimate is high by about 0.1583.
Step 3 — put the gaps in scale. One extra matching trial out of 12 moves the empirical estimate by \(\frac{1}{12} \approx 0.0833\). Both gaps here are smaller than two of those steps. With only 12 trials the empirical value can land on just 13 possible numbers — 0, \(\frac{1}{12}\), \(\frac{2}{12}\), and so on — so it physically cannot sit close to 0.1750 no matter how carefully the experiment was run.
Answer: both empirical values run high, by 0.1100 and about 0.1583, and both gaps are within the resolution that 12 trials can even produce.
The lesson is in Step 3 of that solution. A 12-trial estimate has a step size of \(\frac{1}{12}\), so it cannot be closer to the truth than about half a step even in principle. Run 240 trials instead of 12 and the step size drops to \(\frac{1}{240}\), which lets the empirical value land near the theoretical one — and the Law of Large Numbers from section 3.1 says it will.
A poll of 12 people cannot report 43% support, because 43% is not a multiple of one-twelfth. The resolution of an estimate is set by how many trials paid for it.
Run the 240 trials you cannot hand-pick
The discussion questions ask what changes if you go from 12 picks to 240. You are not going to draw 240 pairs of M&Ms, but you can simulate them. Open rāSHio and choose File → Random Numbers…; the clip below shows the dialog, which asks for a minimum, a maximum, and how many numbers you want. For this lab set the range to 1 through 40, leave No repeats UNCHECKED so a colour can come up twice, and ask for far more than the three the demo generates. Then compare a short column against a long one — the long one is not more random, it is just further along.
Figure 3.6.1 — The rāSHio random-number dialog: set a range and a count, then Generate. Widen the range and the count to simulate a long run.
3.6.10 Discussion Questions
Answer these with your group, using your own two tables. Every one of them is a judgement call, so what matters is whether you can point at your own numbers to defend the answer.
- Why are the "With Replacement" and "Without Replacement" theoretical probabilities different? Name the specific step in the calculation where the two designs stop agreeing, and say what physically causes it.
- Convert \(P(\text{no yellows})\) to decimal form for both the theoretical "With Replacement" value and the empirical "With Replacement" value. Round both to four decimal places.
a. Theoretical "With Replacement": \(P(\text{no yellows}) =\) ________
b. Empirical "With Replacement": \(P(\text{no yellows}) =\) ________
c. Are the two decimals close? Did you expect them to be closer together or farther apart than they turned out? Say what you expected before you converted them, and what changed your mind if anything did.
- If you increased the number of two-M&M picks from 12 to 240, why would the empirical probability values change? Be specific about which part of the fraction moves and which does not.
- Would that change make the empirical and theoretical probabilities closer together or farther apart? How do you know — is your answer a guarantee or a strong expectation, and what is the difference?
- Explain the difference between what \(P(G_1 \text{ AND } R_2)\) and \(P(R_1 \mid G_2)\) represent. Think about the sample space for each one: how many trials does each probability get computed over, and why is it not the same number?
Key Terms
theoretical probability — a probability computed from the known structure of an experiment, as the ratio of favorable outcomes to total outcomes, without running any trials.
empirical probability — a probability estimated from data, as the ratio of the number of trials in which an event occurred to the total number of trials; also called a relative frequency.
trial — one full repetition of the experiment; in this lab, one pick of two M&Ms.
double — an outcome in which both picks in a trial come out the same color.
sampling with replacement — a design in which the first item is returned to the population before the second is drawn, so both picks face the same population.
sampling without replacement — a design in which the first item is not returned before the second is drawn, so the second pick faces a population one item smaller.