6.4 Central Limit Theorem (Pocket Change)

Aligned outcomes:

SLO 3

Describe and apply probability concepts and distributions.

You empty your own pocket change, compute the mean of 36 coin values, and compare it against the theoretical sampling distribution — so the central limit theorem stops being a formula and becomes something you watch happen with real coins.

SLO 4

Demonstrate an understanding of, and ability to use, basic ideas of statistical processes, including hypothesis tests and confidence interval estimation.

This lab builds the intuition that a sample mean from a skewed population still lands roughly normal when n is large enough, which is the whole reason confidence intervals and hypothesis tests work on real data that is never perfectly normal.

Learning Objectives

By the end of this section, you will be able to:

In this section, you will learn to:
  • collect a real, lopsided data set and sketch the shape of its distribution;
  • build the distribution of sample means from that same population by averaging groups of two and groups of five, and sketch each one on the same scale;
  • compare the three pictures and describe what happens to the shape, the center, and the spread as the group size grows.

6.4.1 Stats Lab: Central Limit Theorem (Pocket Change)

Definition 6.4.1: Sampling Distribution of the Sample Mean

Fix a population and a sample size \(n\). Take every possible sample of size \(n\), compute the sample mean \(\bar{x}\) of each one, and collect those means into a data set of their own. The distribution of that collection is the sampling distribution of the sample mean, written \(\bar{X}\).

The thing you are building is not the data

Each classmate hands you one number. Averaging groups of them makes a brand new set of numbers, and that new set has its own shape, its own center, and its own spread. This lab is about the second set, not the first.

It is a distribution of averages, not of individuals. One value in it is built from \(n\) values of the original population, so it is a different object with a different shape, and every question you ask of it is a question about averages.

This section is a lab, not a reading. In §6.1 and §6.2 you were handed a mean and a standard deviation and told what the central limit theorem promises about sample means and sums. In §6.3 you used that promise to compute probabilities. Here nobody hands you anything and nobody promises anything. You are going to build the sampling distribution of the sample mean with your own hands, out of coins in your classmates' pockets, and watch a lopsided pile of numbers turn into a bell.

Class Time:

Names:

You will draw three histograms of the same underlying quantity — the change in a person's pocket. The first is built from single people, the second from the averages of pairs, the third from the averages of groups of five. Scale all three axes the same way, because the whole point of the lab is invisible if the pictures are drawn at different sizes.

Write down what you actually did, including anything that went sideways. If someone counted a dollar coin, if a group of five was really a group of four because someone left, if you rounded the average of a pair to the nearest nickel — record it. A lab report that hides its own irregularities cannot be checked by anyone, and being checkable is most of what makes a result worth anything.

This lab is hungry for people

Thirty pairs is sixty people and thirty groups of five is a hundred and fifty. No single class holds that many, so classes pool their data. If your section runs it alone, say so in your write-up and expect a rougher picture at \(n = 5\).

The arithmetic of that headcount is worth doing before you start collecting, because it decides how the class organizes itself. It also explains a departure you are likely to make from the written instructions: most classes survey each person once and then reuse those same numbers to form the pairs and the groups of five, rather than going out and finding a hundred and fifty separate people. That reuse is a compromise, not the ideal experiment, and it is worth naming in the report. The averages you build from reused values are not quite independent of one another the way the theory assumes. In practice the picture still comes out the way the theorem predicts, which is itself a small piece of evidence about how forgiving the central limit theorem is — but a reader deserves to know which version of the experiment you ran.

Try It Now 6.4.1

Caleb is planning his section's run of the lab and wants to know how many people he has to line up. Work out the total if every value is collected fresh: 30 single people, 30 pairs, and 30 groups of five. Then say what that number tells him about the instruction to combine data from several classes.

Solution

Step 1 — count each round separately.

$$ 30 \times 1 = 30 \qquad 30 \times 2 = 60 \qquad 30 \times 5 = 150 $$

Step 2 — add them up.

$$ 30 + 60 + 150 = 240 $$

Answer: Caleb needs 240 people if nothing is reused. His class of 30 cannot produce that on its own, and even the third round alone needs 150. That is why the lab is written to be run across several sections with the data combined. It also tells you what the fallback has to be if you are working alone: reuse the same surveyed values to build the pairs and the groups, and say in your write-up that you did. The number 240 is not a detail of the setup, it is the reason the instructions are written the way they are.

6.4.2 Collect the Data

  1. Count the change in your pocket. Do not include bills.
  2. Randomly survey 30 classmates. Record the values of the change in Table 6.4.1.
Table 6.4.1 — The change carried by each of your 30 surveyed classmates, in dollars.
1–56–1011–1516–2021–2526–30

Try it in rāSHio

Once the grid above is full, get the 30 amounts into rāSHio in one step: choose File → Delimited List… and type them separated by commas, spaces, or one per line, in whatever order you surveyed them. Every other tool below reads from that one column, so keying it in cleanly once saves re-entering it for the histogram and for \(\bar{x}\) and \(s\).

Figure 6.4.1 — Getting your 30 collected amounts into rāSHio: File → Delimited List… The walkthrough imports its own demo column; the steps are the ones you run on your own 30 amounts.

  1. Construct a histogram. Make five to six intervals. Sketch the graph using a ruler and pencil, and scale the axes. Put the amount of change on the horizontal axis and label it Value of the Change, and put the count on the vertical axis and label it Frequency.

Try it in rāSHio

rāSHio’s Graph → Histogram will draw your five or six intervals: set the bin start below your smallest amount and the bin width to whatever cuts your range into that many pieces. Draw the pencil version first — choosing the intervals yourself is the point of step 3 — then use this to check you counted each bar right. Write down the bin start and width you used, because the next two rounds have to reuse them.

  1. Calculate the following, with \(n = 1\) — you are surveying one person at a time:

a. \(\bar{x} =\) _________

b. \(s =\) _________

  1. Draw a smooth curve through the tops of the bars of the histogram. Use one to two complete sentences to describe the general shape of the curve.

Try it in rāSHio

Stats → Summary Statistics in rāSHio returns the mean and standard deviation of your column in one step. Use it to CHECK the \(\bar{x}\) and \(s\) you worked out by hand, not to replace them — if the tool and your arithmetic disagree you have either mis-keyed an amount or dropped a term out of the sum of squares, and both are worth finding now.

Figure 6.4.2 — Checking your hand-computed \(\bar{x}\) and \(s\) in rāSHio: Stats → Summary Statistics. The walkthrough runs on rāSHio's built-in sample data, so the numbers it returns are not yours; the steps are.

Pocket change is lopsided on purpose

Most people carry under a dollar and a few carry a small pile, so the picture you draw here will lean to the right rather than sit in a neat bell. That is exactly why the lab uses it. A population that was already bell-shaped would prove nothing.

Keep that shape in mind, because it is the baseline everything else is measured against. The curve you draw in this first round is the population you are sampling from, and it is emphatically not normal — it is bunched up against zero on the left, because nobody can carry a negative amount of change, and it trails off to the right, because there is no upper limit on how much someone might have. If your histogram comes out looking symmetric, check whether you rounded aggressively or whether your class happens to be unusually uniform, and say which. Then choose your intervals before you look at the counts, not after. It is tempting to slide the boundaries around until the picture looks tidy, and a histogram built that way tells you about your own preferences rather than about your classmates' pockets.

Throughout the rest of this lab, one class's collected data is used to show every calculation. That class surveyed 30 people and recorded the amounts in Table 6.4.2, in ascending order. Use it to check that you can reproduce each step; your own numbers will differ.

Table 6.4.2 — One class's 30 pocket-change amounts, in dollars, sorted from smallest to largest.
1–56–1011–1516–2021–2526–30
0.050.250.420.610.851.23
0.110.280.450.630.911.38
0.130.310.470.680.971.52
0.170.340.530.721.061.76
0.220.360.560.791.142.10
Try It Now 6.4.2

Ximena, who runs the Pride Center's study group, has the 30 amounts in Table 6.4.2 written up on her whiteboard: the values total \(\$21.00\) and the squares of the values total \(22.3092\). Work alongside her to find \(\bar{x}\) and \(s\), then find the median and use it to describe the shape of the distribution.

Solution

Step 1 — divide the total by the count.

$$ \bar{x} = \frac{21.00}{30} = 0.70 $$

Step 2 — work out the sample standard deviation.

$$ s^2 = \frac{\sum x^2 - n\bar{x}^2}{n-1} = \frac{22.3092 - 30(0.70)^2}{29} = \frac{22.3092 - 14.7}{29} = \frac{7.6092}{29} \approx 0.26238 $$ $$ s = \sqrt{0.26238} \approx 0.5122 $$

Step 3 — locate the median. With 30 sorted values the median averages positions 15 and 16, which hold \(\$0.56\) and \(\$0.61\).

$$ \text{median} = \frac{0.56 + 0.61}{2} = 0.585 $$

Answer: Ximena gets \(\bar{x} = \$0.70\) and \(s \approx \$0.51\), with a median of \(\$0.585\). The mean sits about twelve cents above the median, which is the arithmetic signature of a right skew: the handful of people carrying more than a dollar drag the mean upward while leaving the middle value alone. Eighteen of the 30 people carry less than the mean. Divide by \(n-1\) rather than \(n\) — these 30 are a sample of a much larger population of pockets, not the whole of it.

6.4.3 Collecting Averages of Pairs

Definition 6.4.2: Central Limit Theorem for Sample Means

Draw simple random samples of size \(n\) from a population with mean \(\mu\) and standard deviation \(\sigma\). As \(n\) grows, the sampling distribution of the sample mean approaches a normal distribution:

$$ \bar{X} \sim N\!\left(\mu,\ \frac{\sigma}{\sqrt{n}}\right) $$
Averaging pulls the extremes in

A single \(\$2.10\) pocket sits far out on the right of the first picture. Pair that person with anyone typical and the average lands much closer to the middle. Every extreme value gets diluted by its partner, which is why the second histogram is narrower than the first.

The conclusion holds whatever the shape of the original population — skewed, flat, or full of gaps. Only the mean and the standard deviation of that population enter the statement.

Read what the theorem does and does not change. The center does not move: the mean of the averages is the same \(\mu\) you started with, so your \(\bar{x}\) for the pairs should land close to the \(\bar{x}\) you got for the individuals. The spread shrinks by a factor of \(\sqrt{n}\), so the pair histogram should be about \(\sqrt{2} \approx 1.41\) times narrower than the first one. And the shape moves toward a bell, even though the population it came from was nowhere near one. Those three claims are exactly what your three drawings are testing, which is why the instruction to reuse the same axis scaling is not fussiness — a narrower distribution redrawn on a narrower axis looks identical to the one before it, and the finding disappears.

Repeat steps one through five of the section titled Collect the Data, with one exception. Instead of recording the change of 30 classmates, record the average change of 30 pairs.

  1. Randomly survey 30 pairs of classmates.
  2. Record the values of the average of their change in Table 6.4.3.
Table 6.4.3 — The average change of each of your 30 surveyed pairs, in dollars.
1–56–1011–1516–2021–2526–30
  1. Construct a histogram. Scale the axes using the same scaling you used for the section titled Collect the Data. Sketch the graph using a ruler and a pencil.

Try it in rāSHio

Put the 30 pair averages in a second column and run the same two tools on it — Graph → Histogram with the bin start and width you wrote down in the first round, then Stats → Summary Statistics. Reusing the bin settings is what makes the two pictures comparable; re-binning to fit the new column is the one thing that would hide the result. Do it a third time for the groups of five.

  1. Calculate the following, with \(n = 2\) — you are surveying two people at a time:

a. \(\bar{x} =\) _________

b. \(s =\) _________

  1. Draw a smooth curve through the tops of the bars of the histogram. Use one to two complete sentences to describe the general shape of the curve.
Try It Now 6.4.3

Mai surveyed pairs for her group, and two of them held \(\$0.13\) and \(\$1.52\), and \(\$0.45\) and \(\$0.72\). Find those two pair averages with her. Then use the population figures from Try It Now 6.4.2 to predict the standard deviation of the whole distribution of pair averages.

Solution

Step 1 — average the first pair.

$$ \frac{0.13 + 1.52}{2} = \frac{1.65}{2} = 0.825 $$

Step 2 — average the second pair.

$$ \frac{0.45 + 0.72}{2} = \frac{1.17}{2} = 0.585 $$

Step 3 — predict the spread of all such averages. With \(\sigma \approx 0.5122\) and \(n = 2\),

$$ \frac{\sigma}{\sqrt{n}} = \frac{0.5122}{\sqrt{2}} \approx \frac{0.5122}{1.414} \approx 0.3622 $$

Answer: Mai's pair averages are \(\$0.825\) and \(\$0.585\), and the distribution of pair averages should have a standard deviation near \(\$0.36\), down from \(\$0.51\). Watch what happened to the \(\$1.52\) in her first pair. On its own it sat well out in the right tail; paired with a thirteen-cent pocket it produced an average of \(\$0.825\), barely above the middle of the picture. Nothing was thrown away and nothing was corrected — the extreme value is still in there, it is just sharing its slot with a partner. That dilution is the whole mechanism, and doing it with two people already recovers about 29% of the spread.

6.4.4 Collecting Averages of Groups of Five

Definition 6.4.3: Standard Error of the Mean

The standard deviation of the sampling distribution of the sample mean:

$$ \sigma_{\bar{x}} = \frac{\sigma}{\sqrt{n}} $$
Five times the work buys less than half the spread

Going from single people to groups of five divides the spread by \(\sqrt{5} \approx 2.24\), not by 5. To halve it again you would need groups of 20. Precision gets expensive fast, and knowing the exchange rate is what stops you over-promising.

It measures how far a typical sample mean lands from the population mean \(\mu\). It is not the spread of the individuals — that is \(\sigma\) — and it shrinks toward zero as \(n\) grows, which is why a larger sample gives a more reliable estimate.

Repeat steps one through five of the section titled Collect the Data, with one exception. Instead of recording the change of 30 classmates, record the average change of 30 groups of five.

  1. Randomly survey 30 groups of five classmates.
  2. Record the values of the average of their change in Table 6.4.4.
Table 6.4.4 — The average change of each of your 30 surveyed groups of five, in dollars.
1–56–1011–1516–2021–2526–30
  1. Construct a histogram. Scale the axes using the same scaling you used for the section titled Collect the Data. Sketch the graph using a ruler and a pencil.
  2. Calculate the following, with \(n = 5\) — you are surveying five people at a time:

a. \(\bar{x} =\) _________

b. \(s =\) _________

  1. Draw a smooth curve through the tops of the bars of the histogram. Use one to two complete sentences to describe the general shape of the curve.
Try It Now 6.4.4

Grant ran the third round with his husband Peter and three classmates, and their five pockets held \(\$0.05\), \(\$0.28\), \(\$0.63\), \(\$1.06\), and \(\$2.10\). Find the group's average. Then find the standard error for \(n = 5\), and use it to say how unusual that group average is compared with how unusual the \(\$2.10\) was on its own.

Solution

Step 1 — average the five values.

$$ \frac{0.05 + 0.28 + 0.63 + 1.06 + 2.10}{5} = \frac{4.12}{5} = 0.824 $$

Step 2 — find the standard error. With \(\sigma \approx 0.5122\) and \(n = 5\),

$$ \sigma_{\bar{x}} = \frac{0.5122}{\sqrt{5}} \approx \frac{0.5122}{2.236} \approx 0.2291 $$

Step 3 — measure each value against the spread that applies to it. The \(\$2.10\) is an individual, so it is measured against \(\sigma\):

$$ \frac{2.10 - 0.70}{0.5122} \approx 2.73 $$

The \(\$0.824\) is an average of five, so it is measured against \(\sigma_{\bar{x}}\):

$$ \frac{0.824 - 0.70}{0.2291} \approx 0.54 $$

Answer: Grant's group average is \(\$0.824\), the standard error is about \(\$0.23\), and the two \(z\)-scores are 2.73 and 0.54. His group is a genuine outlier when you look at its largest member and thoroughly ordinary when you look at its average. Neither reading is wrong — they answer different questions, and the number you divide by is what says which question you asked. Using \(\sigma\) where \(\sigma_{\bar{x}}\) belongs is the single most common error in the rest of this course, and it always makes a result look less surprising than it really is.

6.4.5 Discussion Questions

Answer these with your group, in complete sentences, pointing at your own three histograms rather than at what the theorem says ought to have happened.

  1. Why did the shape of the distribution of the data change, as \(n\) changed? Use one to two complete sentences to explain what happened.
  2. In the section titled Collect the Data, what was the approximate distribution of the data?
  1. In the section titled Collecting Averages of Groups of Five, what was the approximate distribution of the averages?
  1. In one to two complete sentences, explain any differences in your answers to the previous two questions.

Questions 2 and 3 are asking for the same three pieces of information each time — a family name, a center, and a spread — and the interesting part is that only one of the two can be filled in confidently. Do not force a familiar family onto the first blank just because the blank is there. Saying plainly that the population does not belong to any distribution you have a name for is a correct and complete answer, and it is the honest starting point for question 4.

Try It Now 6.4.5

Nayeli and their girlfriend Ceci are writing up the discussion for their group. Answer questions 2, 3, and 4 with them, for the class data in Table 6.4.2, using \(\bar{x} = \$0.70\) and \(s \approx \$0.51\).

Solution

Step 1 — describe the population, question 2. From Try It Now 6.4.2 the mean is \(\$0.70\), the standard deviation is about \(\$0.51\), and the mean sits above the median, so the distribution leans right. There is no named family here — it is not normal, not uniform, not exponential. The honest answer is that

$$ X \sim (\text{right-skewed, no named family}),\quad \mu \approx 0.70,\ \sigma \approx 0.51 $$

Step 2 — describe the averages, question 3. The central limit theorem supplies the family, and the two parameters come from Step 1 with \(n = 5\).

$$ \bar{X} \sim N\!\left(0.70,\ \frac{0.51}{\sqrt{5}}\right) = N(0.70,\ 0.23) $$

Step 3 — compare them, question 4. Same center, different family, and a spread cut by a factor of \(\sqrt{5}\).

Answer: \(X\) is right-skewed with mean \(\$0.70\) and standard deviation \(\$0.51\) and belongs to no named family; \(\bar{X}\) is approximately \(N(0.70, 0.23)\). The two answers share a center and differ in everything else, which is the finding the lab was built to produce — averaging changed the shape and the spread without moving the middle. One caution belongs in their write-up and in yours: \(n = 5\) is a small sample to lean the central limit theorem on when the population is this lopsided, so expect your third histogram to be more bell-shaped than the first two without being convincingly bell-shaped on its own. If your class pooled data with other sections and drew groups of ten or twenty as well, the picture would tighten noticeably, and that improvement with larger \(n\) is itself the theorem showing up in your data.

Key Terms

central limit theorem — the result that the sampling distribution of the sample mean approaches \(N(\mu, \sigma/\sqrt{n})\) as the sample size \(n\) grows, whatever the shape of the population.

sample mean — the average of the values in one sample, written \(\bar{x}\); it estimates the population mean \(\mu\).

sampling distribution of the sample mean — the distribution formed by the sample means of all possible samples of a fixed size \(n\) drawn from a population, written \(\bar{X}\).

standard error of the mean — the standard deviation of the sampling distribution of the sample mean, \(\sigma_{\bar{x}} = \sigma/\sqrt{n}\).