2.6 Skewness and the Mean, Median, and Mode

Aligned outcomes:

SLO 2

Identify appropriate graphs and summary statistics for variables and relationships between them and correctly interpret information from graphs and summary statistics.

Shape and summary statistics are two readings of the same data. Naming a histogram symmetrical, skewed left, or skewed right lets you predict the order of mean, median, and mode before computing any of them - and tells you which one to quote so the number matches the picture.

Learning Objectives

By the end of this section, you will be able to:

In this section, you will learn to:
  • identify a symmetrical distribution from its histogram or dot plot and explain what the mirror-image test checks;
  • describe a distribution as skewed to the left or skewed to the right from the shape of its graph;
  • predict the typical ordering of the mean, the median, and the mode for a symmetrical, a left-skewed, and a right-skewed distribution;
  • interpret a graph of real data by naming its shape and saying what that shape implies about which measure of the center to quote.

Section 2.5 gave you three numbers for the center of a data set — the mean, the median, and the mode — and left you with a question it could not fully answer: when they disagree, which one is telling the truth?

The answer is written on the graph. The shape of a distribution and the ordering of its three centers are two views of the same fact, and once you can read one you can predict the other. This section teaches that connection. From here on, "what does this data look like?" and "which average should I quote?" are the same question.

2.6.1 Symmetrical Distributions

Definition 2.6.1: Symmetrical Distribution

A distribution is symmetrical if a vertical line can be drawn at some point in the histogram such that the shape to the left of the line and the shape to the right of the line are mirror images of each other.

The mean is a balance point

Picture the histogram as blocks on a seesaw. The mean is where you would put the pivot so it balances. Fold a symmetrical shape down the middle and the weight on each side is equal — so the pivot and the fold line land in the same place.

Definition 2.6.2: Unimodal Distribution

A distribution is unimodal if it has exactly one mode — a single value (or single interval) that occurs more often than any other.

Here is what both definitions look like on real numbers. Start with a data set of sixteen values:

4; 5; 6; 6; 6; 7; 7; 7; 7; 7; 7; 8; 8; 8; 9; 10

We can draw this as a histogram. Each interval has width one, and each value sits in the middle of its interval, so the height of a bar is just a count of how many times that value occurs.

Figure 2.6.1 — A symmetrical distribution: the tall bar at 7 is the fold line, and the bars on either side match. Figure 2.6.1 — A symmetrical distribution: the tall bar at 7 is the fold line, and the bars on either side match.

Figure 2.6.1 — A symmetrical distribution: the tall bar at 7 is the fold line, and the bars on either side match.

Look at Figure 2.6.1 and imagine dropping a vertical line straight down through the tall bar at 7. The left half and the right half are mirror images — a bar of height 3 on each side, then a bar of height 1 on each side. That mirror-image property is what Definition 2.6.1 describes, and the tall bar at 7 is the only place the fold line works.

Definition 2.6.1 — Drop the fold line and every bar on the left has a partner of the same height the same distance to the right.

For these sixteen values the mean, the median, and the mode are each seven. That is not a coincidence, and it is the first half of the rule this whole section is built on: in a perfectly symmetrical distribution the mean and the median are the same. The fold line is the balance point and the halfway point at once.

The data set also has a single tallest bar, so it has one mode — it is unimodal, and that mode is seven, the same value as the mean and the median. All three centers land together. That agreement depends on the distribution being unimodal: a symmetrical distribution can instead have two modes — the bimodal case from Section 2.5 — and then the two modes sit out on either side of the fold line while the mean and the median stay at the center. The distribution is still symmetrical; the modes simply are not at the middle.

Try It Now 2.6.1

Kyle runs the campus tutoring center with his husband, and he scores a 5-point quiz for the twelve students who came in this week. His twelve scores are

1; 2; 2; 3; 3; 3; 3; 4; 4; 4; 5; 5

Is this distribution symmetrical? Support your answer with the mean and the median.

Solution

Step 1 — Build the frequency picture. The value 1 occurs once, 2 occurs twice, 3 occurs four times, 4 occurs three times, and 5 occurs twice. Bar heights: 1, 2, 4, 3, 2.

Step 2 — Apply the mirror test. Fold at the tall bar (3). To the left the heights are 1, 2; to the right they are 3, 2. Those are not mirror images, so the shape is not symmetrical.

Step 3 — Check with the numbers.

$$\overline{x} = \frac{1 + (2)(2) + (3)(4) + (4)(3) + (5)(2)}{12} = \frac{39}{12} = 3.25$$

With twelve values the median sits between the 6th and 7th ordered values, which are 3 and 3, so \(M = 3\).

Answer: Not symmetrical. The mean (3.25) is larger than the median (3), which tells you the shape is pulled slightly toward the high end — the mirror test and the numbers agree.

2.6.2 Distributions Skewed to the Left

Definition 2.6.3: Skewed to the Left (Negative Skew)

A distribution is skewed to the left (also called negatively skewed) if it has a longer, thinner tail extending toward the lower values, with the bulk of the data concentrated at the higher values.

The name points at the tail, not at the pile

Students reliably get this backwards, because the eye is drawn to the tall bars. The label always names the direction the tail stretches. Tail on the left means skewed left, no matter where the bars are stacked.

Now shorten that data set on the right-hand side. Here are ten values:

4; 5; 6; 6; 6; 7; 7; 7; 7; 8

Figure 2.6.2 — A distribution skewed to the left: the right side looks chopped off and the tail stretches toward the low values.

The right-hand side of Figure 2.6.2 seems chopped off compared to the left side. The bulk of the data is packed against the high end, and a thin tail trails away to the left — the distribution is pulled out to the left, which is what earns it the name.

For these ten values the mean is 6.3, the median is 6.5, and the mode is seven. Notice the ordering: the mean is less than the median, and both are less than the mode. Both the mean and the median react to the low tail by sliding down away from the mode, but the mean slides further. That is because the mean has to be the balance point, so a single value stranded far out on the left tugs the pivot toward itself; the median only counts positions, so one distant value moves it by at most one slot.

Try It Now 2.6.2

Jun is the grader for a section of eleven students, and they record how many of the 11 assigned homework sets each student turned in:

11; 11; 11; 11; 10; 10; 10; 9; 8; 6; 3

Sketch the shape in words, name the skew, and order the mean, the median, and the mode.

Solution

Step 1 — Order the data and count. Ordered: 3; 6; 8; 9; 10; 10; 10; 11; 11; 11; 11. Bar heights from 3 up to 11: one at 3, one at 6, one at 8, one at 9, three at 10, four at 11.

Step 2 — Name the shape. The tall bars are bunched at the top of the scale (10 and 11) and a thin tail of single values runs down to 3. The tail points left, so the distribution is skewed to the left.

Step 3 — Compute the three centers.

$$\overline{x} = \frac{3 + 6 + 8 + 9 + (10)(3) + (11)(4)}{11} = \frac{100}{11} \approx 9.1$$

With eleven values the median is the 6th ordered value, so \(M = 10\). The mode is 11 (four occurrences).

Answer: Skewed left, and mean (9.1) < median (10) < mode (11) — exactly the ordering Figure 2.6.2 shows.

2.6.3 Distributions Skewed to the Right

Definition 2.6.4: Skewed to the Right (Positive Skew)

A distribution is skewed to the right (also called positively skewed) if it has a longer, thinner tail extending toward the higher values, with the bulk of the data concentrated at the lower values.

Flip the situation. Take these ten values, which bunch against the low end instead:

6; 7; 7; 7; 7; 8; 8; 8; 9; 10

Figure 2.6.3 — A distribution skewed to the right: the pile sits at the low end and the tail stretches toward the high values.

Figure 2.6.3 is the mirror of Figure 2.6.2. The left-hand side is the chopped-off one now, and the tail runs out to the right, so the distribution is skewed to the right.

For these ten values the mean is 7.7, the median is 7.5, and the mode is seven. Of the three statistics the mean is now the largest and the mode is the smallest — the ordering from the left-skewed case, reversed. Again the mean reflects the skewing the most.

Right skew is the shape you meet most often in real data, and the reason is simple: many quantities have a hard floor and no ceiling. Nobody earns a negative salary, waits a negative number of minutes, or owns a negative number of cars, so those distributions cannot have a long left tail — but a handful of very large values can always stretch the right one. This is why a news report about "average income" is worth a second thought, and why the median is the number housing statistics almost always quote.

Try It Now 2.6.3

Marisol owns a café near campus, and she times how many minutes each of nine customers waited for a table:

2; 3; 3; 3; 4; 4; 5; 9; 15

Name the skew and predict, before computing, whether the mean or the median will be larger. Then check.

Solution

Step 1 — Name the shape. Seven of the nine waits sit between 2 and 5 minutes; two stragglers at 9 and 15 stretch far out to the high side. Tail on the right, so the distribution is skewed to the right.

Step 2 — Predict. In a right-skewed distribution the high tail drags the balance point up, so the mean should be the larger of the two.

Step 3 — Check.

$$\overline{x} = \frac{2 + 3 + 3 + 3 + 4 + 4 + 5 + 9 + 15}{9} = \frac{48}{9} \approx 5.3$$

With nine values the median is the 5th ordered value, so \(M = 4\).

Answer: Skewed right, and the mean (about 5.3 minutes) is larger than the median (4 minutes), as predicted. The single 15-minute wait moves the mean by more than a full minute and moves the median not at all.

2.6.4 What the Shape Tells You About the Center

Everything above comes down to one mechanism: the mean is affected by outliers that do not influence the median. An extreme value enters the mean's sum at full strength, so it pulls the balance point toward itself. The median only cares about where values sit in the order, so replacing the largest value with one ten times larger does not move it at all.

That gives you three working rules, and they run in both directions — shape to centers, or centers to shape:

Try it in rāSHio

These three rules are quick to test on any data set you have. Open rāSHio, paste a list into File → Delimited List…, then choose Stats → Summary Statistics. The mean and the median come back side by side, so you can read the ordering off one panel and predict the shape before you ever draw the graph — then draw it and check. Try it on the ten values from §2.6.2 and again on the ten from §2.6.3; the two panels disagree in opposite directions.

Read the word "often" carefully. These are strong tendencies, not theorems — they are not true for every data set. The most common exceptions occur in sets of discrete data, where a few repeated values can hold the median in place while the shape leans one way. Problem 2.6.16 in the problem set is exactly such a case: a visibly lopsided data set whose mean and median are identical. So use the rules to form an expectation, then look at the graph to confirm it.

Skewness and symmetry come back in a serious way when we get to probability distributions in later chapters, where the shape of a curve decides which formulas you are allowed to use at all. For now, the payoff is practical: naming the shape tells you which measure of the center to quote, and quoting the wrong one is how a true set of numbers gets used to tell a false story.

Try It Now 2.6.4

Harper is building a slide deck for her campus queer-in-STEM study group, and she wants one shape-reading example from three different sources. Discuss the mean, median, and mode for each of the following. Is there a pattern between the shape and the measure of the center?

a. The number of gold medals won by each of the top 20 medal-winning countries at the 2010 Winter Olympics:

Figure 2.6.4 — Gold medal wins by the top 20 medal-winning countries, 2010 Winter Olympics. Figure 2.6.4 — Gold medal wins by the top 20 medal-winning countries, 2010 Winter Olympics.

Figure 2.6.4 — Gold medal wins by the top 20 medal-winning countries, 2010 Winter Olympics.

b. The ages at which former U.S. presidents died:

Table 2.6.1 — The ages former U.S. presidents died. Key: 8|0 means 80.
StemLeaves
46 9
53 6 7 7 7 8
60 0 3 3 4 4 5 6 7 7 7 8
70 1 1 2 3 4 7 8 8 9
80 1 3 5 8
90 0 3 3

c. The number of hours a group of students spent playing video games over a weekend:

Figure 2.6.5 — Hours spent playing video games on weekends, for 25 students. Figure 2.6.5 — Hours spent playing video games on weekends, for 25 students.

Figure 2.6.5 — Hours spent playing video games on weekends, for 25 students.

Solution

Part a — the medal counts. Reading the dots: 0, 0, 1, 1, 1, 1, 2, 2, 2, 3, 4, 4, 5, 5, 6, 6, 9, 9, 10, 14. Most countries won only a medal or two and one country won 14, so the pile sits at the low end with a long tail to the right — skewed to the right.

$$\overline{x} = \frac{85}{20} = 4.25$$

With 20 values the median is the average of the 10th and 11th ordered values, \(\frac{3 + 4}{2} = 3.5\). The mode is 1 (four countries). So mode (1) < median (3.5) < mean (4.25) — the right-skew ordering. If you wanted one number for "a typical country's gold haul", the mean of 4.25 overstates it; no country actually won between 6 and 9.

Part b — the presidents' ages. The stem plot holds 39 ages. The tallest row is the 60s (12 ages), the 70s row holds 10, and a thinner tail runs up through the 80s and 90s, so the peak sits left of center with the longer tail on the high side — the shape is skewed to the right.

With 39 values the median is the 20th ordered value, which is 68, so \(M = 68\). Adding all 39 ages gives

$$\overline{x} = \frac{2733}{39} \approx 70.1$$

Two leaf values tie for most frequent — 57 and 67, three each — so this data set is bimodal, and both modes sit below the median. Mean (about 70.1) > median (68), matching the right-skew rule.

Part c — the video game hours. Bar heights rise steadily left to right: 2, 3, 4, 7, 9. The pile is at the high end and the thin tail runs to the low end, so the distribution is skewed to the left.

The modal interval is 20–24.99 hours. With 25 students the median is the 13th value, which lands in the 15–19.99 interval. Estimating the mean from interval midpoints gives about 16.1 hours. So mean (about 16.1) < median (in 15–19.99) < mode (20–24.99) — the left-skew ordering.

Answer: Yes, and it is the same pattern all three times. The mode sits at the tall part of the graph, the mean gets dragged out toward the tail, and the median lands between them. Naming the skew is enough to predict the order of the three centers before computing any of them.

Example 2.6.1: Comparing Three Authors' Letter Counts

Statistics are used to compare and sometimes identify authors. Three writers work the same beat for the same local paper: Terry files the crime column and he favors short words, Delgado writes the Spanish-language edition and she keeps her sentences plain, and Raman covers the arts and they reach for longer vocabulary. The following lists show a simple random sample comparing the letter counts of the words each of them used.

Terry: 7; 9; 3; 3; 3; 4; 1; 3; 2; 2

Delgado: 3; 3; 3; 4; 1; 4; 3; 2; 3; 1

Raman: 2; 3; 4; 4; 4; 6; 6; 6; 8; 3

a. Make a dot plot for the three authors and compare the shapes.

b. Calculate the mean for each.

c. Calculate the median for each.

d. Describe any pattern you notice between the shape and the measures of the center.

Solution

Step 1 — Dot plots (part a). For each author, mark one X above each value on a common number line.

Terry's letter count:

Figure 2.6.6 — Terry's distribution has a right (positive) skew.

Delgado's letter count:

Figure 2.6.7 — Delgado's distribution has a left (negative) skew.

Raman's letter count:

Figure 2.6.8 — Raman's distribution is symmetrically shaped.

Step 2 — Means (part b).

$$\overline{x}_{\text{Terry}} = \frac{1 + (2)(2) + (3)(4) + 4 + 7 + 9}{10} = \frac{37}{10} = 3.7$$ $$\overline{x}_{\text{Delgado}} = \frac{(1)(2) + 2 + (3)(5) + (4)(2)}{10} = \frac{27}{10} = 2.7$$ $$\overline{x}_{\text{Raman}} = \frac{2 + (3)(2) + (4)(3) + (6)(3) + 8}{10} = \frac{46}{10} = 4.6$$

So Terry's mean is 3.7, Delgado's mean is 2.7, and Raman's mean is 4.6.

Step 3 — Medians (part c). Each list has ten values, so the median is the average of the 5th and 6th ordered values.

Terry ordered: 1; 2; 2; 3; 3; 3; 3; 4; 7; 9 — the 5th and 6th are both 3, so \(M = 3\).

Delgado ordered: 1; 1; 2; 3; 3; 3; 3; 3; 4; 4 — the 5th and 6th are both 3, so \(M = 3\).

Raman ordered: 2; 3; 3; 4; 4; 4; 6; 6; 6; 8 — the 5th and 6th are 4 and 4, so \(M = 4\).

So Terry's median is three, Delgado's median is three, and Raman's median is four.

Step 4 — The pattern (part d). Line the three authors up against the rules above. Terry is skewed right and has mean (3.7) above median (3). Delgado is skewed left and has mean (2.7) below median (3). Raman is symmetrical and has mean (4.6) and median (4) close together.

Answer: It appears that the median is always closest to the high point of the graph (the mode), while the mean tends to be farther out on the tail. In a symmetrical distribution the mean and the median are both centrally located, close to the high point of the distribution.

Problem Set 2.6

State whether the data are symmetrical, skewed to the left, or skewed to the right.

Problem 1. 1; 1; 1; 2; 2; 2; 2; 3; 3; 3; 3; 3; 3; 3; 3; 4; 4; 4; 5; 5

Solution

Step 1 — Count each value. 1 occurs 3 times, 2 occurs 4 times, 3 occurs 8 times, 4 occurs 3 times, and 5 occurs 2 times, for \(n = 20\) values.

Step 2 — Find the mean.

$$\overline{x} = \frac{(1)(3) + (2)(4) + (3)(8) + (4)(3) + (5)(2)}{20} = \frac{3 + 8 + 24 + 12 + 10}{20} = \frac{57}{20} = 2.85$$

Step 3 — Find the median. With 20 values the median sits between the 10th and 11th ordered values. Counting up, values 8 through 15 are all 3, so both the 10th and 11th are 3 and \(M = 3\).

Step 4 — Compare. The mean (2.85) and the median (3) are close, and the mode (3) sits near the middle of the data rather than out at one end.

Answer: The data are symmetrical.

Problem 2. 16; 17; 19; 22; 22; 22; 22; 22; 23

Solution

Step 1 — Count each value. 16, 17, and 19 occur once each; 22 occurs five times; 23 occurs once. That is \(n = 9\) values.

Step 2 — Find the mean.

$$\overline{x} = \frac{16 + 17 + 19 + (22)(5) + 23}{9} = \frac{185}{9} \approx 20.6$$

Step 3 — Find the median. With nine values the median is the 5th ordered value, which is 22, so \(M = 22\). The mode is also 22.

Step 4 — Read the shape from the ordering. The bulk of the data is packed at 22, and the values 16, 17, and 19 form a thin tail trailing away to the low side. The mean (about 20.6) has been dragged below the median (22).

Answer: The data are skewed to the left.

Problem 3. 87; 87; 87; 87; 87; 88; 89; 89; 90; 91

Solution

Step 1 — Count each value. 87 occurs five times; 88 once; 89 twice; 90 once; 91 once. That is \(n = 10\) values.

Step 2 — Find the mean.

$$\overline{x} = \frac{(87)(5) + 88 + (89)(2) + 90 + 91}{10} = \frac{882}{10} = 88.2$$

Step 3 — Find the median. With ten values the median is the average of the 5th and 6th ordered values, which are 87 and 88, so \(M = 87.5\).

Step 4 — Read the shape. Even though the mean and median are close, the mode (87) lies to the left of the middle of the data, and there are many more 87s than any other value, so the pile is at the low end with a tail running to the high end.

Answer: The data are skewed to the right.

Problem 4. When the data are skewed left, what is the typical relationship between the mean and median?

Solution

In a left-skewed distribution the long thin tail points toward the low values. The mean is a balance point, so those few low values pull it down; the median only counts positions, so it barely moves.

Answer: When the data are skewed left, the mean is typically less than the median.

Problem 5. When the data are symmetrical, what is the typical relationship between the mean and median?

Solution

A symmetrical distribution folds into two matching halves, so the balance point and the halfway point land in the same place.

Answer: When the data are symmetrical, the mean and the median are close to each other or exactly the same.

Problem 6. What word describes a distribution that has two modes?

Solution

Two values tie for the highest frequency, so the graph has two equally tall peaks.

Answer: Bimodal.

Problem 7. Describe the shape of this distribution.

Problem 2.6.7 — the distribution this exercise asks about. Problem 2.6.7 — the distribution this exercise asks about.

Problem 2.6.7 — the distribution this exercise asks about.

Solution

Step 1 — Read the bar heights. From left to right the bars stand at 8, 4, 2, 2, and 1 over the values 3, 4, 5, 6, and 7.

Step 2 — Locate the pile and the tail. The tall bars are stacked at the low end (3 and 4) and the bars shrink steadily as you move right, so the thin tail points toward the high values.

Answer: The distribution is skewed right, because it looks pulled out to the right.

Problem 8. Describe the relationship between the mode and the median of this distribution.

Problem 2.6.8 — the distribution this exercise asks about. Problem 2.6.8 — the distribution this exercise asks about.

Problem 2.6.8 — the distribution this exercise asks about.

Solution

Step 1 — Find the mode. The tallest bar is over 3 with a height of 8, so the mode is 3.

Step 2 — Find the median. The bar heights 8, 4, 2, 2, 1 total \(n = 17\) values, so the median is the 9th ordered value. The first 8 values are all 3 and the next 4 are all 4, so the 9th value is 4 and \(M = 4\).

Answer: The mode (3) is less than the median (4). That is the ordering you expect in a right-skewed distribution: the mode sits at the peak on the left, and the median has already been pushed toward the tail.

Problem 9. Describe the relationship between the mean and the median of this distribution.

Problem 2.6.9 — the distribution this exercise asks about. Problem 2.6.9 — the distribution this exercise asks about.

Problem 2.6.9 — the distribution this exercise asks about.

Solution

Step 1 — Read the counts. Heights 8, 4, 2, 2, 1 over the values 3, 4, 5, 6, 7, so \(n = 17\).

Step 2 — Find the mean.

$$\overline{x} = \frac{(3)(8) + (4)(4) + (5)(2) + (6)(2) + (7)(1)}{17} = \frac{24 + 16 + 10 + 12 + 7}{17} = \frac{69}{17} \approx 4.1$$

Step 3 — Find the median. The 9th ordered value is 4, so \(M = 4\).

Answer: The mean is 4.1 and is slightly greater than the median, which is four.

Problem 10. Describe the shape of this distribution.

Problem 2.6.10 — the distribution this exercise asks about. Problem 2.6.10 — the distribution this exercise asks about.

Problem 2.6.10 — the distribution this exercise asks about.

Solution

Step 1 — Read the bar heights. From left to right: 2, 4, 8, 3, 2 over the values 3, 4, 5, 6, and 7.

Step 2 — Apply the mirror test. The one tall bar is over 5. To its left the heights are 2, 4; to its right they are 3, 2. Those are nearly mirror images but not exactly — the left shoulder carries one extra value.

Answer: The distribution is close to symmetrical, with a single peak at 5 and matching shoulders. It leans very slightly left, because the left shoulder is a little heavier than the right one.

Problem 11. Describe the relationship between the mode and the median of this distribution.

Problem 2.6.11 — the distribution this exercise asks about. Problem 2.6.11 — the distribution this exercise asks about.

Problem 2.6.11 — the distribution this exercise asks about.

Solution

Step 1 — Find the mode. The tallest bar is over 5 with a height of 8, so the mode is 5.

Step 2 — Find the median. The heights 2, 4, 8, 3, 2 total \(n = 19\) values, so the median is the 10th ordered value. The first 2 values are 3, the next 4 are 4 (through position 6), and the next 8 are 5 (positions 7 through 14). Position 10 falls in that run, so \(M = 5\).

Answer: The mode and the median are the same. In this case they are both five.

Problem 12. Are the mean and the median the exact same in this distribution? Why or why not?

Problem 2.6.12 — the distribution this exercise asks about. Problem 2.6.12 — the distribution this exercise asks about.

Problem 2.6.12 — the distribution this exercise asks about.

Solution

Step 1 — Find the mean. The heights are 2, 4, 8, 3, 2 over the values 3 through 7, so \(n = 19\) and

$$\overline{x} = \frac{(3)(2) + (4)(4) + (5)(8) + (6)(3) + (7)(2)}{19} = \frac{6 + 16 + 40 + 18 + 14}{19} = \frac{94}{19} \approx 4.9$$

Step 2 — Find the median. The 10th ordered value is 5, so \(M = 5\).

Step 3 — Compare. They differ by about a tenth of a unit.

Answer: No, they are not exactly the same. The mean is about 4.9 and the median is 5. They are very close because the distribution is nearly symmetrical, but they are not equal, because the shape is not a perfect mirror image — the left shoulder holds one more value than the right one, which pulls the balance point a little below the halfway point.

Problem 13. Describe the shape of this distribution.

Problem 2.6.13 — the distribution this exercise asks about. Problem 2.6.13 — the distribution this exercise asks about.

Problem 2.6.13 — the distribution this exercise asks about.

Solution

Step 1 — Read the bar heights. From left to right: 1, 1, 2, 4, 7 over the values 3, 4, 5, 6, and 7.

Step 2 — Locate the pile and the tail. The bars grow taller as you move right, so the data is packed against the high end and the thin tail of single values runs off to the left.

Answer: The distribution is skewed left, because it looks pulled out to the left.

Problem 14. Describe the relationship between the mode and the median of this distribution.

Problem 2.6.14 — the distribution this exercise asks about. Problem 2.6.14 — the distribution this exercise asks about.

Problem 2.6.14 — the distribution this exercise asks about.

Solution

Step 1 — Find the mode. The tallest bar is over 7 with a height of 7, so the mode is 7.

Step 2 — Find the median. The heights 1, 1, 2, 4, 7 total \(n = 15\) values, so the median is the 8th ordered value. Counting up: position 1 is 3, position 2 is 4, positions 3 and 4 are 5, positions 5 through 8 are 6. So \(M = 6\).

Answer: The mode (7) is greater than the median (6). That is the ordering you expect in a left-skewed distribution: the mode sits at the peak on the right, and the median has been pulled toward the low tail.

Problem 15. Describe the relationship between the mean and the median of this distribution.

Problem 2.6.15 — the distribution this exercise asks about. Problem 2.6.15 — the distribution this exercise asks about.

Problem 2.6.15 — the distribution this exercise asks about.

Solution

Step 1 — Read the counts. Heights 1, 1, 2, 4, 7 over the values 3, 4, 5, 6, 7, so \(n = 15\).

Step 2 — Find the mean.

$$\overline{x} = \frac{(3)(1) + (4)(1) + (5)(2) + (6)(4) + (7)(7)}{15} = \frac{3 + 4 + 10 + 24 + 49}{15} = \frac{90}{15} = 6$$

Step 3 — Find the median. The 8th ordered value is 6, so \(M = 6\).

Answer: The mean and the median are both six.

Problem 16. The mean and median for the data are the same.

3; 4; 5; 5; 6; 6; 6; 6; 7; 7; 7; 7; 7; 7; 7

Is the data perfectly symmetrical? Why or why not?

Solution

Step 1 — Confirm that the mean and the median really are equal. There are 15 values, so

$$\overline{x} = \frac{3 + 4 + (5)(2) + (6)(4) + (7)(7)}{15} = \frac{90}{15} = 6$$

and the median is the 8th ordered value, which is 6. So both are 6.

Step 2 — Apply the mirror test to the shape. The bar heights are 1 at 3, 1 at 4, 2 at 5, 4 at 6, and 7 at 7. Those heights climb steadily to the right — there is no vertical line that makes the left half a mirror image of the right half.

Answer: No, the data are not perfectly symmetrical. The graph is piled against 7 with a thin tail running down to 3, which is a left-skewed shape. The mean and the median happening to agree at 6 does not make a distribution symmetrical — equal mean and median is something symmetry guarantees, not something that guarantees symmetry. Notice the mode (7) is nowhere near the other two, which is the giveaway.

Problem 17. Which is the greatest, the mean, the mode, or the median of the data set?

11; 11; 12; 12; 12; 12; 13; 15; 17; 22; 22; 22

Solution

Step 1 — Find the mode. 12 occurs four times, more than any other value, so the mode is 12.

Step 2 — Find the median. There are 12 values, so the median is the average of the 6th and 7th ordered values, which are 12 and 13:

$$M = \frac{12 + 13}{2} = 12.5$$

Step 3 — Find the mean.

$$\overline{x} = \frac{(11)(2) + (12)(4) + 13 + 15 + 17 + (22)(3)}{12} = \frac{181}{12} \approx 15.1$$

Answer: The mode is 12, the median is 12.5, and the mean is about 15.1. The mean is the greatest. The three 22s form a high tail, and only the mean feels them.

Problem 18. Which is the least, the mean, the mode, and the median of the data set?

56; 56; 56; 58; 59; 60; 62; 64; 64; 65; 67

Solution

Step 1 — Find the mode. 56 occurs three times, more than any other value, so the mode is 56.

Step 2 — Find the median. There are 11 values, so the median is the 6th ordered value, which is 60.

Step 3 — Find the mean.

$$\overline{x} = \frac{(56)(3) + 58 + 59 + 60 + 62 + (64)(2) + 65 + 67}{11} = \frac{667}{11} \approx 60.6$$

Answer: The mode is 56, the median is 60, and the mean is about 60.6. The mode is the least. The data pile up at the low end and thin out toward 67, which is the right-skewed pattern.

Problem 19. Of the three measures, which tends to reflect skewing the most, the mean, the mode, or the median? Why?

Solution

The mode is just the location of the tallest bar, so it does not move at all when a far-out value is added. The median only counts positions in the ordered list, so an extreme value shifts it by at most one slot. The mean has to be the balance point of the whole data set, so every value enters its sum at full strength and a single far-out value drags it.

Answer: The mean tends to reflect skewing the most, because it is affected the most by outliers.

Problem 20. In a perfectly symmetrical distribution, when would the mode be different from the mean and median?

Solution

In a symmetrical distribution the mean and the median both land on the fold line. The mode is a different kind of number — it marks the tallest bar, and nothing forces the tallest bar to sit on the fold line.

Answer: When the distribution is bimodal. Two values tie for the highest frequency, and in a symmetrical shape they sit out on either side of the fold line, so neither one equals the mean or the median. (A perfectly symmetrical distribution in which every value occurs equally often has no single mode either, for the same reason.)

Problem 21. The median age of the U.S. population in 1980 was 30.0 years. In 1991, the median age was 33.1 years.

a) What does it mean for the median age to rise?

b) Give two reasons why the median age could rise.

c) For the median age to rise, is the actual number of children less in 1991 than it was in 1980? Why or why not?

Solution

a) The median age is the age that splits the population into two equal halves. Saying it rose from 30.0 to 33.1 years means the person standing in the middle of the ordered list of every American got about three years older — in 1991 half the population was older than 33.1, while in 1980 half was older than only 30.0.

b) Two reasons, either of which will do it:

  • People are living longer. Better medicine and public health push more of the population into the older ages, which moves the middle of the list up.
  • Families are having fewer children. Fewer births means fewer new values entering at the bottom of the ordered list, so the middle position shifts toward the older ages.

c) No, not necessarily. The median depends on positions in the ordered list, not on counts at either end. The number of children in 1991 could have been the same as in 1980, or even larger, and the median would still rise as long as the older part of the population grew faster. The rise tells you the balance of the age distribution shifted upward; it does not tell you any particular group shrank.

Key Terms

symmetrical distribution — a distribution whose histogram can be split by a vertical line into two halves that are mirror images of each other.

unimodal — having exactly one mode, a single value or interval that occurs more often than any other.

skewed to the left — having a longer, thinner tail toward the lower values, with the data piled at the higher values; also called negatively skewed.

skewed to the right — having a longer, thinner tail toward the higher values, with the data piled at the lower values; also called positively skewed.

skewness — the departure of a distribution's shape from symmetry, named for the direction its longer tail points.