2.4 Box Plots
SLO 2
Identify appropriate graphs and summary statistics for variables and relationships between them and correctly interpret information from graphs and summary statistics.
The box plot is the graph that draws a summary statistic rather than the raw data: five numbers, one number line. You learn to build one and to read spread straight off it — which quarter is packed, which is stretched — and to compare two groups by stacking their boxes on a shared axis.
Learning Objectives
By the end of this section, you will be able to:
- list the five numbers a box plot is built from and read each one off a finished plot;
- construct a box plot on a scaled number line from a data set or from a five-number summary;
- explain what the box, the whiskers, and the median line each say about where the data sits;
- interpret a box plot in which two or more of the five numbers are equal;
- compare two or more data sets by drawing their box plots on one number line and comparing the spreads.
The last section gave us a set of tools for saying where a single value sits inside a data set: the median, the quartiles, the percentiles. This section takes those same numbers and turns them into a picture.
A box plot — you will also see it called a box-and-whisker plot, or a box-whisker plot — is one of the most efficient graphs in statistics. It is built from only five numbers, it fits on a single line, and it answers two questions at a glance: where is the data bunched up? and how far out do the extremes reach? Once you can read one, you can compare two or three data sets side by side in a couple of seconds, which is something no list of numbers will ever let you do.
2.4.1 What a Box Plot Shows
The five-number summary of a data set is the list of its minimum value, first quartile \(Q_1\), median, third quartile \(Q_3\), and maximum value, given in that order. Together these five values divide the ordered data into four groups, each holding approximately 25 percent of the observations.
Those five numbers are all a box plot needs. Nothing else about the data set — not the individual values, not how many observations there are — goes into the drawing.
Definition 2.4.1 — A box plot needs only five numbers out of the whole ordered list, and one of them — the median — need not be a value in the list at all.
A box plot (also called a box-and-whisker plot) is a graph of a data set's five-number summary drawn against a scaled number line. A rectangular box spans from \(Q_1\) to \(Q_3\), so approximately the middle 50 percent of the data falls inside it; the median is marked inside the box by a line; and two line segments called whiskers extend from the ends of the box out to the smallest and largest data values.
Definition 2.4.2 — A box plot is the five-number summary drawn on a scaled number line: box from Q1 to Q3, median inside, whiskers to the extremes.
A whisker is one of the two line segments in a box plot that runs from an end of the box out to an extreme value — from \(Q_1\) down to the minimum, and from \(Q_3\) up to the maximum. The whiskers show how far the extreme values reach away from the middle half of the data.
Picture the data as people queued along a number line. The box fences in the middle half — the crowd. The whiskers are the two people who wandered farthest in each direction. A short box with long whiskers says "most of them are packed together, but a few are way out there."
To construct a box plot, use a horizontal or vertical number line and a rectangular box. The smallest and largest data values label the endpoints of the axis. The first quartile marks one end of the box and the third quartile marks the other end. The whiskers then extend from the ends of the box out to those smallest and largest values. The median can sit anywhere between the first and third quartiles — and, as we will see shortly, it can even land exactly on one of them, or on both.
Here is the idea on the fourteen-value data set from the previous section.
1; 1; 2; 2; 4; 6; 6.8; 7.2; 8; 8.3; 9; 10; 10; 11.5
For this data set the first quartile is two, the median is seven, and the third quartile is nine. The smallest value is one and the largest is 11.5, so the five-number summary is \(1,\ 2,\ 7,\ 9,\ 11.5\). Drawn against a number line, it looks like this.
Figure 2.4.1 — Box plot of the fourteen-value data set, with the box running from 2 to 9, the median at 7, and whiskers out to 1 and 11.5.
The two whiskers extend from the first quartile down to the smallest value and from the third quartile up to the largest value. The median is shown with a dashed line inside the box.
A box plot only means something if the axis underneath it is drawn to scale — equal distances standing for equal amounts. Sketch the box first and pencil in the numbers afterward and you get a picture that looks like data but reports the wrong spreads entirely.
One more convention is worth knowing before you meet box plots in the wild. Some software, and some textbooks, draw box plots that mark unusually extreme values as individual dots outside the whiskers. In those versions the whiskers stop short of the true minimum and maximum — they run only as far as the most extreme value that is not flagged as an outlier, and everything beyond gets its own dot. Both conventions are common, and both are correct as long as the graph tells you which one it is using. In this section every whisker runs all the way out to the true smallest and largest values.
A plot whose whiskers stop early and whose extremes appear as separate dots is using the outlier convention from the last section, where a value more than \(1.5 \times IQR\) beyond a quartile gets flagged. Same five numbers underneath — different drawing rule for the ends.
Every box plot starts from the same five numbers, and those five numbers have a name of their own.
Mai Xiong runs a peer tutoring group and keeps a record of how their students do. The following fifteen quiz scores from their group have already been ordered from smallest to largest.
4; 5; 5; 6; 7; 7; 8; 8; 8; 9; 9; 10; 10; 10; 10
a) Write down the five-number summary.
b) Describe in words where the box and the two whiskers would sit if you drew the box plot.
Solution
Step 1 — Read off the extremes. The list is already ordered, so the minimum is the first value and the maximum is the last: minimum \(= 4\), maximum \(= 10\).
Step 2 — Find the median. There are 15 values, an odd count, so the median is the single middle value — the eighth one. Counting in: 4, 5, 5, 6, 7, 7, 8, 8. The median is 8.
Step 3 — Find the quartiles. \(Q_1\) is the median of the seven values below the median: 4, 5, 5, 6, 7, 7, 8, so \(Q_1 = 6\). \(Q_3\) is the median of the seven values above it: 8, 9, 9, 10, 10, 10, 10, so \(Q_3 = 10\).
Answer (a): The five-number summary is \(4,\ 6,\ 8,\ 10,\ 10\).
Answer (b): The box runs from 6 to 10 with the median line at 8, so the box sits well to the right of the axis. The left whisker reaches down from 6 to 4. There is no visible right whisker at all: \(Q_3\) and the maximum are both 10, so the right end of the box is the maximum. That is the situation the next-but-one subsection is about.
2.4.2 Constructing a Box Plot from the Five-Number Summary
When someone hands you the five numbers, drawing the plot is mechanical: scale the axis so it covers the minimum through the maximum, mark the two quartiles and box them in, mark the median inside the box, and run a whisker out to each extreme. The interesting work is in what you can then say about the data.
The following data are the number of pages in 40 books on a shelf. Construct a box plot and state the interquartile range.
136; 140; 178; 190; 205; 215; 217; 218; 232; 234; 240; 255; 270; 275; 290; 301; 303; 315; 317; 318; 326; 333; 343; 349; 360; 369; 377; 388; 391; 392; 398; 400; 402; 405; 408; 422; 429; 450; 475; 512
Solution
Step 1 — The data is already ordered, so read the extremes straight off the ends: minimum \(= 136\), maximum \(= 512\).
Step 2 — Median. There are 40 values, an even count, so the median is the average of the 20th and 21st values, which are 318 and 326.
$$ \text{Median} = \frac{318 + 326}{2} = \frac{644}{2} = 322 $$Step 3 — First quartile. \(Q_1\) is the median of the lower 20 values, so it is the average of the 10th and 11th values, 234 and 240.
$$ Q_1 = \frac{234 + 240}{2} = \frac{474}{2} = 237 $$Step 4 — Third quartile. \(Q_3\) is the median of the upper 20 values (positions 21 through 40), so it is the average of the 30th and 31st values, 392 and 398.
$$ Q_3 = \frac{392 + 398}{2} = \frac{790}{2} = 395 $$Step 5 — Interquartile range.
$$ IQR = Q_3 - Q_1 = 395 - 237 = 158 $$Step 6 — The plot. Scale a number line from about 100 to about 550. Box from 237 to 395, median line at 322, left whisker out to 136 and right whisker out to 512.
Answer: The five-number summary is \(136,\ 237,\ 322,\ 395,\ 512\) and \(IQR = 158\) pages. The median line sits left of centre inside the box, so the third quarter of the books is more spread out than the second.
Valeria Ocampo measured the heights, in inches, of the 40 students in her statistics class. Her data are shown below.
59; 60; 61; 62; 62; 63; 63; 64; 64; 64; 65; 65; 65; 65; 65; 65; 65; 65; 65; 66; 66; 67; 67; 68; 68; 69; 70; 70; 70; 70; 70; 71; 71; 72; 72; 73; 74; 74; 75; 77
Construct a box plot with the following properties, then answer the questions below.
- Minimum value \(= 59\)
- \(Q_1\): First quartile \(= 64.5\)
- \(Q_2\): Second quartile, or median \(= 66\)
- \(Q_3\): Third quartile \(= 70\)
- Maximum value \(= 77\)
a. What percentage of the data does each quarter hold?
b. Find the spread of each of the four quarters. Which quarter is the most tightly packed, and which is the most spread out?
c. Find the range of the data.
d. Find the interquartile range.
e. Which holds more data — the interval from 59 to 65, or the interval from 66 to 70?
f. What is the range of the middle 50 percent of the data?
Solution
Figure 2.4.2 — Box plot of the heights of 40 statistics students, with the spread of each of the four quarters measured out.
a. Each quarter has approximately 25 percent of the data. That is what the quartiles are for — they are the three cut points that split the ordered data into four equal-sized groups.
b — Spread of each quarter. Subtract the two numbers bounding each quarter:
$$ 64.5 - 59 = 5.5 \quad\text{(first quarter)} $$ $$ 66 - 64.5 = 1.5 \quad\text{(second quarter)} $$ $$ 70 - 66 = 4 \quad\text{(third quarter)} $$ $$ 77 - 70 = 7 \quad\text{(fourth quarter)} $$So the second quarter has the smallest spread and the fourth quarter has the largest. Notice what that means physically: the ten students between 64.5 and 66 inches are packed into an inch and a half of the axis, while the ten students above 70 inches are strung out over seven inches.
c — Range.
$$ \text{Range} = \text{maximum} - \text{minimum} = 77 - 59 = 18 $$d — Interquartile range.
$$ IQR = Q_3 - Q_1 = 70 - 64.5 = 5.5 $$e. The interval from 59 to 65 stretches past \(Q_1 = 64.5\), so it contains more than 25 percent of the data. The interval from 66 to 70 runs from the median to \(Q_3\), which is exactly the third quarter — 25 percent. So 59 to 65 holds more.
f. The middle 50 percent of the data is exactly the box, and the box is the interquartile range: a range of 5.5 inches.
Answer: Each quarter holds about 25%; spreads are 5.5, 1.5, 4, and 7 inches; range \(= 18\); \(IQR = 5.5\); the interval 59–65 holds more data than 66–70; and the middle half spans 5.5 inches.
Try it in rāSHio
Open rāSHio and paste the forty heights into File → Delimited List… — the semicolons parse as-is. Choose Stats → Summary Statistics and the five numbers this example hands you come back measured rather than given: smallest value 59, first quartile 64.5, median 66, third quartile 70, largest value 77. Then choose Graph → Box Plot to draw them. What appears is the same picture as above — box from 64.5 to 70, median line at 66, whiskers out to 59 and 77 — so the sketch you made by hand can be checked against the one the tool builds from the same forty numbers.
Figure 2.4.3 — Drawing a box plot in rāSHio: Graph → Box Plot, which builds the box, the median line and both whiskers from a column of data. The walkthrough narrates a different data set, and rāSHio applies the outlier convention from the previous section, so one extreme value in the clip's demo data is drawn as its own dot rather than reached by a whisker. The forty heights have no value beyond their fences, so running the same steps on them gives whiskers out to 59 and 77, as described above.
2.4.3 When Some of the Five Numbers Coincide
For some data sets, two or more of the five numbers turn out to be equal. You might have a data set in which the median and the third quartile are the same. In that case the drawing has no dashed line inside the box — the right side of the box is the median as well as the third quartile.
Nothing is broken when this happens. A box plot with a missing piece is telling you something specific: a stack of identical values is sitting at that spot, big enough to swallow one of the cut points.
Suppose the smallest value and the first quartile are both one, the median and the third quartile are both five, and the largest value is seven. The box plot looks like this.
Figure 2.4.4 — Box plot in which the minimum and the first quartile are both 1 and the median and the third quartile are both 5, so the left whisker and the median line both disappear.
Read the picture back into words and it says a great deal. At least 25 percent of the values are equal to one — that is why there is no left whisker; the box starts at the smallest value. Twenty-five percent of the values are between one and five, inclusive. At least 25 percent of the values are equal to five, which is why there is no dashed median line inside the box. And the top 25 percent of the values fall between five and seven, inclusive.
When the box ends flush against the minimum, the graph is telling you that a quarter of the data is piled on that one value. The absence of a line is the finding.
Nathan Whitfield finishes a lab and reports the five-number summary of his data set as \(3,\ 3,\ 6,\ 9,\ 15\).
a) Sketch the box plot in words: where does the box sit, and which of the usual pieces are missing?
b) What percentage of the data is equal to 3? What percentage falls between 9 and 15, inclusive?
Solution
Step 1 — Match each number to its role. In order, the five-number summary is minimum \(= 3\), \(Q_1 = 3\), median \(= 6\), \(Q_3 = 9\), maximum \(= 15\).
Step 2 — Spot the coincidence. The minimum and \(Q_1\) are both 3, so the left end of the box lands exactly on the smallest value.
Answer (a): The box runs from 3 to 9 with the median line drawn at 6. There is no left whisker — the box begins at the minimum. The right whisker runs from 9 out to 15, and it is long: six units of whisker against six units of box.
Answer (b): Because \(Q_1\) and the minimum are the same value, at least 25 percent of the data is equal to 3. The interval from \(Q_3\) to the maximum is the fourth quarter, so about 25 percent of the data falls between 9 and 15 inclusive — spread thinly, since that quarter covers as much of the axis as the entire box does.
2.4.4 Comparing Two Data Sets on One Number Line
Box plots earn their keep when you draw more than one of them against the same number line. Two lists of numbers side by side tell you almost nothing; two boxes stacked on a shared axis tell you immediately which group is higher, which is more spread out, and which has the longer reach.
The following data set shows the heights, in inches, for the boys in a class of 40 students.
66; 66; 67; 67; 68; 68; 68; 68; 68; 69; 69; 69; 70; 71; 72; 72; 72; 73; 73; 74
Construct a box plot for this data set and state the spread of the middle 50 percent of the data.
Solution
Step 1 — Count and read the extremes. The 20 values are already in order, so minimum \(= 66\) and maximum \(= 74\).
Step 2 — Median. With 20 values the median is the average of the 10th and 11th, both of which are 69.
$$ M = \frac{69 + 69}{2} = 69 $$Step 3 — First quartile. \(Q_1\) is the median of the lower ten values, the average of the 5th and 6th, both 68.
$$ Q_1 = \frac{68 + 68}{2} = 68 $$Step 4 — Third quartile. \(Q_3\) is the median of the upper ten values (positions 11 through 20), the average of the 15th and 16th, both 72.
$$ Q_3 = \frac{72 + 72}{2} = 72 $$Step 5 — Spread of the middle half.
$$ IQR = Q_3 - Q_1 = 72 - 68 = 4 $$Step 6 — The plot. Scale the axis from about 65 to 75. Box from 68 to 72 with the median line at 69, left whisker to 66, right whisker to 74.
Answer: The five-number summary is \(66,\ 68,\ 69,\ 72,\ 74\), and the middle 50 percent of the boys' heights spans just 4 inches. The median sits left of the centre of the box, so the boys in the third quarter are more spread out than those in the second.
Test scores for a college statistics class held during the day are:
99; 56; 78; 55.5; 32; 90; 80; 81; 56; 59; 45; 77; 84.5; 84; 70; 72; 68; 32; 79; 90
Test scores for a college statistics class held during the evening are:
98; 78; 68; 83; 81; 89; 88; 76; 65; 45; 98; 90; 80; 84.5; 85; 79; 78; 98; 90; 79; 81; 25.5
a. Find the smallest and largest values, the median, and the first and third quartiles for the day class.
b. Find the smallest and largest values, the median, and the first and third quartiles for the evening class.
c. For each data set, what percentage of the data is between the smallest value and the first quartile? The first quartile and the median? The median and the third quartile? The third quartile and the largest value? What percentage of the data is between the first quartile and the largest value?
d. Create a box plot for each set of data. Use one number line for both box plots.
e. Which box plot has the widest spread for the middle 50 percent of the data — the data between the first and third quartiles? What does that mean for that set of data compared with the other?
Solution
a — Day class. Ordered, the 20 day scores are 32, 32, 45, 55.5, 56, 56, 59, 68, 70, 72, 77, 78, 79, 80, 81, 84, 84.5, 90, 90, 99. With 20 values the median is the average of the 10th and 11th, and each quartile is the median of a half.
$$ \text{Min} = 32, \quad Q_1 = \frac{56 + 56}{2} = 56, \quad M = \frac{72 + 77}{2} = 74.5, \quad Q_3 = \frac{81 + 84}{2} = 82.5, \quad \text{Max} = 99 $$b — Evening class. Ordered, the 22 evening scores are 25.5, 45, 65, 68, 76, 78, 78, 79, 79, 80, 81, 81, 83, 84.5, 85, 88, 89, 90, 90, 98, 98, 98. With 22 values the median is the average of the 11th and 12th, and each quartile is the middle value of an 11-value half.
$$ \text{Min} = 25.5, \quad Q_1 = 78, \quad M = \frac{81 + 81}{2} = 81, \quad Q_3 = 89, \quad \text{Max} = 98 $$c — Percentages. By definition the quartiles cut each data set into four quarters holding about 25 percent each, so roughly 25 percent falls in each of the four intervals, and the first quartile to the largest value covers three of those quarters — about 75 percent.
Counting the day class by hand gets close to those figures, with the boundary values inflating the first two counts a little because a score sitting exactly on a cut point gets counted on both sides of it. There are six day scores from 32 to 56 (30 percent), six from 56 to 74.5 (30 percent), five from 74.5 to 82.5 (25 percent), and five from 82.5 to 99 (25 percent); 16 of the 20 scores sit between the first quartile, 56, and the largest value, 99, which is 75 percent. The repeated 56 is the culprit for the two 30-percent counts — it is both the fifth and sixth score, so it lands on the cut point itself.
d — The two box plots on one axis.
Figure 2.4.5 — Day-class and evening-class test scores on one number line: the day class has the wider middle 50 percent, the evening class the longer left whisker.
e — Which middle 50 percent is wider? Compare the two interquartile ranges.
$$ IQR_{\text{day}} = 82.5 - 56 = 26.5 \qquad IQR_{\text{evening}} = 89 - 78 = 11 $$The day class has the wider spread for the middle 50 percent of its data. Its \(IQR\) is more than twice the evening class's, which means there is far more variability among the middle half of the day scores. The evening class's middle half is packed into eleven points — most of that class scored within a narrow band — but its long left whisker down to 25.5 shows that a couple of students trailed a long way behind.
Answer: Day class \(32,\ 56,\ 74.5,\ 82.5,\ 99\); evening class \(25.5,\ 78,\ 81,\ 89,\ 98\). The day class has the wider middle 50 percent, \(IQR = 26.5\) against \(IQR = 11\), so there is more variability in the middle half of the day scores.
Try it in rāSHio
Two data sets on one axis is the comparison box plots exist for, and it is where the tool saves more than arithmetic. Open rāSHio, paste the twenty day-class scores into File → Delimited List…, then paste the twenty-two evening-class scores into a second column. Run Stats → Summary Statistics to get both five-number summaries side by side, then choose Graph → Box Plot. The dialog opens in Single mode with one column dropdown, so switch it to Stacked — that is what lets you pick a second column, and it draws both plots against a single shared scale. That shared scale is the whole point: the evening class sitting higher through its middle half while reaching farther down at the bottom is visible at a glance, and it is not visible at all from two lists of numbers.
2.4.5 Box Plots for Strongly Skewed Data
Not every data set is bunched neatly in the middle. When a few values are enormous compared with the rest, the box plot shows it in a shape you learn to recognise instantly: a box shoved hard to one side of the axis with one very long whisker trailing away from it.
Camila Reyes wants to check that they can run the whole procedure unaided on a second batch of numbers. Work through the same five steps — order the data, find the median, then find each quartile, then draw the box and whiskers on a scaled number line — to graph a box-and-whisker plot for the data values they recorded.
0; 5; 5; 15; 30; 30; 45; 50; 50; 60; 75; 110; 140; 240; 330
Solution
Step 1 — Extremes. The 15 values are ordered, so minimum \(= 0\) and maximum \(= 330\).
Step 2 — Median. With 15 values, an odd count, the median is the 8th value: 0, 5, 5, 15, 30, 30, 45, 50. The median is 50.
Step 3 — First quartile. \(Q_1\) is the median of the seven values below the median (0, 5, 5, 15, 30, 30, 45), which is the 4th of them: \(Q_1 = 15\).
Step 4 — Third quartile. \(Q_3\) is the median of the seven values above the median (50, 60, 75, 110, 140, 240, 330), which is the 4th of them: \(Q_3 = 110\).
Step 5 — Draw it. Scale the axis from 0 to about 350. Box from 15 to 110 with the median line at 50; left whisker from 15 down to 0; right whisker from 110 out to 330.
Answer: The five-number summary is \(0,\ 15,\ 50,\ 110,\ 330\), with \(IQR = 95\). Like the other data set in this subsection, this one is right-skewed — the right whisker alone is longer than the entire box, so the largest values reach much farther from the middle than the smallest ones do.
Daniel Boyd and his husband, Reid, have been recording values for a class project. Graph a box-and-whisker plot for the data values they collected.
10; 10; 10; 15; 35; 75; 90; 95; 100; 175; 420; 490; 515; 515; 790
Solution
Step 1 — Find the five numbers. There are 15 values, already ordered.
- Minimum: 10
- \(Q_1\): the median of the lower seven values (10, 10, 10, 15, 35, 75, 90) is the 4th of them, 15
- Median: the 8th of the 15 values, 95
- \(Q_3\): the median of the upper seven values (100, 175, 420, 490, 515, 515, 790) is the 4th of them, 490
- Maximum: 790
Step 2 — Draw it. Scale a number line that reaches from 10 to 790, mark the box from 15 to 490, put the median line at 95, and run the whiskers out to 10 and 790.
Figure 2.4.6 — A strongly right-skewed data set: the median sits at 95, near the left end of a box that stretches from 15 to 490.
Step 3 — Read the shape back. The median line is jammed up against the left end of a very wide box, and the right whisker runs 300 units past \(Q_3\). Half of these values are 95 or less, yet the top quarter climbs all the way to 790. That lopsidedness is the signature of a right-skewed data set.
Answer: The five numbers used to create the box-and-whisker plot are min \(= 10\), \(Q_1 = 15\), median \(= 95\), \(Q_3 = 490\), max \(= 790\).
Problem Set 2.4
Problem 1. In a survey of 20-year-olds in China, Germany, and the United States, people were asked the number of foreign countries they had visited in their lifetime. The following box plots display the results.
Figure 2.4.7 — Box plots of the number of foreign countries visited by 20-year-olds in China, Germany, and the United States, drawn on a shared 0-to-11 number line.
a) In complete sentences, describe what the shape of each box plot implies about the distribution of the data collected.
b) Have more U.S. citizens or more Germans surveyed been to over eight foreign countries?
c) Compare the three box plots. What do they imply about the foreign travel of 20-year-old residents of the three countries when compared to each other?
Solution
Reading the three plots. Every value below is read off the shared number line in Figure 2.4.7.
- China — the row is a bare line from 0 to 5 with no box at all. A box of zero width means \(Q_1\), the median and \(Q_3\) are all the same value: 0. Five-number summary \(0,\ 0,\ 0,\ 0,\ 5\).
- Germany — box from 4 to 8 with the median drawn on the box's right edge, whiskers out to 0 and 11. Five-number summary \(0,\ 4,\ 8,\ 8,\ 11\).
- United States — box from 0 to 5 with the median at 2 and the right whisker out to 11. The box starts flush against the minimum, so there is no left whisker. Five-number summary \(0,\ 0,\ 2,\ 5,\ 11\).
a — What each shape implies.
The Chinese plot is collapsed at zero: at least 75 percent of the 20-year-olds surveyed had visited no foreign country at all, and the whole upper quarter stretches from 0 out to only 5. The distribution is piled hard against the left end of the scale.
The German plot sits highest of the three. Its box runs from 4 to 8 and the median lands on the box's right edge, which tells us at least a quarter of German respondents had visited exactly eight countries. The middle half is up in the 4-to-8 range — travel is common and fairly evenly spread.
The U.S. plot is spread wide but low. Its box runs from 0 to 5 with the median at 2, so half the American respondents had visited two or fewer countries; the long right whisker to 11 says a small group had travelled a great deal. This is a right-skewed distribution.
b — Over eight countries. For Germany, \(Q_3 = 8\), so about 25 percent of Germans surveyed had been to more than eight countries. For the United States, \(Q_3 = 5\) and the maximum is 11, so the entire top quarter spans 5 to 11 and only part of it lies beyond 8 — fewer than 25 percent.
Answer (b): More Germans surveyed had been to over eight foreign countries.
c — The three compared. German 20-year-olds had travelled the most: their box sits farthest right and three quarters of them had visited at least four countries. Americans reach just as high at the top end (both maxima are 11) but their middle half is far lower, so the American figures are dragged up by a well-travelled minority rather than being typical. Chinese respondents had travelled the least by a wide margin — most had not left the country at all, and the most-travelled respondent had visited five countries, fewer than the German or American median-to-max reach.
Problem 2. Given the following box plot, answer the questions.
Figure 2.4.8 — Box plot with the minimum and first quartile both at 0, the median at 20, the third quartile at about 95, and the maximum at 150.
a) Think of an example (in words) where the data might fit into the above box plot. In two to five sentences, write down the example.
b) What does it mean to have the first and second quartiles so close together, while the second to third quartiles are far apart?
Solution
Reading the plot. In Figure 2.4.8 the box runs from 0 to about 95, the median line sits at 20, and the right whisker reaches 150. There is no left whisker: the box starts flush against the minimum, so the smallest value and \(Q_1\) are both 0. The five-number summary is roughly \(0,\ 0,\ 20,\ 95,\ 150\).
a — An example that would fit. Suppose we ask 200 people how many minutes they spent commuting by bus yesterday. A large share did not take a bus at all, so their answer is 0 — that is why at least a quarter of the data sits at zero and there is no left whisker. Among the people who did ride, most had short trips of ten to thirty minutes, which pulls the median down to 20. A smaller group commutes across the county and reports an hour, ninety minutes, or in one case two and a half hours, which is what stretches the box out to 95 and the whisker all the way to 150. Any variable with a big pile-up at zero and a long thin upper tail — rainfall by day, dollars donated per person, minutes of overtime worked — would produce this same shape.
b — Quartiles close together, then far apart. The gap between two of the five numbers tells you how much of the number line one quarter of the data is spread across, not how many values it holds. Every quarter holds about 25 percent no matter how wide it looks.
So \(Q_1 = 0\) and the median \(= 20\) sitting close together means that a full quarter of the values are squeezed into the narrow band from 0 to 20 — those observations are tightly bunched and nearly identical to one another. The median at 20 and \(Q_3\) at 95 being far apart means the next quarter of the values — the same number of observations — is smeared across 75 units of the scale.
Answer (b): The data below the median is densely packed near zero, while the data just above the median is thinly spread over a wide range. Equal counts, wildly unequal spread — the signature of a right-skewed distribution.
Problem 3. Given the following box plots, answer the questions.
Figure 2.4.9 — Box plot labelled Data 1: box from 2 to about 4.7 with the median at 4, whiskers out to 0 and 7.
Figure 2.4.10 — Box plot labelled Data 2: a narrow box centred near 2 with the median at 2, whiskers out to 0 and 7.
a) In complete sentences, explain why each statement is false.
i. Data 1 has more data values above two than Data 2 has above two.
ii. The data sets cannot have the same mode.
iii. For Data 1, there are more data values below four than there are above four.
b) For which group, Data 1 or Data 2, is the value of "7" more likely to be an outlier? Explain why in complete sentences.
Solution
Reading the two plots. From Figures 2.4.8 and 2.4.9:
- Data 1 — box from 2 to about 4.7 with the median at 4, whiskers out to 0 and 7. Five-number summary about \(0,\ 2,\ 4,\ 4.7,\ 7\).
- Data 2 — a narrow box from about 1.5 to about 2.4 with the median at 2, whiskers out to 0 and 7. Five-number summary about \(0,\ 1.5,\ 2,\ 2.4,\ 7\).
a i — "Data 1 has more data values above two than Data 2 has above two."
False, because a box plot reports percentages, not counts. It tells us that 75 percent of Data 1 lies above 2 (since \(Q_1 = 2\)) and that 50 percent of Data 2 lies above 2 (since the median is 2) — but neither plot says how many observations are in its data set. If Data 1 holds 20 values and Data 2 holds 200, then 75 percent of Data 1 is 15 values while 50 percent of Data 2 is 100. The graph gives us no way to rule that out, so the comparison of counts cannot be made at all.
a ii — "The data sets cannot have the same mode."
False, because a box plot does not show the mode. The mode is the most frequently occurring value, and none of the five numbers a box plot is built from carries frequency information. Both data sets could easily have a mode of 2 — Data 2's median is 2, and Data 1's box covers 2 — and the two pictures would look exactly as they do.
a iii — "For Data 1, there are more data values below four than there are above four."
False, because 4 is the median of Data 1. The median is the value that splits the ordered data into halves: about 50 percent below and about 50 percent above. They are equal, not unbalanced.
b — For which group is 7 more likely to be an outlier?
Use the \(1.5 \times IQR\) rule from the previous section on each set.
$$ IQR_{\text{Data 1}} \approx 4.7 - 2 = 2.7 \qquad Q_3 + 1.5 \times IQR \approx 4.7 + 4.05 = 8.75 $$ $$ IQR_{\text{Data 2}} \approx 2.4 - 1.5 = 0.9 \qquad Q_3 + 1.5 \times IQR \approx 2.4 + 1.35 = 3.75 $$For Data 1 the upper fence lands at about 8.75, so a value of 7 is comfortably inside it and would not be flagged. For Data 2 the fence lands at about 3.75, and 7 is nearly twice that.
Answer (b): Data 2. Its middle half is packed into a band less than one unit wide, so the fences sit close in and a value of 7 is far outside them. In Data 1 the same value of 7 is just the top of a genuinely wide distribution. Note that this is exactly why the plots are drawn with whiskers reaching 7 in both cases — the same number can be ordinary in one data set and extraordinary in another, and it is the \(IQR\) that decides which.
Problem 4. A survey was conducted of 130 purchasers of new BMW 3 series cars, 130 purchasers of new BMW 5 series cars, and 130 purchasers of new BMW 7 series cars. In it, people were asked the age they were when they purchased their car. The following box plots display the results.
Figure 2.4.11 — Box plots of purchaser age for the BMW 3, 5, and 7 series, drawn on a shared number line running from 25 to 80.
a) In complete sentences, describe what the shape of each box plot implies about the distribution of the data collected for that car series.
b) Which group is most likely to have an outlier? Explain how you determined that.
c) Compare the three box plots. What do they imply about the age of purchasing a BMW from the series when compared to each other?
d) Look at the BMW 5 series. Which quarter has the smallest spread of data? What is the spread?
e) Look at the BMW 5 series. Which quarter has the largest spread of data? What is the spread?
f) Look at the BMW 5 series. Estimate the interquartile range (\(IQR\)).
g) Look at the BMW 5 series. Are there more data in the interval 31 to 38 or in the interval 45 to 55? How do you know this?
h) Look at the BMW 5 series. Which interval has the fewest data in it? How do you know this?
i. 31–35
ii. 38–41
iii. 41–64
Solution
Reading the three plots. All values below are estimates read off the shared number line in Figure 2.4.11.
- BMW 3 series — about \(23,\ 29,\ 33,\ 41,\ 66\)
- BMW 5 series — about \(31,\ 41,\ 42,\ 55,\ 64\)
- BMW 7 series — about \(36,\ 41,\ 46,\ 59,\ 67\)
a — What each shape implies. The 3 series distribution is the youngest and the most right-skewed: its box sits between 29 and 41, its median is 33, and yet the right whisker runs all the way to 66. Most 3 series buyers were in their late twenties and thirties, with a thin tail of much older buyers. The 5 series is shifted up the scale, and its median sits almost on top of \(Q_1\) — a dense cluster of buyers right around 41 or 42 with a wide, thinly populated upper half. The 7 series is the oldest and the most symmetric of the three: the median sits near the middle of a box running from 41 to 59, with whiskers of similar length on both sides.
b — Most likely to have an outlier. Apply the \(1.5 \times IQR\) fence to each.
$$ \text{3 series: } IQR \approx 41 - 29 = 12, \quad Q_3 + 1.5 \times IQR \approx 41 + 18 = 59 $$ $$ \text{5 series: } IQR \approx 55 - 41 = 14, \quad Q_3 + 1.5 \times IQR \approx 55 + 21 = 76 $$ $$ \text{7 series: } IQR \approx 59 - 41 = 18, \quad Q_3 + 1.5 \times IQR \approx 59 + 27 = 86 $$Only the 3 series has a maximum (about 66) beyond its own upper fence (about 59).
Answer (b): The BMW 3 series. Its middle half is narrow, which pulls the fences in close, while its maximum reaches 66 — the combination is what makes a value extreme, not the raw size of the number.
c — The three compared. Buyer age climbs steadily across the series: median about 33 for the 3 series, about 42 for the 5 series, about 46 for the 7 series. The spread of the middle half climbs too, from an \(IQR\) of about 12 years up to about 18. Read together, the plots say the cheaper 3 series draws a younger and more tightly clustered group of buyers, while the more expensive 7 series draws older buyers across a broader age range — and all three overlap heavily in the forties, so knowing a buyer's age would not let you guess their series with any confidence.
d — 5 series, smallest spread. The four quarters of the 5 series span about \(41 - 31 = 10\), \(42 - 41 = 1\), \(55 - 42 = 13\), and \(64 - 55 = 9\) years.
Answer (d): The second quarter, spanning roughly 1 year. A quarter of all 5 series buyers were essentially the same age, right around 41 or 42.
Answer (e): The third quarter, spanning roughly 13 years (about 42 to 55).
f — Interquartile range.
$$ IQR \approx Q_3 - Q_1 = 55 - 41 = 14 \text{ years} $$g — 31 to 38 or 45 to 55? Both intervals sit inside a single quarter, and each quarter holds about 25 percent of the buyers, so compare how much of its quarter each interval covers. The interval 45 to 55 covers 10 of the 13 years of the third quarter — most of it. The interval 31 to 38 covers 7 of the 10 years of the first quarter, but it stops 3 years short of \(Q_1\) at both the busy end and is spread over a slightly wider share of empty scale.
Answer (g): More data lies in the interval 45 to 55, because it takes in nearly the whole third quarter, while 31 to 38 leaves out the part of the first quarter nearest \(Q_1\) where those buyers are most densely packed.
h — Fewest data. Option iii, 41 to 64, runs from \(Q_1\) all the way to essentially the maximum, so it holds about 75 percent of the buyers — the most, not the fewest. Options i and ii both sit inside the first quarter, which holds about 25 percent in total, spread across the 10 years from 31 to 41. Option i, 31 to 35, takes in 4 of those 10 years; option ii, 38 to 41, takes in 3.
Answer (h): ii, the interval 38 to 41 holds the fewest — it is the narrowest slice of the single quarter that both candidate intervals fall inside. All of these are estimates read off a graph; with the underlying ages we could count exactly.
Problem 5. Grace Whitfield surveyed 25 randomly selected students, asking each of them the number of movies they watched the previous week. Her results are as follows.
| # of movies | Frequency |
|---|---|
| 0 | 5 |
| 1 | 9 |
| 2 | 6 |
| 3 | 4 |
| 4 | 1 |
Construct a box plot of the data.
Solution
Step 1 — Turn the frequency table back into an ordered list. There are \(5 + 9 + 6 + 4 + 1 = 25\) students. Listing every response in order gives five 0s, then nine 1s, then six 2s, then four 3s, then one 4 — so positions 1–5 hold 0, positions 6–14 hold 1, positions 15–20 hold 2, positions 21–24 hold 3, and position 25 holds 4.
Step 2 — Extremes. Minimum \(= 0\), maximum \(= 4\).
Step 3 — Median. With 25 values, an odd count, the median is the 13th value. Position 13 falls inside the run of 1s, so the median is 1.
Step 4 — First quartile. \(Q_1\) is the median of the 12 values below the median (positions 1 through 12), so it is the average of the 6th and 7th values. Both are 1.
$$ Q_1 = \frac{1 + 1}{2} = 1 $$Step 5 — Third quartile. \(Q_3\) is the median of the 12 values above the median (positions 14 through 25), so it is the average of the 19th and 20th values. Both are 2.
$$ Q_3 = \frac{2 + 2}{2} = 2 $$Step 6 — Draw it. Scale a number line from 0 to 4. The box runs from 1 to 2 — just one unit wide. The median is 1, which is also \(Q_1\), so the median line lands exactly on the left edge of the box and no dashed line appears inside it. The left whisker runs from 1 down to 0 and the right whisker from 2 out to 4.
Answer: The five-number summary is \(0,\ 1,\ 1,\ 2,\ 4\) with \(IQR = 1\). This is the coinciding-values case from earlier in the section: because \(Q_1\) and the median are both 1, at least a quarter of the students watched exactly one movie, and the missing median line inside the box is what tells you so.
Key Terms
box plot — a graph of a data set's five-number summary drawn on a scaled number line: a box from \(Q_1\) to \(Q_3\) with the median marked inside it, plus a whisker out to each extreme value.
five-number summary — the minimum, first quartile, median, third quartile, and maximum of a data set, listed in that order.
whisker — a line segment in a box plot running from an end of the box out to an extreme value, showing how far the extremes reach from the middle half of the data.