Introduction to Statistics · Chapter 1 · Sampling and Data

Frequency, Frequency Tables, and Levels of Measurement

Counting is the easy part. The work is deciding how far to carry an answer, what arithmetic the data will even tolerate, and how to turn a raw list into a table someone can read.


bookSHelf  ·  Introduction to Statistics  ·  §1.3  ·  a self-paced section

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Learning objectives — by the end of this section you will be able to

Objectives

  1. Round a computed answer to the right number of decimal places for the data you started with §1.3.1
  2. Classify data as nominal, ordinal, interval, or ratio — and say which arithmetic each level allows Definitions 1.3.1–1.3.5
  3. Build a frequency table with relative and cumulative relative frequency columns from a raw list Definitions 1.3.6–1.3.8
  4. Read counts and percentages off a frequency table, including grouped-interval tables Tables 1.3.4–1.3.5
  5. Spot errors in a published table and explain how the mistake changes what it appears to say Example 1.3.4
1.3

§1.3.1 — before any of the counting starts

Organizing data means dividing, and division produces long decimals. So the section opens with the one rule that decides how far to carry an answer — and it does not ask you to judge how precise a number "looks."

Carry your final answer one more decimal place than was present in the original data.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

§1.3.1 — three quiz scores, one rounding decision

Round once, at the end

  1. The data are 4, 6, 9 — whole numbers, so zero decimal places.
  2. Zero places in, so the answer carries one.
  3. Add and divide; the calculator shows 6.3333333. Do not touch it.
  4. Only now trim to one place: 6.3.

Had the scores been recorded as 4.0, 6.0, 9.0 — one decimal place — the same calculation would be reported to two: 6.33. Also: most fractions in this course need not be reduced. An unreduced fraction usually shows where the numbers came from.

The rule — data precision, plus one

4+6+93=193=6.3333333 \frac{4 + 6 + 9}{3} = \frac{19}{3} = 6.3333333\ldots

Reported as 6.3. Round off only the final answer; if an intermediate result must be rounded, carry it to at least twice as many places as the final answer.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Insight Note — why the rule says "only the final answer"

Rounding early is like trimming a board before you measure twice.

Every time you round in the middle of a calculation you throw away a sliver of the answer, and those slivers pile up. Keep the full decimal on your calculator until the very last step, then trim once.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Your turn — commit to an answer before the reveal

Try It Now 1.3.1

Try It Now 1.3.1 — five logged study times

Five students in a study group recorded how many minutes they spent on homework last night. Alex Delgado logged 42 minutes; Hannah Wolcott logged 55; Marcus Bell logged 38; Kayla Nguyen logged 61; Ethan Shaw logged 47. Find the group's mean study time and round it off correctly.

48.6 minutes.

42+55+38+61+47=243 42 + 55 + 38 + 61 + 47 = 243 , and 2435=48.6 \frac{243}{5} = 48.6 . The data are whole numbers, so the answer carries one decimal place. No rounding was actually needed here — the division came out exactly; had it come out 48.62, we would have reported 48.6.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Try it in rāSHio

Read the mean off directly

Paste the five logged times (42; 55; 38; 61; 47) into File → Delimited List…, then choose Stats → Summary Statistics. The mean comes back as 48.6 minutes — no adding and dividing by hand.

Worth reopening that panel after the next subsection: a mean is only meaningful once the data reach the interval or ratio level — which is exactly the distinction Levels of Measurement draws.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Figure 1.3.1 — the tool, in one pass

Stats → Summary Statistics

Load a column of values, open the Summary Statistics panel, and the mean arrives alongside the median, the standard deviation, the quartiles, and the range.

The rounding rule still applies to whatever you copy out of that panel — the tool will happily hand you seven decimal places.

Figure 1.3.1: reading the mean directly in rāSHio: Stats → Summary Statistics.

1.3

§1.3.2 — the question you answer before you calculate

The way a set of data is measured is called its level of measurement, and not every statistical operation can be used with every set of data. There are four levels, ordered from least mathematical structure to most: nominal, ordinal, interval, ratio.

Each level adds one power to the level below it.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

§1.3.2 — the classification itself

Level of measurement

Definition 1.3.1 — Level of measurement

The level of measurement of a data set is the classification — nominal, ordinal, interval, or ratio — that describes how much mathematical structure the measurements carry, and therefore which calculations are meaningful on them.

Definition 1.3.1 — Measurement-scale ladder Four rounded boxes rise in a staircase from lower-left to upper-right, each box exactly one step higher than the box before it. Every box holds three stacked lines: the level name in bold ink at top, the power that level adds over the one before it in rust accent beneath, and a tiny real-world example in dim grey at the bottom. Left to right: Nominal (named categories; jersey colors), Ordinal (+ order; loyalty tiers), Interval (+ even spacing; temperature degrees F), Ratio (+ true zero; number of machines). The rising boxes are the visual argument -- each scale keeps every property of the one below it and adds exactly one more. Nominal named categories jersey colors Ordinal + order loyalty tiers Interval + even spacing temperature °F Ratio + true zero number of machines

Figure: the four levels as a ladder — each step adds one power to the step below it.

Correct statistical procedures depend on knowing which rung you are standing on. Get it wrong and the arithmetic still runs — it just stops meaning anything.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Level 1 — labels only

Nominal scale level

Definition 1.3.2 — Nominal scale level

Data measured on a nominal scale are qualitative: categories, colors, names, labels, favorite foods, yes-or-no responses. Nominal data are not ordered and cannot be used in calculations.

Putting pizza first and sushi second is not meaningful — that is just the order you happened to write them down in. Smartphone brands are the same: no agreed-upon order, however strong anyone's preferences.

Definition 1.3.2 — Nominal scale level Four rounded chips labelled pizza, sushi, tacos and salad sit in a row of four fixed slots. Bold rank tags 1st, 2nd, 3rd and 4th fade in above the slots, as if ranking the chips. The pizza and tacos chips then swap slots along a curved path -- pizza dips below the row, tacos arcs up near the tags -- while the tags stay fixed in place. Next the sushi and salad chips swap the same way. Finally the rank tags fade back out, leaving the chips in their new order: tacos, salad, pizza, sushi. The ranking never meant anything -- nominal categories have no inherent order, so any arrangement is equally valid. 1st 2nd 3rd 4th pizza sushi tacos salad

Figure: nominal data are labels — any arrangement is equally valid.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Context Pause — why this classification earns a subsection

The level decides what math is legal

You can compute an average shoe size. You cannot compute an average favorite pizza topping.

Level of measurement is the rule that tells you which of those two situations you are in — before you start calculating, not after the answer looks strange.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Level 2 — order, but no measurable gaps

Ordinal scale level

Definition 1.3.3 — Ordinal scale level

Data measured on an ordinal scale are categorical like nominal data, with one difference: they can be ordered. Differences between two ordinal values cannot be measured, and ordinal data cannot be used in calculations.

The top five national parks can be ranked 1 to 5, but the distance from first to second is unknowable. A cruise survey reading "excellent, good, satisfactory, unsatisfactory" is ordered — and there is no reason the gap from excellent to good equals the gap from good to satisfactory.

Definition 1.3.3 — Ordinal scale level An order arrow runs left to right beneath four fixed-order chips reading unsatisfactory, satisfactory, good, excellent. A double-headed accent arrow and a question mark sit in each of the three gaps between neighbouring chips. The three gaps breathe in and out independently and out of phase, so their widths are never equal at the same moment and never settle -- only the order of the chips is fixed; the size of each gap is unknowable. ? ? ? unsatisfactory satisfactory good excellent

Figure: ordinal data can be ranked — the size of each gap is unknowable.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Level 3 — differences work, ratios do not

Interval scale level

Definition 1.3.4 — Interval scale level

Data measured on an interval scale have a definite ordering and are numerical, so differences can be calculated. But an interval scale has no true zero — its zero point is a convention, not an absence — so ratios are not meaningful.

Celsius and Fahrenheit: 40=10060 40^\circ = 100^\circ - 60^\circ , so differences make sense. Yet 0 is not a minimum — −10 °F and −15 °C exist, and they are colder than zero.

Definition 1.3.4 — Interval scale level A bare number line running -20 to 100, labelled in degrees Fahrenheit, with a tick every 20 degrees. A curved brace under the label "40deg" spans 20 to 60, then slides right to span 60 to 100 with the same width and the same label, showing that a 40-degree difference is a 40-degree difference no matter where it falls on the scale -- because the scale's zero is a convention, not a true minimum. -20 0 20 40 60 80 100 °F 40°

Figure: subtraction works; ratios do not, because zero is a convention.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Insight Note — the whole difference between interval and ratio

A true zero means "none of it."

Zero dollars means you have no money. Zero degrees Celsius means the water is freezing — it does not mean there is no temperature.

That is why "80° is twice as hot as 40°" is a sentence that sounds fine and means nothing.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Level 4 — the most information a measurement can carry

Ratio scale level

Definition 1.3.5 — Ratio scale level

Ratio scale data are like interval data — ordered, numerical, with meaningful differences — but there is a true minimum value of zero, so ratios between values are meaningful.

Four machine-graded exam scores out of 100: 80, 68, 20, 92. They order (20, 68, 80, 92); they subtract (92 beats 68 by 24 points); and because the minimum possible score is 0, the score of 80 really is four times the score of 20.

Definition 1.3.5 — Ratio scale level A vertical Exam-score axis, ticked every 20 from 0 to 100, with a heavy baseline at 0 tagged "true 0". Four bars for scores 20, 68, 80 and 92 rise from that baseline -- 20 and 80 in the accent (rust) tone, 68 and 92 in the curve (blue) tone -- each labelled with its score at the bar top. Dashed guide lines extend right from the tops of the 20 and 80 bars to a curly brace, labelled "80 = 4 x 20" in bold accent text: because the scale has a true zero, 80 really is four times 20. 0 20 40 60 80 100 Exam score true 0 20 68 80 92 80 = 4 × 20

Figure: anchored at a true zero, so 80 really is four times 20.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Definitions 1.3.1–1.3.5, on one page

What each level lets you do

LevelOrdered?Differences meaningful?True zero (ratios meaningful)?Example
NominalNoNoNoCrayon colors, smartphone brands
OrdinalYesNoNoSurvey ratings, park rankings
IntervalYesYesNoTemperature in °C or °F, calendar years
RatioYesYesYesExam scores, distance, income

Table: the four levels of measurement, lowest to highest.

Read it as a ladder, not a list: every column that says "Yes" stays "Yes" all the way down. Ratio data are the only kind that survive every operation in the table.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Your turn — name the level, and say why

Try It Now 1.3.2

Try It Now 1.3.2 — Rachel Whitfield's field-day records

  • a) The jersey numbers worn by the players on a soccer team.
  • b) The finishing places in a 5K race: 1st, 2nd, 3rd, …
  • c) The number of minutes each runner took to finish that same 5K.

a) nominal  ·  b) ordinal  ·  c) ratio

A jersey number is a label — player 10 is not "twice" player 5. Places are genuinely ordered, but the gap from 1st to 2nd is not a measurable amount. Finishing times are ordered, subtract sensibly, and zero minutes is a true zero — a 40-minute runner took twice as long as a 20-minute runner. Only part c lets Rachel do real arithmetic with her records.

1.3

§1.3.3 — from a raw list to a table you can read

A frequency table takes a jumble of responses and answers three questions at once: how many landed on each value, what share of the sample that is, and how much of the sample you have accounted for by the time you reach a given row.

Count. Divide by the total. Then keep a running total.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Column one — the count

Frequency

Definition 1.3.6 — Frequency

A frequency is the number of times a value of the data occurs.

Twenty students were asked how many hours they worked per day:
5; 6; 3; 3; 2; 4; 7; 5; 2; 3; 5; 6; 5; 4; 4; 3; 5; 2; 5; 3

Definition 1.3.6 -- Frequency is a count of how many responses land on each value Twenty response numerals, laid out in the order the definition lists them, sit above six bins for the values 2-7. One value at a time -- 2, then 3, then 4, then 5, then 6, then 7 -- its matching numerals turn accent orange and fly down into that bin, landing as a stack of small squares; the bin's total count fades in above the completed stack. The bins finish at heights 3, 5, 3, 6, 2, 1, matching Table 1.3.1. 2 3 4 5 6 7 5 6 3 3 2 4 7 5 2 3 5 6 5 4 4 3 5 2 5 3 3 5 3 6 2 1

Figure: each dot is one student; the stack height is the frequency.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

The same twenty responses, sorted and counted

Table 1.3.1 — student work hours

Data valueFrequency
23
35
43
56
62
71

Table 1.3.1: frequency table of student work hours.

Three students work two hours, five work three hours, and so on. The frequency column sums to 20 — the total number of students in the sample. If it does not sum to your sample size, you have miscounted.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Column two — the share

Relative frequency

Definition 1.3.7 — Relative frequency

A relative frequency is the ratio of the number of times a value occurs to the total number of outcomes. Divide each frequency by the sample total — here, 20. Write it as a fraction, a percent, or a decimal.

The relative frequency column of Table 1.3.2 sums to 2020 \frac{20}{20} , or 1.

ValueFreq.Relative freq.
233/20 or 0.15
355/20 or 0.25
433/20 or 0.15
566/20 or 0.30
622/20 or 0.10
711/20 or 0.05

Table 1.3.2: work hours with relative frequencies.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Column three — the running total

Cumulative relative frequency

Definition 1.3.8 — Cumulative relative frequency

Cumulative relative frequency is the accumulation of the previous relative frequencies. To find it, add all the previous relative frequencies to the relative frequency for the current row.

The last entry is 1 — one hundred percent of the data has been accumulated.

Definition 1.3.8 — Cumulative relative frequency An axis from 0.00 to 1.00 with a dashed line at 1.00 and x-axis values 2 through 7. Six steps build left to right: each riser rises by that value's relative frequency (0.15, 0.25, 0.15, 0.30, 0.10, 0.05) and its tread draws across the value's slot, with the running cumulative total printed above it. The totals are 0.15, 0.40, 0.55, 0.85, 0.95, 1.00 -- the last step reaches exactly one. 0.00 0.25 0.50 0.75 1.00 2 3 4 5 6 7 0.15 0.40 0.55 0.85 0.95 1.00

Figure: a running total — the last step must reach one.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

One table, three questions answered

Table 1.3.3 — the finished table

Data valueFrequencyRelative frequencyCumulative relative frequency
233/20 or 0.150.15
355/20 or 0.250.15 + 0.25 = 0.40
433/20 or 0.150.40 + 0.15 = 0.55
566/20 or 0.300.55 + 0.30 = 0.85
622/20 or 0.100.85 + 0.10 = 0.95
711/20 or 0.050.95 + 0.05 = 1.00

Table 1.3.3: work hours with relative and cumulative relative frequencies.

Every cumulative entry carries everything above it. That is what makes the column worth building: a single cell answers "how much of the sample is at this value or below?" without any adding.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Insight Note — the shape of the cumulative column

A running total, not a fresh count.

Think of the cumulative column as a bucket you keep pouring into. Each row adds its own relative frequency to everything already in the bucket, so the last row must come out full — one whole, or 100%.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Context Pause — when the last entry reads 0.99

Why the column may not land exactly on 1

Because of rounding, the relative frequency column may not always sum to one, and the last cumulative entry may not be exactly one.

Each should be close to one. If yours is far off, you have an arithmetic error — not a rounding artifact.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

100 measurements of a continuous quantity — grouped into intervals

Table 1.3.4 — soccer player heights

Heights (in.)Freq.Rel. freq.Cum. rel. freq.
59.95–61.9550.050.05
61.95–63.9530.030.08
63.95–65.95150.150.23
65.95–67.95400.400.63
67.95–69.95170.170.80
69.95–71.95120.120.92
71.95–73.9570.070.99
73.95–75.9510.011.00
Total1001.00

Table 1.3.4: heights of 100 semiprofessional soccer players.

Figure 1.3.2 — Frequency histogram of heights for 100 semiprofessional soccer players Eight bars of equal width, touching with no gaps, over the interval boundaries 59.95, 61.95, 63.95, 65.95, 67.95, 69.95, 71.95, 73.95, 75.95. Bar heights (frequency): 5, 3, 15, 40, 17, 12, 7, 1 -- summing to 100 and peaking in the 65.95 to 67.95 interval. Y axis: Frequency, 0 to 45 by 5. X axis: Heights (inches). 59.95 61.95 63.95 65.95 67.95 69.95 71.95 73.95 75.95 0 5 10 15 20 25 30 35 40 45 Heights (inches) Frequency

Figure 1.3.2: the same table as a histogram, in 2-inch intervals.

The boundaries end in .95 on purpose: every height falls between the endpoints of an interval and never at one, so no measurement can land in two intervals at once.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Context Pause — a table with a sequel

You will meet this table again

The soccer-height data returns in Descriptive Statistics, where the method used to compute the intervals is explained.

For now, take the intervals as given and focus on reading the counts off them. Where the interval boundaries came from is a later question.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Your turn — read a percentage off the cumulative column

Try It Now 1.3.3

Rainfall (in.)Freq.Rel. freq.Cum. rel. freq.
2.95–4.9760.120.12
4.97–6.9970.140.26
6.99–9.01150.300.56
9.01–11.0380.160.72
11.03–13.0590.180.90
13.05–15.0750.101.00
Total501.00

Table 1.3.5: annual rainfall in a sample of 50 towns.

Try It Now 1.3.3 — the question

From Table 1.3.5, find the percentage of rainfall that is less than 9.01 inches.

56%

The row ending at 9.01 is the third one. Its cumulative relative frequency already contains every interval below it: 0.56=56% 0.56 = 56\% .

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Worked example — the cumulative column had already done the adding

Example 1.3.1 · Reading a percentage off the cumulative column

Example 1.3.1 — Reading a percentage off the cumulative column

From Table 1.3.4, find the percentage of heights that are less than 65.95 inches.

23%

The first three rows are all under 65.95 inches, so 5+3+15=23 5 + 3 + 15 = 23 players qualify, and 23100=0.23=23% \frac{23}{100} = 0.23 = 23\% . That is exactly the cumulative relative frequency in the third row — the column had already done the adding.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Try it in rāSHio

Build all three columns at once

Paste the twenty students' work-hours list (5; 6; 3; 3; 2; …) into File → Delimited List… — the semicolons parse as-is — then choose Graph → Frequency Table with Discrete values checked.

Table 1.3.1's frequency column comes back, along with the relative and cumulative relative frequency columns that follow it.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Figure 1.3.3 — Graph → Frequency Table

The same table, built for you

One row per distinct value, with the counts, the shares, and the running total already filled in — the arithmetic of Tables 1.3.1 through 1.3.3, done in one pass.

Use it to check work you did by hand, not to skip the hand version. The reason the columns mean anything is the definition behind each one.

Figure 1.3.3: building Table 1.3.1's columns in rāSHio.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Your turn — a band in the middle, not a cutoff

Try It Now 1.3.4

Try It Now 1.3.4 — how common is a middling rainfall year?

Daniel Okada and his husband are deciding which town to move to. From Table 1.3.5, find the percentage of rainfall that is between 6.99 and 13.05 inches.

64%

The band covers three rows — 6.99–9.01, 9.01–11.03, 11.03–13.05 — so add their relative frequencies: 0.30+0.16+0.18=0.64 0.30 + 0.16 + 0.18 = 0.64 . Most of the towns they are considering sit in that middle band.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Worked example — when to add instead of read

Example 1.3.2 · Adding relative frequencies for a middle band

Example 1.3.2 — Adding relative frequencies for a middle band

From Table 1.3.4, find the percentage of heights that fall between 61.95 and 65.95 inches.

0.18, or 18%

The band is covered by the second row (61.95–63.95) and the third (63.95–65.95): 0.03+0.15=0.18 0.03 + 0.15 = 0.18 . Because we want a band in the middle rather than everything below a cutoff, we add the individual relative frequencies instead of reading a single cumulative entry.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Your turn — read the question's units carefully

Try It Now 1.3.5

Try It Now 1.3.5 — a count, not a percentage

From Table 1.3.5, find the number of towns that have rainfall between 2.95 and 9.01 inches.

28 towns

The band covers the first three rows. The question asks for a number of towns, not a percentage, so add the frequencies rather than the relative frequencies: 6+7+15=28 6 + 7 + 15 = 28 .

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Worked example — every column of a grouped table

Example 1.3.3 · Working every column

Example 1.3.3 — Use the 100 soccer player heights in Table 1.3.4

  • a. The percentage of heights from 67.95 to 71.95 inches is ____.
  • b. The percentage of heights from 67.95 to 73.95 inches is ____.
  • c. The percentage of heights more than 65.95 inches is ____.
  • d. The number of players between 61.95 and 71.95 inches tall is ____.
  • e. What kind of data are the heights?
  • f. How could you gather these heights so the data are characteristic of all semiprofessional soccer players?

Remember: you count frequencies. Relative frequency is the frequency divided by the total. Cumulative relative frequency adds all the previous relative frequencies to the current row's.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Example 1.3.3 — worked through

Which column the question is really asking for

Percentages — add relative frequencies

a. 0.17+0.12=0.29=29% 0.17 + 0.12 = 0.29 = 29\%
b. 0.29+0.07=0.36=36% 0.29 + 0.07 = 0.36 = 36\%
c. Everything above the third row is the rest of the whole: 10.23=0.77=77% 1 - 0.23 = 0.77 = 77\%

Counts and context

d. Rows two through six, adding frequencies: 3+15+40+17+12=87 3 + 15 + 40 + 17 + 12 = 87
e. Height lands anywhere on a continuous scale — quantitative continuous.
f. Get rosters from every team and take a simple random sample from each.

Part f is the reason for sampling every team: one team's unusual roster cannot dominate, and randomizing inside each team keeps you from unconsciously picking the tallest players.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

In class — build the table from your own data

Collaborative Exercise

Collaborative Exercise — survey the class and build a frequency table

Have someone conduct a survey of the number of siblings each student has. Create a frequency table, then add a relative frequency column and a cumulative relative frequency column. Answer the following:

  1. What percentage of the students in your class have no siblings?
  2. What percentage of the students have from one to three siblings?
  3. What percentage of the students have fewer than three siblings?
Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Your turn — answer as a fraction

Try It Now 1.3.6

Try It Now 1.3.6 — a single interval

Table 1.3.5 gives the annual rainfall in a sample of towns. What fraction of towns surveyed get between 11.03 and 13.05 inches of rainfall each year?

950 \frac{9}{50}

The interval 11.03–13.05 is a single row with a frequency of 9, and there are 50 towns in the sample. As a decimal that is 0.18, or 18%.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Worked example — the table is wrong; find out how

Example 1.3.4 · Finding the errors in a published table

Example 1.3.4 — nineteen commuters

Nineteen people were asked how many miles they commute to work each day:
2; 5; 7; 3; 2; 10; 18; 15; 20; 7; 10; 18; 5; 12; 13; 12; 4; 5; 10

  • a. Is Claire's table correct? If not, what is wrong?
  • b. True or false: three percent of those surveyed commute three miles or less.
  • c. What fraction commute five or seven miles?
  • d. What fraction commute 12 miles or more? Less than 12? Between five and 13, exclusive?
DataFreq.Rel. freq.Cum. rel. freq.
222/190.1053
311/190.1579
411/190.2105
533/190.3684
722/190.4737
1033/190.6316
1222/190.7368
1311/190.7895
1511/190.8421
1811/190.8948
2011/191.0000

Table 1.3.6: commuting distances, as published — with errors.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Example 1.3.4 — worked through

One missing person, and a count read as a percent

a. The frequency column sums to 18, not 19.

Go back to the raw list: 18 appears twice, but the table lists a frequency of 1. Because the frequency column is wrong, the cumulative relative frequencies are wrong too. Corrected, that column reads 219, 319, 419, 719, 919, 1219, 1419, 1519, 1619, 1819, 1919 \frac{2}{19},\ \frac{3}{19},\ \frac{4}{19},\ \frac{7}{19},\ \frac{9}{19},\ \frac{12}{19},\ \frac{14}{19},\ \frac{15}{19},\ \frac{16}{19},\ \frac{18}{19},\ \frac{19}{19} .

b. False

Three people commute three miles or less — that is 31915.8% \frac{3}{19} \approx 15.8\% , not 3%. The "3" is a count being reported as a percent.

c & d. Fractions

c. 3+219=519 \frac{3+2}{19} = \frac{5}{19} .   d. 12 miles or more: 719 \frac{7}{19} ; less than 12: 1219 \frac{12}{19} ; between 5 and 13 exclusive: 719 \frac{7}{19} .

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Your turn — four readings off one table

Try It Now 1.3.7

Years of serviceEmployees
242
251
263
270
284
296
3011
3112
327
338
346
3510

Table 1.3.7: years of service for 70 federal employees.

Try It Now 1.3.7 — the questions

  • a. Cumulative frequency for 30 to 35 years, inclusive?
  • b. Relative frequency for 30 years of service?
  • c. Relative frequency for 30 years or less?
  • d. Relative frequency for 25 years or more?

a. 54  ·  b. 15.7%  ·  c. 38.6%  ·  d. 97.1%

11+12+7+8+6+10=54 11+12+7+8+6+10 = 54 ; 11700.157 \frac{11}{70} \approx 0.157 ; 27700.386 \frac{27}{70} \approx 0.386 ; and "25 years or more" is everything except the two employees at 24 years, 68700.971 \frac{68}{70} \approx 0.971 .

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Worked example — a two-column table that runs down the page

Example 1.3.5 · Eighteen years of fatal crashes

Example 1.3.5 — Table 1.3.8, 653,782 crashes over 18 years

  • a. Frequency of deaths from Year 7 through Year 11?
  • b. What percentage occurred after Year 13?
  • c. Relative frequency of those in Year 7 or before?
  • d. Percentage of deaths in Year 18?
  • e. Cumulative relative frequency for Year 13 — and what does it tell you?

Nothing new is required here — only deciding, question by question, whether the answer is a count, a share, or a running total.

YearCrashesYearCrashes
136,2541038,477
237,2411138,444
337,4941239,252
437,3241338,648
537,1071437,435
637,1401534,172
737,5261630,862
837,8621730,296
938,4911829,757
Total653,782

Table 1.3.8: fatal motor vehicle traffic crashes, 18 years (shown in two column pairs).

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Example 1.3.5 — worked through

Three-quarters of the crashes came first

a – c

a. 37,526+37,862+38,491+38,477+38,444=190,800 37{,}526 + 37{,}862 + 38{,}491 + 38{,}477 + 38{,}444 = 190{,}800
b. Years 14–18 total 162,522 162{,}522 , and 162,522653,78224.9% \frac{162{,}522}{653{,}782} \approx 24.9\%
c. Years 1–7 total 260,086 260{,}086 , so 260,086653,7820.3978 \frac{260{,}086}{653{,}782} \approx 0.3978

d – e

d. 29,757653,7824.6% \frac{29{,}757}{653{,}782} \approx 4.6\%
e. Years 1–13 account for 653,782162,522=491,260 653{,}782 - 162{,}522 = 491{,}260 crashes, so the cumulative relative frequency is 491,260653,7820.7514 \frac{491{,}260}{653{,}782} \approx 0.7514

About 75% of the period's fatal crashes happened in the first 13 years. The totals dropped sharply from Year 14 onward, so the later years contribute far less to the total than an even split would suggest — which is the whole reason to read the cumulative column rather than assume.

Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

Key Terminology — the eight terms this section defined

The vocabulary

Levels of measurement

level of measurement — the classification (nominal, ordinal, interval, ratio) describing how much mathematical structure a data set carries.
nominal — labels only; not ordered, not usable in calculations.
ordinal — rankable, but the differences cannot be measured.
interval — ordered, meaningful differences, no true zero.
ratio — ordered, meaningful differences, true zero, so ratios are meaningful.

Frequency columns

frequency — the number of times a value of the data occurs.
relative frequency — the ratio of the number of times a value occurs to the total number of outcomes.
cumulative relative frequency — the running total of the relative frequencies up to and including the current row.

Each of the three frequency terms is one column of the same table, and each level of measurement is one rung of the same ladder. Neither list is a set of unrelated definitions to memorize.

1.3
Frequency, Tables & Levels of Measurement · bookSHelf Intro Stats§1.3

§1.3 — conclusions

What §1.3 leaves you with

The core idea

Read the precision off your data and add one. Read the level of measurement off your data and let it decide which arithmetic is allowed. Then build the three columns — count, share, running total — and a single cell answers the question you would otherwise add up by hand.

The failure case

A table whose frequency column sums to 18 when 19 people were surveyed still prints clean cumulative percentages all the way down to 1.0000. Nothing in the arithmetic complains. The only defense is checking the total against the sample size before you trust a single number in the table.

Next: §1.4 — Experimental Design and Ethics, where the question shifts from how the data are organized to whether the study that produced them can support the claim being made.