Introduction to Statistics · Chapter 1 · Sampling and Data
Counting is the easy part. The work is deciding how far to carry an answer, what arithmetic the data will even tolerate, and how to turn a raw list into a table someone can read.
bookSHelf · Introduction to Statistics · §1.3 · a self-paced section
Learning objectives — by the end of this section you will be able to
§1.3.1 — before any of the counting starts
Organizing data means dividing, and division produces long decimals. So the section opens with the one rule that decides how far to carry an answer — and it does not ask you to judge how precise a number "looks."
Carry your final answer one more decimal place than was present in the original data.
§1.3.1 — three quiz scores, one rounding decision
Had the scores been recorded as 4.0, 6.0, 9.0 — one decimal place — the same calculation would be reported to two: 6.33. Also: most fractions in this course need not be reduced. An unreduced fraction usually shows where the numbers came from.
The rule — data precision, plus one
34+6+9=319=6.3333333…Reported as 6.3. Round off only the final answer; if an intermediate result must be rounded, carry it to at least twice as many places as the final answer.
Insight Note — why the rule says "only the final answer"
Rounding early is like trimming a board before you measure twice.
Every time you round in the middle of a calculation you throw away a sliver of the answer, and those slivers pile up. Keep the full decimal on your calculator until the very last step, then trim once.
Your turn — commit to an answer before the reveal
Try It Now 1.3.1 — five logged study times
Five students in a study group recorded how many minutes they spent on homework last night. Alex Delgado logged 42 minutes; Hannah Wolcott logged 55; Marcus Bell logged 38; Kayla Nguyen logged 61; Ethan Shaw logged 47. Find the group's mean study time and round it off correctly.
48.6 minutes.
42+55+38+61+47=243, and 5243=48.6. The data are whole numbers, so the answer carries one decimal place. No rounding was actually needed here — the division came out exactly; had it come out 48.62, we would have reported 48.6.
Try it in rāSHio
Paste the five logged times (42; 55; 38; 61; 47) into File → Delimited List…, then choose Stats → Summary Statistics. The mean comes back as 48.6 minutes — no adding and dividing by hand.
Worth reopening that panel after the next subsection: a mean is only meaningful once the data reach the interval or ratio level — which is exactly the distinction Levels of Measurement draws.
Figure 1.3.1 — the tool, in one pass
Load a column of values, open the Summary Statistics panel, and the mean arrives alongside the median, the standard deviation, the quartiles, and the range.
The rounding rule still applies to whatever you copy out of that panel — the tool will happily hand you seven decimal places.
Figure 1.3.1: reading the mean directly in rāSHio: Stats → Summary Statistics.
§1.3.2 — the question you answer before you calculate
The way a set of data is measured is called its level of measurement, and not every statistical operation can be used with every set of data. There are four levels, ordered from least mathematical structure to most: nominal, ordinal, interval, ratio.
Each level adds one power to the level below it.
§1.3.2 — the classification itself
Definition 1.3.1 — Level of measurement
The level of measurement of a data set is the classification — nominal, ordinal, interval, or ratio — that describes how much mathematical structure the measurements carry, and therefore which calculations are meaningful on them.
Figure: the four levels as a ladder — each step adds one power to the step below it.
Correct statistical procedures depend on knowing which rung you are standing on. Get it wrong and the arithmetic still runs — it just stops meaning anything.
Level 1 — labels only
Definition 1.3.2 — Nominal scale level
Data measured on a nominal scale are qualitative: categories, colors, names, labels, favorite foods, yes-or-no responses. Nominal data are not ordered and cannot be used in calculations.
Putting pizza first and sushi second is not meaningful — that is just the order you happened to write them down in. Smartphone brands are the same: no agreed-upon order, however strong anyone's preferences.
Figure: nominal data are labels — any arrangement is equally valid.
Context Pause — why this classification earns a subsection
You can compute an average shoe size. You cannot compute an average favorite pizza topping.
Level of measurement is the rule that tells you which of those two situations you are in — before you start calculating, not after the answer looks strange.
Level 2 — order, but no measurable gaps
Definition 1.3.3 — Ordinal scale level
Data measured on an ordinal scale are categorical like nominal data, with one difference: they can be ordered. Differences between two ordinal values cannot be measured, and ordinal data cannot be used in calculations.
The top five national parks can be ranked 1 to 5, but the distance from first to second is unknowable. A cruise survey reading "excellent, good, satisfactory, unsatisfactory" is ordered — and there is no reason the gap from excellent to good equals the gap from good to satisfactory.
Figure: ordinal data can be ranked — the size of each gap is unknowable.
Level 3 — differences work, ratios do not
Definition 1.3.4 — Interval scale level
Data measured on an interval scale have a definite ordering and are numerical, so differences can be calculated. But an interval scale has no true zero — its zero point is a convention, not an absence — so ratios are not meaningful.
Celsius and Fahrenheit: 40∘=100∘−60∘, so differences make sense. Yet 0 is not a minimum — −10 °F and −15 °C exist, and they are colder than zero.
Figure: subtraction works; ratios do not, because zero is a convention.
Insight Note — the whole difference between interval and ratio
A true zero means "none of it."
Zero dollars means you have no money. Zero degrees Celsius means the water is freezing — it does not mean there is no temperature.
That is why "80° is twice as hot as 40°" is a sentence that sounds fine and means nothing.
Level 4 — the most information a measurement can carry
Definition 1.3.5 — Ratio scale level
Ratio scale data are like interval data — ordered, numerical, with meaningful differences — but there is a true minimum value of zero, so ratios between values are meaningful.
Four machine-graded exam scores out of 100: 80, 68, 20, 92. They order (20, 68, 80, 92); they subtract (92 beats 68 by 24 points); and because the minimum possible score is 0, the score of 80 really is four times the score of 20.
Figure: anchored at a true zero, so 80 really is four times 20.
Definitions 1.3.1–1.3.5, on one page
| Level | Ordered? | Differences meaningful? | True zero (ratios meaningful)? | Example |
|---|---|---|---|---|
| Nominal | No | No | No | Crayon colors, smartphone brands |
| Ordinal | Yes | No | No | Survey ratings, park rankings |
| Interval | Yes | Yes | No | Temperature in °C or °F, calendar years |
| Ratio | Yes | Yes | Yes | Exam scores, distance, income |
Table: the four levels of measurement, lowest to highest.
Read it as a ladder, not a list: every column that says "Yes" stays "Yes" all the way down. Ratio data are the only kind that survive every operation in the table.
Your turn — name the level, and say why
Try It Now 1.3.2 — Rachel Whitfield's field-day records
a) nominal · b) ordinal · c) ratio
A jersey number is a label — player 10 is not "twice" player 5. Places are genuinely ordered, but the gap from 1st to 2nd is not a measurable amount. Finishing times are ordered, subtract sensibly, and zero minutes is a true zero — a 40-minute runner took twice as long as a 20-minute runner. Only part c lets Rachel do real arithmetic with her records.
§1.3.3 — from a raw list to a table you can read
A frequency table takes a jumble of responses and answers three questions at once: how many landed on each value, what share of the sample that is, and how much of the sample you have accounted for by the time you reach a given row.
Count. Divide by the total. Then keep a running total.
Column one — the count
Definition 1.3.6 — Frequency
A frequency is the number of times a value of the data occurs.
Twenty students were asked how many hours they worked per day:
5; 6; 3; 3; 2; 4; 7; 5; 2; 3; 5; 6; 5; 4; 4; 3; 5; 2; 5; 3
Figure: each dot is one student; the stack height is the frequency.
The same twenty responses, sorted and counted
| Data value | Frequency |
|---|---|
| 2 | 3 |
| 3 | 5 |
| 4 | 3 |
| 5 | 6 |
| 6 | 2 |
| 7 | 1 |
Table 1.3.1: frequency table of student work hours.
Three students work two hours, five work three hours, and so on. The frequency column sums to 20 — the total number of students in the sample. If it does not sum to your sample size, you have miscounted.
Column two — the share
Definition 1.3.7 — Relative frequency
A relative frequency is the ratio of the number of times a value occurs to the total number of outcomes. Divide each frequency by the sample total — here, 20. Write it as a fraction, a percent, or a decimal.
The relative frequency column of Table 1.3.2 sums to 2020, or 1.
| Value | Freq. | Relative freq. |
|---|---|---|
| 2 | 3 | 3/20 or 0.15 |
| 3 | 5 | 5/20 or 0.25 |
| 4 | 3 | 3/20 or 0.15 |
| 5 | 6 | 6/20 or 0.30 |
| 6 | 2 | 2/20 or 0.10 |
| 7 | 1 | 1/20 or 0.05 |
Table 1.3.2: work hours with relative frequencies.
Column three — the running total
Definition 1.3.8 — Cumulative relative frequency
Cumulative relative frequency is the accumulation of the previous relative frequencies. To find it, add all the previous relative frequencies to the relative frequency for the current row.
The last entry is 1 — one hundred percent of the data has been accumulated.
Figure: a running total — the last step must reach one.
One table, three questions answered
| Data value | Frequency | Relative frequency | Cumulative relative frequency |
|---|---|---|---|
| 2 | 3 | 3/20 or 0.15 | 0.15 |
| 3 | 5 | 5/20 or 0.25 | 0.15 + 0.25 = 0.40 |
| 4 | 3 | 3/20 or 0.15 | 0.40 + 0.15 = 0.55 |
| 5 | 6 | 6/20 or 0.30 | 0.55 + 0.30 = 0.85 |
| 6 | 2 | 2/20 or 0.10 | 0.85 + 0.10 = 0.95 |
| 7 | 1 | 1/20 or 0.05 | 0.95 + 0.05 = 1.00 |
Table 1.3.3: work hours with relative and cumulative relative frequencies.
Every cumulative entry carries everything above it. That is what makes the column worth building: a single cell answers "how much of the sample is at this value or below?" without any adding.
Insight Note — the shape of the cumulative column
A running total, not a fresh count.
Think of the cumulative column as a bucket you keep pouring into. Each row adds its own relative frequency to everything already in the bucket, so the last row must come out full — one whole, or 100%.
Context Pause — when the last entry reads 0.99
Because of rounding, the relative frequency column may not always sum to one, and the last cumulative entry may not be exactly one.
Each should be close to one. If yours is far off, you have an arithmetic error — not a rounding artifact.
100 measurements of a continuous quantity — grouped into intervals
| Heights (in.) | Freq. | Rel. freq. | Cum. rel. freq. |
|---|---|---|---|
| 59.95–61.95 | 5 | 0.05 | 0.05 |
| 61.95–63.95 | 3 | 0.03 | 0.08 |
| 63.95–65.95 | 15 | 0.15 | 0.23 |
| 65.95–67.95 | 40 | 0.40 | 0.63 |
| 67.95–69.95 | 17 | 0.17 | 0.80 |
| 69.95–71.95 | 12 | 0.12 | 0.92 |
| 71.95–73.95 | 7 | 0.07 | 0.99 |
| 73.95–75.95 | 1 | 0.01 | 1.00 |
| Total | 100 | 1.00 |
Table 1.3.4: heights of 100 semiprofessional soccer players.
Figure 1.3.2: the same table as a histogram, in 2-inch intervals.
The boundaries end in .95 on purpose: every height falls between the endpoints of an interval and never at one, so no measurement can land in two intervals at once.
Context Pause — a table with a sequel
The soccer-height data returns in Descriptive Statistics, where the method used to compute the intervals is explained.
For now, take the intervals as given and focus on reading the counts off them. Where the interval boundaries came from is a later question.
Your turn — read a percentage off the cumulative column
| Rainfall (in.) | Freq. | Rel. freq. | Cum. rel. freq. |
|---|---|---|---|
| 2.95–4.97 | 6 | 0.12 | 0.12 |
| 4.97–6.99 | 7 | 0.14 | 0.26 |
| 6.99–9.01 | 15 | 0.30 | 0.56 |
| 9.01–11.03 | 8 | 0.16 | 0.72 |
| 11.03–13.05 | 9 | 0.18 | 0.90 |
| 13.05–15.07 | 5 | 0.10 | 1.00 |
| Total | 50 | 1.00 |
Table 1.3.5: annual rainfall in a sample of 50 towns.
Try It Now 1.3.3 — the question
From Table 1.3.5, find the percentage of rainfall that is less than 9.01 inches.
56%
The row ending at 9.01 is the third one. Its cumulative relative frequency already contains every interval below it: 0.56=56%.
Worked example — the cumulative column had already done the adding
Example 1.3.1 — Reading a percentage off the cumulative column
From Table 1.3.4, find the percentage of heights that are less than 65.95 inches.
23%
The first three rows are all under 65.95 inches, so 5+3+15=23 players qualify, and 10023=0.23=23%. That is exactly the cumulative relative frequency in the third row — the column had already done the adding.
Try it in rāSHio
Paste the twenty students' work-hours list (5; 6; 3; 3; 2; …) into File → Delimited List… — the semicolons parse as-is — then choose Graph → Frequency Table with Discrete values checked.
Table 1.3.1's frequency column comes back, along with the relative and cumulative relative frequency columns that follow it.
Figure 1.3.3 — Graph → Frequency Table
One row per distinct value, with the counts, the shares, and the running total already filled in — the arithmetic of Tables 1.3.1 through 1.3.3, done in one pass.
Use it to check work you did by hand, not to skip the hand version. The reason the columns mean anything is the definition behind each one.
Figure 1.3.3: building Table 1.3.1's columns in rāSHio.
Your turn — a band in the middle, not a cutoff
Try It Now 1.3.4 — how common is a middling rainfall year?
Daniel Okada and his husband are deciding which town to move to. From Table 1.3.5, find the percentage of rainfall that is between 6.99 and 13.05 inches.
64%
The band covers three rows — 6.99–9.01, 9.01–11.03, 11.03–13.05 — so add their relative frequencies: 0.30+0.16+0.18=0.64. Most of the towns they are considering sit in that middle band.
Worked example — when to add instead of read
Example 1.3.2 — Adding relative frequencies for a middle band
From Table 1.3.4, find the percentage of heights that fall between 61.95 and 65.95 inches.
0.18, or 18%
The band is covered by the second row (61.95–63.95) and the third (63.95–65.95): 0.03+0.15=0.18. Because we want a band in the middle rather than everything below a cutoff, we add the individual relative frequencies instead of reading a single cumulative entry.
Your turn — read the question's units carefully
Try It Now 1.3.5 — a count, not a percentage
From Table 1.3.5, find the number of towns that have rainfall between 2.95 and 9.01 inches.
28 towns
The band covers the first three rows. The question asks for a number of towns, not a percentage, so add the frequencies rather than the relative frequencies: 6+7+15=28.
Worked example — every column of a grouped table
Example 1.3.3 — Use the 100 soccer player heights in Table 1.3.4
Remember: you count frequencies. Relative frequency is the frequency divided by the total. Cumulative relative frequency adds all the previous relative frequencies to the current row's.
Example 1.3.3 — worked through
Percentages — add relative frequencies
a. 0.17+0.12=0.29=29%
b. 0.29+0.07=0.36=36%
c. Everything above the third row is the rest of the whole: 1−0.23=0.77=77%
Counts and context
d. Rows two through six, adding frequencies: 3+15+40+17+12=87
e. Height lands anywhere on a continuous scale — quantitative continuous.
f. Get rosters from every team and take a simple random sample from each.
Part f is the reason for sampling every team: one team's unusual roster cannot dominate, and randomizing inside each team keeps you from unconsciously picking the tallest players.
In class — build the table from your own data
Collaborative Exercise — survey the class and build a frequency table
Have someone conduct a survey of the number of siblings each student has. Create a frequency table, then add a relative frequency column and a cumulative relative frequency column. Answer the following:
Your turn — answer as a fraction
Try It Now 1.3.6 — a single interval
Table 1.3.5 gives the annual rainfall in a sample of towns. What fraction of towns surveyed get between 11.03 and 13.05 inches of rainfall each year?
509
The interval 11.03–13.05 is a single row with a frequency of 9, and there are 50 towns in the sample. As a decimal that is 0.18, or 18%.
Worked example — the table is wrong; find out how
Example 1.3.4 — nineteen commuters
Nineteen people were asked how many miles they commute to work each day:
2; 5; 7; 3; 2; 10; 18; 15; 20; 7; 10; 18; 5; 12; 13; 12; 4; 5; 10
| Data | Freq. | Rel. freq. | Cum. rel. freq. |
|---|---|---|---|
| 2 | 2 | 2/19 | 0.1053 |
| 3 | 1 | 1/19 | 0.1579 |
| 4 | 1 | 1/19 | 0.2105 |
| 5 | 3 | 3/19 | 0.3684 |
| 7 | 2 | 2/19 | 0.4737 |
| 10 | 3 | 3/19 | 0.6316 |
| 12 | 2 | 2/19 | 0.7368 |
| 13 | 1 | 1/19 | 0.7895 |
| 15 | 1 | 1/19 | 0.8421 |
| 18 | 1 | 1/19 | 0.8948 |
| 20 | 1 | 1/19 | 1.0000 |
Table 1.3.6: commuting distances, as published — with errors.
Example 1.3.4 — worked through
a. The frequency column sums to 18, not 19.
Go back to the raw list: 18 appears twice, but the table lists a frequency of 1. Because the frequency column is wrong, the cumulative relative frequencies are wrong too. Corrected, that column reads 192, 193, 194, 197, 199, 1912, 1914, 1915, 1916, 1918, 1919.
b. False
Three people commute three miles or less — that is 193≈15.8%, not 3%. The "3" is a count being reported as a percent.
c & d. Fractions
c. 193+2=195. d. 12 miles or more: 197; less than 12: 1912; between 5 and 13 exclusive: 197.
Your turn — four readings off one table
| Years of service | Employees |
|---|---|
| 24 | 2 |
| 25 | 1 |
| 26 | 3 |
| 27 | 0 |
| 28 | 4 |
| 29 | 6 |
| 30 | 11 |
| 31 | 12 |
| 32 | 7 |
| 33 | 8 |
| 34 | 6 |
| 35 | 10 |
Table 1.3.7: years of service for 70 federal employees.
Try It Now 1.3.7 — the questions
a. 54 · b. 15.7% · c. 38.6% · d. 97.1%
11+12+7+8+6+10=54; 7011≈0.157; 7027≈0.386; and "25 years or more" is everything except the two employees at 24 years, 7068≈0.971.
Worked example — a two-column table that runs down the page
Example 1.3.5 — Table 1.3.8, 653,782 crashes over 18 years
Nothing new is required here — only deciding, question by question, whether the answer is a count, a share, or a running total.
| Year | Crashes | Year | Crashes |
|---|---|---|---|
| 1 | 36,254 | 10 | 38,477 |
| 2 | 37,241 | 11 | 38,444 |
| 3 | 37,494 | 12 | 39,252 |
| 4 | 37,324 | 13 | 38,648 |
| 5 | 37,107 | 14 | 37,435 |
| 6 | 37,140 | 15 | 34,172 |
| 7 | 37,526 | 16 | 30,862 |
| 8 | 37,862 | 17 | 30,296 |
| 9 | 38,491 | 18 | 29,757 |
| Total | 653,782 | ||
Table 1.3.8: fatal motor vehicle traffic crashes, 18 years (shown in two column pairs).
Example 1.3.5 — worked through
a – c
a. 37,526+37,862+38,491+38,477+38,444=190,800
b. Years 14–18 total 162,522, and 653,782162,522≈24.9%
c. Years 1–7 total 260,086, so 653,782260,086≈0.3978
d – e
d. 653,78229,757≈4.6%
e. Years 1–13 account for 653,782−162,522=491,260 crashes, so the cumulative relative frequency is 653,782491,260≈0.7514
About 75% of the period's fatal crashes happened in the first 13 years. The totals dropped sharply from Year 14 onward, so the later years contribute far less to the total than an even split would suggest — which is the whole reason to read the cumulative column rather than assume.
Key Terminology — the eight terms this section defined
Levels of measurement
level of measurement — the classification (nominal, ordinal, interval, ratio) describing how much mathematical structure a data set carries.
nominal — labels only; not ordered, not usable in calculations.
ordinal — rankable, but the differences cannot be measured.
interval — ordered, meaningful differences, no true zero.
ratio — ordered, meaningful differences, true zero, so ratios are meaningful.
Frequency columns
frequency — the number of times a value of the data occurs.
relative frequency — the ratio of the number of times a value occurs to the total number of outcomes.
cumulative relative frequency — the running total of the relative frequencies up to and including the current row.
Each of the three frequency terms is one column of the same table, and each level of measurement is one rung of the same ladder. Neither list is a set of unrelated definitions to memorize.
§1.3 — conclusions
The core idea
Read the precision off your data and add one. Read the level of measurement off your data and let it decide which arithmetic is allowed. Then build the three columns — count, share, running total — and a single cell answers the question you would otherwise add up by hand.
The failure case
A table whose frequency column sums to 18 when 19 people were surveyed still prints clean cumulative percentages all the way down to 1.0000. Nothing in the arithmetic complains. The only defense is checking the total against the sample size before you trust a single number in the table.
Next: §1.4 — Experimental Design and Ethics, where the question shifts from how the data are organized to whether the study that produced them can support the claim being made.