Introduction to Statistics · Chapter 1 · Sampling and Data
Six words carry this whole course. Learn to tell the number you can compute from the number you actually want.
bookSHelf · Introduction to Statistics · §1.1 · a self-paced section
Learning objectives — by the end of this section you will be able to
The first of the two halves of statistics
Definition 1.1.1 — Descriptive statistics
Descriptive statistics is the practice of organizing and summarizing data — through graphs, and through numerical summaries such as an average.
Definition 1.1.1: descriptive statistics turns a pile of values into a picture and a number.
Nothing here reaches beyond the data in front of you. Descriptive statistics never claims anything about anyone you did not measure.
The second half — and the reason for the rest of the course
Definition 1.1.2 — Inferential statistics
Inferential statistics consists of formal methods for drawing conclusions about a population from "good" data. It uses probability to say how confident we can be that those conclusions are correct.
The science of statistics is the collection, analysis, interpretation and presentation of data — batting averages, opinion polls, weather forecasts and your own gradebook are all statistics at work.
Definition 1.1.2: take a sample from the population, then estimate the parameter from the statistic.
The goal is not to run the formulas. It is to understand your data — the calculator can do the arithmetic, the understanding has to come from you.
In class — everybody contributes one number
Collaborative Exercise — build a dot plot from your own class
Have every class member write down the average time (in hours, to the nearest half-hour) they sleep per night. Your instructor records the data. Then build a dot plot — a number line with a dot stacked above each value. For example:
5; 5.5; 6; 6; 6; 6.5; 6.5; 6.5; 6.5; 7; 7; 8; 8; 9
Descriptive statistics, in one picture
Fourteen students, fourteen dots. Each value gets a dot stacked above its place on the number line, so the height of a stack is how many people gave that answer — and the shape of the stacks is where the class clusters.
Figure 1.1.1: average time (in hours) spent sleeping per night, for a class of 14 students.
You have not concluded anything about students in general — only summarized these fourteen. That is exactly the line between descriptive and inferential.
Try it in rāSHio — do it with your own class list
Open rāSHio and paste your class's sleep-hours list (5; 5.5; 6; 6; …) straight into File → Delimited List… — the semicolons parse as-is. Then choose Graph → Frequency Table with Discrete values checked to count how many classmates gave each value.
Figure 1.1.2: counting each sleep-hours value in rāSHio — Graph → Frequency Table, Discrete values checked.
The frequency table's counts are the same numbers as the heights of the stacks in Figure 1.1.1 — one summary in a table, the other in a picture.
§1.1.1 — the tool that makes inference possible
Inference claims something about people you never measured. What licenses that leap?
If a single outcome is unpredictable, how can many outcomes be dependable enough to bet on?
§1.1.1 — the mathematics of randomness
Definition 1.1.3 — Probability
Probability is a mathematical tool used to study randomness. It deals with the chance — the likelihood — of an event occurring.
Toss a fair coin four times and you may well not get two heads and two tails. Toss the same coin 4,000 times and the outcome will be close to half and half. The theoretical probability of heads on any one toss is 21.
Definition 1.1.3: the fraction of heads is unpredictable over a few tosses, but settles near 0.5 over thousands.
Probability is also how the rest of the course gets its warrants: earthquakes, rainfall, a vaccine's side effects, a portfolio's rate of return — every one is a prediction stated as a probability.
Insight Note — unpredictable in the short run, dependable in the long run
| Who tossed | Tosses | Heads | Fraction heads |
|---|---|---|---|
| a few flips | 4 | anything | unpredictable |
| one of the authors | 2,000 | 996 | 0.498 |
| Karl Pearson | 24,000 | 12,012 | 0.5005 |
| theory | — | — | 0.5 |
Coin-toss results: 2000996=0.498, and Pearson's 24,00012,012=0.5005.
One toss tells you nothing. Four thousand settle reliably near half heads. Probability is the mathematics of that long-run pattern — it cannot call the next flip, yet it predicts the crowd of flips with remarkable accuracy.
Your turn — commit to an answer before the reveal
Try It Now 1.1.1 — seven heads in ten tosses
You toss a fair coin 10 times and get 7 heads. Your friend Emma claims this proves the coin is unfair. Is she right? About what fraction of heads would you expect if you tossed the same coin 10,000 times?
Emma is not right.
Ten tosses is a small number of repetitions, and short runs of a random process are naturally uneven — 7 heads in 10 is not unusual for a fair coin. Over 10,000 tosses you would expect the fraction of heads very close to 0.5 (about 5,000 heads), just as Pearson's 24,000 tosses gave 12,012 heads, a fraction of 0.5005.
§1.1.2 — the vocabulary the whole course leans on
Every study you meet from here on can be taken apart with the same six labels. Read the study, then name each one:
§1.1.2 — the group the study is about
Definition 1.1.4 — Population
A population is the entire collection of persons, things, or objects under study.
Definition 1.1.4: the entire collection of persons, things, or objects under study.
The population is set by the question, not by convenience. "All athletes at the university" and "all students at the university" are different populations, and picking the wrong one invalidates everything downstream.
§1.1.2 — the part you actually examine
Definition 1.1.5 — Sample
A sample is a portion (or subset) of the larger population, selected so that studying it yields information about the whole population.
Examining a whole population costs time and money, so sampling is the practical technique: a sample of students for a campus-wide GPA, an opinion poll of 1,000–2,000 people for a country, a handful of cans to check whether a 16-ounce can really holds 16 ounces.
Definition 1.1.5: a subset of the population, studied to learn about the whole.
§1.1.2 — the number you can actually compute
Definition 1.1.6 — Statistic
A statistic is a number that represents a property of the sample.
Definition 1.1.6: a number computed from the sample.
If you can compute it from the data you collected, it is a statistic. That is the whole test.
§1.1.2 — the number you actually want
Definition 1.1.7 — Parameter
A parameter is a numerical characteristic of the whole population that can be estimated by a statistic.
Treat one math class as a sample of all math classes. The average points earned in that one class is a statistic. The average points earned across all math classes is the parameter it estimates.
Notice the two are described with the same phrase, differing only in whether it points at the sample or at the population. That parallel wording is not an accident — spotting it makes the labeling almost automatic.
The headline result of §1.1
A statistic estimates a parameter only as well as the sample represents the population.
The accuracy of the estimate does not depend on how carefully you did the arithmetic. It depends on whether the sample carries the characteristics of the population — a representative sample.
How accurately a statistic estimates its parameter is one of the main concerns of the entire field. Later chapters use the sample statistic to test a claimed population parameter — and every one of those tests assumes this.
Context Pause — bad samples make confident lies. A poll that only reaches landline phones, or a survey only of volunteers, produces precise-looking numbers that badly mislead. Whether a sample truly represents its population is as much an ethical question as a mathematical one — it decides whose voices the conclusions actually reflect.
§1.1.2 — the one thing you measure per member
Definition 1.1.8 — Variable
A variable, usually notated by a capital letter such as X or Y, is a characteristic or measurement that can be determined for each member of a population.
A numerical variable takes values with equal units — weight in pounds, time in hours. A categorical variable places each member in a category.
Definition 1.1.8: one characteristic that can be measured for each individual.
Let X = points earned by one math student: numerical, and averaging it means something. Let Y = a person's party affiliation: categorical, and an "average party affiliation" means nothing at all.
§1.1.2 — what you actually write down
Definition 1.1.9 — Data
Data are the actual values of the variable. They may be numbers or they may be words. A single value is a datum.
Definition 1.1.9: the actual recorded values of the variable.
Data are the result of sampling from a population. Everything else in this course — every graph, every average, every test — is computed from them.
Two words that come up constantly
Three exam scores of 86, 75 and 92 give a mean of
In a class of 40 students, 22 of them men, the proportion of men is 4022 and the proportion of women is 4018.
Context Pause — "mean" vs. "average." The two words are used interchangeably, and among non-statisticians that substitution is common practice. The technical term is arithmetic mean, while average technically refers to any center location — but in everyday use, "average" is accepted shorthand for the arithmetic mean.
Insight Note — why sampling works at all
Taste a spoonful, judge the pot.
A cook doesn't drink the whole pot of soup to check the seasoning — one well-stirred spoonful speaks for the pot.
Sampling works the same way: a carefully chosen subset can tell you about the entire population — as long as the spoonful truly resembles the pot. "Well-stirred" is doing all the work in that sentence.
Your turn — name all six before the reveal
Try It Now 1.1.2 — school uniforms at Knoll Academy
We want to know the average (mean) amount of money spent on school uniforms each year by families with children at Knoll Academy. We randomly survey 100 families with children in the school. Three of the families spent $65, $75 and $95.
Worked example — the same six labels
Example 1.1.1 — school supplies at ABC College
We want to know the average (mean) amount of money first year students at ABC College spend on school supplies that do not include books. We randomly surveyed 100 first year students. Three of them spent $150, $200 and $225.
Your turn — match each term to a phrase
A survey of athletes in a university studied the heights of athletes, in meters. Match each term to the phrase that describes it.
Terms
Phrases
1 → c · 2 → b · 3 → a · 4 → f · 5 → e · 6 → g
Not d for the population — the study is about athletes, not all students. The statistic (b) is the survey's average; the parameter (a) is the university's.
Worked example — matching, with the reasoning shown
A study at a local college analyzed the average cumulative GPAs of students who graduated last year.
Terms
Phrases
1 → f · 2 → g · 3 → e · 4 → d · 5 → b · 6 → c
Find the population first: the study is about graduates (f), not everyone who attended (a). Then ask of each phrase — population or sample? number or description?
Your turn — a proportion, not a mean
Try It Now 1.1.4 — charging from 50% to 100%
A survey checks how long a mobile phone's battery takes to charge from 50% to 100%, collected under one criterion: a 30 W charger on an Android phone. We want the proportion of Android phones charged to 100% within 30 minutes, starting from a simple random sample of 200 phones.
Worked example — when the "individuals" aren't people
Example 1.1.3 — testing the safety of electric automobiles
The National Transportation Safety Board crashed cars into a wall at 35 miles per hour with dummies in the front seat. We want the proportion of driver dummies that would have suffered head injuries, had they been real drivers. We start from a simple random sample of 75 cars.
Your turn — watch the population's boundary
Try It Now 1.1.5 — points on a license
Alex Fuentes, a reporter, wants the proportion of all truck drivers with no points on their license. They randomly select 1,000 truck drivers from the directory of truck drivers and count how many in the sample have no points.
The population is the directory, not "all truck drivers alive" — a sample can only speak for the list it was drawn from.
In groups of up to four — all six terms, one small study
Collaborative Exercise — find all six terms
Find a population, a sample, the parameter, the statistic, a variable and data for this study: you want the average (mean) number of glasses of milk college students drink per day. Yesterday, in your English class, you asked five students how many glasses they drank the day before. The answers were:
1, 0, 1, 3, 4
Argue about the population before you compute anything — "college students" is a much larger group than "students in one English class," and which one you claim decides whether your mean is worth anything.
Try it in rāSHio — compute your group's number
Open rāSHio, enter your group's five milk answers (1, 0, 1, 3, 4) in the spreadsheet, then choose Stats → Summary Statistics to read off the mean — that number is the statistic your group just identified, not the parameter.
Figure 1.1.3: Stats → Summary Statistics turns the five milk answers into the statistic.
The tool will happily give you a mean for any five numbers. It cannot tell you whether those five people represent college students — that judgement is still yours.
Worked example — the last of the four
Example 1.1.4 — an insurance company's question
An insurance company wants the proportion of all medical doctors who have been involved in one or more malpractice lawsuits. It selects 500 doctors at random from a professional directory and counts how many in the sample have been involved in such a suit.
Your turn — where data actually live
Your instructor keeps every score for the course in one spreadsheet — every student, every graded item. Before any of it can be analysed it has to be laid out so each piece of information has one unambiguous home. How would you arrange the rows, the columns and the cells?
| Student | Assignment 1 | Quiz 1 | Exam 1 |
|---|---|---|---|
| Student 1 | 92 | 85 | 88 |
| Student 2 | 78 | 90 | 81 |
| Student 3 | 85 | 72 | 94 |
Reading across a row summarizes one student; reading down a column summarizes the class on one item.
§1.1.3 — statistics that arrive as a story
A great deal of what you read about statistics does not arrive as a data set. It arrives as a story about one hospital, one school district, one company, one town.
How can a story be entirely true and still be useless for the question you're asking?
§1.1.3 — what one case can and cannot do
Definition 1.1.10 — Case study
A case study is a detailed examination of a single case — one person, one organization, one event — usually reported in depth rather than summarized as numbers.
It is good at two things: showing that something is possible (one documented rare drug reaction proves the reaction can happen), and suggesting what to measure next.
Definition 1.1.10: a case study proves the thing is possible; it cannot say how often it happens.
What it cannot do is tell you how often something happens, or whether it would happen again — with one case there is nothing to compare against. The honest use is as the first step: it generates the question; the study that answers it needs a sample, a comparison group, and a plan made before the data arrive.
Your turn — read the story like a statistician
Try It Now 1.1.7 — a local news segment
"After the city installed a new bike lane on Rowan Street, one cafe owner says her weekday morning sales rose by nearly forty percent."
§1.1 — conclusions
The core idea
A statistic is the number you can compute, from the sample you examined. A parameter is the number you actually want, about the population you care about. Descriptive statistics stops at the first; inferential statistics uses probability to reach the second.
The failure case
A sample that does not represent its population still produces a perfectly precise statistic — and a badly wrong estimate. Volunteers, landlines, one cafe on Rowan Street: the arithmetic is fine, the conclusion is not.
Next: §1.2 — Data, Sampling, and Variation in Data and Sampling, where "representative" stops being an adjective and becomes a procedure.