1.1 Definitions of Statistics, Probability, and Key Terms

Aligned outcomes:

SLO 1

Assess how data were collected and recognize how data collection affects what conclusions can be drawn from the data.

This section hands you the vocabulary that every question about data collection is asked in: population, sample, parameter, statistic, variable, and data. Working through the studies in Examples 1.1.1-1.1.4 — surveys of college spending, crash tests, malpractice records — you practice naming exactly who was studied and what was measured, which is the first move in assessing how data were collected. Once you can tell a sample's statistic from a population's parameter, you can start asking whether the sample really supports the conclusion being drawn from it.

SLO 6

Evaluate ethical issues in statistical practice.

This section plants the seed the course's ethics work grows from: a statistic is only trustworthy when the sample behind it represents the population. The discussion of representative samples — and of how polls or surveys that reach only part of a population can produce precise-looking but misleading numbers — sets up the ethical judgment you will exercise later when weighing study designs. It does not yet ask you to evaluate ethical dilemmas; it gives you the population-versus-sample distinction those evaluations rest on.

Learning Objectives

By the end of this section, you will be able to:

In this section, you will learn to:
  • distinguish between descriptive statistics and inferential statistics, and describe what each one does with data;
  • identify the population, sample, parameter, statistic, variable, and data in a real study;
  • explain how probability measures the long-run regularity of chance events, even when short-run outcomes are unpredictable;
  • explain why a sample must be representative of its population before we can trust conclusions drawn from it.
Definition 1.1.1: Descriptive Statistics

Descriptive statistics is the practice of organizing and summarizing data — through graphs and through numerical summaries such as an average.

Definition 1.1.1 -- Descriptive Statistics A loose scattered row of seven small dots is labelled "data" beneath it. Two arrows fan out downward from the data, one to a left panel and one to a right panel. The left panel is a rounded rectangle holding a small five-bar chart standing on a baseline, labelled "graphs" underneath. The right panel is a rounded rectangle holding the numerical expression x-bar equals 84.3, labelled "numerical summaries" underneath. A caption below reads "two ways to summarize the same data" -- the same raw data organized into a graph on one side and summarized as a number on the other. Descriptive Statistics data graphs
numerical summaries two ways to summarize the same data

Definition 1.1.1 — Descriptive statistics organizes and summarizes data through graphs and numerical summaries.

Definition 1.1.2: Inferential Statistics

Inferential statistics consists of formal methods for drawing conclusions about a population from "good" data. Statistical inference uses probability to determine how confident we can be that our conclusions are correct.

Effective interpretation of data (inference) is based on good procedures for producing data and thoughtful examination of the data. You will encounter what will seem to be too many mathematical formulas for interpreting data. The goal of statistics is not to perform numerous calculations using the formulas, but to gain an understanding of your data. The calculations can be done using a calculator or a computer. The understanding must come from you. If you can thoroughly grasp the basics of statistics, you can be more confident in the decisions you make in life.

The science of statistics deals with the collection, analysis, interpretation, and presentation of data. We see and use data in our everyday lives — batting averages, opinion polls, weather forecasts, and the grades in your own gradebook are all statistics at work.

Definition 1.1.2 — Population, sample, and inference Left: a large circle labeled Population holding 20 blue dots. Right: a smaller circle labeled Sample, empty at first. A solid arrow labeled take a sample grows left to right between them. Five of the population dots briefly pulse larger and rust-colored, then five dots drift from those same positions into the sample circle, turning rust-colored (the sample). A solid arrow labeled estimate the parameter then draws right to left, back from the sample into the population (simplified from the manim original's dashed line — see the CSS comment on .est-shaft for why). Finally captions fade in below each circle: Parameter, a fact about the whole population under Population; Statistic, a fact about the sample under Sample. Population Sample take a sample estimate the parameter Parameter a fact about the whole population Statistic a fact about the sample

Definition 1.1.2 — Inferential statistics: take a sample from the population, then estimate the parameter from the statistic.

In your classroom, try this exercise. Have class members write down the average time (in hours, to the nearest half-hour) they sleep per night. Your instructor will record the data. Then create a simple graph (called a dot plot) of the data. A dot plot consists of a number line and dots (or points) positioned above the number line. For example, consider the following data:

5; 5.5; 6; 6; 6; 6.5; 6.5; 6.5; 6.5; 7; 7; 8; 8; 9

The dot plot for this data is shown in Figure 1.1.1 below. Does your dot plot look the same as or different from the example? Why? If you did the same example in an English class with the same number of students, do you think the results would be the same? Why or why not? Where do your data appear to cluster? How might you interpret the clustering?

Figure 1.1.1 — Dot plot of average nightly sleep hours A semicolon-separated data row above a bare number line from 5 to 9. Each listed value in turn brightens, then travels down and settles as an open circle stacked on its value on the number line, while the list entry dims to mark it used. Ends with the finished dot plot and every list entry dimmed. Data: 5; 5.5; 6; 6; 6; 6.5; 6.5; 6.5; 6.5; 7; 7; 8; 8; 9 5 6 7 8 9

Figure 1.1.1 — Dot plot showing the frequency of average time (in hours) spent sleeping per night for a class of 14 students.

Try it in rāSHio

Open rāSHio, paste your class's sleep-hours list (5; 5.5; 6; 6; …) straight into File → Delimited List… — the semicolons parse as-is — then choose Graph → Frequency Table with Discrete values checked to count how many classmates gave each value, which is exactly what the stacked dots in Figure 1.1.1 show.

Figure 1.1.2 — Counting each sleep-hours value in rāSHio: Graph → Frequency Table with Discrete values checked.

The questions in the exercise above ask you to analyze and interpret your data. With this example, you have already begun your study of statistics.

In this course, you will learn how to organize and summarize data. Two ways to summarize data are by graphing and by using numbers (for example, finding an average). After you have studied probability and probability distributions, you will use formal methods for drawing conclusions from "good" data.

1.1.1 Probability

Definition 1.1.3: Probability

Probability is a mathematical tool used to study randomness. It deals with the chance (the likelihood) of an event occurring.

Unpredictable in the short run, dependable in the long run

One coin toss tells you nothing — but 4,000 tosses settle reliably near half heads. Probability is the mathematics of that long-run pattern: it can't call the next flip, yet it predicts the crowd of flips with remarkable accuracy.

For example, if you toss a fair coin four times, the outcomes may not be two heads and two tails. However, if you toss the same coin 4,000 times, the outcomes will be close to half heads and half tails. The expected theoretical probability of heads in any one toss is \(\frac{1}{2}\) or 0.5. Even though the outcomes of a few repetitions are uncertain, there is a regular pattern of outcomes when there are many repetitions. After reading about the English statistician Karl Pearson, who tossed a coin 24,000 times with a result of 12,012 heads, one of the authors tossed a coin 2,000 times. The results were 996 heads. The fraction \(\frac{996}{2000}\) is equal to 0.498, which is very close to 0.5, the expected probability.

The theory of probability began with the study of games of chance such as poker. Predictions take the form of probabilities. To predict the likelihood of an earthquake, of rain, or whether you will get an A in this course, we use probabilities. Doctors use probability to determine the chance of a vaccination causing the disease the vaccination is supposed to prevent. A stockbroker uses probability to determine the rate of return on a client's investments. You might use probability to decide to buy a lottery ticket or not. In your study of statistics, you will use the power of mathematics through probability calculations to analyze and interpret your data.

Definition 1.1.3 — Probability: the long-run frequency of heads settles near 0.5 An oscillating curve for the fraction of heads across up to 2,000 coin tosses. It starts near 0.8, swings above and below a dashed reference line at 0.5 with shrinking amplitude, and settles onto that line as the toss count grows. Probability: the long run settles 0.5 1.0 0.5 number of tosses fraction of heads one toss is unpredictable — 2,000 tosses are not

Definition 1.1.3 — Probability: the fraction of heads is unpredictable in a few tosses but settles near 0.5 over thousands.

Try It Now 1.1.1

You toss a fair coin 10 times and get 7 heads. Your friend Emma claims this proves the coin is unfair. Is she right? About what fraction of heads would you expect if you tossed the same coin 10,000 times?

Solution

Step 1 — Think about the short run: Ten tosses is a small number of repetitions, and small runs of a random process are naturally uneven. Getting 7 heads in 10 tosses is not unusual for a fair coin, so Emma's conclusion is not justified.

Step 2 — Think about the long run: Probability describes what happens over many repetitions. For a fair coin, the theoretical probability of heads on any one toss is \(\frac{1}{2} = 0.5\).

Answer: Emma is not right — a short run cannot prove the coin is unfair. Over 10,000 tosses, you would expect the fraction of heads to be very close to 0.5 (about 5,000 heads), just as Karl Pearson's 24,000 tosses produced 12,012 heads — a fraction of 0.5005.

1.1.2 Populations, Samples, and Variables

Definition 1.1.4: Population

A population is the entire collection of persons, things, or objects under study.

Definition 1.1.4 -- Population A bold heading reads Population. Below it, a large circle outlined and lightly tinted in the CURVE tone is labelled "every member under study" in the ACCENT tone above it. The circle is filled with a grid-with-jitter cloud of 34 small solid dots in the CURVE tone, representing every member of the population. A caption below reads "the entire collection of persons, things, or objects under study". The frame is deliberately a circle, not a box, matching the same population/sample visual convention used in Definitions 1.1.2 and 1.1.5. Population every member under study the entire collection of persons, things, or objects under study

Definition 1.1.4 — A population: the entire collection of persons, things, or objects under study.

Definition 1.1.5: Sample

A sample is a portion (or subset) of the larger population, selected so that studying it yields information about the whole population.

Two numbers, two worlds

Every study juggles a number you can actually compute (from the sample) and a number you usually can't (from the whole population). Keeping the pair straight — statistic vs. parameter — is the single most-used vocabulary distinction in this entire course.

Because it takes a lot of time and money to examine an entire population, sampling is a very practical technique. If you wished to compute the overall grade point average at your school, it would make sense to select a sample of students who attend the school. The data collected from the sample would be the students' grade point averages. In presidential elections, opinion poll samples of 1,000–2,000 people are taken. The opinion poll is supposed to represent the views of the people in the entire country. Manufacturers of canned carbonated drinks take samples to determine if a 16-ounce can contains 16 ounces of carbonated drink.

Definition 1.1.5 — Sample A bold heading reads Sample. A large circle on the left is labelled Population and filled with a grid-clipped-to-a-disc cloud of 29 small dots in the CURVE tone. Five of those dots, scattered rather than neighbouring, light up one at a time in the ACCENT tone and grow slightly, showing a random draw rather than a convenient contiguous cluster. A black arrow labelled "choose a subset" grows from the Population circle toward a smaller circle on the right labelled Sample. Five accent-coloured copies of the chosen dots then drift out of the Population and into the Sample circle -- the originals stay behind, since sampling observes the population rather than emptying it. A caption reads "a portion (subset) of the population that we actually examine". Sample Population Sample choose a subset a portion (subset) of the population that we actually examine

Definition 1.1.5 — A sample: a subset of the population, studied to learn about the whole.

Definition 1.1.6: Statistic

A statistic is a number that represents a property of the sample.

Definition 1.1.6 — Statistic Left: a rounded rectangle outlined in CURVE, labelled "Sample" above it, holding the three stacked values 86, 75, 92. A thick compute arrow in INK crosses to a rounded rectangle outlined in ACCENT on the right, holding the expression x-bar equals 84.3 in ACCENT, labelled "one number" beneath it. Caption below: "a number that describes the sample". Static figure, no motion. Statistic Sample 86 75 92 compute
one number a number that describes the sample

Definition 1.1.6 — A statistic: a number computed from the sample.

Definition 1.1.7: Parameter

A parameter is a numerical characteristic of the whole population that can be estimated by a statistic.

Bad samples make confident lies

A poll that only reaches landline phones, or a survey only of volunteers, can produce precise-looking numbers that badly mislead. Whether a sample truly represents its population is as much an ethical question as a mathematical one — it decides whose voices the conclusions actually reflect.

For example, if we consider one math class to be a sample of the population of all math classes, then the average number of points earned by students in that one math class at the end of the term is an example of a statistic. The statistic is an estimate of a population parameter. Since we considered all math classes to be the population, the average number of points earned per student over all the math classes is an example of a parameter.

One of the main concerns in the field of statistics is how accurately a statistic estimates a parameter. The accuracy really depends on how well the sample represents the population. The sample must contain the characteristics of the population to be a representative sample. We are interested in both the sample statistic and the population parameter in inferential statistics. In a later chapter, we will use the sample statistic to test the validity of the established population parameter.

A variable, usually notated by capital letters such as \(X\) and \(Y\), is a characteristic or measurement that can be determined for each member of a population. Variables may be numerical or categorical. Numerical variables take on values with equal units such as weight in pounds and time in hours. Categorical variables place the person or thing into a category. If we let \(X\) equal the number of points earned by one math student at the end of a term, then \(X\) is a numerical variable. If we let \(Y\) be a person's party affiliation, then some examples of \(Y\) include Republican, Democrat, and Independent. \(Y\) is a categorical variable. We could do some math with values of \(X\) (calculate the average number of points earned, for example), but it makes no sense to do math with values of \(Y\) (calculating an average party affiliation makes no sense).

Definition 1.1.8: Variable

A variable, usually notated by a capital letter such as \(X\) or \(Y\), is a characteristic or measurement that can be determined for each member of a population. A numerical variable takes on values with equal units, such as weight in pounds; a categorical variable places each member into a category, such as a party affiliation.

Definition 1.1.8 — Variable A four-row, two-column roster table: header row "Individual" / "Score, X", then Ana/x_1, Ben/x_2, Cid/x_3, the score column tinted ACCENT to mark it as the variable column. An arrow from a note reading "one variable" points at the "Score, X" header cell. Static figure, no motion. Variable Individual Score, X Ana x1 Ben x2 Cid x3 one variable one characteristic that can be measured for each individual

Definition 1.1.8 — A variable: one characteristic that can be measured for each individual.

Definition 1.1.9: Data

Data are the actual values of the variable. They may be numbers or they may be words. A datum is a single value.

"Mean" vs. "average."

The words "mean" and "average" are often used interchangeably, and among non-statisticians that substitution is common practice. The technical term is "arithmetic mean," while "average" technically refers to any center location — but in everyday use, "average" is accepted shorthand for the arithmetic mean.

Two words that come up often in statistics are mean and proportion. If you were to take three exams in your math classes and obtain scores of 86, 75, and 92, you would calculate your mean score by adding the three exam scores and dividing by three (your mean score would be 84.3 to one decimal place). If, in your math class, there are 40 students and 22 are men, then the proportion of men students is \(\frac{22}{40}\) and the proportion of women students is \(\frac{18}{40}\). Mean and proportion are discussed in more detail in later chapters.

The examples that follow all rehearse the same skill: reading a study and naming its population, sample, parameter, statistic, variable, and data. This is exactly the vocabulary the Learning Objectives promise, and it is worth slowing down here — nearly every chapter that follows leans on your ability to tell the sample's numbers from the population's numbers. In each example, read the study description first and try to label the six key terms yourself before you open the solution. Notice as you go how the parameter and the statistic are always described with the same phrase, differing only in whether that phrase points at the population or at the sample — that parallel wording is not an accident, and spotting it makes the labeling almost automatic.

In statistics, we generally want to study a population. You can think of a population as a collection of persons, things, or objects under study. To study the population, we select a sample. The idea of sampling is to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. Data are the result of sampling from a population.

Taste a spoonful, judge the pot

A cook doesn't drink the whole pot of soup to check the seasoning — one well-stirred spoonful speaks for the pot. Sampling works the same way: a carefully chosen subset can tell you about the entire population, as long as the spoonful truly resembles the pot.

Definition 1.1.9 — Data A four-row, two-column roster table: header row "Individual" / "Score, X", then Ana/86, Ben/75, Cid/92, the score column tinted to mark it as the recorded-values column. The cell holding 92 is circled in ACCENT and an arrow leads to a note reading "a single value = a datum". Static figure, no motion. Data Individual Score, X Ana 86 Ben 75 Cid 92 a single value = a datum the actual recorded values of the variable

Definition 1.1.9 — Data: the actual recorded values of the variable.

Try It Now 1.1.2

Determine what the key terms refer to in the following study. We want to know the average (mean) amount of money spent on school uniforms each year by families with children at Knoll Academy. We randomly survey 100 families with children in the school. Three of the families spent $65, $75, and $95, respectively.

Solution

The population is all families with children attending Knoll Academy.

The sample is the 100 families with children in the school that were randomly surveyed.

The parameter is the average (mean) amount of money spent on school uniforms each year by all families with children at Knoll Academy.

The statistic is the average (mean) amount of money spent on school uniforms each year by the 100 families in the sample.

The variable is the amount of money spent on school uniforms by one family with children at Knoll Academy. Let \(X\) = the amount of money spent on school uniforms each year by one family with children at Knoll Academy.

The data are the dollar amounts spent by the families. Examples of the data are $65, $75, and $95.

Example 1.1.1: School Supplies at ABC College

Determine what the key terms refer to in the following study. We want to know the average (mean) amount of money first year college students spend at ABC College on school supplies that do not include books. We randomly surveyed 100 first year students at the college. Three of those students spent $150, $200, and $225, respectively.

Solution

The population is all first year students attending ABC College this term.

The sample could be all students enrolled in one section of a beginning statistics course at ABC College (although this sample may not represent the entire population).

The parameter is the average (mean) amount of money spent (excluding books) by first year college students at ABC College this term.

The statistic is the average (mean) amount of money spent (excluding books) by first year college students in the sample.

The variable could be the amount of money spent (excluding books) by one first year student. Let \(X\) = the amount of money spent (excluding books) by one first year student attending ABC College.

The data are the dollar amounts spent by the first year students. Examples of the data are $150, $200, and $225.

Try It Now 1.1.3

Determine what the key terms refer to in the following study.

A survey of athletes in a university was conducted to study the heights of athletes, in meters. Fill in the letter of the phrase that best describes each of the items below.

  1. Population
  2. Statistic
  3. Parameter
  4. Sample
  5. Variable
  6. Data
  1. the average height of athletes in the university
  2. the average height of athletes in the survey
  3. all athletes in the university
  4. all students in the university
  5. the height of one athlete
  6. a group of athletes randomly selected
  7. 1.82, 1.76, 1.69, 1.93
Solution

Answer: 1 → c; 2 → b; 3 → a; 4 → f; 5 → e; 6 → g.

  • Population: c — all athletes in the university (not d: the study is about athletes, not all students).
  • Statistic: b — the average height of the athletes in the survey (computed from the sample).
  • Parameter: a — the average height of all athletes in the university.
  • Sample: f — the group of athletes randomly selected.
  • Variable: e — the height of one athlete. Let \(X\) = the height of one athlete in the university.
  • Data: g — the actual measured values: 1.82, 1.76, 1.69, 1.93.
Example 1.1.2: Cumulative GPAs

Determine what the key terms refer to in the following study.

A study was conducted at a local college to analyze the average cumulative GPAs of students who graduated last year. Fill in the letter of the phrase that best describes each of the items below.

  1. Population
  2. Statistic
  3. Parameter
  4. Sample
  5. Variable
  6. Data
  1. all students who attended the college last year
  2. the cumulative GPA of one student who graduated from the college last year
  3. 3.65, 2.80, 1.50, 3.90
  4. a group of students who graduated from the college last year, randomly selected
  5. the average cumulative GPA of students who graduated from the college last year
  6. all students who graduated from the college last year
  7. the average cumulative GPA of students in the study who graduated from the college last year
Solution

Step 1 — Find the population: the group the study is about is everyone who graduated from the college last year — choice f. (Choice a, all students who attended, is broader than the study's target.)

Step 2 — Match the remaining terms by asking, for each phrase, "population or sample?" and "number or description?":

Answer: 1 → f; 2 → g; 3 → e; 4 → d; 5 → b; 6 → c.

  • Population: f — all students who graduated from the college last year.
  • Statistic: g — the average cumulative GPA of the students in the study (the sample's number).
  • Parameter: e — the average cumulative GPA of all students who graduated last year (the population's number).
  • Sample: d — the randomly selected group of graduates.
  • Variable: b — the cumulative GPA of one graduate.
  • Data: c — the actual recorded values: 3.65, 2.80, 1.50, 3.90.
Try It Now 1.1.4

Determine what the key terms refer to in the following study.

A survey is conducted to check the time taken by a mobile phone's battery to charge from 50% to 100%. The criteria used to collect the data are:

Table 1.1.1 — Charging survey criteria.
Wattage of charger usedType of mobile used
30 WAndroid

We want to know the proportion of Android mobiles that are charged to 100% within 30 minutes. We start with a simple random sample of 200 mobiles.

Solution

The population is all Android mobiles charged with a 30 W charger.

The sample is the 200 mobiles, selected by a simple random sample.

The parameter is the proportion of Android mobiles in the population that are charged to 100% within 30 minutes.

The statistic is the proportion of Android mobiles in the sample that are charged to 100% within 30 minutes.

The variable is whether a mobile is charged to 100% within 30 minutes. Let \(X\) = whether one Android mobile is charged to 100% within 30 minutes.

The data are either: yes, charged to 100% within 30 minutes, or no, did not.

Example 1.1.3: Crash Test Dummies

Determine what the key terms refer to in the following study.

As part of a study designed to test the safety of electric automobiles, the National Transportation Safety Board collected and reviewed data about the effects of a crash on test dummies. Here is the criterion they used:

Table 1.1.2 — Crash test criterion.
Speed at which cars crashedLocation of "drivers" (i.e., dummies)
35 miles/hourFront seat

Cars with dummies in the front seats were crashed into a wall at a speed of 35 miles per hour. We want to know the proportion of dummies in the driver's seat that would have had head injuries, if they had been actual drivers. We start with a simple random sample of 75 cars.

Solution

The population is all cars containing dummies in the front seat.

The sample is the 75 cars, selected by a simple random sample.

The parameter is the proportion of driver dummies (if they had been real people) who would have suffered head injuries in the population.

The statistic is the proportion of driver dummies (if they had been real people) who would have suffered head injuries in the sample.

The variable is whether a dummy (if it had been a real person) would have suffered head injuries. Let \(X\) = whether a dummy (if it had been a real person) would have suffered head injuries.

The data are either: yes, had head injury, or no, did not.

Try It Now 1.1.5

Determine what the key terms refer to in the following study.

A study is being conducted by Alex Fuentes, a reporter at a news agency, to find the proportion of all truck drivers that have no points on their license. They randomly select 1,000 truck drivers from the directory of truck drivers and determine the number of truck drivers in the sample who have no points on their license.

Solution

The population is all truck drivers listed in the directory of truck drivers.

The sample is the 1,000 truck drivers Alex selected at random from the directory.

The parameter is the proportion of all truck drivers in the population who have no points on their license.

The statistic is the proportion of truck drivers in the sample who have no points on their license.

The variable is whether a truck driver has no points on their license. Let \(X\) = whether one truck driver in the directory has no points on their license.

The data are either: yes, has no points, or no, has points.

Do the following exercise collaboratively with up to four people per group. Find a population, a sample, the parameter, the statistic, a variable, and data for the following study: You want to determine the average (mean) number of glasses of milk college students drink per day. Suppose yesterday, in your English class, you asked five students how many glasses of milk they drank the day before. The answers were 1, 0, 1, 3, and 4 glasses of milk.

Try it in rāSHio

Open rāSHio, enter your group's five milk answers (1, 0, 1, 3, 4) in the spreadsheet, then choose Stats → Summary Statistics to read off the mean — that number is the statistic your group just identified, not the parameter.

Figure 1.1.3 — Reading off the mean in rāSHio: Stats → Summary Statistics turns the five milk answers into the statistic.

Example 1.1.4: Malpractice Lawsuits

Determine what the key terms refer to in the following study.

An insurance company would like to determine the proportion of all medical doctors who have been involved in one or more malpractice lawsuits. The company selects 500 doctors at random from a professional directory and determines the number in the sample who have been involved in a malpractice lawsuit.

Solution

The population is all medical doctors listed in the professional directory.

The parameter is the proportion of medical doctors who have been involved in one or more malpractice suits in the population.

The sample is the 500 doctors selected at random from the professional directory.

The statistic is the proportion of medical doctors who have been involved in one or more malpractice suits in the sample.

The variable is whether an individual doctor has been involved in a malpractice suit. Let \(X\) = whether one doctor listed in the directory has been involved in a malpractice suit.

The data are either: yes, was involved in one or more malpractice lawsuits, or no, was not.

Try It Now 1.1.6

Your instructor keeps every score for the course in a single spreadsheet — every student, every graded item. Before any of it can be analysed it has to be laid out so that each piece of information has one unambiguous home. How would you arrange the rows, the columns, and the cells?

Solution

Step 1 — Choose what the rows represent: each row holds one individual — here, one student in the course.

Step 2 — Choose what the columns represent: each column is one variable — here, one graded item (Assignment 1, Quiz 1, Exam 1, and so on).

Step 3 — Fill in the cells: the cell where a student's row meets a graded item's column holds that student's score on that item. For example:

StudentAssignment 1Quiz 1Exam 1
Student 1928588
Student 2789081
Student 3857294

Answer: Organize the gradebook so each row is a student, each column is a graded item, and each cell holds the score that student earned on that item. Reading across a row summarizes one student's performance; reading down a column summarizes the class's performance on one item.

1.1.3 Reading a Case Study

A great deal of what you will read about statistics does not arrive as a data set. It arrives as a story about one hospital, one school district, one company, one town. That story is a case study, and knowing how to read one is a skill in its own right — because a case study can be simultaneously true and useless for the question you are asking.

Definition 1.1.10: Case Study

A case study is a detailed examination of a single case — one person, one organization, one event — usually reported in depth rather than summarized as numbers.

One case cannot be a rate

"A student at our college raised her exam average twenty points after switching to handwritten notes" may be entirely accurate and still tell you nothing about whether handwriting helps. You have one student, one semester, and no idea what else changed. The sentence describes an outcome; it does not estimate a rate, and reading it as one is the single most common way people are misled by true statements.

A case study is genuinely good at two things. It can show that something is possible: one documented case of a rare drug reaction proves the reaction can happen, and no amount of averaging can un-prove it. And it can suggest what to measure next, because the person closest to a single case usually notices the variables nobody thought to record.

What a case study cannot do is tell you how often something happens, or whether it would happen again. With one case there is nothing to compare against and no way to separate the effect you care about from everything else that was true of that one situation.

The honest use of a case study is as the first step. It generates the question. The study that answers it needs a sample, a comparison group, and a plan made before the data arrive.

Definition 1.1.10 — Case Study The case-study card (the case: one person, one organization, or one event, reported in depth, not as numbers) beside a fifty-dot population grid labeled the population it came from. The matching dot turns rust as shows it can happen appears, then a 0% to 100% rate axis with a question mark and how often does it happen? fades in, arguing a single case cannot supply a rate. the case one person, one organization, or one event reported in depth, not as numbers the population it came from shows it can happen how often does it happen? ? 0% 100% a case study is one case, examined in depth one documented case is enough to prove the thing is possible with one case there is nothing to compare against — no rate can be read off

Definition 1.1.10 — A case study proves the thing is possible; it cannot say how often it happens.

Try It Now 1.1.7

A local news segment reports: "After the city installed a new bike lane on Rowan Street, one cafe owner says her weekday morning sales rose by nearly forty percent."

a. What is the case here?

b. Name one thing this report does establish.

c. Name one thing a reader might wrongly conclude from it.

d. What would you need to estimate the bike lane's effect on cafe sales across the city?

Solution

Part a — The case. One cafe, on one street, over some unstated stretch of weekday mornings. That is a single unit of observation, not a sample.

Part b — What it establishes. That a sales increase alongside a new bike lane is possible, and that at least one business owner attributes one to the other. It is a real observation worth following up.

Part c — The wrong conclusion. That bike lanes raise cafe sales, or that they raise them by about forty percent. Neither follows. Weekday morning sales move with weather, season, staffing, nearby construction, and a dozen other things nobody recorded, and we have no cafe without a new bike lane to compare against.

Part d — What you would need. Sales data from many cafes, both on streets that got bike lanes and on comparable streets that did not, covering the same period, with the comparison decided before the data were examined. That turns one story into an estimate.

Problem Set 1.1

For each of the following eight exercises, identify: a. the population, b. the sample, c. the parameter, d. the statistic, e. the variable, and f. the data. Give examples where appropriate.

Problem 1. A fitness center is interested in the mean amount of time a client exercises in the center each week.

Solution

Step 1 — Name the group the fitness center cares about: it wants to know about its clients, so that whole group is the population, and a manageable subset of them is the sample.

Step 2 — Sort the two "mean" numbers: the mean for all clients is a population number (parameter); the mean computed from the surveyed clients is a sample number (statistic).

Step 3 — Name the measurement and its values: the thing measured on one client is the variable; the recorded numbers are the data.

Answer:

a. Population — all clients of the fitness center.

b. Sample — a group of the fitness center's clients whose exercise time is actually recorded.

c. Parameter — the mean amount of time all clients exercise in the center each week.

d. Statistic — the mean amount of time the clients in the sample exercise in the center each week.

e. Variable — \(X\) = the amount of time one client exercises in the center in a week.

f. Data — values of \(X\), such as 2 hours, 5 hours, 0 hours.

Problem 2. Ski resorts are interested in the mean age that children take their first ski and snowboard lessons. They need this information to plan their ski classes optimally.

Solution

Step 1 — Name the group: the resorts care about children who take their first ski or snowboard lesson, so that is the population; the children whose ages actually get recorded form the sample.

Step 2 — Sort the two "mean age" numbers: the mean age over all such children is the parameter; the mean age for the recorded group is the statistic.

Step 3 — Name the measurement and its values: the age of one child at their first lesson is the variable; the recorded ages are the data.

Answer:

a. Population — all children who take ski or snowboard lessons.

b. Sample — a group of these children whose ages are recorded.

c. Parameter — the population mean age of children who take their first snowboard lesson.

d. Statistic — the sample mean age of children who take their first snowboard lesson.

e. Variable — \(X\) = the age of one child who takes their first ski or snowboard lesson.

f. Data — values for \(X\), such as 3, 7, and so on.

Problem 3. A cardiologist is interested in the mean recovery period of their patients who have had heart attacks.

Solution

Step 1 — Name the group: the cardiologist studies their own heart-attack patients, so all of those patients are the population; the ones whose recovery is actually tracked are the sample.

Step 2 — Sort the two "mean recovery period" numbers: across all the patients it is a parameter; across the tracked group it is a statistic.

Step 3 — Name the measurement and its values: the recovery period of one patient is the variable; the recorded lengths are the data.

Answer:

a. Population — all of this cardiologist's patients who have had heart attacks.

b. Sample — a group of those patients whose recovery is tracked.

c. Parameter — the mean recovery period of all of the cardiologist's heart-attack patients.

d. Statistic — the mean recovery period of the patients in the sample.

e. Variable — \(X\) = the recovery period of one heart-attack patient.

f. Data — values for \(X\), such as 14 days, 30 days, 45 days.

Problem 4. Insurance companies are interested in the mean health costs each year of their clients, so that they can determine the costs of health insurance.

Solution

Step 1 — Name the group: the insurance companies care about their clients, so all clients form the population and a surveyed subset forms the sample.

Step 2 — Sort the two "mean health cost" numbers: the mean cost over all clients is the parameter; the mean cost over the surveyed clients is the statistic.

Step 3 — Name the measurement and its values: the yearly health cost of one client is the variable; the recorded dollar amounts are the data.

Answer:

a. Population — the clients of the insurance companies.

b. Sample — a group of those clients.

c. Parameter — the mean health costs of all the clients.

d. Statistic — the mean health costs of the clients in the sample.

e. Variable — \(X\) = the health costs of one client.

f. Data — values for \(X\), such as $34, $9, $82, and so on.

Problem 5. A politician is interested in the proportion of voters in their district who think the politician is doing a good job.

Solution

Step 1 — Name the group: the politician cares about the voters in their district, so all of those voters are the population; the voters actually polled are the sample.

Step 2 — Notice this study measures a proportion, not a mean: the proportion approving among all district voters is the parameter; the proportion approving among the polled voters is the statistic.

Step 3 — Name the measurement and its values: here the variable is categorical — whether one voter thinks the politician is doing a good job — and the data are yes/no answers, not numbers.

Answer:

a. Population — all voters in the politician's district.

b. Sample — a group of voters in the district who are polled.

c. Parameter — the proportion of all district voters who think the politician is doing a good job.

d. Statistic — the proportion of the polled voters who think the politician is doing a good job.

e. Variable — \(X\) = whether one voter thinks the politician is doing a good job.

f. Data — yes, no.

Problem 6. A marriage counselor is interested in the proportion of clients they counsel who stay married.

Solution

Step 1 — Name the group: the counselor studies their own clients, so all of them are the population; a subset of those clients is the sample.

Step 2 — Sort the two proportions: the proportion staying married among all clients is the parameter; the proportion among the sampled clients is the statistic.

Step 3 — Name the measurement and its values: whether one client stayed married is the variable, and the data are yes/no answers.

Answer:

a. Population — all the clients of this counselor.

b. Sample — a group of clients of this marriage counselor.

c. Parameter — the proportion of all their clients who stay married.

d. Statistic — the proportion of the sample of the counselor's clients who stay married.

e. Variable — \(X\) = whether a client stayed married.

f. Data — yes, no.

Problem 7. Political pollsters may be interested in the proportion of people who will vote for a particular cause.

Solution

Step 1 — Name the group: the pollsters care about everyone who could vote on the cause, so all of those people are the population; the people actually contacted are the sample.

Step 2 — Sort the two proportions: the proportion supporting the cause among all voters is the parameter; the proportion among those contacted is the statistic.

Step 3 — Name the measurement and its values: whether one person will vote for the cause is the variable; the data are yes/no answers.

Answer:

a. Population — all people eligible to vote on the particular cause.

b. Sample — a group of those people who are contacted by the pollsters.

c. Parameter — the proportion of all those people who will vote for the cause.

d. Statistic — the proportion of the sampled people who will vote for the cause.

e. Variable — \(X\) = whether one person will vote for the particular cause.

f. Data — yes, no.

Problem 8. A marketing company is interested in the proportion of people who will buy a particular product.

Solution

Step 1 — Name the group: the marketing company cares about all potential buyers (perhaps in a certain geographic area), so that is the population; the people actually surveyed are the sample.

Step 2 — Sort the two proportions: the proportion who would buy among all people is the parameter; the proportion among the surveyed people is the statistic.

Step 3 — Name the measurement and its values: whether one person buys the product is the variable, and the data are buy/not buy.

Answer:

a. Population — all people (maybe in a certain geographic area, such as the United States).

b. Sample — a group of those people.

c. Parameter — the proportion of all people who will buy the product.

d. Statistic — the proportion of the sample who will buy the product.

e. Variable — \(X\) = whether a person bought the product.

f. Data — buy, not buy.

Use the following information to answer the next three exercises: A Lake Tahoe Community College instructor, Professor Vang, is interested in the mean number of days Lake Tahoe Community College math students are absent from class during a quarter.

Problem 9. What is the population Professor Vang is interested in?

a) all Lake Tahoe Community College students

b) all Lake Tahoe Community College English students

c) all Lake Tahoe Community College students in his classes

d) all Lake Tahoe Community College math students

Solution

Step 1 — Read what the study is about: Professor Vang wants the mean number of days Lake Tahoe Community College math students are absent. The population is always the whole group the question is about.

Step 2 — Rule out the near-misses: (a) is too broad — it covers every student at the college, not just math students. (b) names the wrong subject entirely. (c) describes only Professor Vang's own classes, which is a sample of the math students, not the population.

Answer: d. all Lake Tahoe Community College math students.

Problem 10. Consider the following: \(X\) = number of days a Lake Tahoe Community College math student is absent. In this case, \(X\) is an example of a:

a) variable.

b) population.

c) statistic.

d) data.

Solution

Step 1 — Read what \(X\) is: \(X\) = the number of days one math student is absent. It is a measurement that can be determined for each member of the population.

Step 2 — Match that to the vocabulary: a characteristic measured on each member is a variable. It is not the population (that is the group of students), not a statistic (that would be a single summary number computed from a sample), and not the data (those are the actual recorded values of \(X\), such as 0, 3, 7).

Answer: a. variable.

Problem 11. Professor Vang's sample produces a mean number of days absent of 3.5 days. This value is an example of a:

a) parameter.

b) data.

c) statistic.

d) variable.

Solution

Step 1 — Ask where the 3.5 came from: it was computed from Professor Vang's sample, not from every math student at the college.

Step 2 — Match that to the vocabulary: a number that summarizes a sample is a statistic. If the same mean had been computed over all math students it would be a parameter; a single student's absence count would be a datum; and the variable is the measurement itself, not this summary number.

Answer: c. statistic.

Key Terms

average — also called mean; a number that describes the central tendency of the data.

categorical variable — a variable that places the person or thing being studied into a category.

data — the actual values of a variable; they may be numbers or words. A single value is a datum.

descriptive statistics — the practice of organizing and summarizing data with graphs and numerical summaries.

inferential statistics — formal methods that use probability to draw conclusions about a population from sample data.

mean — the arithmetic average: the sum of all values divided by how many values there are.

numerical variable — a variable that takes on values with equal units, such as weight in pounds or time in hours.

parameter — a numerical characteristic of the whole population that can be estimated by a statistic.

population — the entire collection of persons, things, or objects under study.

probability — a mathematical tool used to study randomness; the chance of an event occurring.

proportion — the number of successes divided by the total number in the group.

representative sample — a sample that contains the characteristics of the population it was drawn from.

sample — a subset of the larger population, studied to gain information about the population.

statistic — a number that represents a property of the sample; used to estimate a parameter.

variable — a characteristic or measurement, notated \(X\) or \(Y\), that can be determined for each member of a population.

case study — a detailed examination of a single case; it can show that something is possible, but cannot establish how often it happens.