Introduction to Statistics · Chapter 1 · Sampling and Data

Definitions of Statistics, Probability, and Key Terms

Six words carry this whole course. Learn to tell the number you can compute from the number you actually want.


bookSHelf  ·  Introduction to Statistics  ·  §1.1  ·  a self-paced section

Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

Learning objectives — by the end of this section you will be able to

Objectives

  1. Distinguish descriptive from inferential statistics, and say what each does with data Definitions 1.1.1–1.1.2
  2. Identify the population, sample, parameter, statistic, variable and data in a real study Definitions 1.1.4–1.1.9
  3. Explain how probability measures long-run regularity, even when short-run outcomes are unpredictable Definition 1.1.3
  4. Explain why a sample must be representative before we can trust its conclusions Examples 1.1.1–1.1.4
Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

The first of the two halves of statistics

Organizing what you already have

Definition 1.1.1 — Descriptive statistics

Descriptive statistics is the practice of organizing and summarizing data — through graphs, and through numerical summaries such as an average.

Definition 1.1.1 -- Descriptive Statistics A loose scattered row of seven small dots is labelled "data" beneath it. Two arrows fan out downward from the data, one to a left panel and one to a right panel. The left panel is a rounded rectangle holding a small five-bar chart standing on a baseline, labelled "graphs" underneath. The right panel is a rounded rectangle holding the numerical expression x-bar equals 84.3, labelled "numerical summaries" underneath. A caption below reads "two ways to summarize the same data" -- the same raw data organized into a graph on one side and summarized as a number on the other. Descriptive Statistics data graphs
numerical summaries two ways to summarize the same data

Definition 1.1.1: descriptive statistics turns a pile of values into a picture and a number.

Nothing here reaches beyond the data in front of you. Descriptive statistics never claims anything about anyone you did not measure.

Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

The second half — and the reason for the rest of the course

Reaching past what you measured

Definition 1.1.2 — Inferential statistics

Inferential statistics consists of formal methods for drawing conclusions about a population from "good" data. It uses probability to say how confident we can be that those conclusions are correct.

The science of statistics is the collection, analysis, interpretation and presentation of data — batting averages, opinion polls, weather forecasts and your own gradebook are all statistics at work.

Definition 1.1.2 — Population, sample, and inference Left: a large circle labeled Population holding 20 blue dots. Right: a smaller circle labeled Sample, empty at first. A solid arrow labeled take a sample grows left to right between them. Five of the population dots briefly pulse larger and rust-colored, then five dots drift from those same positions into the sample circle, turning rust-colored (the sample). A solid arrow labeled estimate the parameter then draws right to left, back from the sample into the population (simplified from the manim original's dashed line — see the CSS comment on .est-shaft for why). Finally captions fade in below each circle: Parameter, a fact about the whole population under Population; Statistic, a fact about the sample under Sample. Population Sample take a sample estimate the parameter Parameter a fact about the whole population Statistic a fact about the sample

Definition 1.1.2: take a sample from the population, then estimate the parameter from the statistic.

The goal is not to run the formulas. It is to understand your data — the calculator can do the arithmetic, the understanding has to come from you.

Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

In class — everybody contributes one number

How much do we sleep?

Collaborative Exercise — build a dot plot from your own class

Have every class member write down the average time (in hours, to the nearest half-hour) they sleep per night. Your instructor records the data. Then build a dot plot — a number line with a dot stacked above each value. For example:

5; 5.5; 6; 6; 6; 6.5; 6.5; 6.5; 6.5; 7; 7; 8; 8; 9

  • Does your dot plot look the same as this one, or different — and why?
  • Would an English class of the same size give the same result? Why or why not?
  • Where do your data cluster, and how would you interpret that clustering?
Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

Descriptive statistics, in one picture

The dot plot

Fourteen students, fourteen dots. Each value gets a dot stacked above its place on the number line, so the height of a stack is how many people gave that answer — and the shape of the stacks is where the class clusters.

Figure 1.1.1 — Dot plot of average nightly sleep hours A semicolon-separated data row above a bare number line from 5 to 9. Each listed value in turn brightens, then travels down and settles as an open circle stacked on its value on the number line, while the list entry dims to mark it used. Ends with the finished dot plot and every list entry dimmed. Data: 5; 5.5; 6; 6; 6; 6.5; 6.5; 6.5; 6.5; 7; 7; 8; 8; 9 5 6 7 8 9

Figure 1.1.1: average time (in hours) spent sleeping per night, for a class of 14 students.

You have not concluded anything about students in general — only summarized these fourteen. That is exactly the line between descriptive and inferential.

Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

Try it in rāSHio — do it with your own class list

Counting the values

Open rāSHio and paste your class's sleep-hours list (5; 5.5; 6; 6; …) straight into File → Delimited List… — the semicolons parse as-is. Then choose Graph → Frequency Table with Discrete values checked to count how many classmates gave each value.

Figure 1.1.2: counting each sleep-hours value in rāSHio — Graph → Frequency Table, Discrete values checked.

The frequency table's counts are the same numbers as the heights of the stacks in Figure 1.1.1 — one summary in a table, the other in a picture.

1.1

§1.1.1 — the tool that makes inference possible

Inference claims something about people you never measured. What licenses that leap?

If a single outcome is unpredictable, how can many outcomes be dependable enough to bet on?

Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

§1.1.1 — the mathematics of randomness

What probability measures

Definition 1.1.3 — Probability

Probability is a mathematical tool used to study randomness. It deals with the chance — the likelihood — of an event occurring.

Toss a fair coin four times and you may well not get two heads and two tails. Toss the same coin 4,000 times and the outcome will be close to half and half. The theoretical probability of heads on any one toss is 12\frac{1}{2}.

Definition 1.1.3 — Probability: the long-run frequency of heads settles near 0.5 An oscillating curve for the fraction of heads across up to 2,000 coin tosses. It starts near 0.8, swings above and below a dashed reference line at 0.5 with shrinking amplitude, and settles onto that line as the toss count grows. Probability: the long run settles 0.5 1.0 0.5 number of tosses fraction of heads one toss is unpredictable — 2,000 tosses are not

Definition 1.1.3: the fraction of heads is unpredictable over a few tosses, but settles near 0.5 over thousands.

Probability is also how the rest of the course gets its warrants: earthquakes, rainfall, a vaccine's side effects, a portfolio's rate of return — every one is a prediction stated as a probability.

Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

Insight Note — unpredictable in the short run, dependable in the long run

What actually happens when you keep tossing

Who tossedTossesHeadsFraction heads
a few flips4anythingunpredictable
one of the authors2,0009960.498
Karl Pearson24,00012,0120.5005
theory0.5

Coin-toss results: 9962000=0.498\frac{996}{2000}=0.498, and Pearson's 12,01224,000=0.5005\frac{12{,}012}{24{,}000}=0.5005.

One toss tells you nothing. Four thousand settle reliably near half heads. Probability is the mathematics of that long-run pattern — it cannot call the next flip, yet it predicts the crowd of flips with remarkable accuracy.

Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

Your turn — commit to an answer before the reveal

Try It Now 1.1.1

Try It Now 1.1.1 — seven heads in ten tosses

You toss a fair coin 10 times and get 7 heads. Your friend Emma claims this proves the coin is unfair. Is she right? About what fraction of heads would you expect if you tossed the same coin 10,000 times?

Emma is not right.

Ten tosses is a small number of repetitions, and short runs of a random process are naturally uneven — 7 heads in 10 is not unusual for a fair coin. Over 10,000 tosses you would expect the fraction of heads very close to 0.50.5 (about 5,000 heads), just as Pearson's 24,000 tosses gave 12,012 heads, a fraction of 0.5005.

Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1
1.1

§1.1.2 — the vocabulary the whole course leans on

Six words, one study

Every study you meet from here on can be taken apart with the same six labels. Read the study, then name each one:

  1. Population — everyone or everything the study is about Definition 1.1.4
  2. Sample — the subset you actually examine Definition 1.1.5
  3. Statistic — a number computed from the sample Definition 1.1.6
  4. Parameter — the matching number for the whole population Definition 1.1.7
  5. Variable — the one characteristic measured per member Definition 1.1.8
  6. Data — the actual recorded values Definition 1.1.9
Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

§1.1.2 — the group the study is about

Population

Definition 1.1.4 — Population

A population is the entire collection of persons, things, or objects under study.

Definition 1.1.4 -- Population A bold heading reads Population. Below it, a large circle outlined and lightly tinted in the CURVE tone is labelled "every member under study" in the ACCENT tone above it. The circle is filled with a grid-with-jitter cloud of 34 small solid dots in the CURVE tone, representing every member of the population. A caption below reads "the entire collection of persons, things, or objects under study". The frame is deliberately a circle, not a box, matching the same population/sample visual convention used in Definitions 1.1.2 and 1.1.5. Population every member under study the entire collection of persons, things, or objects under study

Definition 1.1.4: the entire collection of persons, things, or objects under study.

The population is set by the question, not by convenience. "All athletes at the university" and "all students at the university" are different populations, and picking the wrong one invalidates everything downstream.

Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

§1.1.2 — the part you actually examine

Sample

Definition 1.1.5 — Sample

A sample is a portion (or subset) of the larger population, selected so that studying it yields information about the whole population.

Examining a whole population costs time and money, so sampling is the practical technique: a sample of students for a campus-wide GPA, an opinion poll of 1,000–2,000 people for a country, a handful of cans to check whether a 16-ounce can really holds 16 ounces.

Definition 1.1.5 — Sample A bold heading reads Sample. A large circle on the left is labelled Population and filled with a grid-clipped-to-a-disc cloud of 29 small dots in the CURVE tone. Five of those dots, scattered rather than neighbouring, light up one at a time in the ACCENT tone and grow slightly, showing a random draw rather than a convenient contiguous cluster. A black arrow labelled "choose a subset" grows from the Population circle toward a smaller circle on the right labelled Sample. Five accent-coloured copies of the chosen dots then drift out of the Population and into the Sample circle -- the originals stay behind, since sampling observes the population rather than emptying it. A caption reads "a portion (subset) of the population that we actually examine". Sample Population Sample choose a subset a portion (subset) of the population that we actually examine

Definition 1.1.5: a subset of the population, studied to learn about the whole.

Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

§1.1.2 — the number you can actually compute

Statistic

Definition 1.1.6 — Statistic

A statistic is a number that represents a property of the sample.

Definition 1.1.6 — Statistic Left: a rounded rectangle outlined in CURVE, labelled "Sample" above it, holding the three stacked values 86, 75, 92. A thick compute arrow in INK crosses to a rounded rectangle outlined in ACCENT on the right, holding the expression x-bar equals 84.3 in ACCENT, labelled "one number" beneath it. Caption below: "a number that describes the sample". Static figure, no motion. Statistic Sample 86 75 92 compute
one number a number that describes the sample

Definition 1.1.6: a number computed from the sample.

If you can compute it from the data you collected, it is a statistic. That is the whole test.

Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

§1.1.2 — the number you actually want

Parameter

Definition 1.1.7 — Parameter

A parameter is a numerical characteristic of the whole population that can be estimated by a statistic.

Treat one math class as a sample of all math classes. The average points earned in that one class is a statistic. The average points earned across all math classes is the parameter it estimates.

Notice the two are described with the same phrase, differing only in whether it points at the sample or at the population. That parallel wording is not an accident — spotting it makes the labeling almost automatic.

1.1

The headline result of §1.1

A statistic estimates a parameter only as well as the sample represents the population.

The accuracy of the estimate does not depend on how carefully you did the arithmetic. It depends on whether the sample carries the characteristics of the population — a representative sample.

How accurately a statistic estimates its parameter is one of the main concerns of the entire field. Later chapters use the sample statistic to test a claimed population parameter — and every one of those tests assumes this.

Context Pause — bad samples make confident lies. A poll that only reaches landline phones, or a survey only of volunteers, produces precise-looking numbers that badly mislead. Whether a sample truly represents its population is as much an ethical question as a mathematical one — it decides whose voices the conclusions actually reflect.

Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

§1.1.2 — the one thing you measure per member

Variable

Definition 1.1.8 — Variable

A variable, usually notated by a capital letter such as XX or YY, is a characteristic or measurement that can be determined for each member of a population.

A numerical variable takes values with equal units — weight in pounds, time in hours. A categorical variable places each member in a category.

Definition 1.1.8 — Variable A four-row, two-column roster table: header row "Individual" / "Score, X", then Ana/x_1, Ben/x_2, Cid/x_3, the score column tinted ACCENT to mark it as the variable column. An arrow from a note reading "one variable" points at the "Score, X" header cell. Static figure, no motion. Variable Individual Score, X Ana x1 Ben x2 Cid x3 one variable one characteristic that can be measured for each individual

Definition 1.1.8: one characteristic that can be measured for each individual.

Let XX = points earned by one math student: numerical, and averaging it means something. Let YY = a person's party affiliation: categorical, and an "average party affiliation" means nothing at all.

Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

§1.1.2 — what you actually write down

Data

Definition 1.1.9 — Data

Data are the actual values of the variable. They may be numbers or they may be words. A single value is a datum.

Definition 1.1.9 — Data A four-row, two-column roster table: header row "Individual" / "Score, X", then Ana/86, Ben/75, Cid/92, the score column tinted to mark it as the recorded-values column. The cell holding 92 is circled in ACCENT and an arrow leads to a note reading "a single value = a datum". Static figure, no motion. Data Individual Score, X Ana 86 Ben 75 Cid 92 a single value = a datum the actual recorded values of the variable

Definition 1.1.9: the actual recorded values of the variable.

Data are the result of sampling from a population. Everything else in this course — every graph, every average, every test — is computed from them.

Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

Two words that come up constantly

Mean and proportion

Three exam scores of 86, 75 and 92 give a mean of

86+75+923=84.3 (to one decimal place).\frac{86+75+92}{3}=84.3\ \text{(to one decimal place).}

In a class of 40 students, 22 of them men, the proportion of men is 2240\frac{22}{40} and the proportion of women is 1840\frac{18}{40}.

Context Pause — "mean" vs. "average." The two words are used interchangeably, and among non-statisticians that substitution is common practice. The technical term is arithmetic mean, while average technically refers to any center location — but in everyday use, "average" is accepted shorthand for the arithmetic mean.

1.1

Insight Note — why sampling works at all

Taste a spoonful, judge the pot.

A cook doesn't drink the whole pot of soup to check the seasoning — one well-stirred spoonful speaks for the pot.

Sampling works the same way: a carefully chosen subset can tell you about the entire population — as long as the spoonful truly resembles the pot. "Well-stirred" is doing all the work in that sentence.

Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

Your turn — name all six before the reveal

Try It Now 1.1.2

Try It Now 1.1.2 — school uniforms at Knoll Academy

We want to know the average (mean) amount of money spent on school uniforms each year by families with children at Knoll Academy. We randomly survey 100 families with children in the school. Three of the families spent $65, $75 and $95.

Population — all families with children attending Knoll Academy.
Sample — the 100 families randomly surveyed.
Parameter — mean uniform spending across all those families.
Statistic — mean uniform spending across the 100 surveyed.
VariableXX = uniform spending by one family.
Data — the dollar amounts: $65, $75, $95, …
Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

Worked example — the same six labels

Example 1.1.1 · School supplies at ABC College

Example 1.1.1 — school supplies at ABC College

We want to know the average (mean) amount of money first year students at ABC College spend on school supplies that do not include books. We randomly surveyed 100 first year students. Three of them spent $150, $200 and $225.

Population — all first year students attending ABC College this term.
Sample — the 100 first year students surveyed.
Parameter — mean spending (excluding books) by all first year students this term.
Statistic — mean spending (excluding books) by the students in the sample.
VariableXX = money spent (excluding books) by one first year student.
Data — the dollar amounts: $150, $200, $225, …
Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

Your turn — match each term to a phrase

Try It Now 1.1.3

A survey of athletes in a university studied the heights of athletes, in meters. Match each term to the phrase that describes it.

Terms

  1. Population
  2. Statistic
  3. Parameter
  4. Sample
  5. Variable
  6. Data

Phrases

  1. the average height of athletes in the university
  2. the average height of athletes in the survey
  3. all athletes in the university
  4. all students in the university
  5. the height of one athlete
  6. a group of athletes randomly selected
  7. 1.82, 1.76, 1.69, 1.93

1 → c  ·  2 → b  ·  3 → a  ·  4 → f  ·  5 → e  ·  6 → g

Not d for the population — the study is about athletes, not all students. The statistic (b) is the survey's average; the parameter (a) is the university's.

Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

Worked example — matching, with the reasoning shown

Example 1.1.2 · Cumulative GPAs

A study at a local college analyzed the average cumulative GPAs of students who graduated last year.

Terms

  1. Population
  2. Statistic
  3. Parameter
  4. Sample
  5. Variable
  6. Data

Phrases

  1. all students who attended the college last year
  2. the cumulative GPA of one student who graduated last year
  3. 3.65, 2.80, 1.50, 3.90
  4. a randomly selected group of last year's graduates
  5. the average cumulative GPA of students who graduated last year
  6. all students who graduated from the college last year
  7. the average cumulative GPA of the students in the study

1 → f  ·  2 → g  ·  3 → e  ·  4 → d  ·  5 → b  ·  6 → c

Find the population first: the study is about graduates (f), not everyone who attended (a). Then ask of each phrase — population or sample? number or description?

Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

Your turn — a proportion, not a mean

Try It Now 1.1.4

Try It Now 1.1.4 — charging from 50% to 100%

A survey checks how long a mobile phone's battery takes to charge from 50% to 100%, collected under one criterion: a 30 W charger on an Android phone. We want the proportion of Android phones charged to 100% within 30 minutes, starting from a simple random sample of 200 phones.

Population — all Android phones charged with a 30 W charger.
Sample — the 200 phones in the simple random sample.
Parameter — the proportion in the population charged to 100% within 30 min.
Statistic — the proportion in the sample charged to 100% within 30 min.
VariableXX = whether one phone charges to 100% within 30 min.
Data — yes / no for each phone.
Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

Worked example — when the "individuals" aren't people

Example 1.1.3 · Crash test dummies

Example 1.1.3 — testing the safety of electric automobiles

The National Transportation Safety Board crashed cars into a wall at 35 miles per hour with dummies in the front seat. We want the proportion of driver dummies that would have suffered head injuries, had they been real drivers. We start from a simple random sample of 75 cars.

Population — all cars containing dummies in the front seat.
Sample — the 75 cars in the simple random sample.
Parameter — the proportion of driver dummies in the population that would have had head injuries.
Statistic — the same proportion, in the sample.
VariableXX = whether one dummy would have suffered a head injury.
Data — yes, had head injury / no, did not.
Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

Your turn — watch the population's boundary

Try It Now 1.1.5

Try It Now 1.1.5 — points on a license

Alex Fuentes, a reporter, wants the proportion of all truck drivers with no points on their license. They randomly select 1,000 truck drivers from the directory of truck drivers and count how many in the sample have no points.

Population — all truck drivers listed in the directory.
Sample — the 1,000 drivers Alex selected at random.
Parameter — the proportion with no points in the population.
Statistic — the proportion with no points in the sample.
VariableXX = whether one driver has no points on their license.
Data — yes, no points / no, has points.

The population is the directory, not "all truck drivers alive" — a sample can only speak for the list it was drawn from.

Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

In groups of up to four — all six terms, one small study

Glasses of milk

Collaborative Exercise — find all six terms

Find a population, a sample, the parameter, the statistic, a variable and data for this study: you want the average (mean) number of glasses of milk college students drink per day. Yesterday, in your English class, you asked five students how many glasses they drank the day before. The answers were:

1, 0, 1, 3, 4

Argue about the population before you compute anything — "college students" is a much larger group than "students in one English class," and which one you claim decides whether your mean is worth anything.

Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

Try it in rāSHio — compute your group's number

Reading off the mean

Open rāSHio, enter your group's five milk answers (1, 0, 1, 3, 4) in the spreadsheet, then choose Stats → Summary Statistics to read off the mean — that number is the statistic your group just identified, not the parameter.

Figure 1.1.3: Stats → Summary Statistics turns the five milk answers into the statistic.

The tool will happily give you a mean for any five numbers. It cannot tell you whether those five people represent college students — that judgement is still yours.

Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

Worked example — the last of the four

Example 1.1.4 · Malpractice lawsuits

Example 1.1.4 — an insurance company's question

An insurance company wants the proportion of all medical doctors who have been involved in one or more malpractice lawsuits. It selects 500 doctors at random from a professional directory and counts how many in the sample have been involved in such a suit.

Population — all medical doctors listed in the professional directory.
Sample — the 500 doctors selected at random from it.
Parameter — the proportion involved in a malpractice suit in the population.
Statistic — the same proportion, in the sample.
VariableXX = whether one listed doctor has been involved in a malpractice suit.
Data — yes, was involved / no, was not.
Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

Your turn — where data actually live

Try It Now 1.1.6

Your instructor keeps every score for the course in one spreadsheet — every student, every graded item. Before any of it can be analysed it has to be laid out so each piece of information has one unambiguous home. How would you arrange the rows, the columns and the cells?

  1. Each row is one individual — one student in the course.
  2. Each column is one variable — one graded item.
  3. Each cell holds that student's datum for that item.
StudentAssignment 1Quiz 1Exam 1
Student 1928588
Student 2789081
Student 3857294

Reading across a row summarizes one student; reading down a column summarizes the class on one item.

1.1

§1.1.3 — statistics that arrive as a story

A great deal of what you read about statistics does not arrive as a data set. It arrives as a story about one hospital, one school district, one company, one town.

How can a story be entirely true and still be useless for the question you're asking?

Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

§1.1.3 — what one case can and cannot do

Case study

Definition 1.1.10 — Case study

A case study is a detailed examination of a single case — one person, one organization, one event — usually reported in depth rather than summarized as numbers.

It is good at two things: showing that something is possible (one documented rare drug reaction proves the reaction can happen), and suggesting what to measure next.

Definition 1.1.10 — Case Study The case-study card (the case: one person, one organization, or one event, reported in depth, not as numbers) beside a fifty-dot population grid labeled the population it came from. The matching dot turns rust as shows it can happen appears, then a 0% to 100% rate axis with a question mark and how often does it happen? fades in, arguing a single case cannot supply a rate. the case one person, one organization, or one event reported in depth, not as numbers the population it came from shows it can happen how often does it happen? ? 0% 100% a case study is one case, examined in depth one documented case is enough to prove the thing is possible with one case there is nothing to compare against — no rate can be read off

Definition 1.1.10: a case study proves the thing is possible; it cannot say how often it happens.

What it cannot do is tell you how often something happens, or whether it would happen again — with one case there is nothing to compare against. The honest use is as the first step: it generates the question; the study that answers it needs a sample, a comparison group, and a plan made before the data arrive.

Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

Your turn — read the story like a statistician

Try It Now 1.1.7

Try It Now 1.1.7 — a local news segment

"After the city installed a new bike lane on Rowan Street, one cafe owner says her weekday morning sales rose by nearly forty percent."

  1. What is the case here?
  2. Name one thing this report does establish.
  3. Name one thing a reader might wrongly conclude from it.
  4. What would you need to estimate the bike lane's effect on cafe sales across the city?
1.1
Statistics, Probability & Key Terms · bookSHelf Intro Stats§1.1

§1.1 — conclusions

What §1.1 leaves you with

The core idea

A statistic is the number you can compute, from the sample you examined. A parameter is the number you actually want, about the population you care about. Descriptive statistics stops at the first; inferential statistics uses probability to reach the second.

The failure case

A sample that does not represent its population still produces a perfectly precise statistic — and a badly wrong estimate. Volunteers, landlines, one cafe on Rowan Street: the arithmetic is fine, the conclusion is not.

Next: §1.2 — Data, Sampling, and Variation in Data and Sampling, where "representative" stops being an adjective and becomes a procedure.