Introduction to Statistics · Chapter 1 · Sampling and Data

Data, Sampling, and Variation in Data and Sampling

The first question a statistician asks is not "what is the answer?" but "what kind of data is this?" — because the answer decides everything you are allowed to do with it.


bookSHelf  ·  Introduction to Statistics  ·  §1.2  ·  a self-paced section

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Learning objectives — by the end of this section you will be able to

Objectives

  1. Sort data into qualitative, quantitative discrete, and quantitative continuous categories Definitions 1.2.1–1.2.4
  2. Choose the right display — pie chart, bar graph, or Pareto chart — for qualitative data Definitions 1.2.5–1.2.7
  3. Describe how a simple random, stratified, cluster, systematic, or convenience sample is drawn Definitions 1.2.8–1.2.12
  4. Explain how sampling bias, sample size, self-selection, and confounding can make a study untrustworthy Definitions 1.2.15–1.2.19
  5. Explain why two honest samples from the same population give different results — and why that is expected, not an error §1.2.7–1.2.9
1.2

§1.2.1 — the first question a statistician asks

You cannot average a list of hair colors, and you cannot make a pie chart out of measured weights. Getting the type right up front saves you from an analysis that looks fine and means nothing.

Almost all data fall into one of two big families: qualitative and quantitative.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

§1.2.1 — the family of categories and labels

Qualitative (categorical) data

Definition 1.2.1 — Qualitative data

Qualitative data are the result of categorizing or describing attributes of a population. They are generally described by words or letters rather than numbers.

Hair color
Blood type
Ethnic group

Hair color, blood type, ethnic group, the car a person drives — these are all qualitative. It does not make sense to find an average hair color or an average blood type.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Context Pause — the same data, two lives

Numbers collected, categories reported

Quiz scores are recorded all term as numbers; at the end of the term they are reported as A, B, C, D, or F. The same underlying reality can be quantitative when gathered and qualitative when published — always ask which version you are holding.

A variable's type is not a fixed property of the thing being measured. It depends on how you record it.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

§1.2.1 — the family of numbers

Quantitative data

Definition 1.2.2 — Quantitative data

Quantitative data are always numbers. They are the result of counting or measuring attributes of a population.

Amount of money
Pulse rate
Weight

Amount of money, pulse rate, weight, the number of people in your town — all quantitative. They split further into two types, depending on whether you counted or measured.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Insight Note — the one question that decides

Counting gives you whole things. Measuring gives you a reading on a scale.

"How many?" is discrete. "How much?" is continuous. You can have 3 phone calls, never 3.4. But a call lasting 3.4 minutes is perfectly ordinary.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

§1.2.1 — counting lands on certain values

Quantitative discrete data

Definition 1.2.3 — Quantitative discrete data

All data that are the result of counting are called quantitative discrete data. These data take on only certain numerical values.

Definition 1.2.3 -- Quantitative Discrete Data A number line runs from 0 to 6 with arrowheads on both ends. A solid dot sits directly on every integer tick. Between each pair of neighboring dots the whole stretch is drawn as a short line segment struck through by an X, showing that no value in that gap is ever counted. Above the line, "counted: 0, 1, 2, 3, ... whole values only" is set in the accent color; below it, "nothing lands in between" is set in the secondary blue. A caption beneath reads "data that come from counting -- only certain values are possible." This is the deliberate counterpart to Definition 1.2.4's unbroken band: isolated values with gaps, not a continuum. Quantitative Discrete Data 0 1 2 3 4 5 6 counted: 0, 1, 2, 3, … whole values only nothing lands in between data that come from counting — only certain values are possible

Definition 1.2.3: counting gives whole numbers — 0, 1, 2, 3 — never 2.4.

If you count phone calls per day, you might get 0, 1, 2, or 3 — never 2.4. The gaps between possible values are real.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

§1.2.1 — measuring can land anywhere

Quantitative continuous data

Definition 1.2.4 — Quantitative continuous data

Data that are not only made up of counting numbers, but that may include fractions, decimals, or irrational numbers, are called quantitative continuous data.

Definition 1.2.4 — Quantitative continuous data A number line from 0 to 6, arrows on both ends. One unbroken accent-colored band rides above the whole line, edge to edge, contrasted with the paired discrete figure (def_1.2.3), whose dots sit only on the integers with the stretches between them struck out. A single measured reading, 2.638, is marked between 2 and 3 with a short tick, a dot and a numeral label, showing a value that lands off the integers. Static figure, no motion. Quantitative Continuous Data 0 1 2 3 4 5 6 2.638 measured: any value on the scale read the scale finer and you get 2.6381 — there is no next value data that come from measuring — fractions and decimals included

Definition 1.2.4: continuous data come from measurements — lengths, weights, times.

A list of call lengths in minutes — 2.4, 7.5, 11.0 — is continuous. With a more precise timer you could write 2.437 or 7.512. The values can be refined without limit.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

The whole classification in one picture

The data-types tree

Figure 1.2.1 — Data-types classification tree Root node Data splits into two branches: Qualitative (categories), with example "backpack color"; and Quantitative (numbers), which splits again into Discrete (counted), example "number of books", and Continuous (measured), example "weight of a backpack". Branch lines are gentle curves; each terminal box connects to its example leaf by a short dashed stub. Data Qualitative(categories) Quantitative(numbers) Discrete(counted) Continuous(measured) "backpack color" "number of books" "weight ofa backpack"

Figure 1.2.1: data-types classification tree — qualitative vs. quantitative, and quantitative split into discrete and continuous.

Every data set you meet from here on lands in exactly one of these three leaves. The tree is your first diagnostic tool.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Your turn — commit to an answer before the reveal

Try It Now 1.2.1

Try It Now 1.2.1 — counting the machines

The data are the number of machines in a gym. You sample five gyms. One gym has 12 machines, one has 15, one has ten, one has 22, and the other has 20. What type of data is this?

Quantitative discrete.

For each gym we counted the machines. A gym has 12 machines or 13 machines — there is no such thing as 12.5 machines, so only whole numbers can occur.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Worked example — counting books

Example 1.2.1 · Books in backpacks

Example 1.2.1 — Data sample of quantitative discrete data

The data are the number of books students carry in their backpacks. You sample five students. Two students carry three books, one carries four, one carries two, and one carries one book. What type of data is this?

Quantitative discrete.

For each student we wrote down how many books — a count. A backpack holds 3 books or 4 books, never 3.4 books. The values are whole numbers with real gaps between them.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Your turn — commit to an answer before the reveal

Try It Now 1.2.2

Try It Now 1.2.2 — measuring the lawns

The data are the areas of lawns in square feet. You sample five houses. The areas of the lawns are 144, 160, 190, 180, and 210 square feet. What type of data is this?

Quantitative continuous.

Area is obtained by measuring, not by counting. The listed areas happen to be whole numbers, but a lawn could just as easily measure 160.5 or 160.47 square feet — landing on whole numbers here is rounding, not a restriction.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Worked example — weighing backpacks

Example 1.2.2 · Backpack weights

Example 1.2.2 — Data sample of quantitative continuous data

The data are the weights of backpacks with books in them. You sample the same five students. The weights (in pounds) are 6.2, 7, 6.8, 9.1, 4.3. What type of data is this?

Quantitative continuous.

Each backpack was put on a scale — a measurement, not a count. Decimals appear, and with a more precise scale we could write 6.24 or 6.238. Two backpacks carrying three books each can still weigh different amounts, because the books themselves differ.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Your turn — all three types in one purchase order

Try It Now 1.2.3

Try It Now 1.2.3 — a purchase manager's order

Grant Halloway must describe this order to his accountant. Name data sets that are quantitative discrete, quantitative continuous, and qualitative.

  • Two types of nails (2 kg box nails, 3 kg roofing nails)
  • One type of oil (4 L machine oil)
  • Four types of screws (3 kg wood, 5 kg machine, 1 kg set, 2 kg socket)
Quantitative discrete — the counts: two types of nails, one type of oil, four types of screws.
Quantitative continuous — the weights and volume: 2 kg, 3 kg, 4 L, 3 kg, 5 kg, 1 kg, 2 kg.
Qualitative — the names: box nails, roofing nails, machine oil, wood screws, machine screws, set screws, socket screws.
Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Worked example — all three types in one trip

Example 1.2.3 · A shopping trip

Example 1.2.3 — A shopping trip with all three types

You buy three cans of soup (19 oz tomato bisque, 14.1 oz lentil, 19 oz Italian wedding), two packages of nuts (walnuts, peanuts), four kinds of vegetable (broccoli, cauliflower, spinach, carrots), and two desserts (16 oz pistachio ice cream, 32 oz chocolate chip cookies). Name data sets that are quantitative discrete, quantitative continuous, and qualitative.

Quantitative discrete — the counts: 3 cans, 2 packages, 4 kinds, 2 desserts.
Quantitative continuous — the weights: 19 oz, 14.1 oz, 16 oz, 32 oz.
Qualitative — the names: tomato bisque, walnuts, broccoli, pistachio ice cream.
Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Your turn — commit to an answer before the reveal

Try It Now 1.2.4

Try It Now 1.2.4 — walking the block

Sarah Whitfield walks her block and records the colors of five houses. The colors are white, yellow, white, red, and white. What type of data is this?

Qualitative (categorical).

Each observation is a color name — a category, not a measurement or a count. Notice you can count how many houses are white (three); that count is a summary of the qualitative data, not the data itself.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Worked example — colors of backpacks

Example 1.2.4 · Backpack colors

Example 1.2.4 — Data sample of qualitative data

The data are the colors of backpacks. You sample five students. One has a red backpack, two have black, one has green, and one has gray. What type of data is this?

Qualitative (categorical).

For each student we wrote down a color name, not a number. Colors describe an attribute — they place each backpack in a category. "The average of red and green" is meaningless.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Your turn — commit to an answer before the reveal

Try It Now 1.2.5

Try It Now 1.2.5 — reading the phrasing

Determine the correct data type for the number of cars in a parking lot. If it is quantitative, say whether it is discrete or continuous.

Quantitative discrete.

“The number of” is the classic signal of a count. A lot holds 40 cars or 41 cars, never 40.7.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Worked example — sorting a mixed list

Example 1.2.5 · Sorting a mixed list

Determine the correct data type for each item. Hint: data that are discrete often start with "the number of."

Items

  1. pairs of shoes you own
  2. type of car you drive
  3. distance to nearest grocery store
  4. classes per school year
  5. type of calculator you use
  6. weights of dogs at a shelter
  7. correct answers on a quiz
  8. IQ scores

Answers

Discrete: a, d, g — counts of shoes, classes, correct answers.
Continuous: c, f, h — distance, weight, IQ (measured on a continuous scale).
Qualitative: b, e — car make, calculator make.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Your turn — reading the data type off a histogram

Try It Now 1.2.6

Try It Now 1.2.6 — credit hours at State University

The registrar keeps records of the number of credit hours students complete each semester, summarized in the histogram. The class boundaries are 10 to less than 13, 13 to less than 16, 16 to less than 19, 19 to less than 22, and 22 to less than 25. What type of data does this graph show?

Figure 1.2.2 — Histogram of the number of credit hours completed per student Five bars of equal width, touching with no gaps, over the interval boundaries 10, 13, 16, 19, 22, 25. Bar heights (number of students): 250, 575, 735, 630, 250 -- peaking in the 16 to 19 interval. Y axis: Number of Students, 0 to 800 by 100. X axis: Credit Hours Completed. 10 13 16 19 22 25 0 100 200 300 400 500 600 700 800 Number of Credit Hours Completed per Student Credit Hours Completed Number of Students

Figure 1.2.2: credit hours completed per student, grouped into five intervals.

Quantitative discrete.

Credit hours sit on the horizontal axis, so the data are numbers, not categories — and they are counted in whole units. The bars are grouped into intervals to make the picture readable; the grouping does not turn the counts into measurements.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Worked example — reading the data type off a graph

Example 1.2.6 · A pie chart of class standing

Example 1.2.6 — Reading the data type off a graph

A professor collects information about the classification of her students as first-year, sophomore, junior, or senior. The data are summarized in a pie chart. What type of data does this graph show?

Figure 1.2.3 -- Classification of Statistics students by class standing Pie chart, four wedges clockwise from 12 o'clock: First-year 67.2%, Sophomore 17.6%, Junior 11.3%, Senior 3.9%. A vertical legend to the right lists all four categories with colour swatches. No percentage labels are printed on the wedges themselves. First-year Sophomore Junior Senior Classification of Statistics Students

Figure 1.2.3: pie chart of student classification — first-year, sophomore, junior, senior.

Qualitative (categorical).

Each wedge is a class standing — a category name. The percentages describe how many students fall in each category; they summarize the data, they are not the data themselves.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Your turn — commit to an answer before the reveal

Try It Now 1.2.7

Try It Now 1.2.7 — variable and parameter

Priya Raman studies every apprentice enrolled in a statewide electrician programme and records, for each one, the number of hours logged before certification. Is that variable numerical (quantitative) or categorical (qualitative)? What is the parameter of interest?

Numerical — and the parameter is the population mean.

Hours are measured on a continuous scale, so the variable is numerical. A parameter describes the whole population: here that is the mean hours to certification across every apprentice in the programme, written μ\mu.

1.2

§1.2.2 — displaying what you cannot average

Qualitative data need their own kind of display. Three charts do the job — and each one answers a different question.

A pie chart promises that every slice is a piece of one indivisible pie. A bar graph makes no such claim.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

§1.2.2 — categories as wedges of one whole

Pie chart

Definition 1.2.5 — Pie chart

In a pie chart, categories of data are represented by wedges in a circle, and the wedges are proportional in size to the percent of individuals in each category.

Definition 1.2.5 -- Pie Chart Pie chart, four wedges clockwise from 12 o'clock: Walk 40.0%, Bus 25.0%, Car 20.0%, Bike 15.0%, each percentage printed inside its own wedge. A vertical legend to the right lists all four categories with colour swatches. Caption below: every individual falls in exactly one wedge, and the wedges close to 100%. 40.0% 25.0% 20.0% 15.0% Walk Bus Car Bike Pie Chart every individual falls in exactly one wedge, and the wedges close to 100%

Definition 1.2.5: each wedge is a piece of one indivisible whole.

A pie chart works only when every individual belongs to exactly one category and all categories are shown. Overlapping categories or missing data break the promise.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Insight Note — the fundamental difference

A pie is a whole. A bar is a ruler.

A pie chart promises that every slice is a piece of one indivisible pie — each individual belongs to exactly one wedge. Overlapping categories break that promise. Bars make no such claim; each bar is just a measurement against the same ruler.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

§1.2.2 — each category measured against one ruler

Bar graph

Definition 1.2.6 — Bar graph

In a bar graph, the length of the bar for each category is proportional to the number or percent of individuals in that category. Bars may be vertical or horizontal.

Definition 1.2.6 -- Bar graph: commute method by percent Four bars in source order (Walk, Bus, Car, Bike) on a 0% to 50% (by 10%) axis: Walk 40%, Bus 25%, Car 20%, Bike 15%. Caption: same data as the pie -- bars claim no whole, they just measure. Same four values as the def_1.2.5 pie and the def_1.2.7 Pareto chart, unsorted here (the def_1.2.7 Pareto sorts them descending). Bar Graph 0% 10% 20% 30% 40% 50% 40% 25% 20% 15% Walk Bus Car Bike same data as the pie — bars claim no whole, they just measure

Definition 1.2.6: each bar is measured independently against the same scale.

A bar graph does not require the categories to add to 100%. That makes it the honest choice when categories overlap or data are missing.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

§1.2.2 — bars sorted by size

Pareto chart

Definition 1.2.7 — Pareto chart

A Pareto chart consists of bars that are sorted into order by category size, from largest to smallest.

Definition 1.2.7 — Pareto chart: the same bars, sorted largest to smallest Four bars over the categories Walk, Bus, Car, Bike in strictly descending order -- 40%, 25%, 20%, 15% -- illustrating the Pareto-chart rule that a bar graph's categories are ranked tallest first. Y axis 0% to 50% by 10%, no axis title (percentages). Pareto Chart sorted: tallest first 0% 10% 20% 30% 40% 50% 40% 25% 20% 15% Walk Bus Car Bike a bar graph with one extra rule — order by size, so the ranking is readable

Definition 1.2.7: the same bars, sorted so the biggest categories stand out.

Sorting by size does not change the data — it changes what you notice first. The Pareto chart answers "which categories matter most?" at a glance.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

§1.2.2 — the pies or the bars: which makes the comparison clearer?

Three displays of the same enrollment

Figure 1.2.4 -- De Anza College enrollment by full-time / part-time status Pie chart, two wedges clockwise from 12 o'clock: full time 40.9% (blue), part time 59.1% (rust). Percentages are printed inside each wedge. A vertical legend to the right lists Part time above Full time with matching colour swatches. Part time Full time 40.9% 59.1% De Anza College

Figure 1.2.4: De Anza, as one pie.

Figure 1.2.5 — Foothill College enrollment, full time vs. part time Pie chart, wedges clockwise from 12 o'clock: full time 28.6 percent, then part time 71.4 percent for the remainder of the circle. Legend lists part time then full time. Foothill College 28.6% 71.4% Part time Full time

Figure 1.2.5: Foothill, as one pie.

Figure 1.2.6 — Student status at De Anza and Foothill, grouped bar graph Two groups of two bars each, De Anza and Foothill, each split into full-time and part-time counts. De Anza full time 9200, part time 13296. Foothill full time 4059, part time 10124. Y axis 0 to 14000 by 2000, no axis title (raw counts). Student Status 0 2000 4000 6000 8000 10000 12000 14000 9200 13296 4059 10124 De Anza Foothill Full time Part time

Figure 1.2.6: both colleges, side by side.

Each display answers a different question. A single pie answers “how is this college split?” The bar graph answers “how do the two compare?” — and it shows sizes while hiding percentages, exactly the trade-off the pies make in reverse.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Pies vs. bars — which tells the clearer story?

Two colleges, one comparison

StatusDe Anza #De Anza %Foothill #Foothill %
Full-time9,20040.9%4,05928.6%
Part-time13,29659.1%10,12471.4%
Total22,496100%14,183100%

Table 1.2.1: full-time and part-time enrollment at De Anza College and Foothill College.

A single pie answers "how is this college split?" The bar graph answers "how do the two compare?" — and reveals that De Anza is the bigger school, which the percentages alone hide.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Your turn — commit to an answer before the reveal

Try It Now 1.2.8

Try It Now 1.2.8 — reading Table 1.2.1 two ways

Using Table 1.2.1, explain why comparing the percent of part-time students at the two colleges tells a different story than comparing the number of part-time students. Which comparison would you use to argue that Foothill serves a mostly part-time student body?

Use the percentages — 71.4%.

By raw count De Anza has more part-time students (13,296 to 10,124). By share, Foothill is far more part-time (71.4% against 59.1%). The colleges have different totals, so counts are pulled around by the size of the school; percentages remove that effect and describe the make-up of each student body.

1.2

§1.2.3 — when the wedges cannot add up

Sometimes percentages add to more than 100% — or less. When categories overlap, a pie chart is not just unhelpful, it is dishonest.

A student can be full-time, under 25, and intending to transfer — one person counted in three categories.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

When one person counts in multiple categories

Overlapping characteristics

Characteristic%
Full-time40.9%
Intend to transfer48.6%
Under age 2561.0%
TOTAL150.5%

Table 1.2.2: categories overlap, so percentages sum to more than 100%.

Figure 1.2.7 -- Bar graph of De Anza student characteristics (Table 1.2.2) Four bars against a shared 0-100% axis, step 20%: under age 25 61.0%, intend to transfer 48.6%, full-time 40.9%, all students 100.0%. The three characteristic bars sum to 150.5% because a single student can sit in more than one category -- the "overlap" is in group membership, not in the drawing, and the bars themselves do not touch or cross. 0% 20% 40% 60% 80% 100% 61.0% 48.6% 40.9% 100.0% Under age 25 Intend to transfer Full-time All students

Figure 1.2.7: a bar graph handles overlapping categories honestly.

A pie chart cannot represent overlapping categories. A bar graph can — each bar is measured independently against the same 0–100% scale.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Your turn — commit to an answer before the reveal

Try It Now 1.2.9

Try It Now 1.2.9 — would a pie chart be honest?

Katherine Bruce runs a campus survey and reports that 55% of students work at least part-time, 62% commute more than 20 minutes, and 30% are parents. She wants to put all three numbers on one pie chart. Would that be an honest display? Explain, and say what she should use instead.

No — use a bar graph.

55%+62%+30%=147%55\% + 62\% + 30\% = 147\%, well over 100%, because the categories overlap: one student can work part-time, commute 25 minutes, and be a parent. A pie chart requires each individual to land in exactly one wedge. A bar graph measures each bar independently against the same 0–100% scale.

1.2

§1.2.4 — what happens when a category is left out

Nearly one student in ten is invisible in the first graph. Omitting a category does not make the people in it disappear — it makes them invisible to the reader.

"Other/Unknown" is bigger than Native American and Pacific Islander combined, and then some.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

§1.2.4 — the frequencies do not add up to the total

Ethnicity at De Anza, with a category missing

EthnicityFrequencyPercent
Asian8,79436.1%
Black1,4125.8%
Filipino1,2985.3%
Hispanic/Latino4,18017.1%
Native American1460.6%
Pacific Islander2361.0%
White5,97824.5%
TOTAL22,044 of 24,38290.4%

Table 1.2.3: ethnicity of De Anza students, most recent fall term, with “Other/Unknown” omitted.

9.6% of students are not on this table

The omitted “Other/Unknown” category holds students who did not feel they fit the listed categories, or who declined to respond. The frequencies stop 2,338 short of the total enrollment.

When the categories do not account for everyone, build a bar graph and not a pie chart — a pie would silently rescale the listed groups so they appear to be the whole college.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Three bar views of the same data — each tells a different story

Ethnicity at De Anza College

Figure 1.2.8 -- Ethnicity of students, Other/Unknown omitted Bar graph, seven categories in alphabetical order: Asian 36.1%, Black 5.8%, Filipino 5.3%, Hispanic/Latino 17.1%, Native American 0.6%, Pacific Islander 1.0%, White 24.5%. Bars sum to 90.4%, not 100%, because the Other/Unknown category (9.6% of students) is left out of this table. Y axis: percent, 0 to 40 by 5. 36.1% 5.8% 5.3% 17.1% 0.6% 1.0% 24.5% Asian Black Filipino Hispanic/ Latino Native American Pacific Islander White 0.0% 5.0% 10.0% 15.0% 20.0% 25.0% 30.0% 35.0% 40.0% Ethnicity of Students

Fig 1.2.8: Other/Unknown omitted.

Figure 1.2.9 -- Ethnicity of De Anza students, bar graph with Other/Unknown restored Eight bars on a 0.0% to 40.0% (by 5.0%) axis: Asian 36.1%, Black 5.8%, Filipino 5.3%, Hispanic/Latino 17.1%, Native American 0.6%, Pacific Islander 1.0%, White 24.5%, Other/Unknown 9.6%. Same seven categories and same axis as Figure 1.2.8, with the Other/Unknown category (9.6%) restored as an eighth bar -- taller than Native American and Pacific Islander combined. 36.1% 5.8% 5.3% 17.1% 0.6% 1.0% 24.5% 9.6% Asian Black Filipino Hispanic/ Latino Native American Pacific Islander White Other/ Unknown 0.0% 5.0% 10.0% 15.0% 20.0% 25.0% 30.0% 35.0% 40.0% Ethnicity of Students

Fig 1.2.9: Other/Unknown included.

Figure 1.2.10 -- Ethnicity of De Anza students, sorted largest to smallest (Pareto order) Eight bars in strictly descending order: Asian 36.1%, White 24.5%, Hispanic/ Latino 17.1%, Other/ Unknown 9.6%, Black 5.8%, Filipino 5.3%, Pacific Islander 1.0%, Native American 0.6%. Y axis 0.0% to 40.0% by 5.0%. Same table and bar colour as Figure 1.2.8's alphabetical bars -- only the order changed, which is the entire point of a Pareto chart. 0.0% 5.0% 10.0% 15.0% 20.0% 25.0% 30.0% 35.0% 40.0% Asian White Hispanic/Latino Other/Unknown Black Filipino PacificIslander NativeAmerican 36.1% 24.5% 17.1% 9.6% 5.8% 5.3% 1.0% 0.6%

Fig 1.2.10: Pareto chart, sorted.

The omitted category (9.6%) is larger than Native American (0.6%) and Pacific Islander (1.0%) combined. A Pareto chart makes the ranking obvious at a glance.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Same data, same chart type — different arrangement

Pie charts: alphabetical vs. sorted

Figure 1.2.11 -- Ethnicity of Students, pie chart in alphabetical order Eight wedges clockwise from 12 o'clock, alphabetical with Other/Unknown appended last: Asian 36.1%, Black 5.8%, Filipino 5.3%, Hispanic/Latino 17.1%, Native American 0.6%, Pacific Islander 1.0%, White 24.5%, Other 9.6%. Values under 8% are labelled outside the circle on a leader line; the rest are labelled inside their own wedge. 36.1% 17.1% 24.5% 9.6% 5.8% 5.3% 0.6% 1.0% Asian Black Filipino Hispanic/Latino Native American Pacific Islander White Other Ethnicity of Students

Figure 1.2.11: wedges in alphabetical order.

Figure 1.2.12 -- Ethnicity of students, wedges re-ordered by size Pie chart, eight wedges clockwise from 12 o'clock, largest to smallest: Asian 36.1%, White 24.5%, Hispanic/Latino 17.1%, Other 9.6%, Black 5.8%, Filipino 5.3%, Pacific Islander 1.0%, Native American 0.6%. The four wedges at 8% share or larger print their percentage inside the wedge; the four thinner wedges print it outside on a hairline leader line. A vertical legend to the right lists all eight categories in that same size-descending order with colour swatches. Paired with Figure 1.2.11, the alphabetical-order version of this same data. 5.8% 5.3% 0.6% 1.0% 36.1% 17.1% 24.5% 9.6% Asian White Hispanic/Latino Other Black Filipino Pacific Islander Native American Ethnicity of Students

Figure 1.2.12: wedges sorted by size — the more informative arrangement.

The data are identical. Only the order changed — and the second chart is far easier to read. Sorting by size is a design choice, not a data change.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Your turn — commit to an answer before the reveal

Try It Now 1.2.10

Try It Now 1.2.10 — what is wrong with the pie?

Club president Camila Reyes reports the majors of her members as Business 40%, Nursing 25%, Engineering 15%. The club has 200 members but only 160 are accounted for in those three categories. What is wrong with drawing a pie chart of the three reported percentages, and how would you fix the display?

The three wedges cover only 80% of the club.

40%+25%+15%=80%40\% + 25\% + 15\% = 80\%, and 200160=40200 - 160 = 40 members — the same 20% gap. A pie chart claims its wedges fill the circle, so drawing only these three would rescale them and hide 40 people. Either add an “Other/Undeclared” wedge of 20%, or use a bar graph, which never promises the categories are exhaustive.

1.2

§1.2.5 — the practical heart of statistics

Gathering information about an entire population often costs too much or is impossible. Instead, we use a sample. A sample should have the same characteristics as the population it represents.

How do you pick a subset that speaks for the whole?

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

§1.2.5 — the gold standard

Simple random sample

Definition 1.2.8 — Simple random sample

In a simple random sample, any group of nn individuals is equally likely to be chosen as any other group of nn individuals. Each sample of the same size has an equal chance of being selected.

Definition 1.2.8 — Simple random sample Definition 1.2.8 — A simple random sample: every individual had an equal chance of being chosen. Population Sample

Definition 1.2.8: every individual had an equal chance of being chosen.

Lisa wants a four-person study group from her class of 31. She puts all 31 names in a hat, shakes it, and picks three. That is a simple random sample — the simplest version of the gold standard.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Try it in rāSHio — no graphing calculator needed

Drawing Lisa's three IDs

Open rāSHio and choose File → Random Numbers… Set the range to 0 through 30 — the two-digit IDs in the class roster — and ask for 3 numbers. Check No repeats so no ID is drawn twice, then click Generate.

Figure 1.2.13: drawing Lisa's three sample IDs in rāSHio.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

§1.2.5 — a little from every group

Stratified sample

Definition 1.2.9 — Stratified sample

To choose a stratified sample, divide the population into groups called strata and then take a proportionate number from each stratum.

Definition 1.2.9 -- Stratified sample The population is split into four strata -- Stratum 1 through Stratum 4 -- each a band of muted dots inside the Population rectangle. An arrow leads from each band to the matching row of an initially empty Sample box on the right. One stratum at a time, a few of its dots brighten to the accent colour and stay lit in place, while copies of them fly across into that stratum's row of the Sample box, shrinking slightly to fit. By the end all four strata have contributed members, and the Sample box shows one small group drawn from every stratum. Population Sample Stratum 1 Stratum 2 Stratum 3 Stratum 4

Definition 1.2.9: split into strata, draw from every one.

Example: stratify your college by department, then take a simple random sample from each department. Every department is represented.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

§1.2.5 — whole groups at random

Cluster sample

Definition 1.2.10 — Cluster sample

To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from the selected clusters are in the sample.

Definition 1.2.10 — Cluster sample Six cluster boxes are shown in a 3x2 grid, each with a few dots inside, all muted. Cluster 3 and Cluster 6 are highlighted, their dots turn the accent color, and arrows lead from each to an empty Sample box. Each highlighted cluster's dots then fly along its arrow into the Sample box, landing as two small groups, while the original dots stay behind in their clusters. Cluster 1 Cluster 2 Cluster 3 Cluster 4 Cluster 5 Cluster 6 Sample

Definition 1.2.10: whole groups chosen at random, every member taken.

Example: randomly pick four departments from your college and survey every faculty member in those four. Entire departments are left out — that is what clustering does.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Insight Note — the difference in one sentence

Strata slice, clusters scoop.

Stratifying takes a little from every group, so every group is represented. Clustering takes everything from a few groups, so entire groups are left out. Both are random — they just randomize different things.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

§1.2.5 — every kth item

Systematic sample

Definition 1.2.11 — Systematic sample

To choose a systematic sample, randomly select a starting point and take every kkth piece of data from a listing of the population.

Definition 1.2.11 — Systematic sample A rectangle labelled "Population (ordered list)" holds two rows of 15 small grey dots (30 total), the ordered list. A dot near the left end of the first row pulses, grows, and settles into the accent colour, with a small "start" tag appearing beneath it -- the random starting point. Every 5th dot after it (skip-counting by k=5) then lights up in the same accent colour, one at a time, left to right, making the fixed spacing visually obvious. Copies of all six lit dots then travel along an arrow into a rounded "Sample" box on the right, landing in a 3-by-2 grid, while the originals remain lit inside the population strip. Population (ordered list) start Sample

Definition 1.2.11: random start, then every kth item.

Example: a phone book of 20,000 listings, need 400 names. Pick a random start between 1 and 50, then take every 50th name. Simple and fast — but watch for hidden patterns in the list.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

§1.2.5 — the non-random method

Convenience sample

Definition 1.2.12 — Convenience sample

A type of sampling that is non-random is convenience sampling, which involves using results that are readily available.

Definition 1.2.12 — Convenience sample A population of thirty dots sits in a labelled box with a researcher marker at its top-left corner. The seven dots nearest that corner highlight in the accent colour, nearest-first, then fly as copies into a Sample box via an arrow. The rest of the population, away from the corner, is never touched -- the bias is the visual. Population Sample Researcher

Definition 1.2.12: only the readily available members are taken.

Example: a software store interviews customers who happen to be browsing. The results may be very good in some cases and highly biased in others. Convenience sampling is the most common — and the most dangerous.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

§1.2.5 — can the same person be chosen twice?

Sampling with and without replacement

Definition 1.2.13 — Sampling with replacement

Once a member is picked, that member goes back into the population and may be chosen more than once.


Definition 1.2.14 — Sampling without replacement

A member of the population may be chosen only once.

Definition 1.2.13 -- Sampling with replacement A rounded rectangle labelled "population of 5" holds five circles numbered 1-5. To its right, three draws are laid out in a row: 3, 1, 3 -- the same ball (3) appears twice, which is only possible because it went back into the urn. A curved arrow in the accent colour runs from the draws down and left into the bottom-right corner of the urn, labelled "goes back in"; beneath the draws, "the same member, twice" calls out the repeat. A caption along the bottom reads: the population is whole again before every draw, so a member may repeat. Companion figure def_1.2.14 (sampling WITHOUT replacement) shares this exact layout -- balls taken there stay dim and out, and no repeat is possible; the return arrow is the one thing that differs between the two, because it IS the definition. Sampling WITH Replacement population of 5 1 2 3 4 5 goes back in three draws 3 1 3 the same member, twice the population is whole again before every draw, so a member may repeat Definition 1.2.14 — Sampling without replacement A rounded urn box labeled "population of 5" holds five numbered balls, 1 through 5, evenly spaced in a row. Balls 1 and 3 are drawn at low opacity and captioned "taken — cannot be drawn again", showing they have left the population and cannot come up again. To the right, under the heading "three draws", three full-opacity balls read 3, 1, 5, every member different, captioned "every member different". A caption beneath the whole figure reads "the population shrinks with every draw, so no member can repeat". Sampling WITHOUT Replacement population of 5 1 2 3 4 5 taken — cannot be drawn again three draws 3 1 5 every member different the population shrinks with every draw, so no member can repeat

Definitions 1.2.13–1.2.14: with replacement (top) and without (bottom).

In practice, most surveys sample without replacement. When the population is large and the sample is small, the two are nearly equivalent — the chance of picking the same person twice is very low either way.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Context Pause — a fact that feels wrong

Big population, same size sample

A sample of 1,000 can serve a population of 100,000 or 100,000,000 equally well.

In Confidence Intervals you will meet formulas that set sample size from the precision you want — not from the size of the population. It feels wrong, and it is true.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

§1.2.5 — two kinds of error, one fixable

Sampling error and nonsampling error

Definition 1.2.15 — Sampling error

An error caused by the actual process of sampling — for example, the sample may not be large enough.


Definition 1.2.16 — Nonsampling error

An error caused by factors not related to the sampling process — for example, a defective counting device.

Definition 1.2.15 — Sampling error Eight sample-estimate dots scatter both above and below a horizontal line marking the true population value, connected to it by dashed stems, illustrating chance scatter that is not a one-sided push and that shrinks as sample size grows. Paired with Definition 1.2.16 (Nonsampling error), which uses the same axis and dot count but pushes every dot the same direction instead. Sampling Error true population value each dot is one sample's estimate a bigger sample pulls the dots in caused by the sampling process itself — it is always present, and it shrinks with n Definition 1.2.16 — Nonsampling error Titled "Nonsampling Error". A rust-colored horizontal line marks the true population value, labelled at its right end. Eight dots sit on dashed stems rising from the line, all on the SAME side of it -- unlike the paired figure for def_1.2.15 (sampling error), where dots scatter on both sides. Text above reads "every estimate high -- the counter is defective"; text below reads "a bigger sample does NOT pull them in". A caption at the bottom reads "caused by something other than sampling -- the offset survives any sample size". The one-sided offset, not the scatter, is the whole point: a larger sample narrows scatter but cannot pull a one-sided offset back toward the truth. Nonsampling Error every estimate high — the counter is defective true population value a bigger sample does NOT pull them in caused by something other than sampling — the offset survives any sample size

Definitions 1.2.15–1.2.16: sampling error scatters; nonsampling error shifts everything one way.

Sampling error shrinks with a larger sample. Nonsampling error does not — a biased scale or a flawed question stays wrong no matter how many people you ask.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

§1.2.5 — the most dangerous kind of error

Sampling bias

Definition 1.2.17 — Sampling bias

A sampling bias is created when a sample is collected from a population and some members of the population are not as likely to be chosen as others.

Definition 1.2.17 — Sampling bias Two labelled boxes, Group A (accent-coloured) and Group B (blue), each holding a scattered cloud of fifteen dots, sit side by side across a dashed vertical divider. Three accent arrows run from the bottom of Group A's box down into an oval labelled Sample, which holds seven accent dots in a row. A single blue arrow runs from Group B's box toward the same Sample, but a large X crosses it out partway down -- Group B never reaches the sample. A caption below reads: "more data cannot fix it -- the sample cannot speak for Group B." Sampling Bias Group A Group B Sample more data cannot fix it — the sample cannot speak for Group B

Definition 1.2.17: one group has no path into the sample.

Unlike sampling error, more data will not fix bias. A biased method run at ten times the scale produces a bigger, more confident wrong answer.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Context Pause — why we hand the choice to chance

Why randomness is the referee

Left to ourselves, we pick the convenient, the nearby, the willing — and those people are not a cross-section of anybody. Handing the choice to chance removes the researcher's thumb from the scale, which is exactly what lets us generalize from a few hundred people to a few million.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Your turn — draw four different samples from the same 60 scores

Try It Now 1.2.11

Table 1.2.5 holds six sets of quiz scores. Every draw below is already done for you — read the scores straight off the table, then compare the four samples.

#1#2#3#4#5#6
5710983
1059876
9108679
91010989
789574
9991087
7710988
8891088
978778
8810987

Table 1.2.5: six sets of quiz scores, ten rows.

The four draws

1 · Stratified by column — rows 2, 5, 9 from #1; 1, 4, 8 from #2; 3, 3, 7 from #3; 2, 6, 10 from #4; 1, 5, 9 from #5; 4, 7, 8 from #6.
2 · Cluster — take whole columns #2 and #5.
3 · Simple random — number the scores 1–60 down the columns; take 3, 7, 11, 14, 19, 22, 27, 31, 36, 40, 44, 49, 53, 57, 60.
4 · Systematic — random start 4, then every tenth score, wrapping past 60.

18, 20, 15, 12 — and none of them match a classmate's.

Stratified = 3×6=183 \times 6 = 18 scores, every column represented. Cluster = 2×10=202 \times 10 = 20 scores, four columns entirely absent. Simple random = 15 scores, columns over-represented purely by chance. Systematic = 12 scores — and with six columns of ten, a step of ten lands in a predictable pattern, a reminder that a systematic sample can sync up with hidden structure in the list.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Worked example — name that method

Example 1.2.7 · Naming the sampling method

A study determines the average tuition San Jose State undergraduates pay. What type of sampling in each case?

Scenarios

  1. Organize by class standing, select 25 from each.
  2. Random start, then every 50th student.
  3. Completely random — every student has equal chance.
  4. Pick two years at random, take all students in them.
  5. Stand in front of the library, ask the first 100.

Answers

1. Stratified — split by class, draw from each.
2. Systematic — random start, fixed step.
3. Simple random — every group equally likely.
4. Cluster — whole groups chosen, everyone inside taken.
5. Convenience — whoever happens to walk past.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Your turn — commit to an answer before the reveal

Try It Now 1.2.12

Try It Now 1.2.12 — name the sampling method

Principal Diane Kessler polls 50 first-year students, 50 sophomores, 50 juniors, and 50 seniors regarding policy changes for after-school activities. She wants every class year represented. Determine the type of sampling used.

Stratified, with class standing as the strata.

The four class standings divide the student body into non-overlapping groups, and a fixed number is taken from every one. Had she randomly picked two of the four classes and polled everyone in them, it would have been a cluster sample instead.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Worked example — six more studies to name

Example 1.2.8 · Naming the sampling method again

Determine the type of sampling used — simple random, stratified, systematic, cluster, or convenience.

Scenarios

  1. Coach Jordan Achebe picks six players aged 8–10, seven aged 11–12, and three aged 13–14.
  2. A pollster interviews all HR personnel in five high-tech companies.
  3. A researcher interviews 50 public and 50 private high school teachers.
  4. A medical researcher interviews every third patient on a hospital list.
  5. A counselor generates 50 random numbers and picks the matching students.
  6. A student interviews classmates in their own algebra class.

Answers

a. Stratified — age brackets are the groups, some from each.
b. Cluster — five whole companies, everyone inside taken.
c. Stratified — two groups, 50 from each.
d. Systematic — a fixed step down an ordered list.
e. Simple random — random numbers, no grouping.
f. Convenience — the most readily available people.

1.2

§1.2.6 — the most important skill in this course

A study can have a perfect sampling method and still produce worthless conclusions. The question is not just "was it random?" but "was it representative?"

A bigger self-selected sample is a worse one, not a better one.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

In class — are these samples representative?

Collaborative Exercise 1

Collaborative Exercise — determine whether each sample is representative

As a class, determine whether or not the following samples are representative. If they are not, name the kind of bias present.

  • To determine the popularity of a new video game, you ask every third person entering a gaming convention.
  • To determine the average GPA at your college, you survey every 10th student in the campus directory.
  • To find out how students feel about a new parking policy, you hand out surveys at the commuter parking lot.
  • To estimate the proportion of left-handed people, you ask your classmates.
Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Your turn — commit to an answer before the reveal

Try It Now 1.2.13

Try It Now 1.2.13 — name three problems

Héctor Ramos, a spokesperson for a shampoo company, presents a study his own company funded: 4,000 volunteers recruited from the company mailing list, of whom 92% saw improvement. Name at least three items from the list that his study trips over, and say whether the large sample size fixes any of them.

Self-funded, self-selected, and confounded — and 4,000 fixes none of it.

The company has a stake in the answer. Volunteers from its own mailing list are both self-selected and a convenience sample. “Saw improvement” is self-reported with no control group, so any change could be confounded with season or expectation. Sample size fixes sampling error, not sampling bias: a biased method run on 4,000 people gives a more precise estimate of the wrong quantity.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Your turn — commit to an answer before the reveal

Try It Now 1.2.14

Try It Now 1.2.14 — what is left for chance to decide?

In a matched pairs design, what is still left for chance to decide — and what goes wrong if the researcher decides it instead?

Randomize the assignment within each pair.

Pairing removes the person-to-person differences, but something still has to decide which member of each pair gets which treatment, or which treatment comes first. If the researcher decides, that choice can track something else — the healthier twin gets the new drug, the well-rested trial goes first. Flipping a coin inside each pair breaks any link between the treatment and hidden characteristics, including order effects such as practice or fatigue.

1.2

§1.2.7 — data vary. That is the point.

If every can of soda held exactly 12 ounces, there would be no need for statistics. Variation is not a mistake — it is the reason the field exists.

Two honest samples from the same population will give different results. That is expected, not an error.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Your turn — ordinary variation, or a broken machine?

Try It Now 1.2.15

Try It Now 1.2.15 — eight 16-ounce cans, measured

Quality inspector Lucía Herrera reviews eight measured cans (in ounces):

15.8   16.1   15.2   14.8   15.8   15.9   16.0   15.5

Explain in one or two sentences why she would not conclude that her filling machine is broken. What would make you suspect it really was?

Scatter on both sides of the target is normal. A pattern is not.

The values run from 14.8 to 16.1, landing above and below 16 by small amounts — the fingerprint of ordinary variation in both the filling and the measuring. A broken machine would show every can consistently under 16 (a systematic shortfall), or a spread so wide that some held 12 ounces and others 19.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Try it in rāSHio — putting a number on the spread

Measuring variation

Open rāSHio, choose File → Delimited List… and paste the eight can amounts (e.g. 11.9, 12.1, 11.8, 12.0, 12.2, 11.9, 12.1, 12.0). Then choose Stats → Summary Statistics to read off the mean and standard deviation — the standard deviation is the number that measures how much the values spread out from the center.

Figure 1.2.14: Stats → Summary Statistics in rāSHio.

1.2

§1.2.8 — samples vary, and that is okay

If you and a classmate each take a simple random sample of the same size from the same population, your two sample means will almost certainly differ. Neither of you made a mistake.

The question is not whether the samples differ. It is whether the difference is small enough to ignore.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Your turn — commit to an answer before the reveal

Try It Now 1.2.16

Try It Now 1.2.16 — did somebody miscount?

Doreen's 500 students average 6.8 hours of sleep; Jung's 500 students average 7.1 hours. A classmate, Sam Whitlock, says one of them must have made a mistake. Respond to Sam, and explain what would happen to the gap if each of them sampled 5,000 students instead.

Neither made a mistake — that is sampling variability.

Doreen and Jung surveyed different students, and two different random samples from the same population almost never produce the same average. The gap is 0.3 hours — about 18 minutes — on a quantity that varies by hours from student to student. With 5,000 each, both means would sit closer to the population average and therefore closer to each other, but they would still not be identical.

1.2

§1.2.9 — bigger is better, but not in the way you think

A larger sample reduces sampling error — but it does nothing for bias. A biased sample of 10,000 people is just a very confident wrong answer.

Sample size determines precision. Sampling method determines accuracy. Never confuse the two.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Your turn — commit to an answer before the reveal

Try It Now 1.2.17

Try It Now 1.2.17 — is this sample representative?

Station manager Rosa Villalobos wants to know whether her 20,000 listeners prefer more music or more talk. She surveys the first 200 people she meets at one of the station's music concerts: 176 say more music, 24 say more talk. Is this sample representative of the whole fan base?

No — the venue picked the answer.

Attending a music concert is evidence that you like music, so listeners who prefer talk shows had almost no chance of being surveyed. 88% preferring music is exactly what a music-concert crowd would say: the result describes concertgoers, not the fan base. Rosa needs a random sample of the whole listener list.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

In groups — see variation with your own hands

Collaborative Exercise 2 · Dice rolls

Collaborative Exercise — two experiments, 20 rolls each

Divide into groups of two, three, or four. Your instructor will give each group one six-sided die. Roll it 20 times and record each face value. Then roll it another 20 times and record those values separately. Compare the two experiments — are the means the same? The frequency of each face?

Two honest rolls of the same die, same number of rolls, different results. That is sampling variation — and it is the reason we need probability to tell us when a difference is meaningful.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Record your rolls — one table per experiment

Two tallies, twenty rolls each

Face on dieFrequency
1 
2 
3 
4 
5 
6 

Table 1.2.6: first experiment (20 rolls).

Face on dieFrequency
1 
2 
3 
4 
5 
6 

Table 1.2.7: second experiment (20 rolls).

Which experiment had the correct results? They both did. The job of the statistician is to see through the variability and draw appropriate conclusions.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Try it in rāSHio — count the faces

Counting die faces in rāSHio

Once both experiments are rolled, type your 40 face values into a column in rāSHio. Then choose Graph → Frequency Table with Discrete values checked to count how many times each face appeared. Compare the two experiments' frequency tables — the variation between them is sampling variation in action.

Figure 1.2.15: counting each die face in rāSHio.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Worked example — 10,000 part-time students, three attempts to describe them

Example 1.2.9 · Three samples, two bad ones

Quinn Marsden wants the average amount a part-time student at ABC College spends on books.

Sample 1 — convenience

Ten students from a first-term organic chemistry class, many also in calculus.

$128 $87 $173 $116 $130 $204 $147 $189 $93 $153

Sample 2 — systematic

Every fifth senior citizen on a list of those taking P.E. classes.

$50 $40 $36 $15 $50 $100 $40 $53 $22 $22

Sample 3 — stratified by discipline

One student drawn at random from each of ten disciplines.

$180 $50 $150 $85 $260 $75 $180 $200 $200 $150

The first two are biased; the third is not.

Sample 1 is science students buying expensive course books — they pay more than the average part-time student. Sample 2 is senior citizens taking courses for interest — they pay far less. In both, not every student had a chance to be chosen, so neither result describes the population. Sample 3 represents every discipline and draws at random within each, so no group is systematically favored — though a larger sample would still be better. Note the asymmetry: with a biased technique, even a large sample risks not being representative.

1.2

§1.2.10 — after you have chosen it

Even a perfectly drawn sample can fail. The people you cannot reach, and the people who volunteer, both break the randomness that made the sample trustworthy in the first place.

A flawless draw can still return a non-random sample.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

§1.2.10 — the two post-sampling failures

Nonresponse and voluntary response bias

Definition 1.2.18 — Nonresponse bias

The people who do not respond are systematically different from those who do — and the sample no longer represents the population.


Definition 1.2.19 — Voluntary response bias

Voluntary response is nonresponse bias with the dial turned to maximum: nobody was drawn — people chose themselves. The people with the strongest opinions are the most likely to respond.

Definitions 1.2.18 and 1.2.19 -- Nonresponse bias and voluntary response bias Two panels labelled (a) nonresponse bias and (b) voluntary response bias, each showing the same 25-person population arranged on a terrible-to-excellent spectrum line. In panel (a), ten dots -- two per column, covering the whole range -- gain an accent-coloured ring, showing a flawless random draw. Then seven of those ten dim away while three fill solid accent colour, all clustered toward the excellent end: what comes back is not what was drawn. In panel (b), nobody is drawn at all; six dots at the two extreme columns fill solid accent colour on their own, while the large middle never volunteers. A shared caption below both panels narrates all four beats in turn and holds on the final one. No dot ever moves -- only fill and stroke change. (a) nonresponse bias terrible excellent (b) voluntary response bias terrible excellent one population, spread across the whole range of opinion (a) the draw is flawless — every part of the range is represented (a) only the repliers come back — and they are not who you drew (b) nobody is drawn — only the strongest opinions volunteer

Definitions 1.2.18–1.2.19: nonresponse and voluntary response bias.

The practical defence is unglamorous: report your response rate. A study that says "72% of those contacted responded" is far more trustworthy than one that does not say.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Context Pause — the counterintuitive truth

A bigger self-selected sample is worse

With sampling error, more data helps. With bias, more data makes the problem worse.

A self-selected poll that gets 10,000 responses is not more accurate than one that gets 100. It is just a more confident picture of the people who care enough to click — which is not the same as the population.

Data, Sampling & Variation · bookSHelf Intro Stats§1.2

Your turn — name the failure, and say who is over-represented

Try It Now 1.2.18

For each situation, say whether the main problem is nonresponse bias, voluntary response bias, or neither.

Situations

  1. A department mails a printed survey to a random sample of 500 alumni. 90 are returned.
  2. A streaming service pops up “Rate this show!” after each episode and reports 4.6 out of 5.
  3. A researcher draws 200 households at random and visits each in person, returning up to three times. She obtains 191 responses.
  4. A radio host asks listeners to call in about a proposed tax; 78% of callers oppose it.

Answers

a. Nonresponse bias. The draw was random, but only 18% replied — skewing toward alumni with a stronger attachment and more time. Recent graduates and the very busy are under-represented.
b. Voluntary response bias. Nobody was sampled. Viewers who disliked the show enough to stop watching never saw the pop-up, which quietly removes the harshest ratings.
c. Neither, or very little. A 95.5% response rate with deliberate callbacks is about as good as field work gets — persistence, not a larger initial sample.
d. Voluntary response bias. Angry callers dial at far higher rates, so 78% measures intensity of feeling among motivated listeners, not opinion among listeners.

1.2

The headline result of §1.2

Sample size buys precision. Only the sampling method buys accuracy.

Sampling error shrinks as the sample grows. Sampling bias does not — it is a property of how the sample was chosen, and it survives any amount of extra data.

This is why a convenience sample of 10,000 is worse than a random sample of 500, and why every study you read should be judged on how it selected people before it is judged on how many it reached.

† A caution for the rest of the course. Every confidence interval and hypothesis test in later chapters computes precision from the sample size. Not one of them can detect that the sample was drawn badly — that check is yours, and it happens before the arithmetic starts.

1.2
Data, Sampling & Variation · bookSHelf Intro Stats§1.2

§1.2 — conclusions

What §1.2 leaves you with

The core idea

The type of data decides what you can do with it. The sampling method decides whether you can trust the result. Variation is not a bug — it is the reason statistics exists. A sample that is not representative is not saved by being large.

The failure case

A convenience sample of 10,000 people, a self-selected online poll with 50,000 responses, a survey that reaches only landline phones — every one produces a perfectly precise number that may be badly wrong. The arithmetic is fine. The conclusion is not.

Next: §1.3 — Frequency, Frequency Tables, and Levels of Measurement, where the data you have classified gets organized into tables and the measurement scale determines what "average" even means.