Introduction to Statistics · Chapter 1 · Sampling and Data

Experimental Design and Ethics

A study earns the right to say "caused" only by how it was built. This section is the blueprint — and the rules about what you may not do to the people in it.


bookSHelf  ·  Introduction to Statistics  ·  §1.4  ·  a self-paced section

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

Learning objectives — by the end of this section you will be able to

Objectives

  1. Identify the population, sample, experimental units, explanatory variable, response variable and treatments in a described study Definitions 1.4.1–1.4.4
  2. Explain how random assignment removes lurking variables — and recognize a question it cannot answer Definitions 1.4.5–1.4.6
  3. Describe how a control group, a placebo, and blinding protect a study from the power of suggestion Definitions 1.4.9–1.4.12
  4. Assess whether a design supports a cause-and-effect claim or only a weaker one Examples 1.4.1–1.4.3
  5. Evaluate the ethics of a study — informed consent, data honesty, privacy Definitions 1.4.13–1.4.14
1.4

§1.4.1 — four labels, and every study fits them

The purpose of an experiment is to investigate the relationship between two variables: the one you push on, and the one you watch.

Before you look at a single result — can you name what was set, what was measured, and who it was done to?

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4.1 — the variable you set

Explanatory variable

Definition 1.4.1 — Explanatory variable

When one variable causes a change in another, the first is the explanatory variable. In a randomized experiment the researcher deliberately sets its values.

Definition 1.4.1 — Explanatory Variable A dial on the left steps through three settings - none, low dose, full dose - while a response bar on the right rises to match each setting the dial is moved to. A forward arrow labelled "causes a change in" runs from the dial to the meter throughout. After the dial reaches full dose, a second arrow fades in beneath running the other way, labelled "never the other way", and is struck through with an X to show the causal link is one-directional. none low dose full dose explanatory variable you set this causes a change in response variable you measure this never the other way

Definitions 1.4.1 and 1.4.2: the explanatory variable is the one you set; the response is the one you read off.

"Deliberately sets" is the load-bearing phrase. A variable you merely observed people having is not the same thing, and the study built on it cannot make the same claim.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4.1 — the variable you measure

Response variable

Definition 1.4.2 — Response variable

The variable that is affected by the explanatory variable is the response variable. The researcher does not set this one — they measure it, and see whether it moved.

One experiment, one thing pushed and one thing watched. If you cannot say which is which, the study has not been designed yet.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4.1 — the specific values you hand out

Treatment

Definition 1.4.3 — Treatment

The different values of the explanatory variable that the researcher assigns are called the treatments.

Definition 1.4.3 — Treatment A single box, "explanatory variable: oral medication", sits above three empty chips: aspirin, placebo, no pill. Three arrows grow from the box down to the chips, one per chip. Then twelve dots -- the experimental units -- fan out from the box and settle four to a chip, labelled "experimental units". The figure closes on "three treatments = three values of ONE variable", then holds and loops. explanatory variable: oral medication aspirin placebo no pill experimental units three treatments = three values of ONE variable

Definitions 1.4.3 and 1.4.4: the treatments are the values of one explanatory variable, handed out to the experimental units.

One explanatory variable, several treatments. "Aspirin or placebo" is one variable with two treatments — not two variables.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4.1 — who the study is done to

Experimental unit

Definition 1.4.4 — Experimental unit

An experimental unit is a single object or individual to be measured — one person, one rose bush, one rat, one plot of soil.

Insight Note — the light switch and the light bulb

The explanatory variable is the switch you flip; the response variable is the bulb you watch. Treatments are the switch positions you try.

If you never touch the switch, you never learn whether it is wired to that bulb at all.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4.1 — reading a study before you read its result

How the four terms stack

  1. Experimental unitswho is in the study Definition 1.4.4
  2. Treatmentswhat you do to them Definition 1.4.3
  3. Explanatory variable — the general thing the treatments are versions of Definition 1.4.1
  4. Response variable — the one number measured at the end Definition 1.4.2

Naming all four before you look at a single result is the fastest way to find out whether the study can answer the question it claims to answer.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

Your turn — name all six before the reveal

Try It Now 1.4.1

Try It Now 1.4.1 — three medicines, ninety adults

Dr. Lucía Herrera studies the effect of three medicines A, B, and C on the height of adults aged 30 to 45. She selects 90 adults at random and divides them into three equal groups; each group takes one medicine for six months. The average change in height in each group is calculated at the end. Identify: population, sample, experimental units, explanatory variable, response variable, treatments.

The six labels

Population adults aged 30 to 45 · Sample the 90 selected adults · Experimental units the individual adults · Explanatory variable the medicine taken · Treatments medicine A, B, and C · Response variable the change in height over six months.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

Worked example — a real clinical design

Example 1.4.1 — naming the parts of an aspirin study

Example 1.4.1 — does regular aspirin reduce heart-attack risk?

Dr. Karen Whitfield recruits 400 people aged 50 to 84 and divides them randomly into two groups: one takes aspirin, the other a placebo. Each person takes one pill daily for three years without knowing which they received. At the end her team counts heart attacks in each group.

Solution

Population people aged 50 to 84 · Sample the 400 participants · Experimental units the individual people · Explanatory variable oral medication · Treatments aspirin and placebo · Response variable whether the subject had a heart attack.

Note the design already in play: random division, a placebo arm, and subjects kept unaware — the three protections §1.4.2 and §1.4.4 are about.

1.4

§1.4.2 — the difference between finding groups and making them

To prove the explanatory variable caused the change, you must design the study so exactly one difference separates the groups: the treatment they were assigned.

The vitamin-E takers were healthier. Was it the vitamin — or everything else they also did?

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4.2 — the variable nobody wrote down

Lurking variable

Definition 1.4.5 — Lurking variable

A lurking variable is an additional variable, not accounted for in the study, that differs between the groups being compared and can therefore cloud the results — a difference in the response might be caused by it rather than by the explanatory variable.

Definition 1.4.5 — Lurking Variable Two bars, taller for the group that takes vitamin E, shorter for the group that does not, with a double arrow marking the gap between them and a single arrow already running from "vitamin E" to a junction that feeds the gap. Three traits that were always true of the vitamin-E group but were never recorded -- exercises more, eats better, does not smoke -- fade in beneath the bars, then each grows its own arrow into the same junction: four candidate causes now feed one observed gap. The vitamin E arrow and its label then change from the highlight colour to the same blue as the other three, closing on the point that the study cannot tell the four apart. takes vitamin E does not health on average the gap vitamin E never recorded, also true of this group: exercises more eats better does not smoke exercise diet not smoking four candidate causes — the study cannot tell them apart

Definition 1.4.5: in a self-selected comparison the gap has four candidate causes, and the study cannot tell them apart.

Context Pause — this is why "linked to" is not "causes." Headlines saying coffee is "linked to" longer life are almost always reporting a study like the vitamin E one: the coffee drinkers differed from the non-drinkers in a hundred ways nobody controlled.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4.2 — a study that proves nothing, in four lines

Does vitamin E prevent disease?

  1. Recruit subjects and ask whether they regularly take vitamin E.
  2. Observe that the takers are healthier on average than the non-takers.
  3. Notice what else differs: they exercise more, eat better, take other supplements, choose not to smoke.
  4. Any one of those could be the real reason. As described, the study proves nothing about vitamin E.

Nobody assigned anyone to take vitamin E. The two groups formed themselves — and self-selected groups differ in every way at once.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4.2 — the fix, and the reason experiments can prove cause

Random assignment

Definition 1.4.6 — Random assignment

Random assignment is the assignment of experimental units to treatment groups by chance rather than by choice. Because chance knows nothing about the subjects, every potential lurking variable is spread roughly equally across the groups.

Definition 1.4.6 -- Random assignment Twelve dots sit in one pool: six filled (carry a hidden trait) and six hollow (do not). Two empty rounded boxes, "group A" and "group B", fade in below the pool. The twelve units then travel one at a time in a left-to-right wave to their assigned box -- the destination is fixed by the scene's own assignment list, not by which half of the pool a unit started in, so filled and hollow units interleave into both boxes rather than splitting by starting side. Every dot stays visible for the whole trip; none are hidden or merged in transit. Once all twelve have landed, each box holds six units split three filled, three hollow, and a tally line under each box reads "3 with the trait, 3 without". A closing note reads "chance spread the hidden trait evenly -- the only difference left is the treatment", and everything holds through the end of the loop. carries a hidden trait does not one pool of experimental units group A group B 3 with the trait, 3 without 3 with the trait, 3 without chance spread the hidden trait evenly — the only difference left is the treatment

Definition 1.4.6: random assignment spreads a hidden trait evenly, so the treatment is the only difference left.

Once randomization has done its work the only systematic difference left is the one you imposed — so a difference in the response has to be the treatment. This is why experiments prove cause and surveys generally cannot.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4.2 — the uncomfortable question

When the variable cannot be assigned

Some explanatory variables you are not allowed to assign, and some you physically cannot:

  • You cannot randomly assign someone a birth order, a hometown, or a gender.
  • You cannot ethically assign someone to smoke for twenty years.
  • In every such case the experiment simply cannot be run.

The researcher is left comparing groups that formed themselves — exactly the situation random assignment was invented to escape.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

Your turn — design the study yourself

Try It Now 1.4.2

Try It Now 1.4.2 — texting and braking response time

How many seconds does a driver take to respond when the car ahead brakes? Design a study comparing texting with undistracted driving. (a) Name the explanatory and response variables. (b) The treatments. (c) What to consider when selecting participants. (d) Your partner Ryan Caldwell wants two random groups — one texting, one not. Good idea? (e) Lurking variables. (f) How blinding could be used.

Solution

(a) Explanatory = whether the driver is texting; response = braking response time in seconds. (b) Two treatments: texting, and no distraction. (c) Match participants on driving experience, age, and texting comfort — and recruit enough that one outlier cannot swing the result.

(d) Workable, but there is better: have each participant drive under both conditions in randomized order, so every driver is their own comparison. (e) Baseline reaction speed, experience, age, eyesight, tiredness, simulator familiarity. (f) Drivers cannot be blinded — but the person scoring the braking data can be.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

Try it in rāSHio — make the draw a real draw

Randomizing the assignment

Part (d) asks you to split participants at random, and "random" has to mean a real draw — not you deciding who looks like a texter. Open rāSHio, choose File → Random Numbers…, ask for as many draws as you have participants, and assign in the order they come out.

Figure 1.4.1: drawing the assignment in rāSHio — File → Random Numbers.

That is exactly the mechanism Definition 1.4.6 describes: chance, not judgement, decides who lands where — which is what spreads the lurking variables evenly.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

Worked example — the limit of the method

Example 1.4.2 — a question randomization cannot answer

Example 1.4.2 — does birth order affect personality?

A researcher wants to study the effects of birth order on personality. Explain why this could not be conducted as a randomized experiment, and name the main problem in any study that cannot be.

Solution

The explanatory variable is birth order — and it cannot be assigned. Birth order is fixed before the researcher ever shows up. Random assignment is what eliminates lurking variables, so without it the groups differ in family size, parents' age, income at the time, and more. Any personality difference found might be caused by one of those instead, so the study cannot establish cause and effect.

1.4

§1.4.3 — using the information you already have

Random assignment protects you from lurking variables you have not thought of. But when you already know a variable matters, leaving it to chance wastes what you know.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4.3 — sorting before you randomize

Block

Definition 1.4.7 — Block

A block is a group of experimental units that are similar to one another with respect to a variable expected to affect the response.

Two conditions, both required: you can name and measure the variable in advance, and you expect it to move the response on its own.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4.3 — the design in one sentence

Randomized block design

Definition 1.4.8 — Randomized block design

Experimental units are first sorted into blocks, and random assignment to treatments is then carried out separately within each block.

Insight Note — block what you know, randomize what you do not. Blocking handles the nuisance variables you can name; randomization handles the ones you cannot. A good experiment usually uses both — blocking first, randomizing inside the blocks.

Definition 1.4.7 & 1.4.8 — Block and Randomized Block Design Eight experimental units sit in one pool, filled dots working more than twenty hours a week and hollow dots not. Step 1: two block boxes appear and each unit travels to the block matching its own trait, ending four filled and four hollow per block. Step 2: two treatment boxes appear below and each block sends exactly two of its units to each treatment, so every treatment ends up with two filled and two hollow. A tally confirms the count under each treatment box, and the figure closes on "both groups hold the same mix of work schedules -- guaranteed, not hoped for." works more than 20 h a week does not one pool of experimental units block — works > 20 h block — works ≤ 20 h new tutoring current tutoring 2 work > 20 h, 2 do not 2 work > 20 h, 2 do not step 1 — sort into blocks on the variable you already know matters step 2 — randomize to treatments separately within each block both groups hold the same mix of work schedules — guaranteed, not hoped for

Definitions 1.4.7 and 1.4.8: sort on the variable you already know, then randomize separately inside each block.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4.3 — two versions of a tutoring programme

Why blocking pays twice

  1. Students working over twenty hours a week score lower regardless of tutoring — a variable we know matters.
  2. With simple random assignment, bad luck could land most heavy-work students in one group. The comparison becomes a comparison of work schedules in a tutoring costume.
  3. Blocking removes that risk by construction: sort into over-twenty and under-twenty blocks, then randomize half of each block to each version.
  4. The bonus: each block is internally similar, so differences within it are less noisy — a blocked experiment detects a smaller real effect with the same number of subjects.

Blocking is not only insurance against bad luck; it is often the cheapest way to make a study more sensitive.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

Your turn — choose the blocking variable

Try It Now 1.4.3

Try It Now 1.4.3 — a new warm-up routine and injury rates

A campus recreation centre wants to test whether a new warm-up reduces injuries among intramural players. About a third are varsity-level athletes — fitter, and more injury-prone because they play harder. (a) What would you block on, and why? (b) How would you carry out the assignment? (c) What do you risk with simple random assignment across all players? (d) Name one variable you would leave to randomization instead.

Solution

(a) Athlete level, varsity versus non-varsity: known in advance, easy to record, expected to affect injury rate on its own. (b) Sort every player into their block first, then within each block randomly assign half to the new warm-up and half to the current one.

(c) An unlucky draw puts more varsity players in one group; that group shows more injuries whether or not the warm-up worked, and nothing afterwards can tell the explanations apart. (d) Prior injury history, sleep, or competitiveness — plausibly relevant but hard to measure and impractical to sort on. Randomization spreads those without our naming them, which is the one thing blocking cannot do.

1.4

§1.4.4 — control groups, placebos, and blinding

Some of what you measure is the treatment. Some of it is the belief that you received the treatment. Three devices exist to tell them apart.

How much of a drug's effect survives when the subject does not know they took it?

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4.4 — the baseline

Control group

Definition 1.4.9 — Control group

A control group is a treatment group set aside to receive no active treatment. It shows what happens to the response when the treatment is not applied, so the researcher has something to compare the treated groups against.

Definition 1.4.9 — Control group An empty baseline holds two labelled slots, control group (no active treatment) on the left and treatment group (active treatment) on the right. The control bar grows to its measured height first, and a reference line extends from its top across the whole figure -- this is what would have happened without treatment. The treatment bar then grows past that line. A bracket appears marking only the portion of the treatment bar ABOVE the reference line, labelled "the treatment's effect"; a second bracket on the control bar's own height, labelled "would have happened anyway", marks the part every subject -- treated or not -- would have reached regardless. control group no active treatment treatment group active treatment measured response the treatment's effect would have happened anyway

Definition 1.4.9: the control group sets the baseline; only what rises above it is the treatment's effect.

Without a baseline, "the patients improved" is not a finding — people improve for all sorts of reasons, including time.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4.4 — balancing the effect of being in a study

Placebo

Definition 1.4.10 — Placebo

A placebo treatment is an inactive treatment that looks exactly like the active ones but cannot directly influence the response — a sugar pill, a saline injection, an unscented mask. It balances the effect of being in an experiment against the effect of the active treatments.

Definition 1.4.10 — Placebo Two pills the reader cannot tell apart, one over a "sugar (placebo)" bar and one over an "active drug" bar. A belief segment rises to the same height on both bars together. Then only the active-drug bar gains an active-ingredient segment on top. A brace singles out that top segment as what the placebo lets you measure. identical to look at sugar (placebo) active drug measured response belief you are being treated what the placebo lets you measure

Definition 1.4.10: a placebo gives both groups the belief, so whatever is left over is the drug.

Both groups get the ritual, the attention, and the expectation. Only one gets the chemistry — so the gap between them is the chemistry.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4.4 — keeping the secret

Blinding

Definition 1.4.11 — Blinding

Blinding (or masking) means a person involved in a research study does not know who is receiving the active treatment and who is receiving the placebo. Blinding preserves the placebo's power by keeping the participant from knowing which group they are in.

Definition 1.4.11 — Blinding Two arms, aspirin and placebo, each with four experimental-unit dots and a plainly readable label. A subject and a researcher each have a sight line to the arms. A "? ? ?" cover slides down over each label -- the assignment is not deleted, it is restated in a ledger line below saying it is unchanged. The subject's sight line is struck through (single-blind), then the researcher's sight line is struck through too (double-blind), and a closing note reads "the assignment never moved -- only who can read it". aspirin placebo subject researcher what each arm is actually given ? ? ? ? ? ? underneath, unchanged: arm 1 = aspirin · arm 2 = placebo single-blind double-blind the assignment never moved — only who can read it

Definition 1.4.11: blinding hides which arm you are in; the assignment underneath never changes.

Blinding does not change who got what. It changes only who knows — and that turns out to be enough to change the measurement.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4.4 — blinding the researcher too

Double-blind experiment

Definition 1.4.12 — Double-blind experiment

Both the subjects and the researchers working directly with the subjects are blinded. Neither side knows who got what until the study is over.

Insight Note — a placebo only works while the secret holds

If you are in a study and you know your pill is sugar, the power of suggestion switches off — and so does the whole point of the control group. Blinding is what keeps the secret.

Even a scrupulously honest researcher who knows the assignment will time a subject a little differently, prompt them a little differently, read an ambiguous result a little more generously. Blinding the researcher removes that possibility entirely — which is why the double-blind, placebo-controlled randomized experiment is the design medical journals treat as the gold standard.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4.4 — why blinding is not a technicality

Believing you took the drug worked nearly as well as taking it.

In a study of performance-enhancing drugs, believing one had taken the substance produced times almost as fast as consuming the drug itself — while taking the drug without knowledge yielded no significant increment at all.

Some of what you measure is the treatment, and some of it is the belief. Only the design can tell you which.

† Read that twice: the effect survived without the drug, and vanished without the belief. When simply being in a study prompts a physical response, isolating the explanatory variable gets much harder.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

Your turn — where does blinding fit?

Try It Now 1.4.4

Try It Now 1.4.4 — measuring the extent of placebo effects

Dr. Wei Chen's team asked randomly selected men to take a test before and after a pill that induces a mild headache. For half the men, chosen at random, the pill was replaced with a similar pill with no effect. Chen recorded each man's change in completion time. (a) Explanatory and response variables? (b) Treatments? (c) Lurking variables? (d) Is blinding possible?

Solution

(a) Explanatory = which pill he received; response = the change in completion time, before versus after. (b) Two: the active headache pill and the placebo.

(c) Natural test-taking speed, tiredness, caffeine, practice from having seen the test once, expectation. Random assignment spreads these, and measuring each man's change from his own before-score removes most individual differences. (d) Yes — the placebo looks like the active pill, so subjects are blinded; if Chen and the timers are also kept unaware, the study is double-blind.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

Worked example — when you can only get half the blinding

Example 1.4.3 — scent, mazes, and partial blinding

Example 1.4.3 — can smell affect learning?

Dr. Andrea Rivas led a study in which subjects completed pencil-and-paper mazes three times wearing floral-scented masks and three times wearing unscented masks. Participants were randomly assigned to wear the floral mask during the first three trials or the last three. The team recorded completion time and the subject's impression of the scent.

Solution

(a) Explanatory = scent; response = maze completion time. (b) Two treatments: floral mask and unscented mask.

(c) All subjects experienced both treatments and the order was randomly assigned, so there were no differences between treatment groups — randomization handles the lurking variables here. (d) Subjects clearly know whether they can smell flowers, so they cannot be blinded. The assistants timing the mazes can be: set the timing station up so the observer never sees the mask.

1.4

§1.4.5 — "numbers don't lie," but people do

The widespread misuse and misrepresentation of statistical information often gives the field a bad name. Some say that numbers don't lie — the people who use numbers to support their claims often do.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4.5 — fraud on a colossal scale

Diederik Stapel

A three-university investigation of the social psychologist Diederik Stapel found falsified data tainting over 55 papers he authored and 10 Ph.D. dissertations he supervised. The committee named four practices:

  • creating datasets that largely confirmed his prior expectations;
  • altering data in existing datasets;
  • changing measuring instruments without reporting the change;
  • misrepresenting the number of experimental subjects.

"It was a quest for aesthetics, for beauty — instead of the truth," he said, describing a frustration with the messiness of real experimental data.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

Context Pause — where this course comes in

The co-authors are the cautionary tale, not Stapel

Stapel chose to lie. His co-authors just could not read a table well enough to notice. Learning basic statistics is what makes you the person in the room who catches it.

The investigation's own report noted that "statistical flaws frequently revealed a lack of familiarity with elementary statistics."

Harder to spot: researchers who simply stop collecting data once they have just enough to prove what they hoped to prove. retractionwatch.com catalogues the retractions — a quick glance shows the misuse of statistics is a bigger problem than most people realize.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

Your turn — find the ethical failure in each part

Try It Now 1.4.5

Try It Now 1.4.5 — favourite fruit juice among California teens

Grant Halloway is commissioned to run the study. (a) The survey is commissioned by the seller of a popular apple juice. (b) Only two juices are included: apple and cranberry. (c) Participants see the brand as samples are poured for a taste test. (d) Brand X advertises: "Most teens like Brand X as much as or more than Brand Y."

ResponseShare
prefer Brand X25%
prefer Brand Y33%
no preference42%

Table 1: the numbers behind part (d)'s claim.

Solution

(a) A conflict of interest — not automatically unethical, but it must be disclosed. (b) A two-brand menu cannot support a claim about the favourite juice; widen the options or narrow the claim. (c) Visible brands destroy a taste test — pour out of sight, label with neutral codes. (d) Technically true, deeply misleading: it folds the 42% with no preference in with X's 25% to reach 67%, while reporting Y's 33% alone — even though the no-preference group likes Y exactly as much. More teens preferred Brand Y than Brand X.

1.4

§1.4.6 — the rules that come before the data

When a study uses human participants, both ethics and the law require the researcher to be mindful of their safety — and to prove it to someone else before the study begins.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4.6 — the gate a study must clear

Institutional Review Board

Definition 1.4.13 — Institutional Review Board (IRB)

An Institutional Review Board is an oversight committee established by a research institution to review and approve planned studies before they begin, with the purpose of protecting human subjects.

"Before they begin" is the whole design. A review that happens after the data is collected protects nobody.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4.6 — the second gate

Informed consent

Definition 1.4.14 — Informed consent

Informed consent means the risks of participation have been clearly explained to the subjects, and the subjects have agreed to participate in writing. Researchers are required to keep documentation of that consent.

Definition 1.4.13 — Institutional Review Board A closed IRB-review gate and a closed informed-consent gate stand between a planned study and its participants. The study card stops at each gate while that gate's three-item checklist fades in, then the gate's bar lifts before the card continues. Only after both gates clear does the card reach participants. IRB review informed consent participants plannedstudy the IRB checks: risks minimized vs. benefit informed consent required privacy of the data guarded the participant's consent: risks explained in plain language agreement given in writing documentation kept on file approval first, then consent, then anyone is enrolled

Definitions 1.4.13 and 1.4.14: a planned study reaches no one until it clears IRB review and then informed consent.

Two gates in series, in this order. Neither one is optional, and passing the first does not excuse the second.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4.6 — mandated by law

Three protections, and three hard questions

  1. Risks to participants must be minimized and reasonable with respect to projected benefits.
  2. Participants must give informed consent, in writing, with documentation kept.
  3. Data collected from individuals must be guarded carefully to protect their privacy.

Fundamental — and very difficult to verify in practice. Is removing a name enough to protect privacy, or could the identity be recovered from what remains? What happens when unanticipated risks arise mid-study? Once the lab has tested your blood sample, does a researcher have the right to take the remainder for a study?

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

Worked example — three failures on one survey route

Example 1.4.4 — spotting unethical data collection

Example 1.4.4 — Imani Boateng collects survey data in a community

(a) She selects a block where she is comfortable walking because she knows many of the residents. (b) No one is home at four houses; she does not record the addresses and does not return later. (c) She skips four more houses because she is running late, then fills in those forms at home by copying answers from other residents.

Solution

(a) A convenience sample presented as representative — biased, and misleading to claim it represents the community. Select areas at random. (b) Quietly dropping non-responses biases the sample: if the study is about jobs and child care, the people who are out are exactly the working families it is about. Record the addresses and return. (c) It is never acceptable to fake data. Even "real" answers copied from other participants are fraudulent duplication. The only fix is to actually collect the data.

1.4

§1.4.7 — language for whether two variables travel together

Most statistical questions are really questions about whether two variables travel together. Before we can ask why they might, we need language for whether they do.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4.7 — knowing one tells you something

Associated variables

Definition 1.4.15 — Associated variables

Two variables are associated if knowing the value of one tells you something — anything — about the likely value of the other.

Told a student slept three hours, you would revise your guess about their quiz score downward — even though you would sometimes be wrong. That is association.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4.7 — knowing one tells you nothing

Independent variables

Definition 1.4.16 — Independent variables

Two variables are independent if knowing the value of one tells you nothing about the likely value of the other.

Context Pause — these are the only two options. No third category, no partial credit. If knowing one shifts your expectation about the other by any amount, they are associated. "Independent" is the strong claim, not the safe default: it asserts the information content is exactly zero.

Definitions 1.4.15 and 1.4.16 — Associated and independent variables Two side-by-side scatter panels share one sweeping selector. Panel (a), "associated", plots hours of sleep against quiz score: as the selector column sweeps left to right, the accent band showing the admitted quiz scores stays thin and climbs steadily upward -- knowing x narrows y. Panel (b), "independent", plots shoe size against favourite genre (pop, rock, hip-hop): the accent band stays the full height of the panel at every position of the selector -- every genre occurs at every shoe size, so knowing x changes nothing. Dots under the selector turn the accent colour; all others stay the curve colour. The sweep runs 5 seconds and then reverses, looping continuously. hours of sleep quiz score shoe size favourite genre pop rock hip-hop (a) associated (b) independent knowing x narrows y knowing x changes nothing

Definitions 1.4.15 and 1.4.16: knowing x narrows the values y can take, or it does not — that is the whole distinction.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4.7 — worth saying out loud, because everyday language hides it

Association runs both ways

Associated

Hours of sleep and quiz score. Told a student slept three hours, you revise your guess about the score — imperfectly, but you revise it.

Independent

Shoe size and favourite music genre. Told a student wears a size 11, you have learned nothing useful about their playlist.

If sleep tells you something about quiz scores, then quiz scores tell you something about sleep. The arrow you imagine between them is something you brought to the data, not something the association contains — which is precisely why association alone can never establish a direction, let alone a cause.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

Key Terminology — the design vocabulary

The words, part 1

explanatory variable — the variable the researcher manipulates, believed to cause change in another.

response variable — the variable measured to see whether the explanatory variable had an effect.

treatment — one of the specific values of the explanatory variable assigned to a group.

experimental unit — a single object or individual to be measured in the study.

lurking variable — an unaccounted-for variable that differs between groups and can be the real cause of a difference in the response.

random assignment — assigning units to treatment groups by chance, spreading lurking variables equally.

block — a group of units similar with respect to a variable expected to affect the response.

randomized block design — units sorted into blocks first, randomization within each block.

Experimental Design & Ethics · bookSHelf Intro Stats§1.4

Key Terminology — suggestion, ethics, association

The words, part 2

control group — a group that receives no active treatment, used as a baseline for comparison.

placebo — an inactive treatment made to look like the active one, so belief affects both groups equally.

blinding — keeping a person in the study from knowing which treatment was received.

double-blind experiment — both the subjects and the researchers working with them are blinded.

Institutional Review Board (IRB) — the committee that must approve a planned study to protect its human subjects.

informed consent — a subject's written agreement to participate after the risks have been clearly explained.

associated variables — knowing one tells you something about the likely value of the other.

independent variables — knowing one tells you nothing about the likely value of the other.

1.4
Experimental Design & Ethics · bookSHelf Intro Stats§1.4

§1.4 — conclusions

What §1.4 leaves you with

The core idea

A study earns a cause-and-effect claim by construction, not by result size. Random assignment makes the treatment the only systematic difference; blocking spends what you already know; a control group, a placebo, and blinding keep the power of suggestion out of the measurement.

The failure case

Groups the researcher merely found rather than made differ in a hundred uncontrolled ways — so "linked to" never upgrades to "causes." And a design that is sound can still be unethical: undisclosed funding, coerced consent, quietly dropped non-responses, fabricated rows.

Next: §1.5 — Data Collection Experiment, where you run one of these designs yourself instead of reading about it.