1.4 Experimental Design and Ethics
SLO 1
Assess how data were collected and recognize how data collection affects what conclusions can be drawn from the data.
Random assignment, control groups, placebos, and blinding are the machinery that decides whether a study can claim cause and effect at all, so naming which of them a study used tells you how far its conclusion is allowed to reach.
SLO 6
Evaluate ethical issues in statistical practice.
The Stapel fraud case, the IRB's informed-consent rules, and problems on a faked survey route and a bar chart of raw complaint counts give you concrete practice deciding when a study crosses from sloppy into wrong.
Learning Objectives
By the end of this section, you will be able to:
- identify the population, sample, experimental units, explanatory variable, response variable, and treatments in a described study;
- explain how random assignment removes lurking variables, and recognize when a question cannot be answered by a randomized experiment;
- describe how a control group, a placebo, and blinding protect a study from the power of suggestion;
- assess whether a study's design supports a cause-and-effect claim or only a weaker one;
- evaluate the ethics of a statistical study, including informed consent, data honesty, and privacy protections.
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer better at growing roses than another? Is being tired behind the wheel as dangerous as being drunk behind the wheel? Questions like these get answered with randomized experiments. In this section you will learn what has to be built into an experiment before its results mean anything. A study that is designed badly does not produce weak data — it produces data you cannot use at all, which is worse, because the numbers still look like numbers.
1.4.1 The Anatomy of an Experiment
When one variable causes a change in another, the first variable is called the explanatory variable. In a randomized experiment, the researcher deliberately sets or manipulates the values of the explanatory variable.
The variable that is affected by the explanatory variable is called the response variable. The researcher does not set this one — they measure it, and see whether it moved.
Definitions 1.4.1 and 1.4.2 — The explanatory variable is the one you set; the response is the one you read off.
The different values of the explanatory variable that the researcher assigns are called the treatments.
An experimental unit is a single object or individual to be measured — one person, one rose bush, one rat, one plot of soil.
The explanatory variable is the switch you flip; the response variable is the bulb you watch. Treatments are the switch positions you try. If you never touch the switch, you never learn whether it is wired to that bulb at all.
Definitions 1.4.3 and 1.4.4 — The treatments are the values of one explanatory variable, handed out to the experimental units.
Notice how these four terms stack. The experimental units are who is in the study. The treatments are what you do to them. The explanatory variable is the general thing the treatments are versions of, and the response variable is the one number you measure at the end to see if any of it mattered. Naming all four for a study, before you look at a single result, is the fastest way to find out whether the study can answer the question it claims to answer.
The purpose of an experiment is to investigate the relationship between two variables. One variable is the one you push on; the other is the one you watch to see what happens.
Dr. Lucía Herrera, who runs her clinic with her wife, needs to conduct a study of the effect of three medicines A, B, and C on the height of adults aged 30 to 45. She selected 90 adults randomly and divided them into three equal groups. The first group was asked to take medicine A for 6 months. The second group was asked to take medicine B for 6 months. The third group was asked to take medicine C for 6 months. The average change in height in each group is calculated at the end of the study.
Identify the following values for this study: population, sample, experimental units, explanatory variables, response variable, treatments.
Solution
Step 1 — Population. Adults aged 30 to 45 — the group the conclusion is meant to cover.
Step 2 — Sample. The 90 adults who were selected randomly and actually took part.
Step 3 — Experimental units. The individual adults in the study.
Step 4 — Explanatory variable. The medicine taken. Herrera controls which of the three each group receives.
Step 5 — Treatments. Three of them: medicine A, medicine B, and medicine C.
Step 6 — Response variable. The change in height over the six months, which is what gets measured and averaged.
Answer: population = adults aged 30 to 45; sample = the 90 selected adults; experimental units = the individual adults; explanatory variable = medicine; treatments = A, B, and C; response variable = change in height.
Dr. Karen Whitfield, a cardiologist who co-directs her lab with her wife, wants to investigate whether taking aspirin regularly reduces the risk of heart attack. Four hundred people between the ages of 50 and 84 are recruited as participants. The people are divided randomly into two groups: one group will take aspirin, and the other group will take a placebo. Each person takes one pill each day for three years, but they don't know whether they are taking aspirin or the placebo. At the end of the study, Whitfield's team counts the number of people in each group who have had heart attacks.
Identify the following values for this study: population, sample, experimental units, explanatory variable, response variable, treatments.
Solution
Step 1 — Population. The population is the whole group Whitfield wants her conclusion to apply to: people aged 50 to 84.
Step 2 — Sample. The sample is the subset actually studied: the 400 people who participated.
Step 3 — Experimental units. The experimental units are the individual people in the study — one person, measured once.
Step 4 — Explanatory variable. The explanatory variable is oral medication. This is the thing she controls.
Step 5 — Treatments. The treatments are the specific values of that variable that were handed out: aspirin and a placebo.
Step 6 — Response variable. The response variable is whether a subject had a heart attack. This is what gets counted at the end.
Answer: population = people aged 50 to 84; sample = the 400 participants; experimental units = the individual people; explanatory variable = oral medication; treatments = aspirin and placebo; response variable = whether the subject had a heart attack.
1.4.2 Lurking Variables and Random Assignment
A lurking variable is an additional variable, not accounted for in the study, that differs between the groups being compared and can therefore cloud the results. When a lurking variable is present, a difference in the response variable might be caused by the lurking variable rather than by the explanatory variable.
Headlines that say coffee is "linked to" longer life are almost always reporting a study like the vitamin E one. The coffee drinkers differed from the non-drinkers in a hundred ways nobody controlled.
Definition 1.4.5 — In a self-selected comparison the gap has four candidate causes, and the study cannot tell them apart.
To prove that the explanatory variable is what is causing the change in the response variable, you have to isolate it. The researcher must design the experiment so that there is exactly one difference between the groups being compared: the treatment they were assigned. That is accomplished by random assignment.
Random assignment is the assignment of experimental units to treatment groups by chance rather than by choice. Because chance does not know anything about the subjects, every potential lurking variable is spread roughly equally across the groups.
Definition 1.4.6 — Random assignment spreads a hidden trait evenly, so the treatment is the only difference left.
Once random assignment has done its work, the only systematic difference left between the groups is the one the researcher imposed. So if the groups end up with different values of the response variable, that difference has to be a direct result of the different treatments. This is the whole reason experiments can prove a cause-and-effect connection while surveys and observational studies generally cannot.
Which raises an uncomfortable question: what happens when the explanatory variable is something you are not allowed to assign, or physically cannot assign? You cannot randomly assign someone a birth order, a hometown, or a gender. You cannot ethically assign someone to smoke for twenty years. In every one of those cases the experiment simply cannot be run, and the researcher is left comparing groups that formed themselves — which is exactly the situation random assignment was invented to escape. The next example works through one of these, and the reasoning it uses is the reasoning you should reach for any time a study reports a difference between groups the researcher merely found rather than made.
Suppose you want to investigate whether vitamin E prevents disease. You recruit a group of subjects and ask them if they regularly take vitamin E. You notice that the subjects who take vitamin E are healthier on average than those who do not. Does this prove that vitamin E works?
It does not. There are many differences between those two groups besides vitamin E. People who take vitamin E regularly often take other steps to improve their health: they exercise, they eat better, they take other supplements, they choose not to smoke. Any one of those could be the real reason the vitamin E group looks healthier. As described, this study proves nothing about vitamin E.
You are concerned about the effects of texting on driving performance. Design a study to test the response time of drivers while texting and while driving only. How many seconds does it take for a driver to respond when a leading car hits the brakes?
a. Describe the explanatory and response variables in the study.
b. What are the treatments?
c. What should you consider when selecting participants?
d. Your research partner, Ryan Caldwell — who joined the project after his husband was rear-ended by a texting driver — wants to divide participants randomly into two groups: one to drive without distraction and one to text and drive simultaneously. Is this a good idea? Why or why not?
e. Identify any lurking variables that could interfere with this study.
f. How can blinding be used in this study?
Solution
Part a — Variables. The explanatory variable is whether the driver is texting. The response variable is the response time in seconds — how long it takes the driver to hit the brakes after the leading car does.
Part b — Treatments. Two: driving while texting, and driving with no distraction.
Part c — Selecting participants. Participants should be similar in the ways that affect driving on their own: driving experience, age, and how comfortable they are with texting. You also want enough participants that one unusually fast or slow driver cannot swing the result.
Part d — Is Caldwell's split into two groups a good idea? It is workable, but there is a better design. If each participant drives under both conditions, in a randomized order, then each driver acts as their own comparison and the natural differences in reaction speed between people stop mattering. The two-group split he proposes leaves those differences in play.
Part e — Lurking variables. Baseline reaction speed, driving experience, age, eyesight, tiredness, and familiarity with the simulator. Random assignment spreads these across the groups; having each driver do both conditions removes them almost entirely.
Part f — Blinding. The drivers cannot be blinded — they know perfectly well whether they are texting. The researcher scoring the response times can be: they can be shown only the recorded braking data, without being told which condition it came from.
Answer: explanatory variable = texting or not; response variable = braking response time; treatments = texting and no distraction; the strongest design has every driver do both conditions in random order; subjects cannot be blinded but the scorer can be.
Try it in rāSHio
Part (d) asks you to split participants into groups at random, and "random" has to mean a real draw — not you deciding who looks like a texter. Open rāSHio and choose File → Random Numbers…, ask for as many draws as you have participants, and assign in the order they come out. That is exactly the mechanism Definition 1.4.6 describes: chance, not judgement, decides who lands where, which is what spreads the lurking variables evenly.
Figure 1.4.1 — Drawing the assignment in rāSHio: File → Random Numbers.
A researcher wants to study the effects of birth order on personality. Explain why this study could not be conducted as a randomized experiment. What is the main problem in a study that cannot be designed as a randomized experiment?
Solution
Step 1 — Identify the explanatory variable. The explanatory variable is birth order.
Step 2 — Ask whether it can be assigned. It cannot. You cannot randomly assign a person to be a firstborn or a middle child. Birth order is fixed before the researcher ever shows up.
Step 3 — Name the consequence. Random assignment is what eliminates the impact of lurking variables. When you cannot assign subjects to treatment groups at random, there will be differences between the groups other than the explanatory variable — family size, parents' age, income at the time, and so on.
Answer: birth order cannot be randomly assigned, so the groups differ in ways beyond birth order. Any personality difference found might be caused by one of those lurking variables instead, so the study cannot establish cause and effect.
1.4.3 Blocking
Random assignment protects you from lurking variables you have not thought of. But when you already know a variable matters, leaving it to chance is a waste of the information you have.
A block is a group of experimental units that are similar to one another with respect to a variable expected to affect the response.
In a randomized block design, experimental units are first sorted into blocks, and random assignment to treatments is then carried out separately within each block.
This is the whole design principle in one line. Blocking handles the nuisance variables you can name and measure; randomization handles the ones you cannot. Neither replaces the other, and a good experiment usually uses both — blocking first, then randomizing inside the blocks.
Suppose we are comparing two versions of a tutoring programme, and we know that students who work more than twenty hours a week tend to score lower regardless of tutoring. With simple random assignment, we might by bad luck land most of the heavy-work students in one group. The comparison would then be partly a comparison of work schedules wearing a tutoring costume.
Blocking removes that risk by construction. Sort the students into two blocks — those working over twenty hours and those working under — then randomly assign half of each block to each version. Both treatment groups now contain the same mix of work schedules, guaranteed rather than hoped for.
There is a bonus that surprises people. Because each block is internally similar, the differences you see within a block are less noisy, so a blocked experiment can detect a smaller real effect with the same number of subjects. Blocking is not only insurance against bad luck; it is often the cheapest way to make a study more sensitive.
Definitions 1.4.7 and 1.4.8 — Sort on the variable you already know, then randomize separately inside each block.
A campus recreation centre wants to test whether a new warm-up routine reduces injury rates among intramural players. Roughly a third of the players are varsity-level athletes, who are both fitter and more injury-prone than the rest because they play harder.
a. What variable would you block on, and why?
b. Describe how you would carry out the assignment.
c. What would you risk by using simple random assignment across all players instead?
d. Name one variable you would leave to randomization rather than blocking, and say why.
Solution
Part a — The blocking variable. Athlete level: varsity versus non-varsity. It is known in advance, easy to record, and expected to affect the response (injury rate) on its own. That combination is exactly what makes a variable worth blocking on.
Part b — The assignment. Sort every player into the varsity block or the non-varsity block first. Then, within the varsity block, randomly assign half to the new warm-up and half to the current one; do the same, separately, within the non-varsity block. Both treatment groups now hold the same proportion of varsity players by construction.
Part c — The risk of simple random assignment. With about a third of players being varsity, an unlucky draw could put noticeably more of them in one group. That group would show more injuries whether or not the warm-up did anything, and there would be no way afterwards to tell the two explanations apart. The study would not be wrong so much as unable to answer its own question.
Part d — What to leave to chance. Something like prior injury history, sleep, or how competitively each player happens to take the season — plausibly relevant, but hard to measure reliably and impractical to sort on. Randomization spreads such variables evenly across groups on average without our needing to name them, which is the one thing blocking cannot do.
1.4.4 The Power of Suggestion: Control Groups, Placebos, and Blinding
A control group is a treatment group set aside to receive no active treatment. It shows what happens to the response variable when the treatment is not applied, so the researcher has something to compare the treated groups against.
Definition 1.4.9 — The control group sets the baseline; only what rises above it is the treatment's effect.
A placebo treatment is an inactive treatment that looks exactly like the active treatments but cannot directly influence the response variable — a sugar pill, a saline injection, an unscented mask. It balances the effect of being in an experiment against the effect of the active treatments.
Definition 1.4.10 — A placebo gives both groups the belief, so whatever is left over is the drug.
Blinding (or masking) means a person involved in a research study does not know who is receiving the active treatment and who is receiving the placebo. Blinding preserves the placebo's power by keeping the participant from knowing which group they are in.
Definition 1.4.11 — Blinding hides which arm you are in; the assignment underneath never changes.
A double-blind experiment is one in which both the subjects and the researchers working directly with the subjects are blinded. Neither side knows who got what until the study is over.
If you are in a study and you know your pill is sugar, the power of suggestion switches off — and so does the whole point of the control group. Blinding is what keeps the secret.
The double-blind design guards against a second, quieter problem. Even a scrupulously honest researcher who knows which subject got the real treatment will time them a little differently, prompt them a little differently, or read an ambiguous result a little more generously. Blinding the researcher removes that possibility entirely, which is why the double-blind, placebo-controlled randomized experiment is the design that medical journals treat as the gold standard. The next example is a case where full blinding is impossible — the subjects can tell instantly which treatment they are getting — and it shows you what a researcher does when they can only get half of it.
The expectation of the person being studied can be as powerful as the treatment itself. In one study of performance-enhancing drugs, researchers noted:
Results showed that believing one had taken the substance resulted in [performance] times almost as fast as those associated with consuming the drug itself. In contrast, taking the drug without knowledge yielded no significant performance increment.
Read that twice. Believing you took the drug worked nearly as well as taking it. Taking it without knowing did nothing measurable. When simply being in a study prompts a physical response, isolating the effect of the explanatory variable gets much harder — some of what you measure is the treatment, and some of it is the belief.
Dr. Wei Chen and his team at the Placebo Research Group conducted a study to find the extent of placebo effects. A group of men randomly selected were asked to take a test before and after taking a pill that induces a mild headache. The pill in half of the randomly selected men was replaced with a similar pill that has no effect. For each trial, Chen recorded the change in time men took to complete the tests before and after taking the pill.
a. Describe the explanatory and response variable in this study.
b. What are the treatments?
c. Identify any lurking variables that could interfere with this study.
d. Is it possible to use blinding in this study?
Solution
Part a — Variables. The explanatory variable is which pill the man received — the headache-inducing pill or the inactive one. The response variable is the change in the time taken to complete the test, measured before versus after the pill.
Part b — Treatments. Two: the active pill that induces a mild headache, and the placebo pill that has no effect.
Part c — Lurking variables. Natural differences in test-taking speed, tiredness, caffeine, practice effects from having already seen the test once, and how much each man expects the pill to bother him. Random assignment spreads these across the two groups, and measuring each man's change from his own before-score removes most of the individual differences.
Part d — Blinding. Yes. The placebo pill is described as similar to the active one, so a participant cannot tell which he received — the subjects are blinded. If Chen and the assistants timing the tests are also kept from knowing which pill each man got, the study is double-blind.
Answer: explanatory variable = type of pill, response variable = change in test completion time; treatments = active pill and placebo; lurking variables are handled by random assignment plus the before-and-after design; blinding is possible, and double-blinding is possible if the researchers are kept unaware too.
Dr. Andrea Rivas, who uses they/them pronouns, led a study at the Smell & Taste Treatment and Research Foundation, together with their wife Dr. Sofía Rivas, to investigate whether smell can affect learning. Subjects completed mazes multiple times while wearing masks. They completed the pencil and paper mazes three times wearing floral-scented masks, and three times with unscented masks. Participants were assigned at random to wear the floral mask during the first three trials or during the last three trials. For each trial, the Rivases recorded the time it took to complete the maze and the subject's impression of the mask's scent: positive, negative, or neutral.
a. Describe the explanatory and response variables in this study.
b. What are the treatments?
c. Identify any lurking variables that could interfere with this study.
d. Is it possible to use blinding in this study?
Solution
Part a — Variables. The explanatory variable is scent, and the response variable is the time it takes to complete the maze.
Part b — Treatments. There are two treatments: a floral-scented mask and an unscented mask.
Part c — Lurking variables. All subjects experienced both treatments. The order of treatments was randomly assigned, so there were no differences between the treatment groups. Random assignment eliminates the problem of lurking variables here.
Part d — Blinding. Subjects will clearly know whether they can smell flowers or not, so subjects cannot be blinded in this study. The assistants timing the mazes can be blinded, though. Whoever is observing a subject will not know which mask is being worn — Rivas can set the timing station up so that they themselves never see the mask.
Answer: explanatory variable = scent, response variable = maze completion time; treatments = floral mask and unscented mask; random assignment of order handles the lurking variables; subjects cannot be blinded, but the people timing the mazes can be.
1.4.5 Ethics and the Cost of Fraud
The widespread misuse and misrepresentation of statistical information often gives the field a bad name. Some say that "numbers don't lie," but the people who use numbers to support their claims often do.
An investigation of the famous social psychologist Diederik Stapel led to the retraction of his articles from some of the world's top journals. Stapel is a former professor at Tilburg University in the Netherlands. An extensive investigation involving three universities where Stapel had worked concluded that the psychologist was guilty of fraud on a colossal scale. Falsified data taints over 55 papers he authored and 10 Ph.D. dissertations that he supervised.
Stapel did not deny that his deceit was driven by ambition. But it was more complicated than that, he told me. He insisted that he loved social psychology but had been frustrated by the messiness of experimental data, which rarely led to clear conclusions. His lifelong obsession with elegance and order, he said, led him to concoct sexy results that journals found attractive. "It was a quest for aesthetics, for beauty — instead of the truth," he said. He described his behavior as an addiction that drove him to carry out acts of increasingly daring fraud, like a junkie seeking a bigger and better high.
The committee investigating Stapel concluded that he was guilty of several practices, including:
- creating datasets that largely confirmed his prior expectations;
- altering data in existing datasets;
- changing measuring instruments without reporting the change; and
- misrepresenting the number of experimental subjects.
Clearly, it is never acceptable to falsify data the way this researcher did. Sometimes, however, violations of ethics are not as easy to spot.
Researchers have a responsibility to verify that proper methods are being followed. The report describing the investigation of Stapel's fraud states that "statistical flaws frequently revealed a lack of familiarity with elementary statistics." Many of Stapel's co-authors should have spotted irregularities in his data. Unfortunately, they did not know very much about statistical analysis, and they simply trusted that he was collecting and reporting data properly.
Stapel chose to lie. His co-authors just could not read a table well enough to notice. Learning basic statistics is what makes you the person in the room who catches it.
Many types of statistical fraud are difficult to spot. Some researchers simply stop collecting data once they have just enough to prove what they had hoped to prove. They don't want to take the chance that a more extensive study would complicate their lives by producing data that contradicts their hypothesis.
Professional organizations, like the American Statistical Association, clearly define expectations for researchers. There are even laws in the federal code about the use of research data.
It is important that students of statistics take time to consider the ethical questions that arise in statistical studies. How prevalent is fraud in statistical studies? You might be surprised — and disappointed. There is a website (http://www.retractionwatch.com) dedicated to cataloging retractions of study articles that have been proven fraudulent. A quick glance will show that the misuse of statistics is a bigger problem than most people realize.
Vigilance against fraud requires knowledge. Learning the basic theory of statistics will empower you to analyze statistical studies critically.
Describe the unethical behavior, if any, in each example and describe how it could impact the reliability of the resulting data. Explain how the problem should be corrected.
Grant Halloway is commissioned to run a study determining the favorite brand of fruit juice among teens in California.
a. His survey is commissioned by the seller of a popular brand of apple juice.
b. There are only two types of juice included in the study: apple juice and cranberry juice.
c. He allows participants to see the brand of juice as samples are poured for a taste test.
d. Twenty-five percent of participants prefer Brand X, 33% prefer Brand Y and 42% have no preference between the two brands. Brand X references the study in a commercial saying "Most teens like Brand X as much as or more than Brand Y."
Solution
Part a — Who paid for it. Being funded by an interested party is not automatically unethical, but it is a conflict of interest that Halloway must disclose. If the funding is hidden, readers cannot judge how much to trust his design choices. Fix: disclose the sponsor prominently.
Part b — Only two juices. The study claims to find the favorite juice among teens but only offers two options. The result cannot support that claim — the actual favorite may not have been on the table. Fix: include the full range of common brands, or narrow the claim to "apple versus cranberry."
Part c — Visible brands in a taste test. Letting participants see the brand destroys the taste test. Preferences will reflect brand recognition and advertising rather than taste. This is a failure to blind the subjects. Fix: he should pour the samples out of sight and label them only with neutral codes.
Part d — The commercial's claim. This one is technically true and deeply misleading. "As much as or more than" folds the 42% with no preference in with Brand X's 25% to reach 67%, while Brand Y's 33% is reported alone — even though the no-preference group likes Brand Y exactly as much. Fix: report all three numbers, since 33% actually preferred Brand Y and only 25% preferred Brand X.
Answer: (a) undisclosed conflict of interest; (b) a sample of brands too narrow to support the claim; (c) no blinding, so brand recognition contaminates the taste test; (d) a true-but-misleading summary that hides the fact that more teens preferred Brand Y than Brand X.
1.4.6 Protecting Human Subjects
An Institutional Review Board (IRB) is an oversight committee established by a research institution to review and approve planned studies before they begin, with the purpose of protecting human subjects.
Informed consent means the risks of participation have been clearly explained to the subjects of a study, and the subjects have agreed to participate in writing. Researchers are required to keep documentation of that consent.
Definitions 1.4.13 and 1.4.14 — A planned study reaches no one until it clears IRB review and then informed consent.
For this reason, research institutions establish oversight committees, and every planned study must be approved by one in advance. Key protections that are mandated by law include the following:
- Risks to participants must be minimized and reasonable with respect to projected benefits.
- Participants must give informed consent. This means that the risks of participation must be clearly explained to the subjects of the study. Subjects must consent in writing, and researchers are required to keep documentation of their consent.
- Data collected from individuals must be guarded carefully to protect their privacy.
These ideas may seem fundamental, but they can be very difficult to verify in practice. Is removing a participant's name from the data record sufficient to protect privacy? Perhaps the person's identity could be discovered from the data that remains. What happens if the study does not proceed as planned and risks arise that were not anticipated? When is informed consent really necessary? Suppose your doctor wants a blood sample to check your cholesterol level. Once the sample has been tested, you expect the lab to dispose of the remaining blood. At that point the blood becomes biological waste. Does a researcher have the right to take it for use in a study?
When a statistical study uses human participants, as in medical studies, both ethics and the law dictate that researchers should be mindful of the safety of their research subjects. The U.S. Department of Health and Human Services oversees federal regulations of research studies with the aim of protecting participants. When a university or other research institution engages in research, it must ensure the safety of all human subjects.
Describe the unethical behavior in each example and describe how it could impact the reliability of the resulting data. Explain how the problem should be corrected.
Imani Boateng, who lives two neighborhoods over with her wife, is collecting survey data in a community.
a. She selects a block where she is comfortable walking because she knows many of the people living on the street.
b. No one seems to be home at four houses on the route. She does not record the addresses and does not return at a later time to try to find residents at home.
c. She skips four houses on the route because she is running late for an appointment. When she gets home, she fills in the forms by selecting random answers from other residents in the neighborhood.
Solution
Part a — Choosing a comfortable block. By selecting a convenient sample, Boateng is intentionally selecting a sample that could be biased. Claiming that this sample represents the community is misleading. She needs to select areas in the community at random.
Part b — Skipping the houses where no one answered. Intentionally omitting relevant data will create bias in the sample. Suppose she is gathering information about jobs and child care. By ignoring people who are not home, she may be missing data from working families that are exactly the families the study is about. She needs to make every effort to interview all members of the target sample.
Part c — Filling in the forms with other people's answers. It is never acceptable to fake data. Even though the responses she copied are "real" responses provided by other participants, the duplication is fraudulent and can create bias in the data. She needs to work diligently to interview everyone on her route.
Answer: (a) a convenience sample presented as representative — fix it with random selection of blocks; (b) non-responses quietly dropped — fix it by recording the addresses and returning; (c) fabricated responses — never acceptable, and the only fix is to actually collect the data.
1.4.7 Associated and Independent Variables
Most statistical questions are really questions about whether two variables travel together. Before we can ask why they might, we need language for whether they do.
Two variables are associated if knowing the value of one tells you something — anything — about the likely value of the other.
Two variables are independent if knowing the value of one tells you nothing about the likely value of the other.
There is no third category and no partial credit. If knowing one variable shifts your expectation about the other by any amount at all, they are associated — the association may be weak, noisy, and useless in practice, but it exists. "Independent" is the strong claim, not the safe default, because it asserts that the information content is exactly zero.
Hours of sleep and quiz score are plausibly associated: told that a student slept three hours, you would revise your guess about the score downward, even though you would sometimes be wrong. Shoe size and favourite music genre are plausibly independent: told a student wears a size 11, you have learned nothing useful about their playlist.
Association is also symmetric, which is worth saying out loud because everyday language hides it. If sleep tells you something about quiz scores, then quiz scores tell you something about sleep. The arrow you imagine between them is something you brought to the data, not something the association contains — which is precisely why association alone can never establish a direction, let alone a cause.
Definitions 1.4.15 and 1.4.16 — Knowing x narrows the values y can take, or it does not; that is the whole distinction.
Problem Set 1.4
Problem 1. Design an experiment. Identify the explanatory and response variables. Describe the population being studied and the experimental units. Explain the treatments that will be used and how they will be assigned to the experimental units. Describe how blinding and placebos may be used to counter the power of suggestion.
Solution
Step 1 — Pick a question you could actually run. Any well-formed design works; here is one. Does taking a 20-minute walk before class improve quiz scores?
Step 2 — Name the variables. The explanatory variable is whether the student takes the pre-class walk. The response variable is the score on the quiz taken at the end of class.
Step 3 — Population and experimental units. The population is students enrolled in introductory statistics at the college. The experimental units are the individual students who volunteer for the study.
Step 4 — Treatments and how they are assigned. Two treatments: a 20-minute walk before class, and 20 minutes of seated quiet reading before class. Every volunteer is assigned to one of the two by a coin flip or a random number generator — not by letting students pick, which would sort motivated students into the walking group and hand you a lurking variable.
Step 5 — Blinding and placebos. The students obviously know whether they walked, so subjects cannot be blinded. The seated-reading condition is the placebo here: it is an inactive treatment that still puts the student through "being in the study," so the two groups are matched on everything except the walk. The instructor grading the quizzes should be blinded — grade the papers without knowing which group each student was in — so the grader's expectations cannot leak into the scores.
Answer: a complete design names one explanatory variable, one response variable, a defined population and unit, at least two treatments assigned at random, a placebo condition to absorb the power of suggestion, and blinding wherever the study allows it — here, of the grader.
Problem 2. Discuss potential violations of the rule requiring informed consent.
a) People in a correctional facility are offered good behavior credit in return for participation in a study.
b) A research study is designed to investigate a new children's allergy medication.
c) Participants in a study are told that the new medication being tested is highly promising, but they are not told that only a small portion of participants will receive the new medication. Others will receive placebo treatments and traditional treatments.
Solution
Part a — Good behavior credit in a correctional facility. Consent has to be freely given, and it is not free when refusing costs you something. Offering incarcerated people credit toward good behavior makes participation feel compulsory — the "reward" is really a penalty for declining. That is coercion, and incarcerated people are a legally protected population precisely because of it. Fix: remove any benefit tied to release or standing, and get IRB review specifically for a protected population.
Part b — A children's allergy medication. Children cannot give informed consent; they are not legally able to weigh the risks. Consent must come from a parent or guardian, and the child should also be asked for assent in language they can understand. Fix: written guardian consent plus age-appropriate assent from the child.
Part c — Told the medication is promising, not told about the placebo groups. This is informed consent failing on both words. It is not informed, because a material fact — that most participants will receive a placebo or a traditional treatment — has been withheld. And the framing ("highly promising") oversells the benefit while the risk of receiving no active treatment goes unmentioned. Fix: disclose the randomization and every arm of the study up front, in plain language, before anyone signs.
Answer: (a) coercion of a captive population; (b) consent obtained from someone who cannot legally give it; (c) material facts withheld and benefits overstated. In all three, the fix is the same shape — full disclosure, freely given agreement, and IRB review before enrollment.
Problem 3. How does sleep deprivation affect your ability to drive? A recent study measured the effects on 19 professional drivers. Each driver participated in two experimental sessions: one after normal sleep and one after 27 hours of total sleep deprivation. The treatments were assigned in random order. In each session, performance was measured on a variety of tasks including a driving simulation. Use key terms from this module to describe the design of this experiment.
Solution
Step 1 — Identify the variables. The explanatory variable is sleep condition: normal sleep versus 27 hours of total sleep deprivation. The response variable is performance on the tasks, including the driving simulation.
Step 2 — Identify the units and treatments. The experimental units are the 19 professional drivers. The two treatments are the two sleep conditions.
Step 3 — Name the design. Each driver went through both sessions, so every driver is their own control — this is a matched or repeated-measures design. Because the two conditions were assigned in random order, the study is randomized: the order effects (practice on the simulator, time of day) are spread evenly instead of piling onto one condition.
Step 4 — Say what that buys you. Having each driver serve as their own comparison removes the lurking variables that differ between people — baseline reaction speed, experience, eyesight — because the same person appears in both groups. Randomizing the order removes the lurking variable that would otherwise differ within a person: which session came first.
Answer: a randomized, repeated-measures experiment on 19 experimental units, with sleep condition as the explanatory variable, task and driving-simulation performance as the response variable, two treatments (normal sleep, 27 hours of deprivation), and random ordering of the two sessions to control for order effects.
Problem 4. An advertisement for Acme Investments displays the two graphs in Figure 1.4.2 and Figure 1.4.3 to show the value of Acme's product in comparison with the Other Guy's product. Describe the potentially misleading visual effect of these comparison graphs. How can this be corrected?
Figure 1.4.2 — Acme Investments advertisement, graph (a): the value of Acme's product over time.
Figure 1.4.3 — Acme Investments advertisement, graph (b): the value of the Other Guy's product over the same period.
Solution
Step 1 — Look at what is missing from both graphs. Neither Figure 1.4.2 nor Figure 1.4.3 has a single number on either axis. There are gridlines and tick marks, but nothing tells you what one gridline is worth in dollars, or what span of time the horizontal axis covers.
Step 2 — Ask what that lets the advertiser do. Without labeled scales, the two pictures can be drawn to any vertical scale the advertiser likes. Acme's line is drawn to fill its plot from bottom to top; the Other Guy's line is squeezed into the bottom quarter of a plot that is otherwise empty. If the Other Guy's vertical axis runs to a much larger maximum, that flat-looking line could represent the same growth as Acme's — or more.
Step 3 — Notice the shapes are nearly the same. Both lines start low, wobble through the same small dips and bumps, and end higher than they started. The visual difference between "steep climb" and "barely moving" is being produced by the framing, not by the data.
Step 4 — State the correction. Put a labeled, numbered vertical axis on both graphs, using the same scale and the same time range. Better still, plot both companies as two lines on one set of axes, so the reader compares the data instead of comparing two independently stretched pictures.
Answer: the graphs are misleading because neither axis is labeled or scaled, so the reader cannot tell that the two vertical scales may differ — the apparent gap in performance is an artifact of how each plot was framed. Correct it by drawing both series on one set of clearly labeled, identically scaled axes.
Problem 5. The graph in Figure 1.4.4 shows the number of complaints for six different airlines as reported to the US Department of Transportation in February 2013. Alaska, Pinnacle, and Airtran Airlines have far fewer complaints reported than American, Delta, and United. Can we conclude that American, Delta, and United are the worst airline carriers since they have the most complaints?
Figure 1.4.4 — Number of complaints reported to the US Department of Transportation in February 2013, by airline.
Solution
Step 1 — Read the graph. Figure 1.4.4 shows raw complaint counts: roughly 126 for United, 118 for American, 49 for Delta, and about 6 or 7 each for Alaska, Pinnacle, and Airtran.
Step 2 — Ask what the counts are missing. A complaint count is a numerator with no denominator. United and American carry an enormous number of passengers each month; Pinnacle and Airtran carry a small fraction of that. Even if all six airlines annoyed exactly the same share of their passengers, the big carriers would still generate far more complaints in total, simply because they flew far more people.
Step 3 — Say what would settle the question. The comparable figure is a rate: complaints per 100,000 passengers boarded, over the same month. Only that number puts the six airlines on the same footing.
Step 4 — Note the second problem. Complaints are also self-reported. Differences in how easy each airline makes it to file a complaint, and differences in what its passengers expect, both affect the count without saying anything about service quality.
Answer: no. The graph shows raw counts, not rates, so it mostly measures how many passengers each airline carries. Without complaints per passenger, we cannot conclude that American, Delta, and United are the worst carriers.
Problem 6. Seven hundred and seventy-one distance learning students at Long Beach City College responded to surveys in a specific academic year. Highlights of the summary report are listed in Table 1.4.1.
| Survey item | Percent |
|---|---|
| Have computer at home | 96% |
| Unable to come to campus for classes | 65% |
| Age 41 or over | 24% |
| Would like LBCC to offer more DL courses | 95% |
| Took DL classes due to a disability | 17% |
| Live at least 16 miles from campus | 13% |
| Took DL courses to fulfill transfer requirements | 71% |
a) What percent of the students surveyed do not have a computer at home?
b) About how many students in the survey live at least 16 miles from campus?
c) If the same survey were done at Great Basin College in Elko, Nevada, do you think the percentages would be the same? Why?
Solution
Part a — Percent without a computer at home. 96% of the students surveyed have a computer at home, so the rest do not:
$$ 100\% - 96\% = 4\% $$Part b — How many live at least 16 miles from campus. 13% of the 771 respondents:
$$ 0.13 \times 771 = 100.23 \approx 100 \text{ students} $$Part c — Would Great Basin College look the same? Probably not, and the reason is worth stating carefully. These percentages describe the students who responded to a survey about distance learning at Long Beach City College — an urban campus in a dense part of Southern California. Great Basin College serves Elko, Nevada, a rural service area spread over a very large distance. You would expect a higher percentage there to report living at least 16 miles from campus and being unable to come to campus for classes, since distance is the whole reason many rural students enroll in distance learning at all. Other items, like the 96% with a computer at home, might come out lower if home internet access is thinner across the service area.
Answer: (a) 4%; (b) about 100 students; (c) no — the two colleges serve very different populations over very different distances, so a percentage measured at one is not a safe estimate for the other.
Problem 7. Several online textbook retailers advertise that they have lower prices than on-campus bookstores. However, an important factor is whether the Internet retailers actually have the textbooks that students need in stock. Students need to be able to get textbooks promptly at the beginning of the college term. If the book is not available, then a student would not be able to get the textbook at all, or might get a delayed delivery if the book is back ordered.
Avery Whitlock, a college newspaper reporter who uses they/them pronouns, is investigating textbook availability at online retailers. They decide to investigate one textbook for each of the following seven subjects: calculus, biology, chemistry, physics, statistics, geology, and general engineering. They consult textbook industry sales data and select the most popular nationally used textbook in each of these subjects. They visit websites for a random sample of major online textbook sellers and look up each of these seven textbooks to see if they are available in stock for quick delivery through these retailers. Based on this investigation, Whitlock writes an article drawing conclusions about the overall availability of all college textbooks through online textbook retailers.
Write an analysis of their study that addresses the following issues: Is his sample representative of the population of all college textbooks? Explain why or why not. Describe some possible sources of bias in this study, and how it might affect the results of the study. Give some suggestions about what could be done to improve the study.
Solution
Step 1 — Is the sample of textbooks representative? No. Whitlock's conclusion is about all college textbooks, but the seven books were chosen to be the most popular nationally used textbook in each of seven large, common subjects. Popular books in high-enrollment courses are exactly the books an online retailer is most likely to stock. A textbook for a small upper-division seminar, a regional press title, or a custom edition ordered by one department is a completely different case, and none of them had any chance of being selected.
Step 2 — Name the sources of bias.
- Selection bias in the books. Choosing bestsellers guarantees an optimistic estimate of availability. The sample is biased toward the easy cases.
- Selection bias in the subjects. Seven STEM-heavy subjects is not a cross-section of a college catalog. Humanities, languages, nursing, and trade programs are all absent.
- Sample size. Seven books cannot support a claim about a population of tens of thousands of titles, no matter how the retailers were sampled.
- Timing. Availability at the beginning of a term is a different question from availability in the middle of one, and the article's claim is about the moment students actually need the books.
- The one thing he did right. The retailers were a random sample of major online sellers. That part of the design is sound; the book selection is what breaks it.
Step 3 — How the bias moves the result. Every one of these pushes in the same direction: toward overstating how available college textbooks are online. The article's conclusion is likely to be too optimistic.
Step 4 — Suggest improvements. Build a sampling frame of the textbooks actually assigned at a set of colleges that term — the campus bookstore's adoption list is exactly this — and take a random sample from that frame, not the bestseller list. Cover every division of the catalog, not seven subjects. Increase the number of titles substantially. Check availability at the start of the term, when it matters, and repeat the check to see whether stock holds. And if only bestsellers can be sampled, narrow the published claim to match: "the most popular textbooks in seven large subjects," not "all college textbooks."
Answer: the sample is not representative — bestsellers in seven large subjects are the titles most likely to be in stock, so the study systematically overstates availability. Fix it by randomly sampling from an actual list of assigned textbooks across the whole catalog, using far more titles, and checking availability at the start of the term.
Key Terms
explanatory variable — the variable the researcher manipulates, believed to cause change in another variable.
response variable — the variable that is measured to see whether the explanatory variable had an effect.
treatment — one of the specific values of the explanatory variable that gets assigned to a group.
experimental unit — a single object or individual to be measured in the study.
lurking variable — an unaccounted-for variable that differs between groups and can be the real cause of a difference in the response.
random assignment — assigning experimental units to treatment groups by chance, which spreads lurking variables equally across groups.
block — a group of experimental units similar with respect to a variable expected to affect the response.
randomized block design — a design in which units are sorted into blocks first, and random assignment to treatments happens separately within each block.
control group — a group that receives no active treatment, used as a baseline for comparison.
placebo — an inactive treatment made to look like the active one, so that belief affects both groups equally.
blinding — keeping a person in the study from knowing which treatment was received.
double-blind experiment — a study in which both the subjects and the researchers working with them are blinded.
Institutional Review Board (IRB) — the committee that must approve a planned study to protect its human subjects.
informed consent — a subject's written agreement to participate after the risks have been clearly explained.
associated variables — two variables where knowing one tells you something about the likely value of the other.
independent variables — two variables where knowing one tells you nothing about the likely value of the other.