HSS.IC.B.5Common CoreMathStatistics and ProbabilityGrades 9-12
HSS.IC.B.5: Comparing Two Treatments with a Randomized Experiment and Simulation
In plain English: HSS.IC.B.5 is the Common Core statistics standard that asks students to use data from a randomized experiment to compare two treatments, and to use simulation to decide whether the difference between them is significant. Students re-randomize the results many times to see how large a difference chance alone produces. It is usually taught in Algebra II or an introductory statistics course.
Use data from a randomized experiment to compare two treatments; use simulations to decide if differences between parameters are significant.
Common Core State Standards for Mathematics · Domain: Making Inferences and Justifying Conclusions (IC) · Cluster: Make inferences and justify conclusions from sample surveys, experiments, and observational studies Also written as HSS-IC.B.5 or S-IC.5 · Official standard
Students analyze data from randomized experiments that compare two treatments. They compare the groups with dot plots and with the difference in means or in proportions, and they explain why random assignment lets them credit a difference to the treatment rather than to how the groups were formed.
The central question is whether an observed difference is larger than chance alone would produce. Students answer it with a randomization simulation: assume the treatment has no effect, shuffle all the results, deal them into two new groups of the original sizes, and record the difference. After hundreds of shuffles, they see how often chance gives a difference at least as large as the one observed. A rare result (a common cutoff is less than 5% of the shuffles) is called statistically significant.
Learning Objectives
By the end of this lesson, students will be able to:
Use data from a randomized experiment to compare two treatments with dot plots, differences in means and differences in proportions
Explain why random assignment allows a difference in responses to be attributed to the treatments
Carry out a randomization simulation that shows the differences chance alone produces when the treatment has no effect
Estimate from the simulation how often chance gives a difference at least as large as the observed one, and decide whether the difference is significant
State a conclusion in context, including what a result that is not significant does and does not show
Prior Knowledge Required
Students should already be comfortable with:
Comparing two data distributions by their centers and variability 7.SP.B.3
The difference between experiments and observational studies, and the role of random assignment HSS.IC.B.3
Computing means and proportions, and reading dot plots HSS.ID.A.1
Estimating a probability as a relative frequency from a simulation HSS.IC.A.2
Read the scenario aloud and let pairs talk for three minutes:
Warm-Up Prompt
"A basketball coach randomly assigns 8 players to a new free-throw routine and 8 players to the old routine. After a week, the new-routine players make an average of 2 more free throws out of 20. Does the new routine work better, or could the coach have seen a difference like this just by luck? What would you need to know to decide?"
Collect ideas. Students usually suggest that the answer depends on how much individual players vary, and on how many players were in each group. Record the idea "could it be luck?" on the board: the whole lesson is about measuring luck. Ask why the coach used random assignment instead of letting players pick a routine (players who choose the new routine might already be better shooters).
Direct Instruction20-25 minutes
Part 1: Comparing two treatments. In a randomized experiment, chance decides which subjects get which treatment. Chance spreads prior skill, motivation and other differences roughly evenly across the groups, so if the groups end up far apart, the treatment is the likely reason. Compare the groups with dot plots (Diagram 1) and with a single number: the difference in means for a numerical response, or the difference in proportions for a yes-or-no response.
Part 2: Is the difference significant? Even if the treatment did nothing, random assignment would rarely produce two groups with exactly the same mean. A randomization simulation shows how big those chance differences are:
Assume no treatment effect: each subject would have given the same response under either treatment.
Shuffle: write every response on a card (or put them in a spreadsheet column) and mix them.
Re-randomize: deal the cards into two groups with the same sizes as in the experiment, and compute the difference (first group minus second group).
Repeat many times, at least 100 by hand as a class, or 1,000 with technology, and plot the differences.
Compare: count the shuffles that give a difference at least as large as the observed one. If that fraction is small, commonly less than 0.05, chance is not a believable explanation and the difference is called statistically significant.
Typing experiment data, improvement in words per minute after 2 weeks
Group
Improvement (words per minute)
Mean
Game app (A)
12, 9, 15, 11, 8, 14, 10, 13, 7, 16
11.5
Standard drills (B)
8, 6, 10, 9, 5, 11, 7, 12, 4, 9
8.1
Comparing means
Twenty volunteers are randomly assigned, 10 to a typing game app (A) and 10 to standard drills (B). The table shows each person's improvement after 2 weeks.
Equation: Mean A = 115/10 = 11.5, mean B = 81/10 = 8.1, difference = 3.4 words per minute in favor of the app
Deciding significance by simulation
The 20 improvements are shuffled and re-dealt into two groups of 10, 1,000 times (Diagram 2). Only 11 shuffles give a difference of 3.4 or more.
Equation: Estimated probability ≈ 11/1,000 = 0.011, less than 0.05: the difference is statistically significant
Comparing proportions
Sixty students are randomly assigned to study vocabulary with flashcards or by rereading, 30 each. On the quiz, 22 flashcard students and 14 rereading students pass.
Equation: 22/30 ≈ 0.733 and 14/30 ≈ 0.467: a difference of 8/30 ≈ 0.267
Simulation for proportions
Make 36 "pass" cards and 24 "fail" cards, shuffle, and deal 30 to each group. In 1,000 shuffles, 30 give the flashcard group 22 or more passes.
Equation: Estimated probability ≈ 30/1,000 = 0.03, less than 0.05: significant, though less strongly than the typing result
A difference that is not significant
Twenty-four sprinters are randomly assigned to two warm-up routines. Routine 1 is faster by 0.15 seconds on average, and 190 of 500 shuffles give a difference of 0.15 seconds or more.
Equation: Estimated probability ≈ 190/500 = 0.38: chance alone often produces this difference, so it is not significant
Point out that the center of Diagram 2 is near 0, as it must be when the treatment has no effect, and that the observed 3.4 sits far out in the tail. For the last example, stress the careful wording: "not significant" means the data do not give convincing evidence of a difference. It does not prove that the two routines are equally good; a larger experiment might detect a small real effect.
Guided Practice15-20 minutes
Work this small experiment together so every student sees the whole randomization distribution. Six seedlings are randomly assigned, 3 to blue light and 3 to red light. Growth in centimeters: blue 6, 8, 9; red 3, 4, 5. The observed difference is 23/3 - 4 = 11/3 ≈ 3.67 cm. Pairs list all 20 ways to choose 3 of the 6 values for the "blue" group and compute each difference. Only one split, the actual one, gives a difference of 3.67 or more, so the probability is 1/20 = 0.05. Discuss: even the most extreme possible result is only just at the 5% cutoff, because 6 subjects are too few for strong evidence. Listen for pairs who miss some of the 20 splits, and for pairs who compute red minus blue for some splits and blue minus red for others.
Independent Practice15 minutes
Give students the simulation results for three invented experiments and ask for a conclusion in context for each. (1) A new bandage: in 1,000 shuffles, 2 give a healing-time difference at least as large as observed (0.002: significant, strong evidence). (2) A reading app: 70 of 1,000 shuffles (0.07: not significant at the 0.05 cutoff; weak evidence). (3) A new cafeteria menu and food waste: 450 of 1,000 shuffles (0.45: not significant; chance easily explains it). For each, students also say what the experimenter could do next.
Closure5-10 minutes
Exit ticket: In a randomized experiment, the mean score for treatment 1 is 4 points higher than for treatment 2. In 1,000 re-randomizations, 18 give a difference of 4 points or more. (1) Estimate the probability that chance alone gives a difference this large. (Answer: 0.018.) (2) Is the difference significant? Explain in one sentence. (3) What does each shuffle assume about the treatments?
Differentiation Strategies
For Struggling Students
Start with the 6-seedling example and physical cards before moving to 1,000 shuffles with technology
Give a recording sheet with columns for shuffle number, group 1 mean, group 2 mean and difference, with the subtraction order printed at the top
Provide a conclusion frame: "If the treatment had no effect, a difference of ___ or more would happen in about ___ of random assignments, so ___."
For Advanced Students
Ask students to redo the typing test two-sided, counting shuffles with a difference of 3.4 or more in either direction, and explain when a two-sided count is the better choice
Ask students to find the exact probability for the typing experiment by counting all 184,756 possible splits with a short program, and compare it with the simulated 0.011
Ask students to design an experiment with the same observed difference but 5 times as many subjects, predict how the randomization distribution changes, and test the prediction by simulation
Assessment Guidance
What to Look For
Check that students compute the difference in the same order every time and that their simulation keeps the original group sizes. In conclusions, look for three parts: the estimated probability from the simulation, a decision (significant or not), and a sentence in context that names the treatments. Watch for students who say a result that is not significant proves the treatments are the same, and for students who read the estimated probability as the chance that the treatment works.
02
Classroom Activities
3 Activities
1
Shuffle the Memory Cards
25 minPairs
Pairs carry out a randomization test by hand with index cards. The invented data come from an experiment in which 16 volunteers were randomly assigned to study a list of 20 words in silence (8 people) or with music (8 people), then wrote down all the words they remembered.
Observed difference, silence minus music: 2.125 words
Each pair writes the 16 values on 16 index cards
Procedure
Shuffle the 16 cards and deal two piles of 8. Label the first pile "silence" and the second "music"
Compute the mean of each pile and the difference, silence minus music
Repeat 10 times and add each difference to a class dot plot on the board
Count the class's differences that are 2.125 or more and divide by the total number of shuffles. With enough shuffles, the fraction should be close to 0.02
Discussion Questions
Why must each pile have exactly 8 cards?
Where is the class dot plot centered, and why?
Is the difference significant? What does that let us say about music and memory for these volunteers?
Modification for Distance Learning
Students paste the 16 values into a shared spreadsheet, add a column of random numbers, sort by that column and treat the first 8 rows as the silence group. Each student reports 10 differences in a shared table.
2
One Thousand Shuffles with Technology
20 minPairs
Pairs use a spreadsheet or a free online randomization applet to test a difference in proportions. The invented experiment: a library randomly assigns 80 borrowers, 40 to receive a text reminder before the due date and 40 to receive none. In the reminder group, 31 return their books on time; in the other group, 22 do.
Procedure
Compute the two proportions (31/40 = 0.775 and 22/40 = 0.55) and the observed difference, 0.225
Enter 53 values of 1 (on time) and 27 values of 0 (late), one per row
Shuffle, assign the first 40 rows to "reminder", compute the difference in proportions, and repeat 1,000 times
Plot the 1,000 differences and find the fraction that are 0.225 or more. It should come out close to 0.03
Discussion Questions
Why are there 53 ones and 27 zeros, whatever the shuffle?
Is the difference significant at the 0.05 cutoff? Would your decision change with a 0.01 cutoff?
Two pairs got 0.027 and 0.034. Why are their answers different, and does it matter?
Challenge Variation
Keep the same proportions but double the experiment to 80 borrowers per group (62 and 44 on time). Predict whether the fraction of shuffles at 0.225 or more goes up or down, then run the simulation to check.
3
Paper Helicopter Experiment
25 minGroups of 4
The class runs its own randomized experiment and analyzes it with a randomization test. Each group compares paper helicopters with short wings (5 cm) and long wings (8 cm) by measuring flight time from the same height.
Procedure
Each group builds 10 helicopters from the same template and paper clip weight
Flip a coin to decide the wing length for each helicopter until 5 have short wings and 5 have long wings, then cut the wings
Drop each helicopter once from the same height and time its flight with a stopwatch; randomize the order of the drops too
Compute the mean flight time for each wing length and the difference, long minus short
Run at least 200 shuffles with technology, estimate the probability of a difference at least as large as observed, and write a conclusion
Discussion Questions
Why randomize which helicopters get long wings, and the order of the drops?
Other groups got different differences. Did every group reach the same conclusion? Why might they not?
What sources of variation, other than wing length, affected the flight times?
03
Diagrams & Visual Aids
2 diagrams
Diagram 1: Dot Plots of the Two Treatment Groups
Each dot is one volunteer's improvement in words per minute, drawn to scale. The dashed lines mark the group means, 11.5 for the app and 8.1 for standard drills. The groups overlap, so the plots alone do not settle whether the 3.4-point difference could be chance.
Diagram 2: The Randomization Distribution
Differences in means from 1,000 random re-assignments of the 20 improvements into two groups of 10, assuming the app has no effect. Each bar is one possible difference (they come in steps of 0.2). The distribution is centered near 0. Only 11 shuffles, shown in red, reach the observed 3.4, so chance alone rarely produces a difference this large.
04
Homework Assignment
~30 min
HSS.IC.B.5 Homework: Comparing Treatments and Testing Differences
Directions: Show all work. Compute every difference as the first treatment minus the second. For every simulation result, give the estimated probability, say whether the difference is significant at the 0.05 cutoff and write a conclusion in context.
Part 1: Comparing Two Treatments (Problems 1-3)
A teacher randomly assigns 16 volunteer students to review for a test with practice quizzes (8 students) or by rereading notes (8 students). Scores: practice quizzes 84, 90, 78, 88, 92, 81, 86, 85; rereading 80, 76, 85, 79, 82, 74, 83, 77. (a) Find each group's mean and the difference. (b) Sketch dot plots on the same scale and describe how the groups compare.
A clinic randomly assigns 100 patients, 50 to receive a reminder phone call before their appointment and 50 to receive no call. Of the reminder group, 42 keep their appointment; of the other group, 33 do. Find the two proportions and the difference in proportions.
For the clinic in Problem 2, name the two treatments and the response. Explain why random assignment matters if the clinic wants to say that the calls caused the difference, and describe what could go wrong if patients had chosen whether to receive a call.
Part 2: Deciding Significance by Simulation (Problems 4-6)
For the test scores in Problem 1, a student shuffles the 16 scores and re-deals them into two groups of 8, 1,000 times. Nine shuffles give a difference at least as large as the one you found in Problem 1(a). Estimate the probability, decide whether the difference is significant and write a conclusion.
For the clinic in Problem 2, a simulation uses 75 "kept" cards and 25 "missed" cards, dealt into two groups of 50. In 1,000 shuffles, 33 give the reminder group 42 or more kept appointments. Estimate the probability and decide whether the difference is significant. Then explain why the decision is close.
Six tomato plants are randomly assigned, 3 to fertilizer A and 3 to fertilizer B. Yields in kilograms: A 10, 12, 14; B 8, 9, 11. (a) Find the observed difference in means. (b) List all 20 ways to choose the 3 plants for group A and find how many give a difference at least as large as the observed one. (c) What is the probability, and is the difference significant? What does this tell you about experiments with very few subjects?
Rubric
Criterion
Full Credit (2 pts)
Partial Credit (1 pt)
No Credit (0 pts)
Comparing Treatments
Correct means or proportions and difference, with dot plots or a clear comparison
One computational error or comparison incomplete
Missing or incorrect
Simulation
Keeps group sizes, assumes no effect, counts results at least as large as observed
Simulation mostly correct with one flaw
No workable simulation
Significance Decision
Correct probability and decision at the stated cutoff
Correct probability, decision missing or unjustified
Missing or incorrect
Conclusion in Context
Names the treatments, explains the role of random assignment and what the result does not show
Conclusion vague or overstated
No conclusion
05
Quiz: 20 Questions
Interactive, with answers
Instructions
Work through the questions in order. Your score updates as you answer, and Reset quiz clears everything so you or your students can try again.
Multiple choice: pick an option to check it. Short answer: write your answer, then reveal the model answer.
0 of 20 answered · 0 correct
Question 1 of 20 · Multiple Choice
A researcher wants to compare two cold remedies. Why should she use chance to decide which volunteers get which remedy?
Answer: B
Random assignment balances the groups on average, including on traits nobody measured, so a large difference in response points to the treatments. Choice C confuses random assignment with random sampling: generalizing to a population needs a random sample. Choice D is wrong because chance assignment still creates chance differences, which is exactly what the simulation measures.
Question 2 of 20 · Multiple Choice
In a randomized experiment, group X scores 5, 7, 6, 8 and group Y scores 4, 5, 3, 4. What is the difference in means, X minus Y?
Answer: A
Mean X = 26/4 = 6.5 and mean Y = 16/4 = 4, so the difference is 6.5 - 4 = 2.5. Choice B is the difference of the totals (26 - 16) rather than of the means. Choice D subtracts in the wrong order, Y minus X.
Question 3 of 20 · Multiple Choice
A randomized experiment gives 60 people a new sleep routine and 60 people the usual routine. In the new-routine group, 45 report better sleep; in the usual group, 33 do. What is the difference in proportions, new minus usual?
Answer: D
45/60 = 0.75 and 33/60 = 0.55, so the difference is 0.75 - 0.55 = 0.20. Choice A is the difference in counts, not proportions. Choice B divides the count difference by 100 instead of by the group size 60. Choice C is only the usual group's proportion.
Question 4 of 20 · Multiple Choice
In a randomization simulation for an experiment, what does each shuffle of the results assume?
Answer: C
Shuffling the responses between groups only makes sense if a response does not depend on the treatment, so the simulation models the world with no treatment effect. Choice A assumes what the experiment is trying to test. Choice B describes random sampling, which a randomization test does not need.
Question 5 of 20 · Multiple Choice
In 1,000 re-randomizations, 23 give a difference at least as large as the observed difference. What is the estimated probability of a difference this large by chance alone?
Answer: C
The estimated probability is 23/1,000 = 0.023. Choice B divides by 100 instead of 1,000. Choice D is the fraction of shuffles that gave a smaller difference, which is not what we compare with the cutoff.
Question 6 of 20 · Multiple Choice
A simulation shows that chance alone gives a difference at least as large as the observed one in about 31% of re-randomizations. Which conclusion fits?
Answer: D
A result that chance produces 31% of the time is not unusual, so the data do not give convincing evidence of a treatment effect. Choice B overstates the conclusion: "not significant" is not proof of no effect. Choice A reverses the logic: a large probability means the difference is easy to explain by chance.
Question 7 of 20 · Multiple Choice
An experiment has 12 subjects in treatment 1 and 12 in treatment 2. Which procedure correctly produces one simulated difference?
Answer: A
The simulation re-randomizes the actual responses into groups of the original sizes. Choice B never moves any response between groups, so the difference cannot change. Choice C changes the group sizes. Choice D ignores the real data.
Question 8 of 20 · Multiple Choice
A news report says an experiment found a "statistically significant" difference between two treatments. What does that mean?
Answer: B
Significance is about chance: a difference this large would be rare if the treatments had the same effect. Choice A confuses statistical significance with practical importance; a tiny difference can be significant in a very large experiment. Choice C describes groups that do not overlap, which is not required.
Question 9 of 20 · Multiple Choice
Two experiments find the same difference in means. Experiment 1 has 10 subjects per group and experiment 2 has 100 per group, with similar variability. How do their randomization distributions compare?
Answer: C
With more subjects, the means of randomly formed groups stay closer together, so chance differences are smaller and the distribution is narrower. The same observed difference then sits farther out in the tail. Choice D is wrong: a randomization distribution is centered near 0, because it assumes no treatment effect.
Question 10 of 20 · Multiple Choice
Four experiments were analyzed with 1,000 shuffles each. Which result gives the strongest evidence of a real treatment effect?
Answer: B
The fewer shuffles that match or beat the observed difference, the harder it is to explain by chance. 3/1,000 = 0.003 is the smallest. Choice C (0.048) is just under 0.05, so it is significant but weaker evidence. Choices A and D are not significant at the 0.05 cutoff.
Question 11 of 20 · Multiple Choice
The observed difference in an experiment is 2.5. Twenty shuffles give these differences: -3, -2, -2, -1.5, -1, -1, -0.5, -0.5, 0, 0, 0, 0.5, 0.5, 1, 1, 1.5, 2, 2, 2.5, 3. What fraction of the shuffles is at least as large as the observed difference?
Answer: B
Two values, 2.5 and 3, are at least 2.5, so the fraction is 2/20 = 0.10. Choice A counts only the value greater than 2.5 and misses the tie, which counts as "at least as large". Choice D also counts the two values of 2.
Question 12 of 20 · Multiple Choice
A student runs 200 "shuffles" for an experiment whose observed difference is 3.4, and every simulated difference comes out exactly 3.4. What went wrong?
Answer: D
If the responses are never moved between groups, each trial reproduces the original groups and the original difference. A correct simulation mixes all responses and re-deals them, which gives differences spread around 0. Choice A is wrong because more trials of the same mistake still give 3.4 every time. Choice B would change the differences, not freeze them at one value.
Question 13 of 20 · Multiple Choice
In an experiment with 30 subjects, 18 succeed in all. Which set of cards models the simulation for the difference in success proportions between two groups of 15?
Answer: A
The total number of successes stays 18 because, with no treatment effect, the same subjects succeed whatever group they are placed in. Choice B changes the total number of successes. Choice C adds subjects who were not in the experiment.
Question 14 of 20 · Multiple Choice
An experiment's difference is not significant. A student concludes, "So the treatment definitely does not work." What is wrong with this conclusion?
Answer: C
A randomization test can show that chance is an unlikely explanation, but it cannot prove the absence of an effect. A small real effect can easily be hidden by chance in a small experiment. Choice B changes the rule after seeing the data, which is not a sound practice.
Question 15 of 20 · Short Answer
In a randomized experiment, 7 volunteers chew mint gum and 7 chew no gum before a memory test. Words remembered: gum 15, 12, 17, 14, 13, 16, 11; no gum 12, 10, 14, 11, 13, 9, 15. Compute the difference in means (gum minus no gum) and say what else you need to decide whether it is significant.
Mean gum = 98/7 = 14; mean no gum = 84/7 = 12; difference = 2 words. To decide whether it is significant, you need a simulation: shuffle the 14 scores many times, re-deal 7 and 7, and see how often chance alone gives a difference of 2 or more.
Question 16 of 20 · Short Answer
A randomized experiment compares two fertilizers. The observed difference in mean yield is 1.8 kg. In 500 re-randomizations, 6 give a difference of 1.8 kg or more. Estimate the probability, decide whether the difference is significant and write a conclusion in context.
Estimated probability ≈ 6/500 = 0.012. This is less than 0.05, so the difference is statistically significant. If the fertilizers had the same effect, a difference of 1.8 kg or more would happen in only about 1.2% of random assignments, so the data give convincing evidence that the first fertilizer produces a higher mean yield for plants like these.
Question 17 of 20 · Short Answer
A company randomly assigns 200 customers to see website design A or design B, 100 each; 28 buy with A and 17 buy with B. Describe step by step how to use a simulation to decide whether the difference in proportions is significant.
Observed difference: 0.28 - 0.17 = 0.11. (1) Make 45 "buy" cards and 155 "no buy" cards. (2) Shuffle and deal 100 cards to A and 100 to B. (3) Compute the difference in proportions, A minus B. (4) Repeat 1,000 times with technology. (5) Find the fraction of shuffles with a difference of 0.11 or more. If it is below 0.05, the difference is significant and design A likely increases buying; if not, chance is a believable explanation.
Question 18 of 20 · Short Answer
Four volunteers are randomly assigned, 2 to treatment A and 2 to treatment B. Responses: A 7, 9; B 3, 5. List all 6 possible splits, find the probability of a difference (A minus B) at least as large as the observed one, and decide whether the difference is significant.
Observed: 8 - 4 = 4. Splits for A: {7, 9} gives 4; {7, 5} gives 0; {7, 3} gives -2; {9, 5} gives 2; {9, 3} gives 0; {5, 3} gives -4. Only one split reaches 4, so the probability is 1/6 ≈ 0.17. Not significant: with 4 subjects, even the largest possible difference happens 1 time in 6 by chance.
Question 19 of 20 · Short Answer
A student says, "The app group improved by 3 points more than the other group, so the app works. Why do we need a simulation?" Answer the student.
Random assignment alone creates differences between groups, even when the treatment does nothing, because different people land in each group. The simulation shows how big those chance differences typically are for this data. A 3-point difference is convincing only if it is larger than what chance usually produces; if shuffles often give 3 points or more, the app may not be the reason.
Question 20 of 20 · Short Answer
A significant difference was found in an experiment that used 40 volunteers from one high school's robotics club. Can the researchers conclude that the treatment causes the effect? Can they apply the result to all high school students? Explain.
Cause: yes, for these subjects, because the treatments were randomly assigned, which makes chance and the treatment the only differences between the groups, and the simulation made chance unlikely. All high school students: not directly. The volunteers were not a random sample of high school students, and robotics club members may respond differently. Generalizing needs a random sample or more experiments with other groups.
0 of 20 answered · 0 correct
06
Frequently Asked Questions
10 Questions
What does HSS.IC.B.5 mean?
HSS.IC.B.5 means students compare two treatments using data from a randomized experiment and use simulation to decide whether the difference is significant. They compute a difference in means or proportions, then shuffle the results many times, as if the treatments had no effect, to see how often chance alone produces a difference that large.
Is HSS.IC.B.5 part of Algebra 2 or Statistics?
It is usually taught in Algebra II in schools that follow the Common Core course sequence, and in introductory statistics courses. The Common Core places it in the high school Statistics and Probability category, in the cluster on inferences from sample surveys, experiments and observational studies. The standard expects simulation, not formula-based tests.
What is a randomization test?
A randomization test is a simulation that checks whether a difference between two groups is larger than random assignment alone would produce. You pool all the responses, shuffle them, re-deal them into groups of the original sizes and record the difference, many times. The fraction of shuffles with a difference at least as large as the observed one estimates how likely such a result is by chance.
What does "statistically significant" mean in this standard?
It means the observed difference would be unusual if the treatments had the same effect. In practice, students call a difference significant when their simulation shows it, or a larger one, in only a small fraction of the shuffles. A common cutoff is 0.05, or 5% of the shuffles. Significant does not mean large or important; it means hard to explain by chance.
Why is 0.05 used as the cutoff?
The 0.05 cutoff is a convention, not a law of mathematics. It is widely used in science and in statistics courses, but some fields use 0.01 when a false claim would be costly. Students should state the cutoff before looking at the simulation and should report the estimated probability itself, so a reader can judge the strength of the evidence.
Should students count differences in one direction or both?
It depends on the question. If the researcher expected one treatment to be better (the app increases typing speed), count the shuffles with a difference at least as large in that direction. If the question is only whether the treatments differ, count the shuffles at least as far from 0 in either direction, which roughly doubles the fraction. Decide before running the simulation, and always compute the difference in the same order.
What mistakes do students make with randomization simulations?
Common errors include changing the group sizes when re-dealing, subtracting in different orders for different shuffles, and forgetting that a tie with the observed value counts as "at least as large". In conclusions, many students treat a result that is not significant as proof that the treatments are the same, or read the estimated probability as the chance that the treatment works. It is the chance of a difference this large if the treatment had no effect.
How many shuffles are enough?
For a class activity by hand, pooling 100 or more shuffles gives a usable picture. With technology, 1,000 or more is typical and gives estimates that change little from run to run. More shuffles make the estimated probability more precise; they do not change the true probability, which depends only on the data.
Why does HSS.IC.B.5 require a randomized experiment?
Random assignment is what lets the simulation model chance correctly and what allows a cause-and-effect conclusion. If subjects chose their own treatment, the groups could differ in other ways, such as motivation, and a significant difference could come from those differences instead of the treatment. The standard HSS.IC.B.3 introduces this distinction between experiments and observational studies.
How does this standard connect to the SAT and to later statistics?
The digital SAT's Problem-Solving and Data Analysis domain includes evaluating statistical claims from observational studies and experiments, such as whether a conclusion about cause is justified. In later statistics courses, the randomization distribution is replaced by formulas such as the two-sample t-test, which answer the same question. Students who understand the shuffling idea usually find those formulas easier to interpret.
07
Related Standards
6 standards
These standards connect to HSS.IC.B.5: prerequisites to review first, parallel standards at the same level, and next steps that build on it.
Before this lesson
7.SP.B.3Prerequisite
Informally assess the visual overlap of two numerical data distributions