SVHS Website Header

SVHS Website Header Component

Scroll down or resize the browser to test responsive behavior. Hover over the nav items to open mega menus.

My Cart

HSS.IC.B.5Common CoreMathStatistics and ProbabilityGrades 9-12

HSS.IC.B.5: Comparing Two Treatments with a Randomized Experiment and Simulation

In plain English: HSS.IC.B.5 is the Common Core statistics standard that asks students to use data from a randomized experiment to compare two treatments, and to use simulation to decide whether the difference between them is significant. Students re-randomize the results many times to see how large a difference chance alone produces. It is usually taught in Algebra II or an introductory statistics course.

Use data from a randomized experiment to compare two treatments; use simulations to decide if differences between parameters are significant.

Common Core State Standards for Mathematics · Domain: Making Inferences and Justifying Conclusions (IC) · Cluster: Make inferences and justify conclusions from sample surveys, experiments, and observational studies
Also written as HSS-IC.B.5 or S-IC.5 · Official standard

01

Lesson Plan

65-80 min

Overview

Students analyze data from randomized experiments that compare two treatments. They compare the groups with dot plots and with the difference in means or in proportions, and they explain why random assignment lets them credit a difference to the treatment rather than to how the groups were formed.

The central question is whether an observed difference is larger than chance alone would produce. Students answer it with a randomization simulation: assume the treatment has no effect, shuffle all the results, deal them into two new groups of the original sizes, and record the difference. After hundreds of shuffles, they see how often chance gives a difference at least as large as the one observed. A rare result (a common cutoff is less than 5% of the shuffles) is called statistically significant.

Learning Objectives

By the end of this lesson, students will be able to:

  • Use data from a randomized experiment to compare two treatments with dot plots, differences in means and differences in proportions
  • Explain why random assignment allows a difference in responses to be attributed to the treatments
  • Carry out a randomization simulation that shows the differences chance alone produces when the treatment has no effect
  • Estimate from the simulation how often chance gives a difference at least as large as the observed one, and decide whether the difference is significant
  • State a conclusion in context, including what a result that is not significant does and does not show

Prior Knowledge Required

Students should already be comfortable with:

  • Comparing two data distributions by their centers and variability 7.SP.B.3
  • The difference between experiments and observational studies, and the role of random assignment HSS.IC.B.3
  • Computing means and proportions, and reading dot plots HSS.ID.A.1
  • Estimating a probability as a relative frequency from a simulation HSS.IC.A.2

Lesson Procedure

65-80 minutes of class time across 5 phases.

  1. Warm-Up10 minutes

    Read the scenario aloud and let pairs talk for three minutes:

    Warm-Up Prompt

    "A basketball coach randomly assigns 8 players to a new free-throw routine and 8 players to the old routine. After a week, the new-routine players make an average of 2 more free throws out of 20. Does the new routine work better, or could the coach have seen a difference like this just by luck? What would you need to know to decide?"

    Collect ideas. Students usually suggest that the answer depends on how much individual players vary, and on how many players were in each group. Record the idea "could it be luck?" on the board: the whole lesson is about measuring luck. Ask why the coach used random assignment instead of letting players pick a routine (players who choose the new routine might already be better shooters).

  2. Direct Instruction20-25 minutes

    Part 1: Comparing two treatments. In a randomized experiment, chance decides which subjects get which treatment. Chance spreads prior skill, motivation and other differences roughly evenly across the groups, so if the groups end up far apart, the treatment is the likely reason. Compare the groups with dot plots (Diagram 1) and with a single number: the difference in means for a numerical response, or the difference in proportions for a yes-or-no response.

    Part 2: Is the difference significant? Even if the treatment did nothing, random assignment would rarely produce two groups with exactly the same mean. A randomization simulation shows how big those chance differences are:

    1. Assume no treatment effect: each subject would have given the same response under either treatment.
    2. Shuffle: write every response on a card (or put them in a spreadsheet column) and mix them.
    3. Re-randomize: deal the cards into two groups with the same sizes as in the experiment, and compute the difference (first group minus second group).
    4. Repeat many times, at least 100 by hand as a class, or 1,000 with technology, and plot the differences.
    5. Compare: count the shuffles that give a difference at least as large as the observed one. If that fraction is small, commonly less than 0.05, chance is not a believable explanation and the difference is called statistically significant.
    Typing experiment data, improvement in words per minute after 2 weeks
    GroupImprovement (words per minute)Mean
    Game app (A)12, 9, 15, 11, 8, 14, 10, 13, 7, 1611.5
    Standard drills (B)8, 6, 10, 9, 5, 11, 7, 12, 4, 98.1
    • Comparing means

      Twenty volunteers are randomly assigned, 10 to a typing game app (A) and 10 to standard drills (B). The table shows each person's improvement after 2 weeks.

      Equation: Mean A = 115/10 = 11.5, mean B = 81/10 = 8.1, difference = 3.4 words per minute in favor of the app

    • Deciding significance by simulation

      The 20 improvements are shuffled and re-dealt into two groups of 10, 1,000 times (Diagram 2). Only 11 shuffles give a difference of 3.4 or more.

      Equation: Estimated probability ≈ 11/1,000 = 0.011, less than 0.05: the difference is statistically significant

    • Comparing proportions

      Sixty students are randomly assigned to study vocabulary with flashcards or by rereading, 30 each. On the quiz, 22 flashcard students and 14 rereading students pass.

      Equation: 22/30 ≈ 0.733 and 14/30 ≈ 0.467: a difference of 8/30 ≈ 0.267

    • Simulation for proportions

      Make 36 "pass" cards and 24 "fail" cards, shuffle, and deal 30 to each group. In 1,000 shuffles, 30 give the flashcard group 22 or more passes.

      Equation: Estimated probability ≈ 30/1,000 = 0.03, less than 0.05: significant, though less strongly than the typing result

    • A difference that is not significant

      Twenty-four sprinters are randomly assigned to two warm-up routines. Routine 1 is faster by 0.15 seconds on average, and 190 of 500 shuffles give a difference of 0.15 seconds or more.

      Equation: Estimated probability ≈ 190/500 = 0.38: chance alone often produces this difference, so it is not significant

    Point out that the center of Diagram 2 is near 0, as it must be when the treatment has no effect, and that the observed 3.4 sits far out in the tail. For the last example, stress the careful wording: "not significant" means the data do not give convincing evidence of a difference. It does not prove that the two routines are equally good; a larger experiment might detect a small real effect.

  3. Guided Practice15-20 minutes

    Work this small experiment together so every student sees the whole randomization distribution. Six seedlings are randomly assigned, 3 to blue light and 3 to red light. Growth in centimeters: blue 6, 8, 9; red 3, 4, 5. The observed difference is 23/3 - 4 = 11/3 ≈ 3.67 cm. Pairs list all 20 ways to choose 3 of the 6 values for the "blue" group and compute each difference. Only one split, the actual one, gives a difference of 3.67 or more, so the probability is 1/20 = 0.05. Discuss: even the most extreme possible result is only just at the 5% cutoff, because 6 subjects are too few for strong evidence. Listen for pairs who miss some of the 20 splits, and for pairs who compute red minus blue for some splits and blue minus red for others.

  4. Independent Practice15 minutes

    Give students the simulation results for three invented experiments and ask for a conclusion in context for each. (1) A new bandage: in 1,000 shuffles, 2 give a healing-time difference at least as large as observed (0.002: significant, strong evidence). (2) A reading app: 70 of 1,000 shuffles (0.07: not significant at the 0.05 cutoff; weak evidence). (3) A new cafeteria menu and food waste: 450 of 1,000 shuffles (0.45: not significant; chance easily explains it). For each, students also say what the experimenter could do next.

  5. Closure5-10 minutes

    Exit ticket: In a randomized experiment, the mean score for treatment 1 is 4 points higher than for treatment 2. In 1,000 re-randomizations, 18 give a difference of 4 points or more. (1) Estimate the probability that chance alone gives a difference this large. (Answer: 0.018.) (2) Is the difference significant? Explain in one sentence. (3) What does each shuffle assume about the treatments?

Differentiation Strategies

For Struggling Students

  • Start with the 6-seedling example and physical cards before moving to 1,000 shuffles with technology
  • Give a recording sheet with columns for shuffle number, group 1 mean, group 2 mean and difference, with the subtraction order printed at the top
  • Provide a conclusion frame: "If the treatment had no effect, a difference of ___ or more would happen in about ___ of random assignments, so ___."

For Advanced Students

  • Ask students to redo the typing test two-sided, counting shuffles with a difference of 3.4 or more in either direction, and explain when a two-sided count is the better choice
  • Ask students to find the exact probability for the typing experiment by counting all 184,756 possible splits with a short program, and compare it with the simulated 0.011
  • Ask students to design an experiment with the same observed difference but 5 times as many subjects, predict how the randomization distribution changes, and test the prediction by simulation

Assessment Guidance

What to Look For

Check that students compute the difference in the same order every time and that their simulation keeps the original group sizes. In conclusions, look for three parts: the estimated probability from the simulation, a decision (significant or not), and a sentence in context that names the treatments. Watch for students who say a result that is not significant proves the treatments are the same, and for students who read the estimated probability as the chance that the treatment works.

02

Classroom Activities

3 Activities

1

Shuffle the Memory Cards

25 minPairs

Pairs carry out a randomization test by hand with index cards. The invented data come from an experiment in which 16 volunteers were randomly assigned to study a list of 20 words in silence (8 people) or with music (8 people), then wrote down all the words they remembered.

Setup

  • Silence: 14, 12, 15, 11, 13, 16, 12, 14 words (mean 13.375)
  • Music: 11, 10, 13, 9, 12, 10, 14, 11 words (mean 11.25)
  • Observed difference, silence minus music: 2.125 words
  • Each pair writes the 16 values on 16 index cards

Procedure

  • Shuffle the 16 cards and deal two piles of 8. Label the first pile "silence" and the second "music"
  • Compute the mean of each pile and the difference, silence minus music
  • Repeat 10 times and add each difference to a class dot plot on the board
  • Count the class's differences that are 2.125 or more and divide by the total number of shuffles. With enough shuffles, the fraction should be close to 0.02

Discussion Questions

  • Why must each pile have exactly 8 cards?
  • Where is the class dot plot centered, and why?
  • Is the difference significant? What does that let us say about music and memory for these volunteers?

Modification for Distance Learning

Students paste the 16 values into a shared spreadsheet, add a column of random numbers, sort by that column and treat the first 8 rows as the silence group. Each student reports 10 differences in a shared table.

2

One Thousand Shuffles with Technology

20 minPairs

Pairs use a spreadsheet or a free online randomization applet to test a difference in proportions. The invented experiment: a library randomly assigns 80 borrowers, 40 to receive a text reminder before the due date and 40 to receive none. In the reminder group, 31 return their books on time; in the other group, 22 do.

Procedure

  • Compute the two proportions (31/40 = 0.775 and 22/40 = 0.55) and the observed difference, 0.225
  • Enter 53 values of 1 (on time) and 27 values of 0 (late), one per row
  • Shuffle, assign the first 40 rows to "reminder", compute the difference in proportions, and repeat 1,000 times
  • Plot the 1,000 differences and find the fraction that are 0.225 or more. It should come out close to 0.03

Discussion Questions

  • Why are there 53 ones and 27 zeros, whatever the shuffle?
  • Is the difference significant at the 0.05 cutoff? Would your decision change with a 0.01 cutoff?
  • Two pairs got 0.027 and 0.034. Why are their answers different, and does it matter?

Challenge Variation

Keep the same proportions but double the experiment to 80 borrowers per group (62 and 44 on time). Predict whether the fraction of shuffles at 0.225 or more goes up or down, then run the simulation to check.

3

Paper Helicopter Experiment

25 minGroups of 4

The class runs its own randomized experiment and analyzes it with a randomization test. Each group compares paper helicopters with short wings (5 cm) and long wings (8 cm) by measuring flight time from the same height.

Procedure

  • Each group builds 10 helicopters from the same template and paper clip weight
  • Flip a coin to decide the wing length for each helicopter until 5 have short wings and 5 have long wings, then cut the wings
  • Drop each helicopter once from the same height and time its flight with a stopwatch; randomize the order of the drops too
  • Compute the mean flight time for each wing length and the difference, long minus short
  • Run at least 200 shuffles with technology, estimate the probability of a difference at least as large as observed, and write a conclusion

Discussion Questions

  • Why randomize which helicopters get long wings, and the order of the drops?
  • Other groups got different differences. Did every group reach the same conclusion? Why might they not?
  • What sources of variation, other than wing length, affected the flight times?

03

Diagrams & Visual Aids

2 diagrams

Diagram 1: Dot Plots of the Two Treatment Groups

Typing experiment: improvement in words per minute Game app (A) 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 mean 11.5 Standard drills (B) 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 mean 8.1
Each dot is one volunteer's improvement in words per minute, drawn to scale. The dashed lines mark the group means, 11.5 for the app and 8.1 for standard drills. The groups overlap, so the plots alone do not settle whether the 3.4-point difference could be chance.

Diagram 2: The Randomization Distribution

1,000 re-randomizations of the 20 typing results 0 10 20 30 40 50 60 70 -5 -4 -3 -2 -1 0 1 2 3 4 5 observed 3.4 11 of 1,000 at 3.4 or more Difference in means, A minus B (words per minute), if the app had no effect Number of re-randomizations
Differences in means from 1,000 random re-assignments of the 20 improvements into two groups of 10, assuming the app has no effect. Each bar is one possible difference (they come in steps of 0.2). The distribution is centered near 0. Only 11 shuffles, shown in red, reach the observed 3.4, so chance alone rarely produces a difference this large.

04

Homework Assignment

~30 min

HSS.IC.B.5 Homework: Comparing Treatments and Testing Differences

Directions: Show all work. Compute every difference as the first treatment minus the second. For every simulation result, give the estimated probability, say whether the difference is significant at the 0.05 cutoff and write a conclusion in context.

Part 1: Comparing Two Treatments (Problems 1-3)

  1. A teacher randomly assigns 16 volunteer students to review for a test with practice quizzes (8 students) or by rereading notes (8 students). Scores: practice quizzes 84, 90, 78, 88, 92, 81, 86, 85; rereading 80, 76, 85, 79, 82, 74, 83, 77. (a) Find each group's mean and the difference. (b) Sketch dot plots on the same scale and describe how the groups compare.
  2. A clinic randomly assigns 100 patients, 50 to receive a reminder phone call before their appointment and 50 to receive no call. Of the reminder group, 42 keep their appointment; of the other group, 33 do. Find the two proportions and the difference in proportions.
  3. For the clinic in Problem 2, name the two treatments and the response. Explain why random assignment matters if the clinic wants to say that the calls caused the difference, and describe what could go wrong if patients had chosen whether to receive a call.

Part 2: Deciding Significance by Simulation (Problems 4-6)

  1. For the test scores in Problem 1, a student shuffles the 16 scores and re-deals them into two groups of 8, 1,000 times. Nine shuffles give a difference at least as large as the one you found in Problem 1(a). Estimate the probability, decide whether the difference is significant and write a conclusion.
  2. For the clinic in Problem 2, a simulation uses 75 "kept" cards and 25 "missed" cards, dealt into two groups of 50. In 1,000 shuffles, 33 give the reminder group 42 or more kept appointments. Estimate the probability and decide whether the difference is significant. Then explain why the decision is close.
  3. Six tomato plants are randomly assigned, 3 to fertilizer A and 3 to fertilizer B. Yields in kilograms: A 10, 12, 14; B 8, 9, 11. (a) Find the observed difference in means. (b) List all 20 ways to choose the 3 plants for group A and find how many give a difference at least as large as the observed one. (c) What is the probability, and is the difference significant? What does this tell you about experiments with very few subjects?

Rubric

CriterionFull Credit (2 pts)Partial Credit (1 pt)No Credit (0 pts)
Comparing TreatmentsCorrect means or proportions and difference, with dot plots or a clear comparisonOne computational error or comparison incompleteMissing or incorrect
SimulationKeeps group sizes, assumes no effect, counts results at least as large as observedSimulation mostly correct with one flawNo workable simulation
Significance DecisionCorrect probability and decision at the stated cutoffCorrect probability, decision missing or unjustifiedMissing or incorrect
Conclusion in ContextNames the treatments, explains the role of random assignment and what the result does not showConclusion vague or overstatedNo conclusion

05

Quiz: 20 Questions

Interactive, with answers

Instructions

Work through the questions in order. Your score updates as you answer, and Reset quiz clears everything so you or your students can try again.

Multiple choice: pick an option to check it. Short answer: write your answer, then reveal the model answer.

0 of 20 answered · 0 correct

  1. Question 1 of 20 · Multiple Choice

    A researcher wants to compare two cold remedies. Why should she use chance to decide which volunteers get which remedy?

  2. Question 2 of 20 · Multiple Choice

    In a randomized experiment, group X scores 5, 7, 6, 8 and group Y scores 4, 5, 3, 4. What is the difference in means, X minus Y?

  3. Question 3 of 20 · Multiple Choice

    A randomized experiment gives 60 people a new sleep routine and 60 people the usual routine. In the new-routine group, 45 report better sleep; in the usual group, 33 do. What is the difference in proportions, new minus usual?

  4. Question 4 of 20 · Multiple Choice

    In a randomization simulation for an experiment, what does each shuffle of the results assume?

  5. Question 5 of 20 · Multiple Choice

    In 1,000 re-randomizations, 23 give a difference at least as large as the observed difference. What is the estimated probability of a difference this large by chance alone?

  6. Question 6 of 20 · Multiple Choice

    A simulation shows that chance alone gives a difference at least as large as the observed one in about 31% of re-randomizations. Which conclusion fits?

  7. Question 7 of 20 · Multiple Choice

    An experiment has 12 subjects in treatment 1 and 12 in treatment 2. Which procedure correctly produces one simulated difference?

  8. Question 8 of 20 · Multiple Choice

    A news report says an experiment found a "statistically significant" difference between two treatments. What does that mean?

  9. Question 9 of 20 · Multiple Choice

    Two experiments find the same difference in means. Experiment 1 has 10 subjects per group and experiment 2 has 100 per group, with similar variability. How do their randomization distributions compare?

  10. Question 10 of 20 · Multiple Choice

    Four experiments were analyzed with 1,000 shuffles each. Which result gives the strongest evidence of a real treatment effect?

  11. Question 11 of 20 · Multiple Choice

    The observed difference in an experiment is 2.5. Twenty shuffles give these differences: -3, -2, -2, -1.5, -1, -1, -0.5, -0.5, 0, 0, 0, 0.5, 0.5, 1, 1, 1.5, 2, 2, 2.5, 3. What fraction of the shuffles is at least as large as the observed difference?

  12. Question 12 of 20 · Multiple Choice

    A student runs 200 "shuffles" for an experiment whose observed difference is 3.4, and every simulated difference comes out exactly 3.4. What went wrong?

  13. Question 13 of 20 · Multiple Choice

    In an experiment with 30 subjects, 18 succeed in all. Which set of cards models the simulation for the difference in success proportions between two groups of 15?

  14. Question 14 of 20 · Multiple Choice

    An experiment's difference is not significant. A student concludes, "So the treatment definitely does not work." What is wrong with this conclusion?

  15. Question 15 of 20 · Short Answer

    In a randomized experiment, 7 volunteers chew mint gum and 7 chew no gum before a memory test. Words remembered: gum 15, 12, 17, 14, 13, 16, 11; no gum 12, 10, 14, 11, 13, 9, 15. Compute the difference in means (gum minus no gum) and say what else you need to decide whether it is significant.

  16. Question 16 of 20 · Short Answer

    A randomized experiment compares two fertilizers. The observed difference in mean yield is 1.8 kg. In 500 re-randomizations, 6 give a difference of 1.8 kg or more. Estimate the probability, decide whether the difference is significant and write a conclusion in context.

  17. Question 17 of 20 · Short Answer

    A company randomly assigns 200 customers to see website design A or design B, 100 each; 28 buy with A and 17 buy with B. Describe step by step how to use a simulation to decide whether the difference in proportions is significant.

  18. Question 18 of 20 · Short Answer

    Four volunteers are randomly assigned, 2 to treatment A and 2 to treatment B. Responses: A 7, 9; B 3, 5. List all 6 possible splits, find the probability of a difference (A minus B) at least as large as the observed one, and decide whether the difference is significant.

  19. Question 19 of 20 · Short Answer

    A student says, "The app group improved by 3 points more than the other group, so the app works. Why do we need a simulation?" Answer the student.

  20. Question 20 of 20 · Short Answer

    A significant difference was found in an experiment that used 40 volunteers from one high school's robotics club. Can the researchers conclude that the treatment causes the effect? Can they apply the result to all high school students? Explain.

0 of 20 answered · 0 correct

06

Frequently Asked Questions

10 Questions

What does HSS.IC.B.5 mean?

HSS.IC.B.5 means students compare two treatments using data from a randomized experiment and use simulation to decide whether the difference is significant. They compute a difference in means or proportions, then shuffle the results many times, as if the treatments had no effect, to see how often chance alone produces a difference that large.

Is HSS.IC.B.5 part of Algebra 2 or Statistics?

It is usually taught in Algebra II in schools that follow the Common Core course sequence, and in introductory statistics courses. The Common Core places it in the high school Statistics and Probability category, in the cluster on inferences from sample surveys, experiments and observational studies. The standard expects simulation, not formula-based tests.

What is a randomization test?

A randomization test is a simulation that checks whether a difference between two groups is larger than random assignment alone would produce. You pool all the responses, shuffle them, re-deal them into groups of the original sizes and record the difference, many times. The fraction of shuffles with a difference at least as large as the observed one estimates how likely such a result is by chance.

What does "statistically significant" mean in this standard?

It means the observed difference would be unusual if the treatments had the same effect. In practice, students call a difference significant when their simulation shows it, or a larger one, in only a small fraction of the shuffles. A common cutoff is 0.05, or 5% of the shuffles. Significant does not mean large or important; it means hard to explain by chance.

Why is 0.05 used as the cutoff?

The 0.05 cutoff is a convention, not a law of mathematics. It is widely used in science and in statistics courses, but some fields use 0.01 when a false claim would be costly. Students should state the cutoff before looking at the simulation and should report the estimated probability itself, so a reader can judge the strength of the evidence.

Should students count differences in one direction or both?

It depends on the question. If the researcher expected one treatment to be better (the app increases typing speed), count the shuffles with a difference at least as large in that direction. If the question is only whether the treatments differ, count the shuffles at least as far from 0 in either direction, which roughly doubles the fraction. Decide before running the simulation, and always compute the difference in the same order.

What mistakes do students make with randomization simulations?

Common errors include changing the group sizes when re-dealing, subtracting in different orders for different shuffles, and forgetting that a tie with the observed value counts as "at least as large". In conclusions, many students treat a result that is not significant as proof that the treatments are the same, or read the estimated probability as the chance that the treatment works. It is the chance of a difference this large if the treatment had no effect.

How many shuffles are enough?

For a class activity by hand, pooling 100 or more shuffles gives a usable picture. With technology, 1,000 or more is typical and gives estimates that change little from run to run. More shuffles make the estimated probability more precise; they do not change the true probability, which depends only on the data.

Why does HSS.IC.B.5 require a randomized experiment?

Random assignment is what lets the simulation model chance correctly and what allows a cause-and-effect conclusion. If subjects chose their own treatment, the groups could differ in other ways, such as motivation, and a significant difference could come from those differences instead of the treatment. The standard HSS.IC.B.3 introduces this distinction between experiments and observational studies.

How does this standard connect to the SAT and to later statistics?

The digital SAT's Problem-Solving and Data Analysis domain includes evaluating statistical claims from observational studies and experiments, such as whether a conclusion about cause is justified. In later statistics courses, the randomization distribution is replaced by formulas such as the two-sample t-test, which answer the same question. Students who understand the shuffling idea usually find those formulas easier to interpret.