HSS.ID.A.4Common CoreMathStatistics and ProbabilityGrades 9-12
HSS.ID.A.4: Fitting Data to a Normal Distribution and Estimating Percentages
In plain English: HSS.ID.A.4 is the Common Core statistics standard that asks students to use the mean and standard deviation of a data set to fit a normal distribution and estimate the percent of a population in an interval. Students find normal-curve areas with tables, calculators or spreadsheets and recognize skewed or multi-peaked data where the model does not fit. It is usually taught in Algebra II or in a statistics course.
Use the mean and standard deviation of a data set to fit it to a normal distribution and to estimate population percentages. Recognize that there are data sets for which such a procedure is not appropriate. Use calculators, spreadsheets, and tables to estimate areas under the normal curve.
Common Core State Standards for Mathematics · Domain: Interpreting Categorical and Quantitative Data (ID) · Cluster: Summarize, represent, and interpret data on a single count or measurement variable Also written as HSS-ID.A.4 or S-ID.4 · Official standard
Students learn when and how to replace a data set with a normal curve. They compute the mean and standard deviation of a data set, use them to fit a normal distribution, and then use areas under that curve to estimate what percent of the population lies above, below or between given values. The 68-95-99.7 rule gives quick estimates, and z-tables, graphing calculators and spreadsheets give precise ones.
An equally important part of the standard is judgment. Many real data sets, such as waiting times, incomes or counts that pile up at zero, are skewed or have more than one peak. Students learn to look at the shape first, to test a fitted model against the data, and to reject a normal model when it predicts impossible values or clearly wrong percentages.
Learning Objectives
By the end of this lesson, students will be able to:
Fit a normal distribution to a data set by using its mean and standard deviation
Estimate population percentages and counts in an interval with the 68-95-99.7 rule and with z-scores
Find areas under a normal curve with a z-table, a graphing calculator and a spreadsheet
Decide from a histogram, summary statistics or a model check whether a normal model is appropriate for a data set
Prior Knowledge Required
Students should already be comfortable with:
Displaying data in dot plots, histograms and box plots HSS.ID.A.1
Choosing measures of center and spread to fit the shape of a distribution HSS.ID.A.2
Describing the shape of a data distribution, including symmetry, skew and outliers 6.SP.A.2
Computing a mean and reading a standard deviation from a calculator or spreadsheet
Project the prompt and give students three minutes to write a guess for each part before anyone calculates.
Warm-Up Prompt
"An orchard weighs 1,000 of its apples (invented data). The weights make a single, symmetric mound. The mean is 180 g and the standard deviation is 12 g. (1) Roughly what share of the apples would you expect between 168 g and 192 g? (2) Would you expect many apples under 144 g? (3) Could the same method work for the number of worms found in each apple?"
Collect guesses for (1) on the board, then reveal the idea the lesson makes precise: for mound-shaped, symmetric data, about 68% of values lie within one standard deviation of the mean, so about 680 of the 1,000 apples. For (2), 144 g is three standard deviations below the mean, so very few apples (well under 1%) should be that light. Leave (3) open: most apples have 0 worms, a few have 1 or 2, and none can have a negative number, so that distribution is piled up at 0 and skewed. Students come back to this question at the end of Direct Instruction.
Direct Instruction20 minutes
Fitting a normal model. A normal distribution is a symmetric, bell-shaped curve fixed by two numbers: its mean μ (the center) and its standard deviation σ (the spread). To fit a data set to a normal distribution, compute the mean x̄ and standard deviation s of the data and use the curve N(x̄, s). Areas under the curve then estimate the percent of the population in any interval. Present the steps:
Look first. Make a histogram or dot plot. Fit a normal model only if the data are roughly symmetric with a single peak and no strong outliers.
Fit. Use the mean and standard deviation of the data: the model is N(mean, SD).
Standardize. For a value x, the z-score z = (x - mean) / SD says how many standard deviations x is from the mean.
Find the area. Use the 68-95-99.7 rule for whole standard deviations, a z-table for the area to the left of z, a calculator (normalcdf(lower, upper, mean, SD)) or a spreadsheet (=NORM.DIST(x, mean, SD, TRUE) gives the area to the left of x).
Interpret. State the answer as an estimated percent (or count) of the population, in context, and say whether the model is believable.
Fit with the 68-95-99.7 rule
A company weighs 500 bags of trail mix (invented data). The histogram is a symmetric mound with mean 250 g and SD 4 g. Fit a normal model and estimate the percent of bags between 242 g and 258 g.
Equation: N(250, 4); 242 and 258 are 2 SD from the mean, so about 95% of bags (about 475 of 500)
z-score and a z-table
Resting pulse rates of 300 teens (invented) are mound-shaped with mean 72 beats per minute and SD 8. Estimate the percent of teens with a pulse above 84.
Equation: z = (84 - 72) / 8 = 1.5; table area to the left = 0.9332; above 84: 1 - 0.9332 = 0.0668, about 6.7%
Calculator and spreadsheet
Heights of 250 tenth-grade boys (invented) are mound-shaped with mean 68.5 in and SD 3.0 in. Estimate the percent between 65 in and 72 in.
Equation: normalcdf(65, 72, 68.5, 3) ≈ 0.757, or =NORM.DIST(72,68.5,3,TRUE)-NORM.DIST(65,68.5,3,TRUE) ≈ 0.757: about 75.7%
When the procedure is not appropriate
Wait times for 100 patients at a walk-in clinic (invented, Diagram 2) have mean 14 min and SD 11 min, with most waits short and a long right tail. What does N(14, 11) predict for waits under 0 minutes?
Equation: z = (0 - 14) / 11 ≈ -1.27; area ≈ 0.102, so the model predicts about 10% negative waits, which is impossible: do not use a normal model here
Use Diagram 1 with Example 1 so students see why 95% sits between 242 and 258. In Example 2, stress that a z-table gives the area to the left, so "above" needs 1 minus the table value. In Example 3, show both technologies on the projector and point out that the calculator and the spreadsheet give the same area. Then return to Warm-Up question (3) with Example 4 and Diagram 2: a normal curve can be drawn for any mean and SD, but if the data are skewed the areas it gives are wrong, here even putting 10% of the patients at negative wait times.
Guided Practice15 minutes
Pairs fit a normal model to a small data set and test it against the data. These are the lengths of 20 rainbow trout from one hatchery pond (invented data), in centimeters:
Invented lengths (cm) of 20 hatchery trout, in order
Fish 1-5
Fish 6-10
Fish 11-15
Fish 16-20
26.4
29.4
30.2
31.3
27.8
29.6
30.4
31.7
28.3
29.8
30.6
32.2
28.9
30.0
30.8
32.9
29.1
30.1
31.1
33.4
Steps and answers:
A dot plot shows one peak near 30 cm and rough symmetry, so a normal model is reasonable.
Mean x̄ = 604 / 20 = 30.2 cm; standard deviation s ≈ 1.7 cm (calculator: 1-Var Stats; spreadsheet: =STDEV.S). Model: N(30.2, 1.7).
Within 1 SD (28.5 to 31.9 cm) the model predicts 68%; the data have 14 of 20 = 70%. Within 2 SD (26.8 to 33.6 cm) the model predicts 95%; the data have 19 of 20 = 95%. The fit is good.
Estimate the percent of trout in the pond longer than 32 cm: z = (32 - 30.2) / 1.7 ≈ 1.06, and the area to the right is about 14.5%. The sample has 3 of 20 = 15%.
Ask pairs why the model estimate is useful even though the sample already gives 15%: the model describes the whole pond, and it gives answers for lengths that do not appear in a sample of 20.
Independent Practice10-15 minutes
Backpack weights of 400 students at one school (invented) are mound-shaped with mean 8.2 lb and SD 1.9 lb. Students answer each part three ways where they can: with a z-table, a calculator and a spreadsheet formula. (1) What percent of backpacks weigh more than 10 lb? (z ≈ 0.95, about 17%, so about 69 backpacks.) (2) What percent weigh less than 6 lb? (z ≈ -1.16, about 12%.) (3) A student says the model predicts about 0.0008% of backpacks weigh less than 0 lb, so the model is useless. Respond. (That share is tiny, far less than one backpack out of 400, so here the model is still reasonable; compare with Example 4, where the impossible region held 10%.)
Closure10 minutes
Exit ticket: Lengths of 900 new pencils from one factory (invented) are mound-shaped with mean 19.0 cm and SD 0.2 cm. (1) Use a z-table or a calculator to estimate the percent between 18.7 and 19.4 cm. (2) Name one data set about your school for which this method would give poor estimates, and say why. Answers: (1) z = -1.5 and z = 2, so the area is 0.9772 - 0.0668 = 0.9104, about 91.0%. (2) Sample answers: minutes of commute to school or number of siblings, which pile up at small values and are skewed right.
Differentiation Strategies
For Struggling Students
Give a blank normal curve with tick marks already labeled mean, ±1 SD, ±2 SD and ±3 SD, and the region percents 34, 13.5, 2.35 and 0.15 written in, so students fill in only the data values
Always have students sketch and shade the region before computing, and ask whether the answer should be more or less than 50%
Start with values that sit exactly on whole standard deviations, then move to z-scores such as 1.5 and 0.75
For Advanced Students
Give a data set of 30 values and ask students to compare the percent of data within 1, 2 and 3 SD with 68%, 95% and 99.7%, and to decide how far off the percents can be before the model should be rejected
Ask students to explain why a normal model of adult heights that mixes men and women gives worse estimates than two separate models
Have students build a spreadsheet column with =NORM.DIST that gives the expected count for each histogram bin and compare it with the observed counts
Assessment Guidance
What to Look For
Check that students look at the shape of the data before they fit anything, and that they can say in words why a normal model is or is not reasonable. When students compute areas, look for a shaded sketch, a correct z-score with its sign, and the right use of "area to the left" (1 minus the table value for "above", a difference of two values for "between"). Students should report answers as estimates about a population, with a percent or count in context, and should notice when a model predicts impossible values.
02
Classroom Activities
3 Activities
1
Fit It and Test It
20 minGroups of 3-4
Groups measure a quantity for everyone in the class, fit a normal model, and test whether the model matches their own data.
Procedure
Each student measures their hand span (thumb tip to little-finger tip, fingers spread) to the nearest half centimeter with a ruler and posts it on a sticky note on the board
Groups copy the class data, draw a histogram on grid paper and decide whether the shape is roughly symmetric with one peak
Groups compute the mean and standard deviation with a calculator or spreadsheet and write the model N(mean, SD)
Groups count the hand spans within 1 SD and within 2 SD of the mean and compare those percents with 68% and 95%
Each group uses its model to estimate the percent of students at the school with a hand span over 22 cm, and states one reason the estimate could be off
Discussion Questions
How close to 68% and 95% did your class come? What could explain the difference?
Is your class a good sample of all students at the school? Of all adults?
Would the model work better or worse with 300 students? Why?
Modification for Distance Learning
Students enter their measurements in a shared spreadsheet. Each group uses =AVERAGE, =STDEV.S and =COUNTIFS to test the model, and posts a screenshot of its histogram.
2
Three Tools, One Area
20 minGroups of 3
Each group member finds the same normal-curve areas with a different tool, a z-table, a graphing calculator or a spreadsheet, and the group reconciles the answers.
The Four Questions
Invented data: the masses of 2,000 chicken eggs from one farm are mound-shaped with mean 58 g and SD 5 g.
Card 1: the percent of eggs under 50 g (about 5.5%)
Card 2: the percent of eggs over 65 g (about 8.1%)
Card 3: the percent of eggs between 53 g and 63 g (about 68.3%)
Card 4: the number of the 2,000 eggs between 55 g and 70 g (about 1,435)
Procedure
Roles: the table reader computes z-scores and uses the printed z-table, the calculator user types normalcdf, and the spreadsheet user types NORM.DIST formulas
All three work each card alone, then compare. Answers should agree to within about 0.5 percentage points; the z-table is less exact because z is rounded to two decimals
Rotate roles after every card so each student uses every tool at least once
The group writes one sentence per card that interprets the answer for the farm
Discussion Questions
For Card 2, the table reader must subtract from 1. Why do the other two tools not need that step, or do they?
Which tool would you choose for 50 different intervals? Why?
Card 3 matches the 68-95-99.7 rule. Which tool needed the least work for it?
Challenge Variation
Ask the group to find the egg mass that only 10% of eggs exceed, using the z-table backward and then invNorm or NORM.INV, and to explain why the three answers differ slightly.
3
Normal or Not? Card Sort
15 minPairs
Pairs sort 8 data-set cards into "a normal model is reasonable" and "a normal model is not appropriate", and justify each choice with evidence from the card.
The 8 Cards (invented data)
Card A: diameters of 400 ball bearings from one machine; symmetric mound, mean 10.00 mm, SD 0.02 mm
Card B: monthly phone bills for 300 families; mean $85, median $62, long right tail
Card C: heights of 150 people at a family picnic that includes young children and adults; two peaks
Card D: scores of 1,000 students on a long multiple-choice exam; single symmetric peak
Card E: number of pets in 200 homes; mean 1.4, SD 1.5, many zeros
Card F: gestation lengths of 500 dairy calves; symmetric mound, mean 283 days, SD 6 days
Card G: ages of 250 people at a toddler swim class, including the parents; two clusters
Card H: masses of 600 bags of flour filled by one machine; symmetric mound, mean 2.27 kg, SD 0.01 kg
Procedure
Pairs sort the cards and write the evidence for each one: a shape, a mean far from the median, or an impossible prediction
For Card E, pairs compute what N(1.4, 1.5) predicts for fewer than 0 pets (about 18%) and use it as evidence
Two pairs join, compare their sorts and settle any disagreement
Discussion Questions
Which kinds of variables tend to be skewed? Why?
Card C could be split into two groups. Could each group then be modeled as normal?
Is it ever acceptable to use a normal model for data that are only roughly symmetric?
Modification for Distance Learning
Put the 8 cards on a shared slide with two labeled columns. Pairs drag each card into a column and type their evidence in a comment on the card.
03
Diagrams & Visual Aids
2 diagrams
Diagram 1: The 68-95-99.7 Rule on a Fitted Normal Model
The curve is N(250, 4), fitted to the invented trail-mix data in Worked Example 1 and drawn to scale. The darker band holds about 68% of the area (within 1 SD of the mean), the lighter band extends to about 95% (within 2 SD), and nearly all of the area (99.7%) lies within 3 SD. Each band is a percent of the population the model describes.
Diagram 2: A Data Set Where a Normal Model Fails
The bars show 100 invented clinic wait times from Worked Example 4, and the curve shows the counts a normal model with the same mean (14 min) and SD (11 min) would predict. The data pile up at short waits and have a long right tail, while the curve is symmetric and puts about 10% of its area below 0 minutes (shaded red), which is impossible. This is why a normal model is not appropriate for these data.
04
Homework Assignment
~30 min
HSS.ID.A.4 Homework: Normal Models and Percentages
Directions: All data sets are invented. For every problem, sketch the normal curve and shade the region before you compute. Show the z-score or the exact calculator or spreadsheet command you used, round percents to one decimal place, and answer in a sentence about the context.
Part 1: Fitting and the 68-95-99.7 Rule (Problems 1-2)
The lifetimes of 2,000 LED bulbs from one factory are mound-shaped and symmetric, with mean 25,000 hours and standard deviation 1,500 hours. (a) Write the normal model. (b) About what percent of the bulbs last between 23,500 and 26,500 hours? (c) About how many of the 2,000 bulbs last more than 28,000 hours? (d) About what percent last between 23,500 and 29,500 hours?
Twelve students measure their hand spans in centimeters: 18.5, 19.0, 19.5, 20.0, 20.0, 20.5, 20.5, 21.0, 21.5, 21.5, 22.0, 22.0. (a) Find the mean and the standard deviation s, and write the fitted normal model. (b) What percent of the 12 values lie within 1 SD of the mean? Within 2 SD? (c) Compare with 68% and 95%. Does a normal model look reasonable?
Part 2: Tables, Calculators and Spreadsheets (Problems 3-4)
A machine fills cans of juice. The fill amounts of 5,000 cans are mound-shaped with mean 12.10 oz and SD 0.05 oz. Use a z-table. (a) What percent of cans contain less than 12.00 oz? (b) What percent contain between 12.02 oz and 12.18 oz?
Scores on a district math test taken by 3,600 students are mound-shaped with mean 74 and SD 9. (a) Write the calculator command and the spreadsheet formula that give the proportion of students who scored 80 or higher, and give the percent. (b) What percent scored between 60 and 90? (c) About how many of the 3,600 students is that?
Part 3: Is a Normal Model Appropriate? (Problems 5-6)
A coffee kiosk records what 200 customers spend. The mean is $6.40, the SD is $5.10, the median is $4.75 and the smallest amount is $1.95. (a) What percent of customers would the model N(6.40, 5.10) predict spend less than $0? (b) What does the comparison of the mean and the median suggest about the shape? (c) Should the kiosk use a normal model to estimate the percent of customers who spend more than $15? Explain.
The finishing times of 600 runners in a 5K race are mound-shaped with mean 28 minutes and SD 4 minutes. (a) About how many runners finished between 24 and 30 minutes? (b) About how many finished faster than 21 minutes? (c) The race also had 40 walkers, who all finished in 45-60 minutes. If their times were added, would a single normal model still be appropriate? Explain.
Rubric
Criterion
Full Credit (2 pts)
Partial Credit (1 pt)
No Credit (0 pts)
Fitting the Model
Correct mean and SD, model written as N(mean, SD)
Mean or SD computed incorrectly
No model or wrong parameters
Areas and Percents
Correct z-scores, shaded sketch and area for every part
Correct setup with one area error, such as forgetting 1 minus the table value
Areas missing or unrelated to the question
Technology
Calculator commands and spreadsheet formulas written correctly
Right tool with a wrong argument order
No command or formula
Judging the Model
Uses shape, mean versus median or impossible predictions to decide
States a decision with weak evidence
No decision or no reason
05
Quiz: 20 Questions
Interactive, with answers
Instructions
All data sets in this quiz are invented. Answer each question before you open its explanation. The score at the top counts your multiple-choice answers, and Reset quiz starts the whole set again.
Multiple choice: pick an option to check it. Short answer: write your answer, then reveal the model answer.
0 of 20 answered · 0 correct
Question 1 of 20 · Multiple Choice
The masses of 1,200 boxes of cereal (invented data) are mound-shaped and symmetric, with mean 368 g and SD 3 g. About what percent of the boxes weigh between 362 g and 374 g?
Answer: B
362 g and 374 g are each 2 standard deviations (6 g) from the mean, and about 95% of a normal distribution lies within 2 SD. Choice A is the percent within 1 SD, which would be 365 g to 371 g. Choice D is only one side of the mean, from 368 g to 374 g.
Question 2 of 20 · Multiple Choice
Reaction times of 400 students (invented) are mound-shaped with mean 250 milliseconds and SD 30 milliseconds. About what percent of the students have a reaction time slower (greater) than 310 milliseconds?
Answer: A
310 is 2 SD above the mean. About 95% of values lie within 2 SD, and the other 5% is split between the two tails, so about 2.5% lie above 310 (a z-table gives 2.3%). Choice B counts both tails. Choice D is the percent below 310.
Question 3 of 20 · Multiple Choice
On an invented test, the mean is 50 and the SD is 5. What is the z-score of a score of 58?
Answer: D
z = (58 - 50) / 5 = 8 / 5 = 1.6: the score is 1.6 standard deviations above the mean. Choice B stops at the difference 8 and does not divide by the SD. Choice C divides the SD by the difference. Choice A has the wrong sign: 58 is above the mean.
Question 4 of 20 · Multiple Choice
The weights of 500 bags of dog food (invented) are mound-shaped with mean 120 oz and SD 16 oz. A z-table gives 0.1056 as the area to the left of z = -1.25. About what percent of the bags weigh less than 100 oz?
Answer: C
z = (100 - 120) / 16 = -1.25, and the table value 0.1056 is already the area to the left, so about 10.6% weigh less than 100 oz. Choice A is 1 - 0.1056, the percent above 100 oz. Choice D is the 68-95-99.7 estimate for exactly 1 SD below the mean, but 100 oz is 1.25 SD below.
Question 5 of 20 · Multiple Choice
For which data set is a normal model least likely to be appropriate?
Answer: B
Incomes are strongly skewed to the right: most households earn moderate amounts and a few earn far more, so the mean is pulled above the median and the distribution is not a symmetric bell. Heights of one group, egg masses and machine-made lengths are usually mound-shaped and symmetric, so choices A, C and D are data sets where a normal model often works.
Question 6 of 20 · Multiple Choice
Invented scores are mound-shaped with mean 70 and SD 6. Which graphing-calculator command estimates the proportion of scores above 80? (The syntax is normalcdf(lower, upper, mean, SD).)
Answer: A
For "above 80" the lower bound is 80 and the upper bound is a very large number (1E99), with the mean and SD in that order; the result is about 0.048. Choice B gives the area below 80. Choice C swaps the mean and the SD. Choice D gives the area between the mean and 80.
Question 7 of 20 · Multiple Choice
Invented data are modeled by N(70, 6). Which spreadsheet formula gives the proportion of values below 64?
Answer: D
NORM.DIST(x, mean, SD, TRUE) returns the cumulative area to the left of x, here about 0.159 (64 is 1 SD below 70). Choice A uses FALSE, which returns the height of the curve, not an area. Choice B gives the area above 64. Choice C swaps x and the mean.
Question 8 of 20 · Multiple Choice
The diameters of 500 washers (invented) are mound-shaped and symmetric, with mean 12.4 mm and SD 0.8 mm. Which interval contains about the middle 68% of the washers?
Answer: B
The middle 68% of a normal model lies within 1 SD of the mean: 12.4 - 0.8 = 11.6 and 12.4 + 0.8 = 13.2. Choice A goes only half an SD each way. Choice C is the 2 SD interval, which holds about 95%. Choice D goes 1.5 SD each way.
Question 9 of 20 · Multiple Choice
The masses of 800 tomatoes from one greenhouse (invented) are mound-shaped with mean 150 g and SD 20 g. A z-table gives 0.9332 as the area to the left of z = 1.50. About how many of the tomatoes weigh more than 180 g?
Answer: D
z = (180 - 150) / 20 = 1.5, so the area above is 1 - 0.9332 = 0.0668, and 0.0668 × 800 ≈ 53 tomatoes. Choice A (747) is the number below 180 g. Choice B uses 16%, the tail beyond 1 SD, but 180 g is 1.5 SD above the mean.
Question 10 of 20 · Multiple Choice
Invented data on nightly homework time for 300 students have mean 40 minutes and SD 25 minutes, and no student reports less than 0 minutes. The model N(40, 25) puts about 5.5% of students below 0 minutes. What is the best conclusion?
Answer: C
A good model should put almost no area on impossible values. 5.5% of 300 is about 16 students with negative homework time, which signals a right-skewed distribution bounded at 0. Choice A ignores what the 5.5% means. Choice D changes only the tool: any method that uses the same normal curve inherits the same problem.
Question 11 of 20 · Multiple Choice
Scores on an invented reading test are mound-shaped with mean 100 and SD 15. About what percent of scores lie between 85 and 130?
Answer: D
85 is 1 SD below the mean and 130 is 2 SD above, so the area is about 0.9772 - 0.1587 = 0.8186, or 81.9% (normalcdf(85, 130, 100, 15) gives the same). Choice A is only the part from 85 to 100, and choice B only the part from 100 to 130. Choice C is the area below 130.
Question 12 of 20 · Multiple Choice
An invented data set has mean 30, median 22 and a long tail of large values. What does this suggest about using a normal model?
Answer: A
A mean well above the median shows that large values pull the mean up: the distribution is skewed right, not symmetric, so a normal curve with this mean and SD will misplace the percents. Choice B confuses being able to compute a mean and SD with the model fitting. Choice C gets the direction of the skew wrong.
Question 13 of 20 · Multiple Choice
Heights of 700 adult women (invented) are mound-shaped with mean 64 in and SD 2.5 in. A z-table shows an area of about 0.90 to the left of z = 1.28. About what height separates the tallest 10% from the rest?
Answer: C
Read the table backward: 90% of the area is to the left of z ≈ 1.28, so the height is 64 + 1.28 × 2.5 = 67.2 in. Choice A is only 1 SD above the mean, which leaves about 16% above it, not 10%. Choice B is 1.28 SD below the mean, the cutoff for the shortest 10%.
Question 14 of 20 · Multiple Choice
A histogram of the heights of everyone at a school event, including both sixth graders and adults (invented data), shows two clear peaks. Which statement is correct?
Answer: D
Two peaks mean the data come from two groups with different centers. A single normal curve has one peak, so it would put too much area between the two groups and too little near each peak. Choice B ignores the shape: a larger SD only widens the one peak. Choice C is the same error.
Question 15 of 20 · Short Answer
Eight students run 100 meters (invented times in seconds): 14.2, 13.8, 15.1, 14.6, 13.5, 14.9, 14.0, 14.3. (a) Find the mean and the standard deviation s. (b) Write the fitted normal model. (c) Use it to estimate the percent of students in the whole grade who would run slower than (more than) 15.0 seconds.
(a) The sum is 114.4, so the mean is 14.3 s, and s ≈ 0.55 s. (b) N(14.3, 0.55). (c) z = (15.0 - 14.3) / 0.55 ≈ 1.27, and the area to the right is about 10%. With only 8 students, this estimate is rough.
Question 16 of 20 · Short Answer
The run times of 1,500 phone batteries (invented) are mound-shaped with mean 9.5 hours and SD 0.6 hours. A z-table gives 0.0668 for z = -1.50 and 0.8413 for z = 1.00. What percent of the batteries last between 8.6 and 10.1 hours?
z = (8.6 - 9.5) / 0.6 = -1.5 and z = (10.1 - 9.5) / 0.6 = 1.0. The area between is 0.8413 - 0.0668 = 0.7745, so about 77.5% of the batteries.
Question 17 of 20 · Short Answer
A school records the number of days each of its 900 students was absent last year (invented data). The mean is 3.1 days, the SD is 4.2 days and many students have 0 absences. (a) What percent of students would N(3.1, 4.2) predict have fewer than 0 absences? (b) Is a normal model appropriate? Explain.
(a) z = (0 - 3.1) / 4.2 ≈ -0.74, so the model puts about 23% of students below 0 days. (b) No. Negative absences are impossible, and a model that puts nearly a quarter of its area there does not describe the data. The counts pile up at 0 with a long right tail, so the distribution is skewed right.
Question 18 of 20 · Short Answer
The weights of 1,000 apples in a shipment (invented) are mound-shaped with mean 182 g and SD 14 g. Write a spreadsheet formula that gives the proportion heavier than 200 g, and state the percent.
=1-NORM.DIST(200,182,14,TRUE), which is about 0.099, so about 9.9% of the apples weigh more than 200 g. The "1 -" is needed because NORM.DIST with TRUE gives the area to the left.
Question 19 of 20 · Short Answer
A student fits N(20, 4) to 50 invented data values. She finds that 40 of the 50 values (80%) lie between 16 and 24, and all 50 lie between 12 and 28. Does the normal model fit these data well? Explain using the 68-95-99.7 rule.
Not very well. A normal model predicts about 68% within 1 SD (16 to 24), but the data have 80%, and it predicts about 95% within 2 SD, while the data have 100%. The data are more tightly packed around the mean than a normal curve, so percents from the model, especially in the tails, would be off.
Question 20 of 20 · Short Answer
Heights of 2,500 adult men (invented) are mound-shaped with mean 69.2 in and SD 2.9 in. Use a calculator or spreadsheet to estimate how many of the men are taller than 72 in.
normalcdf(72, 1E99, 69.2, 2.9) ≈ 0.167 (z ≈ 0.97), and 0.167 × 2,500 ≈ 418 men, about 16.7%.
0 of 20 answered · 0 correct
06
Frequently Asked Questions
10 Questions
What does HSS.ID.A.4 mean?
HSS.ID.A.4 means students can replace a mound-shaped data set with a normal curve and use it to estimate percentages. They fit the curve with the mean and standard deviation of the data, find areas under it with tables, calculators or spreadsheets, and recognize data sets, such as skewed ones, for which this procedure does not work.
Is HSS.ID.A.4 Algebra 1 or Algebra 2?
It depends on the school. The Common Core appendix on course design places HSS.ID.A.4 in the traditional Algebra II course, and many schools also introduce the 68-95-99.7 rule in Algebra I or in a statistics course. The integrated pathway places it in Mathematics III.
What is the 68-95-99.7 rule?
It is a quick summary of the normal curve: about 68% of the area lies within 1 standard deviation of the mean, about 95% within 2 and about 99.7% within 3. It only gives estimates for whole standard deviations. For a value such as 1.5 SD above the mean, students need a z-table, a calculator or a spreadsheet.
How do you fit a normal distribution to a data set?
Compute the mean and the standard deviation of the data and use them as the center and spread of the curve, written N(mean, SD). Before that, check a histogram or dot plot: the data should be roughly symmetric with one peak. After fitting, compare the percent of data within 1 and 2 SD with 68% and 95% as a check.
When is it not appropriate to use a normal distribution?
A normal model is not appropriate when the data are strongly skewed, have more than one peak, have strong outliers, or pile up against a boundary such as 0. Waiting times, incomes and counts such as the number of pets are typical examples. A warning sign is a model that predicts a real share of impossible values, such as negative times.
How do I find the area under a normal curve on a TI-84 calculator?
Use normalcdf(lower, upper, mean, SD) from the DISTR menu. For "less than x" use a very small lower bound such as -1E99, and for "more than x" use a very large upper bound such as 1E99. The result is a proportion; multiply by 100 for a percent and by the population size for a count.
How do I use a spreadsheet to find normal percentages?
In Excel or Google Sheets, =NORM.DIST(x, mean, SD, TRUE) gives the area to the left of x. For "more than x" use 1 minus that value, and for "between a and b" subtract the two areas. The last argument must be TRUE: FALSE gives the height of the curve, which is not a percent.
What are common mistakes students make with normal distributions?
A frequent error is forgetting that a z-table gives the area to the left, so students report the area below when the question asks for above. Other common mistakes are dropping the minus sign on a z-score, swapping the mean and SD in a calculator command, and fitting a normal model without first looking at the shape of the data.
Is a z-score the same as a percentile?
No. A z-score says how many standard deviations a value is from the mean, and a percentile says what percent of values are below it. For a normal model the two are linked through the curve: z = 0 is the 50th percentile, and z = 1 is about the 84th percentile.
How does HSS.ID.A.4 connect to later statistics courses?
HSS.ID.A.4 is the first contact with the normal curve, which returns throughout statistics. It prepares students for margins of error in sample surveys (HSS.IC.B.4), for deciding whether a model is consistent with data (HSS.IC.A.2), and for the sampling distributions used in college statistics. On the digital SAT, related questions about data distributions appear in the Problem-Solving and Data Analysis domain.
07
Related Standards
5 standards
These standards connect to HSS.ID.A.4: prerequisites to review first, parallel standards at the same level, and next steps that build on it.
Before this lesson
HSS.ID.A.1Prerequisite
Represent data with plots on the number line: dot plots, histograms and box plots