HSS.ID.C.8Common CoreMathStatistics and ProbabilityGrades 9-12
HSS.ID.C.8: Computing and Interpreting the Correlation Coefficient
In plain English: HSS.ID.C.8 is the Common Core statistics standard that asks students to use technology to compute the correlation coefficient r of a linear fit and to explain what it says: its sign gives the direction of a linear association, and its distance from 0 gives the strength, from -1 to 1. It is usually taught in Algebra I and introductory Statistics.
Compute (using technology) and interpret the correlation coefficient of a linear fit.
Common Core State Standards for Mathematics · Domain: Interpreting Categorical and Quantitative Data (ID) · Cluster: Interpret linear models Also written as HSS-ID.C.8 or S-ID.8 · Official standard
Students learn to compute the correlation coefficient r with technology and to read it correctly. The value of r is always between -1 and 1. Its sign gives the direction of the linear association, and its distance from 0 gives the strength: values near 1 or -1 mean the points lie close to a line, and values near 0 mean a line describes the data poorly.
Because the standard asks students to compute r with technology, the lesson spends little time on the formula and more time on what r does and does not say. Students see that r measures only linear association, so a strong curved pattern can have r near 0; that one outlier can change r a great deal; that r has no units and does not change when units change or x and y are swapped; and that r is not the slope.
Learning Objectives
By the end of this lesson, students will be able to:
Compute the correlation coefficient of a linear fit with a graphing calculator, Desmos or a spreadsheet
Interpret the sign and size of r as the direction and strength of a linear association, in the context of the data
Explain why r near 0 does not rule out a strong nonlinear relationship, and check a scatter plot before trusting r
Describe how an outlier can change r, and why r does not change with units or when x and y are swapped
Distinguish r from the slope of the fitted line and from the proportion of points on the line
Prior Knowledge Required
Students should already be comfortable with:
Making scatter plots and describing positive, negative, linear and nonlinear association 8.SP.A.1
Informally fitting a line and judging how close the points are to it 8.SP.A.2
Fitting a linear function to a scatter plot with technology HSS.ID.B.6
Interpreting the slope and intercept of a linear model HSS.ID.C.7
Students sketch a quick scatter plot for each pair of variables and describe the association before any numbers appear.
Warm-Up Prompt
"For each pair, sketch what a scatter plot of 20 people might look like and say whether the association is positive, negative or neither: (1) a person's height and their arm span, (2) the number of absences in a semester and the final grade, (3) a person's height and the last digit of their phone number. Which pattern would be closest to a straight line?"
Expect (1) strong positive, (2) negative with more scatter, and (3) no association. Tell students that statisticians summarize "how close to a line, and which way" in one number, the correlation coefficient r, and that today they will compute it with technology and learn to read it.
Direct Instruction20 minutes
What r measures. The correlation coefficient r describes the direction and strength of the linear association between two quantitative variables. It is always between -1 and 1; r = 1 or r = -1 only when every point is on one line. Use Diagram 1 to connect values of r to pictures. A rough guide used in many textbooks, not a fixed rule: |r| of 0.8 or more is strong, 0.5 to 0.8 is moderate, and under 0.5 is weak.
Computing r with technology.
Graphing calculator (TI-83/84): enter x in L1 and y in L2 (STAT, Edit). Turn on DiagnosticOn once (in the Catalog, or with Stat Diagnostics in the Mode menu on newer models). Then STAT, CALC, LinReg(ax+b) shows a, b, r² and r.
Desmos: type the data in a table with columns x₁ and y₁, then type y₁ ~ mx₁ + b on a new line. Desmos shows m, b, r² and r.
Spreadsheet: with x in A2:A9 and y in B2:B9, type =CORREL(A2:A9, B2:B9). (RSQ gives r², not r.)
Always graph the data too: r only makes sense as a summary when the scatter plot looks roughly linear and has no extreme outliers.
Invented data: arm span and height of 8 students (cm)
Student
1
2
3
4
5
6
7
8
Height
152
158
160
165
168
171
175
181
Arm span
150
160
157
166
165
173
172
183
Computing r: strong positive
Enter the height and arm-span table above into a calculator or Desmos and run a linear regression.
Equation: Technology gives r ≈ 0.97. Interpretation: a strong positive linear association; taller students in this sample tend to have longer arm spans, and the points lie close to a line.
Computing r: negative
Invented data for 10 students, (absences, final grade): (0, 94), (1, 88), (2, 91), (3, 80), (4, 89), (6, 78), (7, 82), (9, 72), (10, 79), (12, 70).
Equation: Technology gives r ≈ -0.88. Interpretation: a strong negative linear association; students with more absences tend to have lower final grades, with more scatter than the arm-span data.
Equation: Technology gives r ≈ -0.10, a very weak linear association. But the scatter plot rises and then falls: fuel economy depends strongly on speed in a curved way. r near 0 means "no linear pattern", not "no relationship".
Effect of an outlier
Invented data for 8 students, (hours studied, test score): (1, 64), (2, 70), (2, 72), (3, 75), (4, 79), (4, 83), (5, 86), (6, 92). A ninth student studied 6 hours but scored 58 after being ill on test day.
Equation: For the 8 students, r ≈ 0.99. Adding the ninth student gives r ≈ 0.39 (Diagram 2). One point changed the summary from very strong to weak, so report r with and without the outlier and explain it.
Close by listing what r is not: it is not the slope (a data set with slope 0.4 and another with slope 35 can have the same r), it is not a percentage of points on the line, and it has no units, so changing the units of either variable or swapping x and y leaves r unchanged.
Guided Practice15-20 minutes
Pairs enter the invented data below into technology. One partner uses a calculator and the other Desmos or a spreadsheet, and they confirm they get the same r.
Invented data: horsepower and 0-60 mph time for 8 cars
Car
1
2
3
4
5
6
7
8
Horsepower
130
150
180
200
240
280
300
350
0-60 time (s)
9.6
9.4
8.3
8.0
6.8
6.9
5.7
5.3
Tasks: (1) Compute r. (Answer: r ≈ -0.98.) (2) Interpret it in context. (A very strong negative linear association: cars with more horsepower tend to reach 60 mph in less time.) (3) Swap the columns so horsepower is y, and compute r again. (Still about -0.98: r does not depend on which variable is x.) (4) Would r change if time were measured in tenths of a second? (No: r has no units.) Circulate and check that students report r, not r², and that the sign matches the scatter plot.
Independent Practice15 minutes
Students work alone. The invented data list the height and rebounds per game of 10 players on a high school basketball team.
Invented data: height and rebounds for 10 players
Player
1
2
3
4
5
6
7
8
9
10
Height (in)
72
74
75
76
77
78
79
80
81
83
Rebounds per game
4.1
6.3
3.8
7.0
5.2
8.1
5.9
6.6
9.4
7.2
Tasks: (1) Make a scatter plot and describe it. (2) Use technology to compute r. (Answer: r ≈ 0.66.) (3) Interpret r in context. (A moderate positive linear association: taller players tend to get more rebounds, but there is a lot of scatter.) (4) A teammate says, "r = 0.66 means 66% of the points are on the line." Explain what is wrong. (r is not a percentage; it measures how tightly the points cluster around a line and in which direction.)
Closure5-10 minutes
Exit ticket: (1) For invented data on 25 households, the correlation between household size and weekly water use is r = 0.83. Interpret it. (2) Explain why a data set can have r = 0 even though y depends strongly on x. (3) Name one thing you should always do before reporting r. (Answers: a strong positive linear association, larger households tend to use more water; a curved pattern such as a U shape can have no linear trend; make a scatter plot and check for curves and outliers.)
Differentiation Strategies
For Struggling Students
Give a one-page technology guide with screenshots of the calculator menus and of the Desmos regression line, and have students circle r in the output
Have students first sort Diagram 1-style pictures into strong, moderate and weak, and positive or negative, before computing any values
Use a two-column frame: "direction: positive/negative because ..." and "strength: strong/moderate/weak because |r| = ..."
For Advanced Students
Ask students to build, in Desmos, a data set of 6 points with r above 0.9 and then add one point that makes r negative
Ask students to compute r by hand for a data set of 4 points using the z-score formula r = Σ(zxzy)/(n - 1) and compare it with the technology result
Ask students to explain, using the z-score formula, why multiplying every x value by the same positive number cannot change r
Assessment Guidance
What to Look For
Check that students report r, not r², and that the sign of r agrees with the direction of the scatter plot. Good interpretations say "linear", name the direction and strength in context, and mention the variables. Look for students who make a scatter plot before trusting r, who say that r near 0 means no linear association rather than no relationship, and who can explain the effect of an outlier. Watch for students who read r as a slope or as a percentage.
02
Classroom Activities
3 Activities
1
Class Data Correlations
20 minWhole class, then pairs
Students measure themselves, pool the class data in a shared spreadsheet, predict the value of r for three pairs of variables, then compute it with technology and compare.
Procedure
In pairs, measure height, arm span and hand span (thumb to little finger) in centimeters with a tape measure, and record the day of the month of each student's birthday
Enter the data in a shared spreadsheet, one row per student
Before computing, each pair writes a predicted r for: height and arm span; height and hand span; height and birthday
Compute each r with =CORREL or Desmos and record it next to the prediction
Discussion Questions
Which pair had the strongest linear association? Did it match the scatter plot?
Why is r for height and birthday not exactly 0, even though the variables are unrelated?
What happens to r for height and arm span if we measure in inches instead?
Modification for Distance Learning
Students measure at home and enter their values in a shared online form that feeds a spreadsheet. The teacher shares the spreadsheet so pairs can compute r on their own copies.
2
Build a Target r
20 minPairs
Pairs use Desmos to build small data sets that hit target values of r. Moving points and watching r change builds a feel for what the number responds to.
Targets
Six points with r between 0.90 and 0.95
Six points with r between -0.55 and -0.45
Seven points that follow an obvious curve but have r between -0.1 and 0.1
Six points with r above 0.9, then one added point that makes r negative
Two data sets with r close to 0.9, one with slope near 1 and one with slope near 20
Procedure
Type the points in a Desmos table, add y₁ ~ mx₁ + b, and drag points while watching r
Record each finished data set and its r in a table
Trade with another pair and check each other's values of r with a calculator
Challenge Variation
Build a data set with r = 0 exactly. Pairs should explain why a pattern that is symmetric about a vertical line, such as a U shape centered on the mean of x, gives r = 0.
3
Claim Check
15 minGroups of 3-4
Groups get five claim cards about correlation coefficients. For each card they decide true or false, write a reason, and give a small example if the claim is false.
Claim Cards
"r = -0.8 shows a stronger linear association than r = 0.6." (True: strength depends on |r|.)
"If r = 0, the variables have nothing to do with each other." (False: a curved pattern can have r = 0.)
"If a line has a steep slope, r must be close to 1." (False: r measures closeness to the line, not steepness.)
"Changing temperatures from degrees Fahrenheit to degrees Celsius changes r." (False: r has no units, and a unit change like this does not change how closely the points follow a line.)
"Removing one point can change r from strong to weak." (True: r is sensitive to outliers.)
Procedure
Each group member reads one card aloud and proposes an answer; the group must agree before writing
For every false claim, the group writes or sketches a small data set that shows why
Groups share one counterexample with the class
Modification for Distance Learning
Post the five claims as a poll. After voting, groups meet in breakout rooms to build a Desmos counterexample for each false claim and paste a screenshot in a shared slide.
03
Diagrams & Visual Aids
2 diagrams
Diagram 1: What Different Values of r Look Like
Four invented data sets drawn to scale. As |r| gets closer to 1, the points hug the line more tightly; the sign of r gives the direction. The fourth set follows a clear curve that is symmetric about x = 5, so r = 0: there is a strong relationship, but no linear one.
Diagram 2: How One Outlier Changes r
Invented scores of 8 students with the least-squares line (solid), r ≈ 0.99. Adding a ninth student who studied 6 hours and scored 58 gives the dashed line and r ≈ 0.39. Report r with and without an outlier, and explain the outlier in context.
04
Homework Assignment
~30 min
HSS.ID.C.8 Homework: The Correlation Coefficient
Directions: Use a graphing calculator, Desmos or a spreadsheet to compute every correlation coefficient, and round r to two decimal places. Make a scatter plot for each data set. Every interpretation must give the direction and strength of the linear association in context. All data are invented.
Part 1: Computing r with Technology (Problems 1-3)
A student records how many days she has practiced a piano piece and the number of mistakes in her run-through that day: (1, 14), (2, 12), (3, 12), (4, 9), (5, 8), (6, 6), (7, 5). Compute r and interpret it in context.
A bike shop lists 7 used bikes as (age in years, price in dollars): (1, 410), (2, 360), (3, 330), (4, 300), (5, 250), (6, 240), (7, 190). (a) Compute r. (b) The shop adds a 30-year-old vintage bike priced at $650. Compute r again. (c) Explain the change and which value better describes typical used bikes.
Eight students list (height in inches, weight in pounds): (61, 112), (63, 120), (64, 131), (66, 128), (67, 142), (69, 150), (70, 146), (72, 165). (a) Compute r. (b) Convert the heights to centimeters (multiply by 2.54) and compute r again. (c) Explain the result.
Part 2: Interpreting r (Problems 4-6)
Match each value of r with one description and explain each choice: r = -0.95, r = -0.40, r = 0.10, r = 0.70. (a) Hours of weekly exercise and resting heart rate: points fall steadily with very little scatter. (b) Shoe size and score on a spelling test: a shapeless cloud. (c) Hours of sleep and minutes to fall asleep: a slight downward drift with a lot of scatter. (d) Price of a meal and size of the tip: an upward trend with moderate scatter.
A running club times 100-meter sprints for members of different ages: (age 10, 15.6 s), (14, 13.9), (18, 13.0), (22, 12.6), (26, 12.5), (30, 12.8), (34, 13.4), (38, 14.3), (42, 15.5). (a) Compute r. (b) A club member concludes, "Age has nothing to do with sprint time." Use a scatter plot to explain why this is wrong.
Explain the error in each statement: (a) "r = -0.85 is a weak correlation because it is negative." (b) "The line has slope 3.2, so r = 3.2." (c) "A data set has r = 0.95, so a straight line must be the best model for it."
Rubric
Criterion
Full Credit (2 pts)
Partial Credit (1 pt)
No Credit (0 pts)
Computing r
Every r correct to two decimals, with the correct sign
One value wrong or r² reported instead of r
Values missing or mostly incorrect
Interpreting r
Direction, strength and the word "linear" stated in context
Direction or strength missing
Treats r as a slope or a percentage
Outliers and Curves
Explains the effect of the outlier and the curved pattern using a scatter plot
Notices the change without explaining it
No explanation
Reasoning
Every error in Problem 6 named and corrected
Most errors corrected
Errors not identified
05
Quiz: 20 Questions
Interactive, with answers
Instructions
Keep a calculator, Desmos or a spreadsheet open for the questions that ask you to compute r. All data are invented. Your score updates as you answer, and Reset quiz clears everything so you or your students can try again.
Multiple choice: pick an option to check it. Short answer: write your answer, then reveal the model answer.
0 of 20 answered · 0 correct
Question 1 of 20 · Multiple Choice
Which correlation coefficient shows the strongest linear association?
Answer: B
Strength depends on the distance from 0, |r|, and |-0.91| = 0.91 is the largest. The sign only gives the direction. Choice A is the largest positive value, a common trap when students ignore negative values of r.
Question 2 of 20 · Multiple Choice
For invented data on 12 months, the correlation between average outdoor temperature and a family's heating bill is r = -0.78. Which interpretation is best?
Answer: C
The negative sign means that as temperature goes up, bills tend to go down, and |r| = 0.78 is close to the strong range. Choice B confuses the sign with the strength. Choice D reads r as a slope; the slope would have units of dollars per degree, and r has no units. Choice A reads r as a percentage.
Question 3 of 20 · Multiple Choice
Use technology to compute r for the invented data (3, 40), (5, 34), (6, 37), (8, 28), (10, 27), (11, 19).
Answer: D
A linear regression gives r ≈ -0.94: the y values tend to fall as x grows, with a little scatter. Choice B is the slope of the least-squares line, not r, and cannot be a correlation because |r| is never more than 1. Choice A has the wrong sign.
Question 4 of 20 · Multiple Choice
The height of a thrown ball, measured every tenth of a second, forms an arch on a scatter plot, and r = 0.03. Which conclusion is correct?
Answer: A
r measures only linear association. An arch rises and then falls, so no single line fits it, and r is close to 0 even though time predicts height very well with a curved model. Choice B is the error of reading r near 0 as "no relationship".
Question 5 of 20 · Multiple Choice
For invented data on 20 taxi rides, r = 0.84 between distance in miles and fare in dollars. What is r if the distances are converted to kilometers (multiply by 1.609)?
Answer: B
r has no units. Multiplying every distance by 1.609 stretches the scatter plot but does not change how closely the points follow a line, so r stays 0.84. Choice A is impossible because r cannot be greater than 1.
Question 6 of 20 · Multiple Choice
Invented data are in cells A2:A13 (x) and B2:B13 (y) of a spreadsheet. Which formula returns the correlation coefficient r?
Answer: C
CORREL returns r. RSQ returns r², which is never negative and so loses the direction of the association. SLOPE returns the slope of the least-squares line, a different number with units.
Question 7 of 20 · Short Answer
Use technology to compute r for the invented data (hours of sunshine, solar panel output in kWh): (2, 3.1), (4, 6.8), (5, 6.0), (7, 10.9), (8, 11.5), (10, 14.2), (11, 17.0). Interpret r in context.
Technology gives r ≈ 0.99. There is a very strong positive linear association: days with more hours of sunshine tend to have higher solar panel output, and the points lie very close to a line.
Question 8 of 20 · Multiple Choice
For invented data, r = 0.85 between hours of daylight and daily visitors to a park. What does r = 0.85 tell you?
Answer: D
r describes direction and strength: positive, and 0.85 is strong. It is not a percentage of points (Choice A) and not a rate of change (Choice B); the slope would be in visitors per hour.
Question 9 of 20 · Multiple Choice
A student computes r = -0.62 for x = weekly screen time and y = hours of sleep. She then swaps the variables so that sleep is x. What is the new r?
Answer: A
The formula for r treats x and y the same way, so swapping them does not change r. The slope of the least-squares line does change when the variables are swapped, which is one more way r and the slope differ. Choice C cannot be a correlation because it is below -1.
Question 10 of 20 · Multiple Choice
In an invented data set, r = 0.92. After one point far from the pattern is added, r = 0.41. What should a student conclude?
Answer: B
A single point far from the others can pull the least-squares line and change r a great deal, as in Diagram 2. The right response is to check the point and report both values. Choice C is wrong because r is still positive, and Choice D is false: adding points that follow the pattern can keep r high.
Question 11 of 20 · Short Answer
A student says, "r = -0.9 shows a weaker association than r = 0.5, because -0.9 is less than 0.5." Explain the error.
Strength is measured by the distance from 0, |r|. Since |-0.9| = 0.9 > 0.5, r = -0.9 shows the stronger linear association. The negative sign only tells the direction: as x increases, y tends to decrease.
Question 12 of 20 · Multiple Choice
Use technology to compute r for the invented data (1, 2), (2, 4), (3, 5), (4, 4), (5, 5).
Answer: D
A linear regression gives y ≈ 2.2 + 0.6x with r ≈ 0.77, a moderate to strong positive linear association. Choices B and C are the slope and the intercept, not r. Choice A has the wrong sign.
Question 13 of 20 · Multiple Choice
Which scatter plot description best matches r = 0.3?
Answer: C
0.3 is positive and weak, so the points should drift up with lots of scatter. Choice A describes r close to -1, and Choice B describes r close to 1. A tight U shape (Choice D) usually gives r near 0.
Question 14 of 20 · Multiple Choice
Invented data set P has fitted slope 2 and r = 0.9. Invented data set Q has fitted slope 50 and r = 0.9. Which statement is true?
Answer: D
r measures how closely the points follow a line, not how steep the line is. The slope depends on the units and spread of x and y. Choices A and B both treat the slope as a measure of strength.
Question 15 of 20 · Short Answer
For invented data on 30 used phones, the correlation between age in months and resale price is r = -0.87. Interpret r in context.
r = -0.87 shows a strong negative linear association: older phones in this sample tend to have lower resale prices, and the points lie fairly close to a line that falls from left to right.
Question 16 of 20 · Multiple Choice
Which of these cannot be a correlation coefficient?
Answer: A
Every correlation coefficient is between -1 and 1, so 1.2 is impossible and signals a computing error, such as reading the slope. r = -1 is possible when all points lie on a falling line, and r = 0 is possible when there is no linear trend.
Question 17 of 20 · Multiple Choice
A calculator shows this LinReg output for invented data: a = -1.5, b = 4.2, r² = 0.64, r = -0.8. What is the correlation coefficient?
Answer: B
The calculator lists r on its own line: r = -0.8. The value 0.64 is r², the square of r, which drops the sign. The values a and b describe the fitted line y = ax + b, so a = -1.5 is the slope, which has the same sign as r.
Question 18 of 20 · Short Answer
For invented data on 8 students, (shoe size, score on a memory game): (6, 14), (7, 9), (7.5, 17), (8, 11), (9, 15), (9.5, 10), (10, 16), (11, 12). Compute r with technology and say whether a linear model is useful here.
Technology gives r ≈ 0.01. There is essentially no linear association, and the scatter plot is a shapeless cloud, so a line is not useful for predicting memory scores from shoe size.
Question 19 of 20 · Multiple Choice
What does r = 1 tell you about a data set?
Answer: C
r = 1 is the strongest possible positive linear association: all points are on one rising line. The slope can be any positive number, so Choice B is wrong, and the line does not have to be y = x (Choice D).
Question 20 of 20 · Short Answer
For 200 invented students, the correlation between high school GPA and first-year college GPA is 0.58, and the correlation between weekly hours of TV and first-year college GPA is -0.58. Compare the two associations.
Both have the same strength, a moderate linear association, because |0.58| = |-0.58|. They have opposite directions: higher high school GPA tends to go with higher college GPA, while more TV tends to go with lower college GPA.
0 of 20 answered · 0 correct
06
Frequently Asked Questions
10 Questions
What does HSS.ID.C.8 mean?
It means students use technology to find the correlation coefficient r of a linear fit and explain what it says about the data. The sign of r gives the direction of the linear association, and the distance from 0 gives its strength.
Do students need to compute r by hand for HSS.ID.C.8?
No. The standard says "using technology", so a graphing calculator, Desmos or a spreadsheet is expected. The formula is a useful enrichment for strong students, but the focus is on interpreting r correctly.
How do you find r on a TI-84?
Enter the data in L1 and L2, turn on DiagnosticOn (from the Catalog, or Stat Diagnostics in the Mode menu on newer models), then choose STAT, CALC, LinReg(ax+b). The output lists a, b, r² and r. If r does not appear, diagnostics are still off.
What is a strong correlation?
There is no official cutoff. Many textbooks use a rough guide: |r| of 0.8 or more is strong, 0.5 to 0.8 is moderate, and under 0.5 is weak. What counts as strong also depends on the field and on the purpose of the model.
Does r = 0 mean there is no relationship?
No. It means there is no linear association. A U-shaped pattern can have r near 0 even though y depends strongly on x. This is why students should always look at the scatter plot before interpreting r.
What is the difference between r and r²?
r is the correlation coefficient, between -1 and 1, and its sign shows the direction. r² is its square, between 0 and 1, and it is usually interpreted as the fraction of the variation in y that the linear model accounts for. Calculators and Desmos show both, so students should read the right line.
Is the correlation coefficient the same as the slope?
No. The slope tells how much the predicted y changes per unit of x and has units. r has no units and tells how closely the points follow the line. The slope and r always have the same sign, but a steep line can have a weak r and a shallow line a strong r.
Why can one outlier change r so much?
r is based on how far each point is from the means of x and y, so a point far from the rest has a large effect, especially in a small data set. Report r with and without the outlier and explain the outlier in context rather than deleting it silently.
Does a strong correlation mean one variable causes the other?
No. A strong r shows only that the variables tend to move together in a linear way. A lurking variable, reverse cause or chance can produce a strong correlation. Telling correlation apart from causation is the focus of the next standard, HSS.ID.C.9.
Is HSS.ID.C.8 Algebra 1 or Statistics?
Both. It is usually taught in Algebra I with scatter plots and lines of best fit, and again in introductory Statistics, where r is used with regression output and residual plots.
07
Related Standards
6 standards
These standards connect to HSS.ID.C.8: prerequisites to review first, parallel standards at the same level, and next steps that build on it.
Before this lesson
8.SP.A.1Prerequisite
Construct and interpret scatter plots and describe patterns of association