DoAssignment.ca

SL 4.4 · Analyse correlation and linear regression with appropriate caution

Learn to analyse correlation and linear regression with appropriate caution through clear examples and targeted practice.

International Baccalaureate (IB) IB AA SL: Mathematics: Analysis and Approaches SL

Statistics and Probability

IB Mathematics: Analysis and Approaches SL — Study topic SL 4.4

Two quantities can vary together without one necessarily causing changes in the other. Correlation describes how closely two numerical variables follow a straight-line pattern; linear regression gives a line that summarises that pattern and can be used for prediction. Both are useful, but both require careful interpretation. This lesson connects scatter plots, numerical summaries, regression equations, and context, with particular attention to the limits of conclusions drawn from data.

What you will learn

1. Prior knowledge: paired data and scatter plots

Bivariate data consist of paired values for two variables. Write a pair as (x,y)(x,y), where xx is the explanatory variable and yy is the response variable. For example, xx might be practice time and yy a test score. The choice of explanatory and response variables should make sense in context; it does not prove that xx causes changes in yy.
A scatter plot places each pair at its corresponding position on axes. First look for direction: as xx increases, do the yy-values generally increase or decrease? Then consider form: is the pattern roughly straight, curved, or unclear? Finally, consider how tightly the points cluster and whether any points lie far from the general pattern. A straight-line model is most appropriate when the overall pattern is approximately linear.
The mean of the xx-values is written xˉ\bar{x} and the mean of the yy-values is written yˉ\bar{y}. These averages describe the centre of the data. Familiarity with substitution and rearranging a linear equation is enough background for interpreting a regression line. When explaining results, include the variables’ units.

2. Correlation: direction and strength

The Pearson correlation coefficient, usually written rr, is a number between negative one and one. A positive value indicates that larger xx-values tend to occur with larger yy-values; a negative value indicates that larger xx-values tend to occur with smaller yy-values. Values nearer to positive or negative one indicate a stronger linear association. A value near zero indicates a weak linear association, but a curved relationship may still be present.
The coefficient has no units, so it describes direction and strength of linear association without depending on the measurement units. It does not describe the slope of a regression line or how much yy changes for each unit of xx; the regression equation provides that information. Avoid assigning a rigid label such as “strong” to a particular value without considering the scatter plot and context.
A point far from the main cluster can substantially affect the calculated correlation. Inspect the plotted data before relying on a calculator’s value. Ask whether the pattern is reasonably linear, whether the data cover a suitable range, and whether an unusual point needs explanation. A numerical summary cannot replace inspection of the data.
−1≤r≤1-1\leq r\leq 1

3. Linear regression: a model for prediction

A linear regression line is written y^=a+bx\hat{y}=a+bx, where y^\hat{y} is the predicted response, aa is the intercept, and bb is the slope. The slope gives the change in predicted yy for an increase of one unit in xx. Its units are units of yy per unit of xx. The intercept is the predicted yy when x=0x=0; it may have little practical meaning if zero is outside the observed range or does not make sense in context.
The least-squares regression line is the line that makes the sum of the squared vertical differences between observed and predicted yy-values as small as possible. For a regression of yy on xx, the line passes through (xˉ,yˉ)(\bar{x},\bar{y}). Graphing technology can calculate the slope and intercept from paired data, but you must still identify the variables correctly, inspect the scatter plot, and interpret the output.
A residual is the observed response minus the predicted response. A positive residual means the observed point lies above the line; a negative residual means it lies below. Residuals describe how far individual observations are from the model. Keep the predicted value y^\hat{y} distinct from the observed value yy.
Use a regression line for interpolation when the chosen xx-value lies within the observed data range and the linear pattern is suitable. A prediction beyond that range is extrapolation; it is less reliable because the relationship may not continue in the same way. Even an interpolation is an estimate, not a guaranteed result. State predictions in context and include appropriate units.
y^=a+bx\hat{y}=a+bx

4. Technology and cautious conclusions

A graphing calculator or statistics tool can plot paired data and find the correlation coefficient and regression equation. Enter paired values in matching lists, display a scatter plot, and check that the plotted points agree with the original data. Then request the regression of yy on xx and the correlation coefficient. Calculator menus differ, so use the device’s instructions rather than assuming a particular sequence.
Technology supports, but does not replace, reasoning. Verify that the regression uses the intended response variable, retain enough digits in intermediate values, and round the final result sensibly. Interpret the slope and intercept using units and context. If the scatter plot is curved or an unusual point strongly affects the result, the linear summary may be misleading.
An association may occur because both variables are related to another factor, because of the way data were selected, or by coincidence. Avoid claiming that changing xx causes a change in yy unless the study design provides evidence for causation. Describe what the data support: for example, “these observations show a positive linear association,” rather than asserting a cause.

Worked example

1. Calculate and interpret a regression line

Four students’ practice hours, xx, and quiz scores, yy, are recorded as (1,2)(1,2), (2,3)(2,3), (3,5)(3,5), and (4,4)(4,4). Find the least-squares regression line and correlation coefficient, then predict the score at x=3x=3.
  1. Find the means
    The mean practice time is 2.52.5 hours and the mean score is 3.53.5. These means locate the centre of the data, and the regression line passes through their coordinate pair.
    xˉ=2.5,yˉ=3.5\bar{x}=2.5,\quad\bar{y}=3.5
  2. Calculate the slope
    The sum of the products of paired deviations from the means is 44, and the sum of squared deviations in xx is 55. Their ratio gives the slope, or predicted score-point change per practice hour.
    b=45=0.8b=\frac{4}{5}=0.8
  3. Find the intercept and line
    The line passes through (xˉ,yˉ)(\bar{x},\bar{y}). Substituting the means and slope into the line gives an intercept of 1.51.5.
    a=3.5−0.8(2.5)=1.5,y^=1.5+0.8xa=3.5-0.8(2.5)=1.5,\quad\hat{y}=1.5+0.8x
  4. Find the correlation and prediction
    The sum of squared deviations in yy is also 55, so the correlation is 4/5⋅5=0.84/\sqrt{5\cdot5}=0.8. This indicates a positive linear association. At x=3x=3, the model predicts 3.93.9 points. The observed score at that practice time was 55, so the prediction is an estimate, not the actual score.
    r=0.8,y^=1.5+0.8(3)=3.9r=0.8,\quad\hat{y}=1.5+0.8(3)=3.9
Answer: The regression line is y^=1.5+0.8x\hat{y}=1.5+0.8x, and r=0.8r=0.8. The predicted score at 33 hours is 3.93.9 points.
Check: The line predicts 3.53.5 at the mean practice time of 2.52.5 hours, matching the mean score. The positive value of rr agrees with the overall upward pattern.

Worked example

2. Use a model within the observed range

A delivery company models travel time yy, in minutes, against route length xx, in kilometres. Its fitted line is y^=6+2.4x\hat{y}=6+2.4x. Observed route lengths were from 22 km to 1010 km. Estimate the time for a 77 km route and interpret the slope.
  1. Check the range
    Seven kilometres is between the shortest and longest observed route lengths, so the estimate is an interpolation, provided the scatter plot supports a linear pattern.
  2. Substitute into the model
    Replace xx by 77 in the fitted equation. The predicted travel time is 22.822.8 minutes, or about 2323 minutes for a practical estimate.
    y^=6+2.4(7)=22.8\hat{y}=6+2.4(7)=22.8
  3. Interpret the slope
    The slope means that the model predicts an increase of 2.42.4 minutes in travel time for each additional kilometre. This is a modelled average change, not a guarantee for every route. b=2.4 minutes per kilometre
Answer: The estimated time is 22.822.8 minutes, or about 2323 minutes. The model predicts an increase of 2.42.4 minutes per additional kilometre.
Check: The prediction is within the observed route-length range. Its reliability still depends on the scatter plot and on whether the routes are reasonably comparable.

Worked example

3. Identify an unsupported causal claim

A school finds that students who bring more books to class tend to receive higher grades. A student concludes that carrying extra books will cause grades to rise. Explain why the conclusion is not justified by correlation alone.
  1. State what the data show
    If the scatter plot and correlation support it, the observations show an association between number of books and grades for the students studied. This does not establish which variable, if either, causes a change in the other.
  2. Consider another explanation
    A third factor, such as time spent studying, could be related to both the number of books brought and grades. The association might also reflect the way students were selected or chance.
  3. Write a cautious conclusion
    Report the observed association without claiming that carrying extra books raises grades. More evidence designed to investigate cause would be needed for that conclusion.
Answer: The data may show an association, but they do not establish that carrying extra books causes higher grades.
Check: The conclusion distinguishes an observed relationship from a causal explanation.

Common mistakes and how to avoid them

Saying that a high positive correlation proves that one variable causes the other.
Correction: Describe the association shown by the data. Correlation alone does not establish cause and effect.
Assuming that r=0r=0 means there is no relationship of any kind.
Correction: It indicates no linear association; a curved pattern may still exist. Inspect the scatter plot.
Using the regression line far beyond the observed xx-values as if its prediction were reliable.
Correction: Identify this as extrapolation and explain why the pattern may not continue outside the data range.
Interpreting the slope as the predicted value of yy at a particular xx.
Correction: The slope is the change in predicted yy for a one-unit increase in xx. Substitute the chosen xx into the full line to obtain a prediction.

Lesson summary

Check your understanding

Question 1

A scatter plot shows a roughly straight downward pattern. Which sign would you expect for rr?
  1. Positive
  2. Negative
  3. Exactly zero
  4. Greater than one
Show answer and explanation
Negative
As xx increases, yy tends to decrease, so the linear association is negative.

Question 2

A regression line is y^=4+3x\hat{y}=4+3x. What does the slope mean?
  1. The predicted response is always 33.
  2. The predicted response increases by 33 units for each one-unit increase in xx.
  3. The response causes xx to increase by 33.
  4. The predicted response at x=0x=0 is 33.
Show answer and explanation
The predicted response increases by 33 units for each one-unit increase in xx.
The coefficient of xx is the slope. It gives the change in predicted yy for each one-unit increase in xx.

Question 3

A linear model is fitted to observations with xx from 55 to 1212. A prediction is requested at x=16x=16. What is the main concern?
  1. This is interpolation, so no caution is needed.
  2. This is extrapolation, so the pattern may not continue.
  3. The correlation coefficient must be zero.
  4. The intercept must equal 1616.
Show answer and explanation
This is extrapolation, so the pattern may not continue.
The requested value is outside the observed range. A linear pattern seen within the data may not continue beyond it.

Key terms

Bivariate data
Paired numerical observations for two variables.
Correlation coefficient
The number rr between −1-1 and 11 that describes the direction and strength of a linear association.
Regression line
A fitted straight line used to summarise a linear pattern and predict response values.
Residual
The observed response minus the response predicted by the regression line.
Interpolation
Using a model to estimate a value within the range of observed explanatory-variable values.
Extrapolation
Using a model to estimate a value outside the range of observed explanatory-variable values.

Continue through IB AA SL

View the complete IB AA SL International Baccalaureate (IB) IB AA SL: Mathematics: Analysis and Approaches SL curriculum and lessons

About this lesson and its review

Published by DoAssignment. This AI-assisted lesson follows International Baccalaureate (IB) IB AA SL: Mathematics: Analysis and Approaches SL, study topic SL 4.4. It is a study resource, not an official curriculum publication.

Before publication, the draft is checked for structure, mathematical or chemical notation, calculations, course boundaries, and readability, and then requires administrator approval. Errors can still occur, so corrections are welcomed.

Official curriculum reference

Report a correction or ask a question