DoAssignment.ca
SL 4.4 · Analyse correlation and linear regression with appropriate caution
Learn to analyse correlation and linear regression with appropriate caution through clear examples and targeted practice.
International Baccalaureate (IB) IB AA SL: Mathematics: Analysis and Approaches SL
Statistics and Probability
IB Mathematics: Analysis and Approaches SL — Study topic SL 4.4
Two quantities can vary together without one necessarily causing changes in the other. Correlation describes how closely two numerical variables follow a straight-line pattern; linear regression gives a line that summarises that pattern and can be used for prediction. Both are useful, but both require careful interpretation. This lesson connects scatter plots, numerical summaries, regression equations, and context, with particular attention to the limits of conclusions drawn from data.
What you will learn
- Describe the direction, form, and strength of a relationship using a scatter plot and the correlation coefficient.
- Interpret a least-squares regression line and use it to make a prediction within the observed data range.
- Explain why correlation and a regression model do not, by themselves, establish causation or guarantee reliable predictions.
1. Prior knowledge: paired data and scatter plots
Bivariate data consist of paired values for two variables. Write a pair as , where is the explanatory variable and is the response variable. For example, might be practice time and a test score. The choice of explanatory and response variables should make sense in context; it does not prove that causes changes in .
A scatter plot places each pair at its corresponding position on axes. First look for direction: as increases, do the -values generally increase or decrease? Then consider form: is the pattern roughly straight, curved, or unclear? Finally, consider how tightly the points cluster and whether any points lie far from the general pattern. A straight-line model is most appropriate when the overall pattern is approximately linear.
The mean of the -values is written and the mean of the -values is written . These averages describe the centre of the data. Familiarity with substitution and rearranging a linear equation is enough background for interpreting a regression line. When explaining results, include the variables’ units.
- Plot paired values, keeping each matched with its observed .
- Describe direction, form, strength, and unusual points before choosing a model.
- An observed association does not establish a cause-and-effect relationship.
2. Correlation: direction and strength
The Pearson correlation coefficient, usually written , is a number between negative one and one. A positive value indicates that larger -values tend to occur with larger -values; a negative value indicates that larger -values tend to occur with smaller -values. Values nearer to positive or negative one indicate a stronger linear association. A value near zero indicates a weak linear association, but a curved relationship may still be present.
The coefficient has no units, so it describes direction and strength of linear association without depending on the measurement units. It does not describe the slope of a regression line or how much changes for each unit of ; the regression equation provides that information. Avoid assigning a rigid label such as “strong” to a particular value without considering the scatter plot and context.
A point far from the main cluster can substantially affect the calculated correlation. Inspect the plotted data before relying on a calculator’s value. Ask whether the pattern is reasonably linear, whether the data cover a suitable range, and whether an unusual point needs explanation. A numerical summary cannot replace inspection of the data.
- The sign of gives the direction of linear association.
- The magnitude of indicates how closely points follow a straight-line pattern.
- Correlation does not establish causation; other variables or chance may help explain an association.
3. Linear regression: a model for prediction
A linear regression line is written , where is the predicted response, is the intercept, and is the slope. The slope gives the change in predicted for an increase of one unit in . Its units are units of per unit of . The intercept is the predicted when ; it may have little practical meaning if zero is outside the observed range or does not make sense in context.
The least-squares regression line is the line that makes the sum of the squared vertical differences between observed and predicted -values as small as possible. For a regression of on , the line passes through . Graphing technology can calculate the slope and intercept from paired data, but you must still identify the variables correctly, inspect the scatter plot, and interpret the output.
A residual is the observed response minus the predicted response. A positive residual means the observed point lies above the line; a negative residual means it lies below. Residuals describe how far individual observations are from the model. Keep the predicted value distinct from the observed value .
Use a regression line for interpolation when the chosen -value lies within the observed data range and the linear pattern is suitable. A prediction beyond that range is extrapolation; it is less reliable because the relationship may not continue in the same way. Even an interpolation is an estimate, not a guaranteed result. State predictions in context and include appropriate units.
- The slope describes change in predicted response per unit increase in explanatory variable.
- Residual equals observed value minus predicted value.
- Check the data and justify the range before using the line to predict.
4. Technology and cautious conclusions
A graphing calculator or statistics tool can plot paired data and find the correlation coefficient and regression equation. Enter paired values in matching lists, display a scatter plot, and check that the plotted points agree with the original data. Then request the regression of on and the correlation coefficient. Calculator menus differ, so use the device’s instructions rather than assuming a particular sequence.
Technology supports, but does not replace, reasoning. Verify that the regression uses the intended response variable, retain enough digits in intermediate values, and round the final result sensibly. Interpret the slope and intercept using units and context. If the scatter plot is curved or an unusual point strongly affects the result, the linear summary may be misleading.
An association may occur because both variables are related to another factor, because of the way data were selected, or by coincidence. Avoid claiming that changing causes a change in unless the study design provides evidence for causation. Describe what the data support: for example, “these observations show a positive linear association,” rather than asserting a cause.
- Use a scatter plot to check whether linear regression is suitable before interpreting its output.
- A calculator reports a model; it does not decide whether that model is sensible in context.
- Describe association, not causation, unless there is separate evidence supporting a causal claim.
Worked example
1. Calculate and interpret a regression line
Four students’ practice hours, , and quiz scores, , are recorded as , , , and . Find the least-squares regression line and correlation coefficient, then predict the score at .
- Find the meansThe mean practice time is hours and the mean score is . These means locate the centre of the data, and the regression line passes through their coordinate pair.
- Calculate the slopeThe sum of the products of paired deviations from the means is , and the sum of squared deviations in is . Their ratio gives the slope, or predicted score-point change per practice hour.
- Find the intercept and lineThe line passes through . Substituting the means and slope into the line gives an intercept of .
- Find the correlation and predictionThe sum of squared deviations in is also , so the correlation is . This indicates a positive linear association. At , the model predicts points. The observed score at that practice time was , so the prediction is an estimate, not the actual score.
Answer: The regression line is , and . The predicted score at hours is points.
Check: The line predicts at the mean practice time of hours, matching the mean score. The positive value of agrees with the overall upward pattern.
Worked example
2. Use a model within the observed range
A delivery company models travel time , in minutes, against route length , in kilometres. Its fitted line is . Observed route lengths were from km to km. Estimate the time for a km route and interpret the slope.
- Check the rangeSeven kilometres is between the shortest and longest observed route lengths, so the estimate is an interpolation, provided the scatter plot supports a linear pattern.
- Substitute into the modelReplace by in the fitted equation. The predicted travel time is minutes, or about minutes for a practical estimate.
- Interpret the slopeThe slope means that the model predicts an increase of minutes in travel time for each additional kilometre. This is a modelled average change, not a guarantee for every route. b=2.4 minutes per kilometre
Answer: The estimated time is minutes, or about minutes. The model predicts an increase of minutes per additional kilometre.
Check: The prediction is within the observed route-length range. Its reliability still depends on the scatter plot and on whether the routes are reasonably comparable.
Worked example
3. Identify an unsupported causal claim
A school finds that students who bring more books to class tend to receive higher grades. A student concludes that carrying extra books will cause grades to rise. Explain why the conclusion is not justified by correlation alone.
- State what the data showIf the scatter plot and correlation support it, the observations show an association between number of books and grades for the students studied. This does not establish which variable, if either, causes a change in the other.
- Consider another explanationA third factor, such as time spent studying, could be related to both the number of books brought and grades. The association might also reflect the way students were selected or chance.
- Write a cautious conclusionReport the observed association without claiming that carrying extra books raises grades. More evidence designed to investigate cause would be needed for that conclusion.
Answer: The data may show an association, but they do not establish that carrying extra books causes higher grades.
Check: The conclusion distinguishes an observed relationship from a causal explanation.
Common mistakes and how to avoid them
Saying that a high positive correlation proves that one variable causes the other.
Correction: Describe the association shown by the data. Correlation alone does not establish cause and effect.
Assuming that means there is no relationship of any kind.
Correction: It indicates no linear association; a curved pattern may still exist. Inspect the scatter plot.
Using the regression line far beyond the observed -values as if its prediction were reliable.
Correction: Identify this as extrapolation and explain why the pattern may not continue outside the data range.
Interpreting the slope as the predicted value of at a particular .
Correction: The slope is the change in predicted for a one-unit increase in . Substitute the chosen into the full line to obtain a prediction.
Lesson summary
- Use a scatter plot to describe direction, form, strength, and unusual points.
- The correlation coefficient describes the direction and strength of a linear association and lies between and .
- The regression line predicts response values; interpret its slope, intercept, units, and limits in context.
- Prefer interpolation to extrapolation, and do not infer causation from correlation alone.
- Use technology to calculate and check, while applying mathematical and contextual reasoning to judge the result.
Check your understanding
Question 1
A scatter plot shows a roughly straight downward pattern. Which sign would you expect for ?
- Positive
- Negative
- Exactly zero
- Greater than one
Show answer and explanation
Negative
As increases, tends to decrease, so the linear association is negative.
Question 2
A regression line is . What does the slope mean?
- The predicted response is always .
- The predicted response increases by units for each one-unit increase in .
- The response causes to increase by .
- The predicted response at is .
Show answer and explanation
The predicted response increases by units for each one-unit increase in .
The coefficient of is the slope. It gives the change in predicted for each one-unit increase in .
Question 3
A linear model is fitted to observations with from to . A prediction is requested at . What is the main concern?
- This is interpolation, so no caution is needed.
- This is extrapolation, so the pattern may not continue.
- The correlation coefficient must be zero.
- The intercept must equal .
Show answer and explanation
This is extrapolation, so the pattern may not continue.
The requested value is outside the observed range. A linear pattern seen within the data may not continue beyond it.
Key terms
- Bivariate data
- Paired numerical observations for two variables.
- Correlation coefficient
- The number between and that describes the direction and strength of a linear association.
- Regression line
- A fitted straight line used to summarise a linear pattern and predict response values.
- Residual
- The observed response minus the response predicted by the regression line.
- Interpolation
- Using a model to estimate a value within the range of observed explanatory-variable values.
- Extrapolation
- Using a model to estimate a value outside the range of observed explanatory-variable values.
Continue through IB AA SL
- SL 4.1 · Distinguish populations, samples, variables, sampling methods, and bias
- SL 4.2 · Organize and display discrete and continuous data
- SL 4.3 · Calculate and interpret measures of centre, position, and dispersion
- SL 4.5 · Use sample spaces, events, complements, and expected frequencies
- SL 4.6 · Solve combined and conditional probability problems
- SL 4.7 · Use discrete random-variable distributions and expected value
About this lesson and its review
Published by DoAssignment. This AI-assisted lesson follows International Baccalaureate (IB) IB AA SL: Mathematics: Analysis and Approaches SL, study topic SL 4.4. It is a study resource, not an official curriculum publication.
Before publication, the draft is checked for structure, mathematical or chemical notation, calculations, course boundaries, and readability, and then requires administrator approval. Errors can still occur, so corrections are welcomed.