Q BankQuestion BankDocsDocuments

5 Regression Analysis

Syllabus
2026
Section
5
Level

Exam analysis

No tagged past-paper evidence yet

Published Concept pages under this syllabus area do not have tagged past-paper appearances in the selected level yet.

Recent 5 years

In this section

Topic 5.1

5.1 Graphical Representations Between Two Quantitative Variables

Objectives in this topic

5.1.A—Construct scatterplots depicting the relationship between two quantitative variables

Construct scatterplots depicting the relationship between two quantitative variables.

  • A bivariate quantitative data set consists of observations of ordered pairs from two quantitative variables, collected from the same individuals in a sample or population, and can be used to construct a scatterplot.
  • A scatterplot shows the relationship between two quantitative variables for each observation, one corresponding to the value on the x-axis and one corresponding to the value on the y-axis. The explanatory variable is placed on the x-axis and is the variable whose values are used to explain or predict the corresponding values for the response variable, which is placed on the y-axis.

5.1.B—Describe the characteristics of a scatterplot

Describe the characteristics of a scatterplot.

  • A description of the association shown in a scatterplot includes form, direction, strength, and unusual features.
  • The form of the association shown in a scatterplot, if any, can be described as linear or non-linear.
  • The direction of the association shown in a scatterplot, if any, can be described as positive or negative. A positive association means that as values of the explanatory variable increase, the values of the response variable tend to increase. A negative association means that as values of the explanatory variable increase, the values of the response variable tend to decrease.
  • The strength of the association shown in a scatterplot is how closely the points follow the general pattern. Strength can be described as strong, moderate, or weak.
  • Unusual features of a scatterplot include clusters of individual points or points that don’t fit in the general pattern of association between the two variables.

5.1.C—Justify a claim using scatterplots depicting the relationship between two quantitative variables

Justify a claim using scatterplots depicting the relationship between two quantitative variables.

  • Scatterplots depicting the relationship between two numeric variables may reveal information that can be used to justify claims about the variable in context. 149 Regression Analysis UNIT 5

Topic 5.2

5.2 Correlation

Objectives in this topic

5.2.A—Interpret the correlation for a linear relationship

Interpret the correlation for a linear relationship.

  • The correlation coefficient, r, summarizes the strength and direction of the linear association between two quantitative variables. The correlation coefficient r is unit-free and always between −1 and 1, inclusive. A negative correlation coefficient v alue indicates a negative association, and a positive correlation coefficient value indicates a positive association.
  • The strength of the linear association is determined by how close the correlation coefficient is to −1 or 1. A value of r =0 indicates that there is no linear association. A value of r =- 1 or r =1 indicates that there is a perfect linear association.
  • A correlation coefficient close to −1 or 1 does not necessarily mean that a linear model is appropriate.
  • A perceived or real relationship between two variables does not mean that changes in one variable cause changes in the other. That is, correlation does not necessarily imply causation.

Topic 5.3

5.3 Linear Regression Models

Objectives in this topic

5.3.A—Calculate a predicted response value using a linear regression model

Calculate a predicted response value using a linear regression model.

  • If the form of the relationship between x and y appears linear, we can approximate the relationship between x and y using a linear regression model, which is a linear equation that uses an explanatory variable, x, to predict the response variable, y.
  • In a linear regression model, the predicted response value, denoted by y, is calculated as y=+ab x, where a is the y-intercept, b is the slope of the regression line, and x is the explanatory variable.
  • Extrapolation is predicting a response value using a value for the explanatory variable that is beyond the interval of x-values used to determine the regression line. The predicted value is less reliable the further the estimate is extrapolated.
  • Interpolation is predicting a response value using a value for the explanatory variable that is within the interval of x-values used to determine the regression line.

Topic 5.4

5.4 Residuals

Objectives in this topic

5.4.A—Calculate the differences between the observed and predicted values

Calculate the differences between the observed and predicted values.

  • A residual is the difference between the observed response value and the predicted response value for the given value of the explanatory variable: residual =-y y or (residual =-observed pyy redicted .)

5.4.B—Interpret the differences between the observed and predicted values

Interpret the differences between the observed and predicted values.

  • If the residual is positive, the model underpredicts (underestimates) the value of the response variable. If the residual is negative, the model overpredicts (overestimates) the value of the response variable.

5.4.C—Describe the form of association of bivariate data using residual plots

Describe the form of association of bivariate data using residual plots.

  • A residual plot is a scatterplot of the residuals versus the predicted response values (or the explanatory variable values).
  • Residual plots can be used to investigate the appropriateness of the linear regression model for the observed data.
  • The linear regression model should only be fit to the data if the data exhibit a linear trend. Apparent randomness in a residual plot for a linear regression model is confirmation of a linear form in the association between the two variables and indicates that the simple linear regression model is an appropriate model for the data.
  • Curvature in the residual plot for a linear regression model suggests that the linear model is not the most appropriate model for the data.

Topic 5.5

5.5 Least-Squares Regression

Objectives in this topic

5.5.A—Calculate the coefficients for the least-squares regression line model

Calculate the coefficients for the least-squares regression line model.

  • The simple linear regression model is fit to the data by minimizing the sum of the squares of the residuals. Because of this, the resulting equation is often called the least-squares regression line (LSRL) and is calculated using technology. This regression line will pass through the point ()xy, .
  • The slope of the regression line, b, is calculated using technology.
  • The y-intercept of the regression line, a, is calculated using technology.
  • In simple linear regression, the correlation coefficient, r, is calculated using technology.
  • In simple linear regression, the square of the correlation coefficient, r2 , is called the coefficient of determination. The value of r2 is the proportion of variation in the response variable that is explained by the linear relationship with the explanatory variable.

5.5.B—Interpret coefficients for the least-squares regression line model

Interpret coefficients for the least-squares regression line model.

  • The coefficients of the least-squares regression line model (line of best fit) are the slope, b, and the y-intercept, a, because they are based on a sample of values.
  • The slope of the least-squares regression line can be interpreted as the predicted increase or decrease in the response variable for a oneunit increase in the explanatory variable, and it should be interpreted in context.
  • The y-intercept in the least-squares regression line is the predicted value of the response variable when the explanatory variable is equal to 0, and it should be interpreted in context. Sometimes, the y-intercept of the line does not have a reasonable interpretation in context because x =0 might be beyond the interval of x-values used to determine the regression line (extrapolation). At other times, the y-intercept of the line does not have a logical interpretation in context because it might be a negative value for a response variable that has no negative values, such as height.
ConceptAP Statistics