This chapter uses the example of whether an auto mechanic’s salary relates to his years of experience to introduce linear regression and correlation. It covers the basic ideas of linear regression and correlation, creating and interpreting a line of best fit, calculating and interpreting the correlation coefficient, and identifying outliers.
Do two variables move together?
The chapter opens with the question of whether two numeric variables are related, using the pairing of exam grades on two tests and of an auto mechanic’s pay against years of experience as running examples, and discusses the basic ideas of linear regression and correlation as tools for answering it.
The line of best fit and the correlation coefficient
Students learn to create and interpret a line of best fit through a scatter of paired data points, and to calculate and interpret the correlation coefficient, a single number summarizing how strong and in what direction the linear relationship between the two variables runs.
Identifying outliers
The chapter also covers calculating and interpreting outliers, points that fall unusually far from the line of best fit, teaching students to recognize when a data point may be distorting the regression results and should be examined rather than automatically included in the analysis.
This chapter grounds linear regression and correlation in economic model-building, using the theory of consumer choice and the demand curve as an example of how assumptions lead to a testable relationship between variables. It covers the correlation coefficient, testing its significance, linear equations, and the regression equation itself.
Models, theories, and testable relationships
The chapter opens by connecting regression to the broader idea of a model, a theorized cause-and-effect relationship, illustrated with the economic model of consumer choice, where assumptions about preferences and utility maximization generate the prediction embodied in the demand curve, an example of how a theory produces a relationship that can then be tested statistically.
The correlation coefficient and its significance
The correlation coefficient r measures the strength and direction of a linear relationship between two numeric variables, and the chapter covers how to calculate and interpret it, followed by a hypothesis test for whether the observed correlation is statistically significant or could plausibly have arisen from an uncorrelated population.
Linear equations and the regression equation
Building on the algebra of a linear equation, the chapter develops the regression equation as the statistical technique for finding the line of best fit through a set of paired data, letting an analyst predict one variable, such as pay for a repair job, from another, such as an initial fee plus an hourly rate.
Linear regression and correlation describe how two numeric variables are related. This chapter reviews linear equations and scatter plots, shows how to fit a line of best fit through data using the regression equation, and explains how the correlation coefficient measures the strength and direction of a linear relationship.
Scatter Plots and Linear Equations
A relationship between two numeric variables is first explored with a scatter plot, which plots paired observations as points. If the points cluster around a straight line, the variables have an approximately linear relationship. A linear equation describes such a line through its slope and intercept: the slope gives the change in one variable for each unit change in the other, and the intercept gives the value where the line meets the vertical axis. The scatter plot shows whether a line is a reasonable summary.
The Regression Equation
The regression equation gives the line of best fit, the straight line that comes closest to all the points by making the total squared vertical distance from the points to the line as small as possible. This least-squares line can be used to predict the value of the dependent variable from a value of the independent variable. Predictions are most reliable within the range of the observed data, and extending the line far beyond that range can give misleading results.
The Correlation Coefficient
The correlation coefficient measures the strength and direction of the linear relationship between two variables. It ranges from negative one to positive one: values near positive one indicate a strong upward relationship, values near negative one a strong downward one, and values near zero little or no linear relationship. A value close to the extremes means the points lie near the line, while outliers, points far from the pattern, can distort both the line and the coefficient.