This chapter teaches how to organize and present a data collection, using real classroom examples such as exam scores from a precalculus class and a basketball team’s game scores. It covers graphical displays including stem-and-leaf plots, histograms, and box plots, along with measures of location, center, and spread.
Organizing raw data into graphs
Once a collection of data has been gathered, the chapter teaches students to display it graphically and interpret several standard chart types, including stem-and-leaf plots, line graphs, bar graphs, frequency polygons, time series graphs, histograms, box plots, and dot plots, using real classroom examples such as a precalculus class’s exam scores.
Measures of location
The chapter covers measures of location, quartiles and percentiles, which describe where a particular value sits within an ordered data set, giving students a way to describe an individual score, such as a basketball team’s game total, relative to the rest of the data set rather than in isolation.
Measures of center and spread
Beyond location, the chapter covers the measures of center, mean, median, and mode, and the measures of spread, variance, standard deviation, and range, teaching students to recognize, describe, and calculate each one so that a data set can be summarized by both its typical value and how much its values vary.
This chapter distinguishes continuous random variables, which are measured, from discrete random variables, which are counted, using examples such as the length of a phone call or a person’s SAT score. It covers continuous probability density functions in general, and studies the uniform and exponential distributions as two specific continuous models.
Measured values versus counted values
The chapter uses paired examples, such as the number of miles driven, which is counted and discrete, against the actual distance driven, which is measured and continuous, to teach students how the same underlying quantity can be treated as either a discrete or a continuous random variable depending on exactly how it is defined.
Continuous probability density functions
Students learn to recognize and understand continuous probability density functions in general, building on the idea that probability corresponds to the area under a curve rather than to the height of a bar, extending the relative-frequency reasoning already familiar from histograms into the continuous setting.
The uniform and exponential distributions
The chapter studies two named continuous distributions in turn: the uniform distribution, where every outcome in an interval is equally likely, and the exponential distribution, and asks students to recognize each one and apply it appropriately to a described situation, building toward the normal distribution introduced later in the course.
This chapter introduces confidence intervals through an everyday example, estimating the average number of candies in a bag, building toward the formal idea of an interval estimate. It covers calculating and interpreting confidence intervals for a population mean and proportion, the Student’s t-distribution, and how sample size affects margin of error.
From a point estimate to an interval estimate
Starting from the familiar problem of estimating a mean, such as the mean rent of an apartment or the average number of candies in a bag, the chapter shows that a single point estimate is unlikely to equal the true population value exactly, motivating the construction of a confidence interval, a range of values calculated from sample data that is likely to contain the unknown parameter.
The Student’s t-distribution
By the end of the chapter students can discriminate between problems that call for the normal distribution and those that call for the Student’s t-distribution, which is used in place of the normal distribution when the population standard deviation is unknown and must be estimated from the sample, with the t-distribution’s shape changing as the sample size changes.
Sample size for a target confidence level
The chapter also covers calculating the sample size required to estimate a population mean or proportion for a given desired confidence level and margin of error, giving students a way to plan a study in advance rather than only analyzing data that has already been collected.
This chapter presents the normal distribution as the most important of all continuous distributions, appearing across psychology, business, economics, and the sciences, while cautioning that it cannot be applied to everything. It covers the distribution’s two parameters, the standard normal distribution, and how to use the normal distribution to find probabilities.
The most important, and most abused, distribution
The chapter opens by calling the normal distribution the most important of all distributions, noting its bell-shaped graph appears across psychology, business, economics, and the sciences and that some instructors even use it to help set grade curves, while explicitly cautioning that it is also widely misapplied and cannot be assumed for every real-world quantity.
Two parameters and the standard normal distribution
A normal distribution is fully described by two numerical parameters, its mean μ and its standard deviation σ, and the chapter introduces the standard normal distribution, the special case with mean zero and standard deviation one, as the reference distribution used to compute probabilities for any normal distribution by converting, or standardizing, values onto it.
Using the normal distribution
Once a quantity is established as normally distributed with a given mean and standard deviation, the chapter shows how to use the distribution to find probabilities for specified ranges of values, representing them as shaded areas under the bell curve, building the practical skill of moving between raw values and the probabilities they correspond to.
This chapter introduces the chi-square distribution through examples such as whether a coffee machine dispenses a consistent amount each time, framing it as the tool behind three major hypothesis tests. It covers the facts and notation of the chi-square distribution, the test of a single variance, and the goodness-of-fit test.
Three applications of one distribution
The chapter introduces the chi-square distribution as the basis for three distinct hypothesis tests: the goodness-of-fit test, which checks whether data fit a particular distribution, such as evenly distributed lottery numbers; the test of independence, which checks whether two categorical variables such as movie preference and age group are related; and the test of a single variance, illustrated with whether a coffee machine’s dispensed amount is consistent.
Facts about the chi-square distribution
The chi-square distribution is nonsymmetrical and skewed to the right, with a different curve for each value of its degrees of freedom, and the chapter notes that its test statistic is always greater than or equal to zero and that, for a sufficiently large number of degrees of freedom, the chi-square curve begins to approximate the normal distribution.
Testing a single variance and goodness of fit
Where earlier chapters focused on means and proportions, the chi-square test of a single variance shifts attention to variability itself, while the goodness-of-fit test determines whether an observed distribution of categorical data is consistent with a claimed or expected distribution, both worked through with the degrees of freedom appropriate to each specific use.