This chapter introduces confidence intervals through an everyday example, estimating the average number of candies in a bag, building toward the formal idea of an interval estimate. It covers calculating and interpreting confidence intervals for a population mean and proportion, the Student’s t-distribution, and how sample size affects margin of error.
From a point estimate to an interval estimate
Starting from the familiar problem of estimating a mean, such as the mean rent of an apartment or the average number of candies in a bag, the chapter shows that a single point estimate is unlikely to equal the true population value exactly, motivating the construction of a confidence interval, a range of values calculated from sample data that is likely to contain the unknown parameter.
The Student’s t-distribution
By the end of the chapter students can discriminate between problems that call for the normal distribution and those that call for the Student’s t-distribution, which is used in place of the normal distribution when the population standard deviation is unknown and must be estimated from the sample, with the t-distribution’s shape changing as the sample size changes.
Sample size for a target confidence level
The chapter also covers calculating the sample size required to estimate a population mean or proportion for a given desired confidence level and margin of error, giving students a way to plan a study in advance rather than only analyzing data that has already been collected.
This chapter uses business examples, such as estimating monthly iTunes downloads from a marketing survey, to introduce confidence intervals. It explains how a point estimate becomes an interval estimate, covers the Student’s t-distribution for unknown standard deviations, and shows how to size a sample for a target confidence level and margin of error.
From point estimate to confidence interval
A sample mean or sample proportion gives a single point estimate of an unknown population parameter, but a confidence interval expresses that estimate as a range together with a stated probability of accuracy, called the confidence level. Using a marketing example of estimating the mean number of songs downloaded per month, the chapter shows how the central limit theorem lets an analyst attach two standard deviations to a sample mean to state, for instance, 95 percent confidence that the true population mean falls within a given interval.
Known standard deviation and large samples
When the population standard deviation is known, or the sample is large, the confidence interval for a population mean is built directly from the normal distribution and the standard error of the sampling distribution of means. The width of the interval depends on the desired confidence level, set by the z-value the analyst chooses, and on the sample size, since a larger sample narrows the interval for the same confidence level.
Business applications of the interval framework
Because managers rarely know a population’s exact standard deviation, the chapter extends the framework to the Student’s t-distribution, which adjusts for the extra uncertainty of estimating the standard deviation from the sample itself, and shows how to work backward from a desired margin of error and confidence level to the sample size a survey or study needs before it is run.
A confidence interval gives a range of plausible values for an unknown population value, built from sample data. This chapter shows how to construct confidence intervals for a population mean using the normal and the Student t distributions, and for a population proportion, and how the confidence level and sample size affect the interval.
What a Confidence Interval Is
A confidence interval is a range of values, calculated from a sample, that is likely to contain an unknown population value such as a mean or a proportion. It is built as a point estimate plus and minus a margin of error. The confidence level, often ninety-five percent, describes how often intervals built this way would contain the true value if sampling were repeated many times. A wider interval gives more confidence, while a narrower one is more precise.
Intervals for a Mean
When the population standard deviation is known, or the sample is large, a confidence interval for the mean uses the normal distribution, with the margin of error equal to a z-value times the standard error. When the population standard deviation is unknown and estimated from the sample, the Student t distribution is used instead. The t distribution is wider than the normal for small samples and approaches the normal as the sample size grows, which accounts for the extra uncertainty of estimating the standard deviation.
Intervals for a Proportion and Sample Size
A confidence interval for a population proportion is built from the sample proportion plus and minus a margin of error based on the normal distribution. The margin of error depends on the confidence level, the variability in the data, and the sample size. Because a larger sample reduces the margin of error, the required sample size can be calculated in advance from a chosen confidence level and a desired margin of error, letting a study be planned to reach a target precision.