AP Statistics 4.7: Difference of Means CIs
Construct a two-sample t-interval for a difference between means after checking independent sampling, 10% limits, and sample shape.
- Syllabus
- Effective Fall 2025
- Course
- AP Statistics
Construct a two-sample t-interval for a difference between means after checking independent sampling, 10% limits, and sample shape.
Cumulative exposure to nitrogen dioxide ( NO2 ) is a major risk factor for lung disease in tunnel construction workers. In one study, researchers compared cumulative exposure to NO2 (in parts per million per year, ppm/yr) for a random sample of drill and blast workers and for an independent random sample of concrete workers. Summary statistics are shown in the following table:
\begin{tabular}[t]{|l|l|l|l|}
\hline Tunnel Worker
Activity & Sample Size & Mean
Cumulative NO2
Exposure
(ppm/yr) & Standard
Deviation in
Cumulative NO2
Exposure
(ppm/yr) \\
\hline Drill and
blast & 115 & 4.1 & 1.8 \\
\hline Concrete & 69 & 4.8 & 2.4 \\
\hline
\end{tabular}

Identify the appropriate inference procedure.
Use a two-sample t-interval for a difference between population means.
Assume conditions hold and use 95 percent confidence.
Calculate the confidence interval.
For drill-and-blast minus concrete mean exposure, the 95 percent confidence interval is (-1.362,-0.0381) ppm/yr.
If both data sets are somewhat skewed, determine whether the analysis remains valid.
Yes. Both sample sizes, 115 and 69, exceed 30, so the central limit theorem makes the analysis valid despite some skewness.
Independent random samples of 500 households were taken from a large metropolitan area in the United States for the years 1950 and 2000. Histograms of household size (number of people in a household) for the years are shown below.

A researcher wants to use these data to construct a confidence interval to estimate the change in mean household size in the metropolitan area from the year 1950 to the year 2000. State the conditions for using a two-sample t-procedure, and explain whether the conditions for inference are met.
Part (b):
The conditions for applying a two-sample t-procedure are:
The data come from independent random samples or from random assignment to two groups;
The populations are normally distributed, or both sample sizes are large;
The population sizes are at least 10 (or 20) times the sample sizes.
The first condition is satisfied because independent random samples were selected for the years 1950 and 2000. The second condition is satisfied because the sample sizes (500 in each group) are quite large, despite the right skewness of the distributions of household sizes in the sample data. The third condition is satisfied because the number of households in the large metropolitan area in both 1950 and 2000 would easily exceed 10×500=5,000.
Scoring
This question is scored in four sections. Part (a) has three components: (1) comparing the centers of the two distributions; (2) comparing variability for the two distributions; (3) identifying the shapes of both distributions and including context related to the variable of interest. Section 1 consists of part (a), component 1; section 2 consists of part (a), component 2; section 3 consists of part (a), component 3. Section 4 consists of part (b). Sections 1 and 2 are scored as essentially correct (E) or incorrect (I). Sections 3 and 4 are scored as essentially correct (E), partially correct (P), or incorrect (I).
Section 1 is scored as follows:
Essentially correct (E) if the response correctly compares center (or location) for both distributions.
Incorrect (I) otherwise.
Section 2 is scored as follows:
Essentially correct (E) if the response correctly compares variability for both distributions.
Incorrect (I) otherwise.
Section 3 is scored as follows:
Essentially correct (E) if the response includes context related to the variable of interest (household size) AND the response correctly identifies the shapes of both distributions.
Partially correct (P) if the response correctly identifies the shapes of both distributions BUT does NOT include context related to the variable of interest (household size),
OR
if the response correctly identifies the shape of only one distribution AND includes context related to the variable of interest (household size).
Incorrect (I) otherwise.
Section 4 is scored as follows:
Essentially correct (E) if the response correctly states and checks the following two conditions.
The data come from independent random samples
Normality/sample size conditions.
Partially correct (P) if the response correctly states and checks only one of the two conditions listed above,
OR
if the response correctly refers to random samples and large sample size, BUT does NOT state and check either condition correctly.
Incorrect (I) otherwise.
Note: The population size condition does not need to be checked to earn E or P.
Each essentially correct (E) section counts as 1 point. Each partially correct (P) section counts as 21 point.
Complete Response
Substantial Response
Developing Response
Minimal Response
If a response is between two scores (for example, 221 points), use a holistic approach to decide whether to score up or down, depending on the overall strength of the response and communication.