AP Statistics 4.5: Mean Test Conclusions
Carry out a mean or mean-difference t-test by comparing p-value with α and stating a contextual, non-definitive conclusion.
- Syllabus
- Effective Fall 2025
- Course
- AP Statistics
Carry out a mean or mean-difference t-test by comparing p-value with α and stating a contextual, non-definitive conclusion.
A top-100, 7.0-rated ten)nis pro wishes to compare a new racket against his current model. He is interested in whether his hitting speed with the new racket is different from that with his old racket. He strings the new racket with the same type of strings at 60 pounds tension that he uses on his old racket. From past testing, the tennis pro knows that the average forehand crosscourt volley with his old racket is 82 miles per hour (mph). On an indoor court, using a ball machine set at 70 mph, which is the same speed he had his old racket tested against, he takes 47 swings with the new racket. (Assume that swings are independent of each other.) An associate with a speed gun records an average of 83.5 mph with a standard deviation of 3.4 mph. At the 0.01 level of significance, complete the appropriate inference procedure to determine if there is convincing statistical evidence that his mean hitting speed with the new racket is different from that with his old racket.
Complete the inference procedure, including calculating the appropriate statistics.
Let μ be the true mean hitting speed with the new racket. Test H0:μ=82 against Ha:μ=82. The swings are independent, and n=47≥30, so the normality condition is satisfied by the central limit theorem. The test statistic is t=3.0246, with df=46 and p=0.00406.
Justify a conclusion in context.
Because 0.00406 is less than 0.01, reject the null hypothesis. There is convincing evidence that the tennis professional’s mean hitting speed with the new racket differs from that with the old racket.
Past studies indicate that the average number of acres of wildland burned per wildfire each year in the United States is 170 acres. A forest ranger believes that the average might now be greater. A random sample of wildfires was surveyed. With H0:μ=170 and Ha:μ>170, the p-value of the test is 0.022 . Which of the following is a correct interpretation of the p-value?
0.022 is the probability the population mean is greater than the one obtained by the ranger.
0.022 is the probability of obtaining a sample mean as large as or larger than the one obtained by the ranger.
If it is true that the average number of acres of wildland burned per wildfire each year in the United States is 170 acres, 0.022 is the probability of obtaining a sample mean as large as or larger than 170.
If it is true that the average number of acres of wildland burned per wildfire each year in the United States is 170 acres, 0.022 is the probability of obtaining a sample mean as large as or larger than the one obtained by the ranger.
D
A researcher conducted a study to investigate whether local car dealers tend to charge women more than men for the same car model. Using information from the county tax collector's records, the researcher randomly selected one man and one woman from among everyone who had purchased the same model of an identically equipped car from the same dealer. The process was repeated for a total of 8 randomly selected car models.
The purchase prices and the differences (woman - man) are shown in the table below. Summary statistics are also shown.


Dotplots of the data and the differences are shown below.

Purchase Price (in thousands of dollars)

Difference in Purchase Price (woman - man, in dollars)
Do the data provide convincing evidence that, on average, women pay more than men in the county for the same car model?
Step 1: States a correct pair of hypotheses.
Let μdiff represent the population mean difference in purchase price (woman - man) for identically equipped cars of the same model, sold to both men and women by the same dealer, in the county.
The hypotheses to be tested are H0:μdiff =0 versus Ha:μdiff >0.
Step 2: Identifies a correct test procedure (by name or by formula) and checks appropriate conditions.
The appropriate procedure is a paired t-test.
The conditions for the paired t-test are:
The sample is randomly selected from the population.
The population of price differences (woman - man) is normally distributed, or the sample size is large.
The first condition is met because the car models and the individuals were randomly selected. The sample size (n=8) is not large, so we need to investigate whether it is reasonable to assume that the population of price differences is normally distributed. The dotplot of sample price differences reveals a fairly symmetric distribution, so we will consider the second condition to be met.
Step 3: Correct mechanics, including the value of the test statistic and p-value (or rejection region).
The test statistic is t=8530.71585−0≈3.12.
The p-value, based on a t-distribution with 8-1=7 degrees of freedom, is 0.008.
Step 4: States a correct conclusion in the context of the study, using the result of the statistical test.
Because the p-value is very small (for instance, smaller than α=0.05 ), we reject the null hypothesis. The data provide convincing evidence that, on average, women pay more than men in the country for the same car model.
Scoring
Each of steps 1, 2, 3, and 4 were scored as essentially correct (E), partially correct (P), or incorrect (I).
Step 1 is scored as follows:
Essentially correct (E) if the response identifies the correct parameter AND states correct hypotheses.
Partially correct (P) if the response identifies the correct parameter OR states correct hypotheses, but not both.
Incorrect (I) if the response does not meet the criteria for E or P.
Note: Defining the parameter symbol in context or simply using common parameter notation is sufficient.
Step 2 is scored as follows:
Essentially correct (E) if the response identifies the correct test procedure (by name or by formula) AND checks both conditions correctly.
Partially correct (P) if the response correctly completes two of the three components (identification of procedure, check of randomness condition, check of normality condition).
Incorrect (I) if the response does not meet the criteria for E or P.
Note: The random sampling condition can be verified by referring to the random selection of car models or to the random selection of male and female car buyers.
Step 3 is scored as follows:
Essentially correct (E) if the response correctly calculates both the test statistic and the p-value.
Partially correct (P) if the response correctly calculates the test statistic but not the p-value; OR if the response calculates the test statistic incorrectly but then calculates the correct p-value for the computed test statistic.
Incorrect (I) if the response does not meet the criteria for E or P.
Note: If the response identifies a z-test for a mean as the correct procedure in step 2, then the response can earn a P in step 3 if both the test statistic and the p-value are calculated correctly.
Step 4 is scored as follows:
Essentially correct (E) if the response provides a correct conclusion in context, also providing justification based on linkage between the p-value and the conclusion.
Partially correct (P) if the response provides a correct conclusion with linkage to the p-value, but not in context;
OR
if the response provides a correct conclusion in context, but without justification based on linkage to the p-value.
Incorrect (I) if the response does not meet the criteria for E or P.
Notes:
- If the conclusion is consistent with an incorrect p-value from step 3 and also in context with justification based on linkage to the p-value, step 4 is scored as E.
- A response that performs a two-sample t-test with correct calculations should fail to reject H0. A conclusion that is equivalent to "accept H0 " (such as "we conclude that women pay the same amount as men, on average"), either as a stated decision or as a conclusion in context, cannot be scored as E. Such a response will be scored as P provided that the conclusion is in context with linkage. Such a response will be scored as I if it lacks either context or linkage.
Each essentially correct (E) step counts as 1 point. Each partially correct (P) step counts as 21 point.
Complete Response
Substantial Response
Developing Response
Minimal Response
If a response is between two scores (for example, 221 points), use a holistic approach to decide whether to score up or down, depending on the overall strength of the response and communication.