AP Statistics 4.10: Two-Mean Test Conclusions
Conclude a two-sample mean test by comparing p-value with α and explaining what the evidence says about both populations.
- Syllabus
- Effective Fall 2025
- Course
- AP Statistics
Conclude a two-sample mean test by comparing p-value with α and explaining what the evidence says about both populations.
An experiment is run to test whether daily stimulation of specific reflexes in young infants will lead to earlier walking. Twenty infants were recruited through a pediatrician's service and were randomly split into 2 groups of 10 . One group received the daily stimulation, while the other was considered a control group. The ages (in months) at which the infants first walked alone were recorded.
With stimulation: 10,12,11,10.5,11,11.5,11.5,11,12,11.5Mean=11.2,SD=0.63246
Control: 10, 13, 12, 11, 11.5, 11.5, 12.5, 12, 11.5, 11 Mean = 11.6, SD=0.84327

Complete the inference procedure, including calculations.
Let μ1 and μ2 be the true mean walking ages for infants receiving stimulation and no stimulation. Test H0:μ1=μ2 against Ha:μ1<μ2. Random assignment satisfies randomization; the dotplots show no outliers or strong skew, supporting normality. The test gives t=-1.20, df=16.69, and p=0.1234.
Justify a conclusion in context.
Because 0.1234>0.05, fail to reject H0. There is not convincing evidence that mean walking age is lower for infants receiving daily stimulation.
A research scientist is conducting an experiment to determine whether a new chemical process creates less toxic byproduct compared to the current chemical process. The scientist took a random sample of products made with the current chemical process and calculated the sample mean amount of toxic byproduct created. The scientist then took a random sample of products made with the new chemical process and calculated the sample mean amount of toxic byproduct created. The difference in the sample means (current minus new) was 2.31 liters of toxic byproduct. A hypothesis test was conducted using the following hypotheses.
Assuming the conditions for inference were met, the scientist calculated the p-value of the test to be 0.072 . Which of the following statements is the best interpretation of the p-value?
The probability that the null hypothesis is true is 0.072 .
The probability that the alternative hypothesis is true is 0.072.
The probability of observing a difference in means of 2.31 liters of toxic byproduct is 0.072.
If the null hypothesis is true, the probability of observing a difference in means of at least 2.31 liters of toxic byproduct is 0.072.
If the null hypothesis is true, the probability of observing a difference in means of at most 2.31 liters of toxic byproduct is 0.072.
D
Stefan, a psychologist, conducted a study to investigate the effect of time of day on reading
comprehension in children. One hundred children volunteered, with their parents' consent, to
participate in the study. Fifty of the children were randomly assigned to read a story at 9 a.m.
and then answer 25 questions about it. The remaining 50 children were assigned to read the
same story at 3 p.m. and answer the same 25 questions. The reading comprehension for each
child was measured by a reading score, which was determined by the number of questions that
were answered correctly about the story. Stefan is interested in comparing the mean reading
scores for the two times of day. Table 1 shows the results of Stefan's study.

Table 1: Summary Statistics of Reading Scores
Stefan found the conditions for inference were met and conducted a two-sample t-test for the
difference in two population means. Let μAM represent the mean reading score for all children,
similar to those in the study, who would read the story at 9 a.m. Let μPM represent the mean
reading score for all children, similar to those in the study, who would read the story at 3 p.m.
Stefan's hypotheses are as shown.
The p-value for Stefan's hypothesis test was 0.002. State an appropriate conclusion, at the
5 percent significance level, for Stefan's test in the context of the investigation. Justify your
answer.
| Model Solution | Scoring | |
|---|---|---|
| A | Because the p-value of 0.002 is less than the level of significance of 0.05, the null hypothesis should be rejected. There is convincing statistical evidence of a difference between the mean reading score for all children, similar to those who participated in the study, who would read the story at 9 a.m. and the mean reading score for all children, similar to those who participated in the study, who would read the story at 3 p.m. | Essentially correct (E) if the response satisfies the following two components: 1. Provides correct comparison of the p-value to alpha ( p-value is less than α ) AND provides a correct decision about the null and/or alternative hypothesis 2. States a conclusion in context, consistent with, and in terms of the stated alternative hypothesis using nondefinitive language Partially correct (P) if the response satisfies only one of the two components required for E. Incorrect (I) if the response does not meet the criteria for E or P. |
Scoring Notes:
- To satisfy the p-value comparison in component 1, the response can compare the value of the test statistic
to an appropriate critical value; for example, |t|>1.985 if d f=97.489, or |t|>2.01 if d f=49.
- An explicit decision about the null hypothesis is not required to satisfy component 1.
- If an explicit decision is stated and the conclusion is inconsistent with the decision, component 1 is not
satisfied.
- The decision part of component 1 may be satisfied by implying the decision within the conclusion
statement (sufficient evidence for the alternative hypothesis).
- To satisfy component 2, the response must include reference to means, groups (e.g., 9 a.m. and 3 p.m.),
the sampling units (e.g., children), and the variable of interest (reading score).
- Examples of nondefinitive language in component 2 include "evidence to accept the alternative," "there is
evidence for the alternative," and "there is not sufficient evidence for the alternative."
- Examples of definitive language in component 2 include "accepts the null," "proves the null," "proves the
alternative," "accepts the alternative," "there is not evidence for the alternative," and "no evidence for the
alternative."
- If components 1 and/or 2 are satisfied and the response provides an incorrect interpretation of the p-value,
the score is lowered from E to P or P to I.
- The quality of communication for responses with score P should be considered if holistic scoring is
required.