AP Statistics 4.10: Two-Mean Test Conclusions
Justify a comparison of two population means by evaluating the p-value, sampling variation, and the alternative claim in context.
- Syllabus
- Effective Fall 2025
- Course
- AP Statistics
Justify a comparison of two population means by evaluating the p-value, sampling variation, and the alternative claim in context.
Stefan, a psychologist, conducted a study to investigate the effect of time of day on reading
comprehension in children. One hundred children volunteered, with their parents' consent, to
participate in the study. Fifty of the children were randomly assigned to read a story at 9 a.m.
and then answer 25 questions about it. The remaining 50 children were assigned to read the
same story at 3 p.m. and answer the same 25 questions. The reading comprehension for each
child was measured by a reading score, which was determined by the number of questions that
were answered correctly about the story. Stefan is interested in comparing the mean reading
scores for the two times of day. Table 1 shows the results of Stefan's study.

Table 1: Summary Statistics of Reading Scores
Stefan found the conditions for inference were met and conducted a two-sample t-test for the
difference in two population means. Let μAM represent the mean reading score for all children,
similar to those in the study, who would read the story at 9 a.m. Let μPM represent the mean
reading score for all children, similar to those in the study, who would read the story at 3 p.m.
Stefan's hypotheses are as shown.
The p-value for Stefan's hypothesis test was 0.002. State an appropriate conclusion, at the
5 percent significance level, for Stefan's test in the context of the investigation. Justify your
answer.
| Model Solution | Scoring | |
|---|---|---|
| A | Because the p-value of 0.002 is less than the level of significance of 0.05, the null hypothesis should be rejected. There is convincing statistical evidence of a difference between the mean reading score for all children, similar to those who participated in the study, who would read the story at 9 a.m. and the mean reading score for all children, similar to those who participated in the study, who would read the story at 3 p.m. | Essentially correct (E) if the response satisfies the following two components: 1. Provides correct comparison of the p-value to alpha ( p-value is less than α ) AND provides a correct decision about the null and/or alternative hypothesis 2. States a conclusion in context, consistent with, and in terms of the stated alternative hypothesis using nondefinitive language Partially correct (P) if the response satisfies only one of the two components required for E. Incorrect (I) if the response does not meet the criteria for E or P. |
Scoring Notes:
- To satisfy the p-value comparison in component 1, the response can compare the value of the test statistic
to an appropriate critical value; for example, |t|>1.985 if d f=97.489, or |t|>2.01 if d f=49.
- An explicit decision about the null hypothesis is not required to satisfy component 1.
- If an explicit decision is stated and the conclusion is inconsistent with the decision, component 1 is not
satisfied.
- The decision part of component 1 may be satisfied by implying the decision within the conclusion
statement (sufficient evidence for the alternative hypothesis).
- To satisfy component 2, the response must include reference to means, groups (e.g., 9 a.m. and 3 p.m.),
the sampling units (e.g., children), and the variable of interest (reading score).
- Examples of nondefinitive language in component 2 include "evidence to accept the alternative," "there is
evidence for the alternative," and "there is not sufficient evidence for the alternative."
- Examples of definitive language in component 2 include "accepts the null," "proves the null," "proves the
alternative," "accepts the alternative," "there is not evidence for the alternative," and "no evidence for the
alternative."
- If components 1 and/or 2 are satisfied and the response provides an incorrect interpretation of the p-value,
the score is lowered from E to P or P to I.
- The quality of communication for responses with score P should be considered if holistic scoring is
required.