4.3 Chi-squared tests
- Syllabus
- 9231–2028–2029
- Topic
- 4.3
- Level
- AS
Fit a probability model by estimating its parameters from data, then compare expected behaviour with observed frequencies or summary features. A fitted model is an approximation, not an explanation by itself.
Keep class intervals consistent, calculate expected counts from model probabilities, and check that expected counts are large enough for the chosen test. Parameters estimated from the same data affect degrees of freedom.
Fit a Poisson model using the sample mean λ̂, then calculate each class probability from λ̂ before forming expected counts.
Matching the mean does not guarantee a good fit; dispersion, tail behaviour and the test statistic still matter.
The chi-squared goodness-of-fit statistic is Σ(O−E)²/E. Under H₀ the proposed distribution is adequate, subject to model and expected-count conditions.
Combine tail classes when expected counts are too small, subtract parameters estimated from the data when determining degrees of freedom, and state the conclusion in context.
A large statistic relative to the critical value gives evidence against the fitted distribution; it does not identify which class caused the mismatch without inspecting contributions.
Rejecting H₀ does not prove every observation is wrong, and a small statistic cannot prove the model true.
For a contingency table, H₀ states that the two categorical variables are independent. Expected count in a cell is row total × column total ÷ grand total, then use the chi-squared statistic.
Use counts rather than percentages, check expected-count conditions, and describe any rejection as evidence of association—not proof of causation.
If a row total is 40, a column total is 30 and the grand total is 100, the independent expected count is 12. Compare the observed cell with 12.
Independence is not mutual exclusivity, and a significant association can be produced by confounding or sampling bias.