2.4.3—Inferential statistics

Syllabus
First assessment 2027
Objective
2.4.3
Level
HL

Inferential statistics quantify how surprising a result would be under chance

HL only

An inferential test compares an observed result with what would be expected if the null hypothesis were true. A p-value is the probability of a result at least this extreme under that model, not the probability that the null is true.

At a chosen significance level such as 0.05, a result below the threshold is evidence against the null; it does not prove a theory or measure the size of an effect. A critical-value decision must match the test and tail.

A Type I error is a false positive: rejecting a true null. A Type II error is a false negative: failing to reject a false null. Tightening the threshold can reduce Type I risk while making Type II errors more likely.

‘Not significant’ is not proof of no effect, and ‘significant’ is not proof of importance. Report the decision, uncertainty and design limits together.

Choose the test from design and data: an unrelated t-test compares two independent means when parametric assumptions are suitable; a related t-test compares paired means; Mann–Whitney and Wilcoxon are corresponding rank-based alternatives; chi-square tests association between categorical variables; a correlation coefficient describes direction and strength of a relationship. Statistical significance addresses compatibility with the null model, while effect size addresses magnitude. Neither repairs bias, confounding or a non-causal design.