A Level only mathematical and statistical requirements
- Syllabus
- 9700–2028–2029
- Topic
- —
- Level
- A2
Probability gives the expected long-run frequency of an outcome. A biological sample is one finite set of observations, so its observed frequency can differ from expectation through chance even when the model is correct.
For 20 offspring under a 3:1 model, expected counts are 15 dominant and 5 recessive. Observing 14 and 6 is not automatically evidence against the genetic model; expectation is not a quota imposed on each family.
A large sample reduces random sampling variation but does not repair systematic bias. Sampling only the easiest organisms to find can remain unrepresentative however many are counted.
Use the provided formula only after identifying what each symbol represents and what biological quantity is being estimated: allele/genotype frequency, total population size, or diversity.
Question | Formula | Interpretation
Hardy–Weinberg | p + q = 1; p² + 2pq + q² = 1 | p and q are allele frequencies; p², 2pq and q² are genotype frequencies. If recessive phenotype frequency is q², take √q² to find q, then p = 1 − q.
Lincoln index | N = (n₁ × n₂) ÷ m₂ | n₁ first capture, n₂ second capture, m₂ marked recaptures; N estimates population size.
Simpson's index | D = 1 − Σ(n/N)² | n is the number of each type and N the total; larger D indicates greater diversity because abundance is spread more evenly among types.
Interpret assumptions with the result. Hardy–Weinberg treats frequencies as stable under its model conditions. Lincoln requires marking not to change survival or recapture, marks to persist, mixing between samples, and little migration, birth or death. Simpson's D depends on representative identification and sampling of types.
Do not set q equal to a recessive genotype frequency: q² is the genotype frequency. In Lincoln estimation, a very small m₂ makes N highly unstable. A diversity index summarises richness and evenness; it does not identify why communities differ.
First inspect the distribution. A roughly symmetric bell-shaped distribution is normal; skew, multiple peaks or strong outliers may make it non-normal. Then choose a centre, spread or uncertainty statistic that answers the question.
Statistic | What it describes
Mean | arithmetic centre; sensitive to extreme values
Median | middle ordered value; often more representative for skewed data
Mode | most frequent value or category
Range | total span from minimum to maximum
Sample SD, s | spread of individual observations around the sample mean
SE = s/√n | uncertainty in the sample mean; decreases as n increases
95% CI ≈ mean ± (2 × SE) | interval estimating the population mean with stated confidence
Using the provided formula, sample SD is s = √[Σ(x − mean)² ÷ (n − 1)]. Then calculate SE and the lower/upper 95% confidence limits. Plot error bars symmetrically from the mean and label the legend or axis note as SD, SE or 95% CI so the reader knows what they represent.
Small SD means observations are clustered; small SE means the mean is estimated precisely. Neither alone proves accuracy or biological significance. Error-bar overlap can inform comparison but is not a substitute for the required statistical test.
Use chi-squared to compare observed and expected frequencies of nominal categories. Use the syllabus t-test to compare the means of two independent samples of continuous data when the populations are approximately normal and their standard deviations are approximately equal.
Chi-squared: χ² = Σ[(O − E)²/E]; degrees of freedom v = c − 1, where c is the number of classes.
t-test: use the provided two-sample formula with each mean, sample SD and sample size; degrees of freedom v = n₁ + n₂ − 2.
Always state H₀ first: there is no significant difference from expectation, or no significant difference between the two population means.
Do not write that H₀ is proved or accepted. Rejecting H₀ does not identify a mechanism, prove the biological hypothesis, or show that the effect is large; failing to reject H₀ does not prove that the groups are identical.
A correlation coefficient describes the direction and strength of association between paired observations from −1 (perfect negative) through 0 (no monotonic or linear correlation) to +1 (perfect positive). It does not by itself show that either variable causes the other.
Pearson's linear correlation | two continuous variables; scatter plot suggests a linear relationship; population data are normally distributed; at least 5 paired observations, ideally 10 or more.
Spearman's rank correlation | ordinal/ranked data or non-normal data; independent points; scatter plot suggests an increasing or decreasing relationship; more than 5 pairs, ideally 10–30; individuals selected randomly with equal selection chance.
Use the provided formula, state H₀ as no significant correlation, and compare the coefficient with the correct critical value for sample size and significance level.
Even a strong significant correlation may arise because a third variable affects both measurements, because selection was biased, or by coincidence. A causal claim needs an appropriate controlled design, plausible biological mechanism, temporal direction and replicated evidence beyond the coefficient.
A coefficient near 0 can occur when a real relationship is curved rather than linear/monotonic, so inspect the scatter plot. Do not use Pearson merely because both columns contain numbers, and do not convert paired observations into ranks unless Spearman's conditions apply.