A Level only mathematical and statistical requirements

Syllabus
9700–2028–2029
Topic
Level
A2

Learning objectives

Probability predicts patterns; samples contain chance variation

Probability gives the expected long-run frequency of an outcome. A biological sample is one finite set of observations, so its observed frequency can differ from expectation through chance even when the model is correct.

  1. Convert a genetic ratio to probabilities: a 3:1 ratio predicts P(dominant) = 3/4 and P(recessive) = 1/4.
  2. Multiply probability by sample size to obtain expected numbers.
  3. Sample randomly and representatively so every eligible individual has an appropriate chance of selection.
  4. Replicate or increase sample size to reduce the proportional effect of chance variation.
  5. Compare observed and expected outcomes with an appropriate test before deciding whether the departure is larger than chance alone would plausibly produce.

For 20 offspring under a 3:1 model, expected counts are 15 dominant and 5 recessive. Observing 14 and 6 is not automatically evidence against the genetic model; expectation is not a quota imposed on each family.

A large sample reduces random sampling variation but does not repair systematic bias. Sampling only the easiest organisms to find can remain unrepresentative however many are counted.

Three formulae answer three different population questions

Use the provided formula only after identifying what each symbol represents and what biological quantity is being estimated: allele/genotype frequency, total population size, or diversity.

Question | Formula | Interpretation
Hardy–Weinberg | p + q = 1; p² + 2pq + q² = 1 | p and q are allele frequencies; p², 2pq and q² are genotype frequencies. If recessive phenotype frequency is q², take √q² to find q, then p = 1 − q.
Lincoln index | N = (n₁ × n₂) ÷ m₂ | n₁ first capture, n₂ second capture, m₂ marked recaptures; N estimates population size.
Simpson's index | D = 1 − Σ(n/N)² | n is the number of each type and N the total; larger D indicates greater diversity because abundance is spread more evenly among types.

Interpret assumptions with the result. Hardy–Weinberg treats frequencies as stable under its model conditions. Lincoln requires marking not to change survival or recapture, marks to persist, mixing between samples, and little migration, birth or death. Simpson's D depends on representative identification and sampling of types.

Do not set q equal to a recessive genotype frequency: q² is the genotype frequency. In Lincoln estimation, a very small m₂ makes N highly unstable. A diversity index summarises richness and evenness; it does not identify why communities differ.

Spread and uncertainty require different statistics

First inspect the distribution. A roughly symmetric bell-shaped distribution is normal; skew, multiple peaks or strong outliers may make it non-normal. Then choose a centre, spread or uncertainty statistic that answers the question.

Statistic | What it describes
Mean | arithmetic centre; sensitive to extreme values
Median | middle ordered value; often more representative for skewed data
Mode | most frequent value or category
Range | total span from minimum to maximum
Sample SD, s | spread of individual observations around the sample mean
SE = s/√n | uncertainty in the sample mean; decreases as n increases
95% CI ≈ mean ± (2 × SE) | interval estimating the population mean with stated confidence

Using the provided formula, sample SD is s = √[Σ(x − mean)² ÷ (n − 1)]. Then calculate SE and the lower/upper 95% confidence limits. Plot error bars symmetrically from the mean and label the legend or axis note as SD, SE or 95% CI so the reader knows what they represent.

Small SD means observations are clustered; small SE means the mean is estimated precisely. Neither alone proves accuracy or biological significance. Error-bar overlap can inform comparison but is not a substitute for the required statistical test.

A test statistic becomes a decision through df and a critical value

Use chi-squared to compare observed and expected frequencies of nominal categories. Use the syllabus t-test to compare the means of two independent samples of continuous data when the populations are approximately normal and their standard deviations are approximately equal.

Chi-squared: χ² = Σ[(O − E)²/E]; degrees of freedom v = c − 1, where c is the number of classes.
t-test: use the provided two-sample formula with each mean, sample SD and sample size; degrees of freedom v = n₁ + n₂ − 2.

Always state H₀ first: there is no significant difference from expectation, or no significant difference between the two population means.

  1. Calculate the statistic and df. 2. Choose the stated significance level, commonly p = 0.05. 3. Read the critical value for that df. 4. If the calculated statistic exceeds the critical value, p is below the threshold: reject H₀ and call the difference statistically significant. Otherwise, do not reject H₀ because the evidence is insufficient to distinguish the result from chance variation.

Do not write that H₀ is proved or accepted. Rejecting H₀ does not identify a mechanism, prove the biological hypothesis, or show that the effect is large; failing to reject H₀ does not prove that the groups are identical.

Choose correlation from distribution, scale and pattern

A correlation coefficient describes the direction and strength of association between paired observations from −1 (perfect negative) through 0 (no monotonic or linear correlation) to +1 (perfect positive). It does not by itself show that either variable causes the other.

Pearson's linear correlation | two continuous variables; scatter plot suggests a linear relationship; population data are normally distributed; at least 5 paired observations, ideally 10 or more.
Spearman's rank correlation | ordinal/ranked data or non-normal data; independent points; scatter plot suggests an increasing or decreasing relationship; more than 5 pairs, ideally 10–30; individuals selected randomly with equal selection chance.

Use the provided formula, state H₀ as no significant correlation, and compare the coefficient with the correct critical value for sample size and significance level.

Even a strong significant correlation may arise because a third variable affects both measurements, because selection was biased, or by coincidence. A causal claim needs an appropriate controlled design, plausible biological mechanism, temporal direction and replicated evidence beyond the coefficient.

A coefficient near 0 can occur when a real relationship is curved rather than linear/monotonic, so inspect the scatter plot. Do not use Pearson merely because both columns contain numbers, and do not convert paired observations into ranks unless Spearman's conditions apply.