Mathematical and statistical skills
- Syllabus
- 9700–2028–2029
- Section
- —
- Level
- A2
A number is meaningful only with its unit and scale. Use the most appropriate unit, then convert all quantities to compatible units before comparing, multiplying or dividing them.
Prefix | Symbol | Multiplier
Giga | G | 10⁹
Mega | M | 10⁶
Kilo | k | 10³
Milli | m | 10⁻³
Micro | µ | 10⁻⁶
Nano | n | 10⁻⁹
To convert, replace the prefix by its power of ten. For example, 2.5 µm = 2.5 × 10⁻⁶ m = 2500 nm.
Use decimal or standard form as the scale demands. Read <, >, ⩽ and ⩾ as comparison limits; ∝ means directly proportional; Σ means sum. In a table heading or graph axis, write a quantity followed by a solidus and unit, for example length / µm, so entries contain numbers only.
Capitalisation matters: M is mega and m is milli. A bare value such as 2.5 cannot be compared with 3000 until both quantities have stated, compatible units.
Calculation quality has three checks: estimate the expected size, calculate without premature rounding, then report a precision justified by the measurements—not by the number of digits on the calculator display.
For 12.4 ÷ 3.2, an estimate of 12 ÷ 3 ≈ 4 makes 38.75 implausible. The calculator gives 3.875. The inputs have 3 and 2 significant figures, so the syllabus permits a final value with 2 or 3 significant figures: 3.9 or 3.88, with the required unit.
Decimal places and significant figures are not interchangeable. In 0.00450, the leading zeros are placeholders and the value has three significant figures. Do not round intermediate steps so aggressively that the final result drifts.
First decide whether the unknown is a length, area, surface area or volume. Convert measurements to compatible units, choose the formula with the correct dimension, and attach the resulting unit: unit, unit² or unit³.
Magnification = image size ÷ actual size; actual size = image size ÷ magnification.
Triangle area = ½bh; rectangle area = lw; circle area = πr².
Rectangle perimeter = 2(l + w); circle circumference = 2πr.
Cuboid surface area = 2(lw + lh + wh); cuboid volume = lwh.
Cylinder surface area = 2πr² + 2πrh; cylinder volume = πr²h.
A cell of actual length 20 µm is shown as 40 mm. Convert 40 mm to 40 000 µm, then magnification = 40 000 ÷ 20 = ×2000. If every linear dimension doubles, surface area becomes 2² = 4 times larger and volume becomes 2³ = 8 times larger.
Do not mix radius and diameter, and do not attach a linear unit to area or volume. Magnification has no physical unit because image and actual size are divided in the same unit.
Choose a summary that answers the biological question. For any ratio or percentage, state what the numerator represents and which reference quantity belongs in the denominator.
Question | Calculation
Typical value using every observation | mean = Σx ÷ n
Middle of ordered observations | median
Most frequent value | mode
Spread from extremes | range = maximum − minimum
Relative amounts | ratio a:b, simplified or scaled consistently
Part of a whole | percentage = part ÷ whole × 100
Change relative to the starting value | percentage change = (final − initial) ÷ initial × 100
Measurement uncertainty relative to the measured value | percentage error = absolute error ÷ measured value × 100
A mass increases from 10 g to 12 g: absolute change = 2 g and percentage change = 2 ÷ 10 × 100 = 20%. If a 10.0 cm reading has an absolute error of ±0.1 cm, percentage error = 0.1 ÷ 10.0 × 100 = 1%. These percentages answer different questions.
Do not divide percentage change by the final value. A mean can also conceal skew or an anomalous value; use median or mode only when the data and question justify them, not as interchangeable labels.
A graph is a transformation of data, not decoration. Choose the representation from the variable and question, place the independent variable on x and dependent variable on y, and preserve units and numerical meaning.
Purpose/data | Representation
Separate categories | bar chart with separated bars
Parts of one whole | pie chart
Frequency distribution of continuous measurements | histogram with touching class intervals
Response across an ordered continuous IV | line graph or scatter plot with an appropriate straight or curved best-fit line
Label each axis as quantity / unit, choose a simple scale that uses the plotting area, plot accurately, and decide from context whether points represent a sequence to join with straight ruled lines or a trend needing best fit.
Rate of change = Δy ÷ Δx, with units from y per unit x. For a straight line, use two widely separated points on the line. For an average rate on a curve, use the relevant interval. For the rate at one instant, draw a tangent at that point and calculate the tangent gradient from a large triangle.
Do not join independent scatter points dot-to-dot or use a bar chart for a continuous IV merely because there are only a few values. A visually steeper line is not necessarily a larger rate: compare gradients only after checking both axis scales and units.
Probability gives the expected long-run frequency of an outcome. A biological sample is one finite set of observations, so its observed frequency can differ from expectation through chance even when the model is correct.
For 20 offspring under a 3:1 model, expected counts are 15 dominant and 5 recessive. Observing 14 and 6 is not automatically evidence against the genetic model; expectation is not a quota imposed on each family.
A large sample reduces random sampling variation but does not repair systematic bias. Sampling only the easiest organisms to find can remain unrepresentative however many are counted.
Use the provided formula only after identifying what each symbol represents and what biological quantity is being estimated: allele/genotype frequency, total population size, or diversity.
Question | Formula | Interpretation
Hardy–Weinberg | p + q = 1; p² + 2pq + q² = 1 | p and q are allele frequencies; p², 2pq and q² are genotype frequencies. If recessive phenotype frequency is q², take √q² to find q, then p = 1 − q.
Lincoln index | N = (n₁ × n₂) ÷ m₂ | n₁ first capture, n₂ second capture, m₂ marked recaptures; N estimates population size.
Simpson's index | D = 1 − Σ(n/N)² | n is the number of each type and N the total; larger D indicates greater diversity because abundance is spread more evenly among types.
Interpret assumptions with the result. Hardy–Weinberg treats frequencies as stable under its model conditions. Lincoln requires marking not to change survival or recapture, marks to persist, mixing between samples, and little migration, birth or death. Simpson's D depends on representative identification and sampling of types.
Do not set q equal to a recessive genotype frequency: q² is the genotype frequency. In Lincoln estimation, a very small m₂ makes N highly unstable. A diversity index summarises richness and evenness; it does not identify why communities differ.
First inspect the distribution. A roughly symmetric bell-shaped distribution is normal; skew, multiple peaks or strong outliers may make it non-normal. Then choose a centre, spread or uncertainty statistic that answers the question.
Statistic | What it describes
Mean | arithmetic centre; sensitive to extreme values
Median | middle ordered value; often more representative for skewed data
Mode | most frequent value or category
Range | total span from minimum to maximum
Sample SD, s | spread of individual observations around the sample mean
SE = s/√n | uncertainty in the sample mean; decreases as n increases
95% CI ≈ mean ± (2 × SE) | interval estimating the population mean with stated confidence
Using the provided formula, sample SD is s = √[Σ(x − mean)² ÷ (n − 1)]. Then calculate SE and the lower/upper 95% confidence limits. Plot error bars symmetrically from the mean and label the legend or axis note as SD, SE or 95% CI so the reader knows what they represent.
Small SD means observations are clustered; small SE means the mean is estimated precisely. Neither alone proves accuracy or biological significance. Error-bar overlap can inform comparison but is not a substitute for the required statistical test.
Use chi-squared to compare observed and expected frequencies of nominal categories. Use the syllabus t-test to compare the means of two independent samples of continuous data when the populations are approximately normal and their standard deviations are approximately equal.
Chi-squared: χ² = Σ[(O − E)²/E]; degrees of freedom v = c − 1, where c is the number of classes.
t-test: use the provided two-sample formula with each mean, sample SD and sample size; degrees of freedom v = n₁ + n₂ − 2.
Always state H₀ first: there is no significant difference from expectation, or no significant difference between the two population means.
Do not write that H₀ is proved or accepted. Rejecting H₀ does not identify a mechanism, prove the biological hypothesis, or show that the effect is large; failing to reject H₀ does not prove that the groups are identical.
A correlation coefficient describes the direction and strength of association between paired observations from −1 (perfect negative) through 0 (no monotonic or linear correlation) to +1 (perfect positive). It does not by itself show that either variable causes the other.
Pearson's linear correlation | two continuous variables; scatter plot suggests a linear relationship; population data are normally distributed; at least 5 paired observations, ideally 10 or more.
Spearman's rank correlation | ordinal/ranked data or non-normal data; independent points; scatter plot suggests an increasing or decreasing relationship; more than 5 pairs, ideally 10–30; individuals selected randomly with equal selection chance.
Use the provided formula, state H₀ as no significant correlation, and compare the coefficient with the correct critical value for sample size and significance level.
Even a strong significant correlation may arise because a third variable affects both measurements, because selection was biased, or by coincidence. A causal claim needs an appropriate controlled design, plausible biological mechanism, temporal direction and replicated evidence beyond the coefficient.
A coefficient near 0 can occur when a real relationship is curved rather than linear/monotonic, so inspect the scatter plot. Do not use Pearson merely because both columns contain numbers, and do not convert paired observations into ranks unless Spearman's conditions apply.
The examination provides p + q = 1 and p² + 2pq + q² = 1. Here p and q are the frequencies of two alleles; p² and q² are homozygous genotype frequencies and 2pq is the heterozygous genotype frequency.
When a recessive phenotype frequency is given: 1. identify it as q²; 2. calculate q = √q²; 3. calculate p = 1 − q; 4. calculate p² and 2pq; 5. check that p + q = 1 and p² + 2pq + q² = 1. Multiply a frequency by population size only if the question asks for an expected number.
If 16% of a population shows the recessive phenotype, q² = 0.16, so q = 0.40 and p = 0.60. Expected genotype frequencies are p² = 0.36, 2pq = 0.48 and q² = 0.16; these sum to 1.00.
Do not assign q = 0.16 in this example. The equations model an equilibrium population under assumptions; a correct substitution does not show that the assumptions hold in the real population.
Lincoln index: N = (n₁ × n₂) ÷ m₂, where N is estimated population size, n₁ is the first captured and marked sample, n₂ is the total second sample, and m₂ is the marked number recaptured in the second sample.
Simpson's index in this syllabus: D = 1 − Σ(n/N)², where n is the number of individuals of each type and N is the total across all types. For each type calculate n/N, square it, sum the squared proportions, then subtract from 1. A larger D means greater diversity under this stated formula.
Lincoln: n₁ = 40, n₂ = 50 and m₂ = 10 gives N = 200. Simpson: type counts 50, 30 and 20 give N = 100 and D = 1 − (0.50² + 0.30² + 0.20²) = 1 − 0.38 = 0.62.
A small recapture m₂ inflates N and makes the estimate sensitive to one animal. Use the exact Simpson formula printed in the question/syllabus: alternative textbooks may name reciprocal or complement indices differently.
The examination provides: χ² = Σ[(O − E)²/E]; sample s = √[Σ(x − mean)²/(n − 1)]; SE = s/√n; and 95% CI ≈ mean ± (2 × SE). Use the symbols exactly as keyed in the supplied formula.
χ²: pair each observed O with its expected E, calculate every (O − E)²/E contribution, then sum.
Sample SD: find deviations from the mean, square and sum them, divide by n − 1, then square-root.
SE: divide sample SD by √n.
95% CI: calculate 2SE, subtract it from and add it to the mean to obtain lower and upper limits.
χ² measures total discrepancy between observed and expected category frequencies. SD measures spread of individual observations. SE measures precision of the estimated mean. The 95% CI expresses uncertainty around that mean. The numerical output must be followed by the relevant biological or inferential interpretation.
Do not lose the summation in χ² or SD, pair an O with the wrong E, or use n instead of n − 1 in the provided sample-SD formula. A 95% CI is not the range expected to contain 95% of individual observations.
The provided t-test formula standardises the difference between two means using sample SDs and sizes. Pearson's r measures linear correlation from paired continuous data. Spearman's rₛ = 1 − [6ΣD²/(n³ − n)] measures monotonic correlation from paired ranks, where D is each rank difference.
Degrees of freedom are not provided as formulae: for chi-squared, v = c − 1; for the syllabus two-sample t-test, v = n₁ + n₂ − 2. Pearson and Spearman critical-value tables use sample size n rather than these df formulae. Confirm whether the table asks for df or n before reading it.
State H₀, calculate the statistic or coefficient with the supplied formula, identify df or n, select the stated probability threshold, then compare the magnitude with the correct critical value. If it crosses the critical threshold, reject H₀ and state the significant difference or correlation in the biological context.
The sign of r or rₛ gives direction; significance uses the coefficient's magnitude and the table. A strong or significant correlation does not prove causation, and no formula can compensate for choosing an invalid test.
Statistical selection begins with the question: difference from expectation, difference between two means, or correlation between paired measurements. Then check data type, distribution, independence, pattern and sample size before calculating.
Chi-squared | observed versus expected nominal frequencies; syllabus questions use one row or one column of classes.
t-test | difference between two sample means; continuous data; populations normal; SDs approximately equal; valid even when each sample has fewer than 30 values.
Pearson | correlation; two continuous variables; normal population; scatter plot suggests linearity; at least 5 pairs, ideally 10+.
Spearman | correlation; ordinal/rankable or non-normal data; independent points; scatter suggests increasing/decreasing monotonic pattern; more than 5 pairs, ideally 10–30; random equal-chance selection.
Selection route: 1. name the biological question; 2. classify each variable; 3. inspect distribution and scatter pattern where relevant; 4. check independence, sample size and test-specific conditions; 5. state why the method fits; 6. state H₀; 7. calculate and interpret using the critical table.
Do not choose a test because its formula is supplied or familiar. Pearson is not a general test for any two numerical columns; a t-test is not valid for category counts; and statistical significance does not repair biased or non-independent sampling.