A.1 Handling data

Syllabus
2021
Topic
Level
A2

Learning objectives

A.1.1—Significant figuresUse an appropriate number of significant figures; report calculated results consistently with the precision of the raw data and the least accurate measurement.A.1.2—Arithmetic meansCalculate arithmetic means from biological data, such as mean stomatal counts.A.1.3—Tables, diagrams, bar charts and histogramsConstruct and interpret frequency tables and diagrams, bar charts and histograms; use clear headings, units and consistent decimal places; select suitable formats and interpret biological tables and graphs such as enzyme-activity graphs and ECG traces.A.1.4—Simple probabilityUnderstand simple probability and use probability and chance appropriately, including in genetic inheritance.A.1.5—Sampling scientific dataUnderstand sampling principles for scientific data, including analysing randomly collected data and calculating an index of diversity to compare habitats.A.1.6—Mean, median and modeUnderstand and calculate or compare the mean, median and mode of biological datasets.A.1.7—Scatter diagrams and correlationUse scatter diagrams to identify correlations between variables, for example between lifestyle factors and health.A.1.8—Order-of-magnitude calculationsMake order-of-magnitude calculations, including manipulating magnification = image size ÷ real-object size.A.1.9—Statistical testsSelect and use statistical tests, including chi-squared tests for observed versus expected results, Student's t-tests and correlation coefficients.A.1.10—Dispersion, standard deviation and rangeUnderstand measures of dispersion, including range and standard deviation; calculate standard deviation and judge when it is more useful, including when data contain an outlier.A.1.11—Measurement uncertaintyIdentify measurement uncertainties and use simple techniques to determine uncertainty when data are combined, including calculating percentage error.

Report significant figures without inventing precision

Significant figures communicate the precision supported by a measurement or calculation. Zeros between non-zero digits count; leading zeros only locate the decimal point.

Keep guard digits during working, then round the final result to match the least precise relevant measurement. State a value such as 2.40 to show its precision.

A mean based on readings recorded to 0.1 s should not be reported as 2.437891 s; 2.4 s or 2.44 s may be appropriate depending on the data.

More digits do not make a result more accurate, and exact counted quantities are not limited by instrument precision.

Use an arithmetic mean only when the data can be combined

The arithmetic mean is the total of the observations divided by their count. It summarises repeated measurements when they represent the same quantity and are on the same scale.

Inspect the raw values first, calculate the mean with full precision and report spread or anomalous values rather than hiding them.

Rates 9, 10 and 11 units min⁻¹ have mean 10; if one reading was taken at a different temperature, combining it may be misleading.

A mean is sensitive to outliers and does not prove that the system is stable or normally distributed.

Choose a display that preserves the pattern in biological data

Tables keep exact values visible; bar charts compare categories; histograms show how continuous measurements are distributed; diagrams clarify structure or process.

Label axes and units, choose equal scales, include a key when needed and make the display match the variable type. Keep raw data available behind any summary.

Use a histogram for a distribution of cell diameters, but a bar chart for separate treatment groups; joining category bars can falsely imply continuity.

A polished graph cannot repair missing units, selective data or a misleading scale.

Use probability in inheritance and biological risk

Probability is a number from 0 to 1 for an event under stated assumptions. It can also be expressed as a fraction, percentage or '1 in nn'. Define the event and denominator before calculating.

For mutually exclusive alternatives, add probabilities. For independent events that must both occur, multiply them. Independence must come from the biological model—for example, separate meioses or independently assorting genes—not from convenience.

In an Aa×AaAa\times Aa cross, the probability of an aaaa child is 1/41/4. Each birth is a new independent event, so two aaaa children in succession have probability (1/4)(1/4)=1/16(1/4)(1/4)=1/16. For independently assorting AaBb×AaBbAaBb\times AaBb, the probability of aabbaabb is (1/4)(1/4)=1/16(1/4)(1/4)=1/16.

Observed frequency estimates probability: 7 cases among 800 000 people gives 7/800000=8.75×1067/800000=8.75\times10^{-6}, about 0.000875%0.000875\% or 1 in 114 000. Sampling variation means the observed proportion is not an exact law.

Probability does not predict which individual outcome must occur. Do not multiply events that are dependent or add alternatives that can overlap without correcting for the overlap.

Link representative sampling to valid biodiversity estimates

A sample supports inference only for the population represented by its sampling frame. Define the habitat, organism, spatial/temporal boundary and inclusion rule before choosing random, systematic or stratified sampling.

Method Best use Main protection and limit
Random Estimate a relatively uniform habitat without deliberate site choice Random coordinates reduce selection bias; rare zones may be missed
Systematic Detect change along a gradient Fixed intervals are reproducible; one transect may not represent the habitat
Stratified Habitat contains known subareas Sample each stratum in proportion or by justified allocation; requires a valid map

From sampled species counts, D=N(N1)n(n1)D=\dfrac{N(N-1)}{\sum n(n-1)}, where NN is total individuals and nn is each species count. Use the same effort, method and identification rules when comparing habitats; a larger DD indicates greater diversity under this index.

Increase independent sampling units, distribute them across time/space where appropriate and report uncertainty. A small sample of 23 chicks with an 8:15 sex ratio may not represent the adult population or a stable 1:1 population ratio.

A large convenience sample can remain biased. The diversity index combines richness and evenness but does not identify why habitats differ or correct misidentification.

Choose mean, median or mode for the shape of the data

The mean uses every value, the median is the middle after ordering, and the mode is the most frequent value. The best summary depends on the measurement and its distribution.

Use the median when an outlier would distort the mean; use the mode for common categories or repeated discrete values. Report the raw context and sample size.

Income-like values 2, 2, 3, 3 and 20 have mean 6 but median 3, so the median better represents a typical observation.

No summary statistic tells you the spread or the mechanism producing the data.

Use a scatter diagram to inspect association, not causation

A scatter diagram pairs two measured variables. Direction, strength and form of the pattern describe association; they do not by themselves identify a causal mechanism.

Plot the independent variable consistently, inspect outliers and restricted ranges, and use a correlation measure only with the assumptions and data type it requires.

Light intensity and photosynthetic rate may rise together before a plateau; the pattern suggests a relationship, while temperature or CO₂ could also influence both.

No correlation does not prove no biological relationship, and correlation never rules out confounding variables.

Use magnification and orders of magnitude to recover real size

Magnification is a dimensionless ratio: M=image sizeactual sizeM=\dfrac{\text{image size}}{\text{actual size}}. Convert image and actual size to the same units before substituting, then rearrange as actual size=image size/M\text{actual size}=\text{image size}/M or image size=M×actual size\text{image size}=M\times\text{actual size}.

A leaf section measures 50 mm in a photograph at ×100\times100. Its actual thickness is 50/100=0.5050/100=0.50 mm =500=500 µm =5.0×102=5.0\times10^2 µm. The unit conversion happens after or before division only if it is applied consistently.

Write sizes in standard form to compare scale. 2×1052\times10^{-5} m is one order of magnitude larger than 2×1062\times10^{-6} m because their powers differ by 1; a power difference of 3 means a thousandfold difference. Coefficients near a power boundary must be considered when rounding to the nearest order.

A measured scale bar can recover actual size even if an image is resized, provided the object and bar were resized together. A printed magnification label may become invalid after resizing.

Magnification enlarges an image but does not guarantee resolution. Do not compare powers of ten until both quantities use the same base unit.

Select and interpret chi-squared, t and correlation tests

Start with a null hypothesis and choose the test before looking for a desirable result. Match the biological question, variable type, pairing/independence and test assumptions; a statistical result addresses chance variation under the null, not biological importance or causation.

Question and data Test Null hypothesis
Do observed categorical counts differ from expected counts? Chi-squared, χ2=(OE)2/E\chi^2=\sum (O-E)^2/E Observed and expected frequencies do not differ beyond chance
Do two sample means differ? Student's t-test with the required paired/independent form The population means do not differ
Are two ranked/continuous variables associated? Appropriate correlation coefficient, often Spearman's rank There is no association/correlation

Calculate the statistic with unrounded data, determine degrees of freedom where required and compare its magnitude with the critical value at the stated significance level. If the result exceeds the relevant critical threshold, reject the null; otherwise do not reject it. Use the course-provided table and state the biological conclusion.

Chi-squared uses frequency counts with suitable expected values; t-tests require quantitative samples and the specified distribution/design assumptions; correlation requires paired observations and tests association, with Spearman based on ranks.

Failing to reject is not proof that the null is true. Statistical significance does not establish causation, effect size or practical importance, and changing the test after seeing the result inflates error risk.

Calculate range and sample standard deviation

Dispersion describes how measurements vary around their centre. Range is maximum minus minimum and depends strongly on two extreme values. Sample standard deviation uses every deviation from the mean and is reported in the original measurement unit.

For nn sample values, s=(xxˉ)2n1s=\sqrt{\dfrac{\sum(x-\bar{x})^2}{n-1}}. Calculate the mean, subtract it from each value, square deviations, sum, divide by n1n-1, then take the square root. Retain calculator precision until the final step.

For 14, 16 and 18 g, xˉ=16\bar{x}=16 g, range =1814=4=18-14=4 g, and s=(4+0+4)/2=2.0s=\sqrt{(4+0+4)/2}=2.0 g. The SD means observations typically vary around the mean; it is not an error bar for the mean unless explicitly defined that way.

SD is usually more informative than range because it uses all values, but it can still be affected by an outlier. Compare groups using both centre and spread and consider sample size and measurement scale.

Do not omit the square root or use nn when the required equation is sample SD with n1n-1. Standard deviation describes variation among observations, not its biological cause.

Carry measurement uncertainty through calculations

Record a measured value with absolute uncertainty in the same unit, such as 10.0±0.110.0\pm0.1 cm. Estimate uncertainty from instrument resolution, repeated readings and the measurement procedure; do not invent extra precision from a calculator.

Percentage uncertainty is absolute uncertaintymeasured value×100\dfrac{\text{absolute uncertainty}}{\text{measured value}}\times100. Percentage error compares a measured value with an accepted/reference value: measuredacceptedaccepted×100\dfrac{|\text{measured}-\text{accepted}|}{\text{accepted}}\times100. They answer different questions.

Calculation Simple uncertainty rule
Addition or subtraction Add absolute uncertainties
Multiplication or division Add percentage uncertainties
Quantity raised to a power pp Multiply percentage uncertainty by p|p|

A diameter 2.0±0.12.0\pm0.1 mm has 5% uncertainty. For area proportional to d2d^2, the simple propagated percentage uncertainty is about 2×5=10%2\times5=10\%. Report the area to precision consistent with that uncertainty.

Uncertainty is not necessarily a mistake and percentage error requires a defensible reference value. Combining many precise digits does not make the original measurements more accurate.