A.1 Handling data
- Syllabus
- 2021
- Topic
- —
- Level
- A2
Significant figures communicate the precision supported by a measurement or calculation. Zeros between non-zero digits count; leading zeros only locate the decimal point.
Keep guard digits during working, then round the final result to match the least precise relevant measurement. State a value such as 2.40 to show its precision.
A mean based on readings recorded to 0.1 s should not be reported as 2.437891 s; 2.4 s or 2.44 s may be appropriate depending on the data.
More digits do not make a result more accurate, and exact counted quantities are not limited by instrument precision.
The arithmetic mean is the total of the observations divided by their count. It summarises repeated measurements when they represent the same quantity and are on the same scale.
Inspect the raw values first, calculate the mean with full precision and report spread or anomalous values rather than hiding them.
Rates 9, 10 and 11 units min⁻¹ have mean 10; if one reading was taken at a different temperature, combining it may be misleading.
A mean is sensitive to outliers and does not prove that the system is stable or normally distributed.
Tables keep exact values visible; bar charts compare categories; histograms show how continuous measurements are distributed; diagrams clarify structure or process.
Label axes and units, choose equal scales, include a key when needed and make the display match the variable type. Keep raw data available behind any summary.
Use a histogram for a distribution of cell diameters, but a bar chart for separate treatment groups; joining category bars can falsely imply continuity.
A polished graph cannot repair missing units, selective data or a misleading scale.
Probability is a number from 0 to 1 for an event under stated assumptions. It can also be expressed as a fraction, percentage or '1 in n'. Define the event and denominator before calculating.
For mutually exclusive alternatives, add probabilities. For independent events that must both occur, multiply them. Independence must come from the biological model—for example, separate meioses or independently assorting genes—not from convenience.
In an Aa×Aa cross, the probability of an aa child is 1/4. Each birth is a new independent event, so two aa children in succession have probability (1/4)(1/4)=1/16. For independently assorting AaBb×AaBb, the probability of aabb is (1/4)(1/4)=1/16.
Observed frequency estimates probability: 7 cases among 800 000 people gives 7/800000=8.75×10−6, about 0.000875% or 1 in 114 000. Sampling variation means the observed proportion is not an exact law.
Probability does not predict which individual outcome must occur. Do not multiply events that are dependent or add alternatives that can overlap without correcting for the overlap.
A sample supports inference only for the population represented by its sampling frame. Define the habitat, organism, spatial/temporal boundary and inclusion rule before choosing random, systematic or stratified sampling.
| Method | Best use | Main protection and limit |
|---|---|---|
| Random | Estimate a relatively uniform habitat without deliberate site choice | Random coordinates reduce selection bias; rare zones may be missed |
| Systematic | Detect change along a gradient | Fixed intervals are reproducible; one transect may not represent the habitat |
| Stratified | Habitat contains known subareas | Sample each stratum in proportion or by justified allocation; requires a valid map |
From sampled species counts, D=∑n(n−1)N(N−1), where N is total individuals and n is each species count. Use the same effort, method and identification rules when comparing habitats; a larger D indicates greater diversity under this index.
Increase independent sampling units, distribute them across time/space where appropriate and report uncertainty. A small sample of 23 chicks with an 8:15 sex ratio may not represent the adult population or a stable 1:1 population ratio.
A large convenience sample can remain biased. The diversity index combines richness and evenness but does not identify why habitats differ or correct misidentification.
The mean uses every value, the median is the middle after ordering, and the mode is the most frequent value. The best summary depends on the measurement and its distribution.
Use the median when an outlier would distort the mean; use the mode for common categories or repeated discrete values. Report the raw context and sample size.
Income-like values 2, 2, 3, 3 and 20 have mean 6 but median 3, so the median better represents a typical observation.
No summary statistic tells you the spread or the mechanism producing the data.
A scatter diagram pairs two measured variables. Direction, strength and form of the pattern describe association; they do not by themselves identify a causal mechanism.
Plot the independent variable consistently, inspect outliers and restricted ranges, and use a correlation measure only with the assumptions and data type it requires.
Light intensity and photosynthetic rate may rise together before a plateau; the pattern suggests a relationship, while temperature or CO₂ could also influence both.
No correlation does not prove no biological relationship, and correlation never rules out confounding variables.
Magnification is a dimensionless ratio: M=actual sizeimage size. Convert image and actual size to the same units before substituting, then rearrange as actual size=image size/M or image size=M×actual size.
A leaf section measures 50 mm in a photograph at ×100. Its actual thickness is 50/100=0.50 mm =500 µm =5.0×102 µm. The unit conversion happens after or before division only if it is applied consistently.
Write sizes in standard form to compare scale. 2×10−5 m is one order of magnitude larger than 2×10−6 m because their powers differ by 1; a power difference of 3 means a thousandfold difference. Coefficients near a power boundary must be considered when rounding to the nearest order.
A measured scale bar can recover actual size even if an image is resized, provided the object and bar were resized together. A printed magnification label may become invalid after resizing.
Magnification enlarges an image but does not guarantee resolution. Do not compare powers of ten until both quantities use the same base unit.
Start with a null hypothesis and choose the test before looking for a desirable result. Match the biological question, variable type, pairing/independence and test assumptions; a statistical result addresses chance variation under the null, not biological importance or causation.
| Question and data | Test | Null hypothesis |
|---|---|---|
| Do observed categorical counts differ from expected counts? | Chi-squared, χ2=∑(O−E)2/E | Observed and expected frequencies do not differ beyond chance |
| Do two sample means differ? | Student's t-test with the required paired/independent form | The population means do not differ |
| Are two ranked/continuous variables associated? | Appropriate correlation coefficient, often Spearman's rank | There is no association/correlation |
Calculate the statistic with unrounded data, determine degrees of freedom where required and compare its magnitude with the critical value at the stated significance level. If the result exceeds the relevant critical threshold, reject the null; otherwise do not reject it. Use the course-provided table and state the biological conclusion.
Chi-squared uses frequency counts with suitable expected values; t-tests require quantitative samples and the specified distribution/design assumptions; correlation requires paired observations and tests association, with Spearman based on ranks.
Failing to reject is not proof that the null is true. Statistical significance does not establish causation, effect size or practical importance, and changing the test after seeing the result inflates error risk.
Dispersion describes how measurements vary around their centre. Range is maximum minus minimum and depends strongly on two extreme values. Sample standard deviation uses every deviation from the mean and is reported in the original measurement unit.
For n sample values, s=n−1∑(x−xˉ)2. Calculate the mean, subtract it from each value, square deviations, sum, divide by n−1, then take the square root. Retain calculator precision until the final step.
For 14, 16 and 18 g, xˉ=16 g, range =18−14=4 g, and s=(4+0+4)/2=2.0 g. The SD means observations typically vary around the mean; it is not an error bar for the mean unless explicitly defined that way.
SD is usually more informative than range because it uses all values, but it can still be affected by an outlier. Compare groups using both centre and spread and consider sample size and measurement scale.
Do not omit the square root or use n when the required equation is sample SD with n−1. Standard deviation describes variation among observations, not its biological cause.
Record a measured value with absolute uncertainty in the same unit, such as 10.0±0.1 cm. Estimate uncertainty from instrument resolution, repeated readings and the measurement procedure; do not invent extra precision from a calculator.
Percentage uncertainty is measured valueabsolute uncertainty×100. Percentage error compares a measured value with an accepted/reference value: accepted∣measured−accepted∣×100. They answer different questions.
| Calculation | Simple uncertainty rule |
|---|---|
| Addition or subtraction | Add absolute uncertainties |
| Multiplication or division | Add percentage uncertainties |
| Quantity raised to a power p | Multiply percentage uncertainty by ∣p∣ |
A diameter 2.0±0.1 mm has 5% uncertainty. For area proportional to d2, the simple propagated percentage uncertainty is about 2×5=10%. Report the area to precision consistent with that uncertainty.
Uncertainty is not necessarily a mistake and percentage error requires a defensible reference value. Combining many precise digits does not make the original measurements more accurate.