2 Handling data
- Syllabus
- 2024
- Topic
- 2
- Level
- —
Significant figures count meaningful digits from the first non-zero digit. A final biological result should not imply finer precision than the measurements used to calculate it.
| Value | Significant figures and reason |
|---|---|
| 0.00450 | 3; leading zeros locate the decimal point, while the final zero is significant |
| 1200 | ambiguous without context; 1.2×103 states 2 significant figures |
| 12.46 to 3 s.f. | 12.5 because the next digit is 6, so 4 rounds up |
| 0.06784 to 2 s.f. | 0.068 because counting starts at 6 and the next digit rounds 7 up |
Keep extra digits during intermediate steps and round once at the end. Match the requested significant figures and retain the unit.
Decimal places count positions after the decimal point; significant figures count meaningful digits. For example, 0.0125 has four decimal places but three significant figures.
The arithmetic mean shares the total of all included observations equally across their number. It summarises the centre of repeated numerical data but does not show their spread.
mean=sumofincludedvalues/numberofincludedvalues
For gelatine volumes 0.55, 0.54 and 0.61 cm³, the mean is (0.55+0.54+0.61)/3=0.5666… cm³, reported as 0.57 cm³ to two decimal places.
Count only the values actually included. Do not discard an anomalous result merely because it changes the mean; exclusion needs a recorded evidence-based reason, and the reported mean should make that decision clear.
A bar chart compares numerical values for distinct categories. Because the categories are separate rather than continuous, the bars are separated by gaps.
| Construct | Interpret |
|---|---|
| place categories on the horizontal axis | identify the named group represented by each bar |
| label the vertical axis with variable and unit | read values from the scale, not from apparent bar area |
| start at zero unless a clearly shown break is justified | compare heights using the same scale |
| use equal bar widths and gaps | state the largest, smallest or difference with values |
If unlikely and likely groups have LH concentrations of 5 and 45 arbitrary units, the likely category is 40 units higher and nine times the unlikely value.
Do not join category bars into a continuous shape. A histogram looks similar but represents continuous grouped intervals, so its bars touch and may use frequency density.
A frequency table records how many observations fall in each value or class interval. A histogram displays grouped continuous data with touching bars whose areas represent frequencies.
| Data move | Rule |
|---|---|
| define classes | make intervals non-overlapping and cover every possible value |
| tally observations | place each observation in exactly one class |
| record frequency | count the tally in each class and check the total equals the sample size |
| draw equal-width histogram classes | bar height may be frequency because equal widths preserve area comparisons |
| draw unequal-width classes | use frequency density = frequency ÷ class width, so bar area equals frequency |
Bar height alone does not represent frequency when class widths differ. Histograms are for continuous grouped measurements; ordinary bar charts are for separate categories.
Sampling measures a manageable subset to infer properties of a larger population. The sample must be selected without systematic bias and be large and repeated enough to capture natural variation.
| Situation | Defensible sampling design |
|---|---|
| organisms across an area | overlay a grid, choose coordinates randomly, use equal-sized quadrats and count consistently |
| change along a gradient | place quadrats at fixed intervals along a transect |
| mobile organisms | use a defined trapping method for equal times and avoid counting the same individual twice where possible |
| microscopic density | count in known equal areas, repeat across randomly selected fields of view, find mean density and scale to total area |
Two stomata in 0.4 mm×0.4 mm=0.0016 cm² give a density of 2/0.0016=1250 stomata cm⁻². Applied to 150 cm², the estimate is 187500 stomata.
A large convenient sample can still be biased. More repeats improve representation only when locations, times or individuals are selected by a method appropriate to the population and question.
Probability measures how likely an outcome is, from 0 for impossible to 1 for certain. For equally likely outcomes, it is the number of favourable outcomes divided by the total number of possible outcomes.
P(event)=favourableoutcomes/totalpossibleoutcomes
For a heterozygous monohybrid cross Aa×Aa, one of four equally likely genotype outcomes is aa, so P(aa)=1/4=0.25=25%. For two independent events, multiply their probabilities.
A probability predicts a long-run proportion, not an exact result in a small family or sample. Observed frequencies may differ by chance, and events must be independent before their probabilities are multiplied.
The median is the middle value after numerical data are ordered; the mode is the value or category that occurs most often. They answer different questions about a dataset.
| Measure | Method | Useful when |
|---|---|---|
| median | order values; choose the middle, or average the two middle values when the count is even | extreme values would pull the mean away from a typical central value |
| mode | count occurrences and select the most frequent value or category | the most common outcome matters, including non-numerical categories |
For 6, 8, 8, 9, 16, the median is 8 and the mode is 8. For 6, 8, 9, 16, the median is (8+9)/2=8.5 and there is no mode because no value repeats.
The median cannot be found from unordered positions, and a dataset may have no mode or more than one mode. Do not call the largest value the mode unless it is also most frequent.
A scatter diagram plots paired measurements for two numerical variables. The overall point pattern can show positive correlation, negative correlation or no clear correlation.
| Pattern | Interpretation |
|---|---|
| points rise from left to right | positive correlation: larger values of one variable tend to accompany larger values of the other |
| points fall from left to right | negative correlation |
| points show no direction | no clear correlation in the observed range |
| points lie close to a trend | stronger correlation than a widely scattered pattern |
| one point lies far from the pattern | possible anomaly; check it without deleting it automatically |
Plot the independent variable on the horizontal axis and its paired dependent value vertically. A best-fit line or curve follows the overall pattern rather than joining each point.
Correlation alone does not prove that one variable causes the other. A third variable, reverse causation or chance may explain the pattern, and extrapolation beyond the observed range is uncertain.
An order of magnitude is a factor of ten. Expressing a quantity in standard form reveals its scale and allows rapid comparisons between very large or very small biological values.
| Move | Example |
|---|---|
| express each value in standard form | a bacterium 2×10−6 m; a cell 2×10−5 m |
| compare powers of ten | exponents differ by −5−(−6)=1 |
| convert exponent difference to a factor | 101=10, so the cell is about ten times longer |
| estimate a product or quotient | combine rounded coefficients and add or subtract exponents |
A difference of two orders of magnitude means a factor of 102=100, not a difference of 2. Coefficients near a power boundary can affect the nearest order, so keep them visible until the comparison is justified.