9. Statistics
- Syllabus
- 0580–2028–2029
- Section
- 9
- Level
- Core

Classifying data means placing each observation into a clearly defined category; tabulating data records the resulting counts in a structure that can be checked.
Choose categories that do not overlap and that cover every possible observation. Read each observation once, place it in exactly one category, and keep an 'other' category only when the question permits one.
| Table | Use | Recording rule |
|---|---|---|
| tally table | one variable has several categories | add one tally per observation, crossing every fifth tally through the previous four; write the numerical count as the frequency |
| two-way table | every observation has one category from each of two variables | place it in the single interior cell where its row and column categories meet |
For shirt colours blue, red, blue, green, red, blue, the tallies give blue 3, red 2 and green 1. The frequency column contains 3,2,1—not the tally symbols themselves. Their sum is 6, matching the six observations.
| Year group / travel | Walk | Cycle | Total |
|---|---|---|---|
| Year 1 | 6 | 4 | 10 |
| Year 2 | 7 | 3 | 10 |
| Total | 13 | 7 | 20 |
In a two-way table, each row total is the sum across its cells and each column total is the sum down its cells. The sum of all row totals and the sum of all column totals must both equal the grand total. Use subtraction from a known total to find a missing cell only after identifying the correct row or column.
A table organises and counts the data; interpreting patterns, comparing datasets, calculating averages or drawing statistical charts belongs to later topics. Never count one observation in two interior cells.
Reading a statistical display means extracting what it shows accurately; drawing an inference means combining those values into a conclusion that the display actually supports.
Before reading a value, check the title, category labels, units, scale intervals and key. Locate the exact category or interval, read from the correct mark or bar, and state the unit. Then calculate any required difference, total, fraction or percentage from those values.
| Month | Rainfall (mm) | Days with rain |
|---|---|---|
| January | 40 | 8 |
| February | 65 | 6 |
| March | 50 | 10 |
The table shows February has the greatest rainfall, while March has the most rainy days. Therefore 'the month with most rainfall also has most rainy days' is false for these data. A useful inference cites the values or pattern that justify it.
Do not replace one measured variable with another: rainfall amount and number of rainy days answer different questions. An inference should say 'for these data' unless the evidence justifies a wider claim.
A fair comparison uses the same feature, unit and statistical measure for both data sets, and comments separately on typical value and variation.
| Data set | Mean | Median | Range |
|---|---|---|---|
| A | 52 | 51 | 12 |
| B | 58 | 57 | 30 |
Set B has the higher typical value because both its mean and median are higher. Name the measure and direction: 'B has a higher median by 6' is stronger than 'B is better'.
Set A is less variable because its range is smaller: 12 compared with 30. A lower range means the observed values are packed into a narrower span, but it does not say that every A value is close to every other value.
When comparing graphs, also use like-for-like features such as the modal category, peaks, gaps or overall pattern. Two comments should describe genuinely different features rather than repeat the same comparison in new words.
Do not compare a mean from one set with a median from the other, or raw frequencies when sample sizes differ and proportions are needed. Check axes and units before comparing bar heights or plotted positions.
A conclusion is only as strong as the data and method behind it. Before generalising, identify what the display or statistical measure leaves uncertain.
| Restriction | Why it matters | Safer conclusion |
|---|---|---|
| small or unrepresentative sample | the sample may not reflect the population | limit the claim to the sampled group |
| extreme value | it can pull the mean away from most values | compare the median or inspect the data |
| one summary measure | different distributions can share the same average | compare centre and spread together |
| different scales, units or time periods | the visual comparison is not like-for-like | standardise before comparing |
| two variables change together | association alone does not prove cause | describe the association, not a cause |
For values 24,25,26,27,98, the mean is 40 but the median is 26. The single value 98 pulls the mean upward, so calling 40 a typical value would misrepresent most observations.
Use evidence-bounded language: the data 'show', 'suggest' or 'support' a pattern. State the group and time covered, and mention the specific restriction when it affects the conclusion.
Finding a limitation does not make the data useless; it sets the boundary of what can be claimed. Do not reject a conclusion without explaining which feature of the data weakens it.
Mean, median and mode describe a typical or central value; range describes spread. The calculation and the purpose of each measure are different.
| Measure | How to find it | When it is useful |
|---|---|---|
| Mean | add all values, then divide by how many values | uses every numerical value; sensitive to extreme values |
| Median | order the data and find the middle; average the two middle values if needed | gives a central value that is less affected by extremes |
| Mode | identify the most frequent value or category | shows what occurs most often; suitable for categorical data |
| Range | maximum minus minimum | measures the total spread, not a typical value |
For the ordered data 3,4,4,6,8, the total is 25. Therefore the mean is 25÷5=5, the median is 4, the mode is 4, and the range is 8−3=5. Always order a list before locating its median.
| Value x | 2 | 3 | 5 |
|---|---|---|---|
| Frequency f | 1 | 3 | 2 |
| Product fx | 2 | 9 | 10 |
For this ungrouped frequency table, ∑f=6 and ∑fx=21, so the mean is 21÷6=3.5. The ordered positions are 2,3,3,3,5,5, so the median and mode are both 3, and the range is 5−2=3. Frequencies count repeated observations; do not average the headings alone.
Choose the measure to match the question: use the median when an extreme value would distort the mean; use the mode for the most common category; use the range only to describe spread. A data set can have no mode or more than one mode. This objective covers lists and ungrouped frequency tables, not grouped-data estimates.
A statistical display represents frequencies so that categories, proportions or the shape of individual data can be compared. A correct display preserves every frequency and uses a clear scale or key.
| Display | Construction rule | What to interpret |
|---|---|---|
| Bar chart | equal-width bars, consistent gaps and a labelled linear frequency scale | compare bar heights; in dual bars compare paired groups, and in composite bars compare both parts and totals |
| Pie chart | sector angle =totalfrequency×360∘ | compare proportions; convert an angle back with 360∘angle×total |
| Pictogram | use one stated key for every symbol, including fractional symbols | multiply symbols by the key before comparing frequencies |
| Stem-and-leaf | split each value into a stem and leaf, order every row and give a key | recover the original values and read their distribution |
| Frequency distribution | list each value or category once with its frequency | check the total frequency and identify common or rare values |
For bars, choose a linear scale that reaches the largest frequency, label both axes and plot each height accurately. A dual bar chart needs a key and the same scale for both groups. A composite bar's segment heights add to its total; read a segment from the difference between its two boundaries. Unequal widths or a non-linear unmarked scale are misleading.
For frequencies 9,14,7, the total is 30. The sector angles are 9÷30×360∘=108∘, 14÷30×360∘=168∘ and 7÷30×360∘=84∘. They sum to 360∘, which checks the chart.
For 13,15,21,21,28,31, use stems 1,2,3 and ordered leaves: 1∣3 5, 2∣1 1 8, 3∣1. The key 1∣3=13 fixes the place value. Every original value must appear exactly once.
Interpret only what the display supports: read the scale and key before calculating, and compare like with like. Bars show frequencies, pie sectors show proportions, and a stem-and-leaf diagram retains individual values. This objective uses simple frequency distributions; grouped-data estimates and scatter-diagram correlation belong elsewhere.
A scatter diagram plots one point for each paired observation, so the overall relationship between two numerical variables can be seen.
Put the explanatory variable on the horizontal axis and the other variable on the vertical axis when the context makes that choice clear. Label both axes with units, use linear scales, and plot each pair as a small clear cross. For (23,31.2), move to 23 on the horizontal scale and 31.2 on the vertical scale; the point must satisfy both coordinates.
Read the whole cloud of points before describing it. State the variables and direction, for example: 'as time in the shop increases, the number of items bought tends to increase.' A point far from the overall pattern is an unusual point; identify it from its coordinates.
Do not join consecutive points. A scatter diagram shows paired observations and a general pattern, not a time sequence or proof that one variable causes the other.
Correlation describes the direction of the overall relationship between two variables in a scatter diagram.
| Type | Pattern from left to right | Meaning |
|---|---|---|
| Positive | points tend to rise | as one variable increases, the other tends to increase |
| Negative | points tend to fall | as one variable increases, the other tends to decrease |
| Zero | no clear upward or downward pattern | changes in one variable give no consistent direction for the other |
If higher temperature is generally paired with more ice creams sold, the correlation is positive. If greater car power is generally paired with a shorter acceleration time, it is negative. Judge the trend of the whole cloud, not the slope between two selected points.
Correlation can be weak or affected by unusual points, and a very small set may not reveal a reliable pattern. Correlation describes association; it does not by itself prove causation.
A line of best fit is a single straight ruled line drawn by inspection to represent the central trend of a scatter diagram.
First identify the direction and centre of the point cloud. Draw one ruled line across the full data set, with roughly even numbers of points above and below it over its whole length. The line need not pass through any point or through the origin, and an isolated unusual point should not pull it away from the main pattern.
To estimate y from a given x, start at the x-axis value, move vertically to the line, then move horizontally to the y-axis and read the scale. Reverse these moves to estimate x from y. An estimate is approximate, so report a precision justified by the scale.
If the best-fit line for two test scores crosses near (40,48), a test 1 score of 40 gives an estimated test 2 score of about 48. The estimate comes from the line, not from choosing the nearest plotted student.
Estimates within the horizontal range of the data are interpolation. Extending beyond that range is extrapolation and is less reliable because the observed pattern may not continue.