9. Statistics

Syllabus
0580–2028–2029
Section
9
Level
Core

C9.1 Classifying statistical data

Syllabus
0580–2028–2029
Topic
C9.1
Level
Core

Classify data in tally and two-way tables

Classifying data means placing each observation into a clearly defined category; tabulating data records the resulting counts in a structure that can be checked.

Choose categories that do not overlap and that cover every possible observation. Read each observation once, place it in exactly one category, and keep an 'other' category only when the question permits one.

Table Use Recording rule
tally table one variable has several categories add one tally per observation, crossing every fifth tally through the previous four; write the numerical count as the frequency
two-way table every observation has one category from each of two variables place it in the single interior cell where its row and column categories meet

For shirt colours blue, red, blue, green, red, blue, the tallies give blue 3, red 2 and green 1. The frequency column contains 3,2,13,2,1—not the tally symbols themselves. Their sum is 6, matching the six observations.

Year group / travel Walk Cycle Total
Year 1 6 4 10
Year 2 7 3 10
Total 13 7 20

In a two-way table, each row total is the sum across its cells and each column total is the sum down its cells. The sum of all row totals and the sum of all column totals must both equal the grand total. Use subtraction from a known total to find a missing cell only after identifying the correct row or column.

A table organises and counts the data; interpreting patterns, comparing datasets, calculating averages or drawing statistical charts belongs to later topics. Never count one observation in two interior cells.

C9.2 Interpreting statistical data

Syllabus
0580–2028–2029
Topic
C9.2
Level
Core

Read statistical displays and make supported inferences

Reading a statistical display means extracting what it shows accurately; drawing an inference means combining those values into a conclusion that the display actually supports.

Before reading a value, check the title, category labels, units, scale intervals and key. Locate the exact category or interval, read from the correct mark or bar, and state the unit. Then calculate any required difference, total, fraction or percentage from those values.

Month Rainfall (mm) Days with rain
January 40 8
February 65 6
March 50 10

The table shows February has the greatest rainfall, while March has the most rainy days. Therefore 'the month with most rainfall also has most rainy days' is false for these data. A useful inference cites the values or pattern that justify it.

Do not replace one measured variable with another: rainfall amount and number of rainy days answer different questions. An inference should say 'for these data' unless the evidence justifies a wider claim.

Compare data sets using centre and spread

A fair comparison uses the same feature, unit and statistical measure for both data sets, and comments separately on typical value and variation.

Data set Mean Median Range
A 52 51 12
B 58 57 30

Set B has the higher typical value because both its mean and median are higher. Name the measure and direction: 'B has a higher median by 6' is stronger than 'B is better'.

Set A is less variable because its range is smaller: 1212 compared with 3030. A lower range means the observed values are packed into a narrower span, but it does not say that every A value is close to every other value.

When comparing graphs, also use like-for-like features such as the modal category, peaks, gaps or overall pattern. Two comments should describe genuinely different features rather than repeat the same comparison in new words.

Do not compare a mean from one set with a median from the other, or raw frequencies when sample sizes differ and proportions are needed. Check axes and units before comparing bar heights or plotted positions.

Judge the limits of conclusions from data

A conclusion is only as strong as the data and method behind it. Before generalising, identify what the display or statistical measure leaves uncertain.

Restriction Why it matters Safer conclusion
small or unrepresentative sample the sample may not reflect the population limit the claim to the sampled group
extreme value it can pull the mean away from most values compare the median or inspect the data
one summary measure different distributions can share the same average compare centre and spread together
different scales, units or time periods the visual comparison is not like-for-like standardise before comparing
two variables change together association alone does not prove cause describe the association, not a cause

For values 24,25,26,27,9824,25,26,27,98, the mean is 4040 but the median is 2626. The single value 9898 pulls the mean upward, so calling 4040 a typical value would misrepresent most observations.

Use evidence-bounded language: the data 'show', 'suggest' or 'support' a pattern. State the group and time covered, and mention the specific restriction when it affects the conclusion.

Finding a limitation does not make the data useless; it sets the boundary of what can be claimed. Do not reject a conclusion without explaining which feature of the data weakens it.

C9.3 Averages and range

Syllabus
0580–2028–2029
Topic
C9.3
Level
Core

Calculate and choose mean, median, mode and range

Mean, median and mode describe a typical or central value; range describes spread. The calculation and the purpose of each measure are different.

Measure How to find it When it is useful
Mean add all values, then divide by how many values uses every numerical value; sensitive to extreme values
Median order the data and find the middle; average the two middle values if needed gives a central value that is less affected by extremes
Mode identify the most frequent value or category shows what occurs most often; suitable for categorical data
Range maximum minus minimum measures the total spread, not a typical value

For the ordered data 3,4,4,6,83,4,4,6,8, the total is 2525. Therefore the mean is 25÷5=525\div5=5, the median is 44, the mode is 44, and the range is 8−3=58-3=5. Always order a list before locating its median.

Value xx 2 3 5
Frequency ff 1 3 2
Product fxfx 2 9 10

For this ungrouped frequency table, ∑f=6\sum f=6 and ∑fx=21\sum fx=21, so the mean is 21÷6=3.521\div6=3.5. The ordered positions are 2,3,3,3,5,52,3,3,3,5,5, so the median and mode are both 33, and the range is 5−2=35-2=3. Frequencies count repeated observations; do not average the headings alone.

Choose the measure to match the question: use the median when an extreme value would distort the mean; use the mode for the most common category; use the range only to describe spread. A data set can have no mode or more than one mode. This objective covers lists and ungrouped frequency tables, not grouped-data estimates.

C9.4 Statistical charts and diagrams

Syllabus
0580–2028–2029
Topic
C9.4
Level
Core

Draw and interpret statistical charts and diagrams

A statistical display represents frequencies so that categories, proportions or the shape of individual data can be compared. A correct display preserves every frequency and uses a clear scale or key.

Display Construction rule What to interpret
Bar chart equal-width bars, consistent gaps and a labelled linear frequency scale compare bar heights; in dual bars compare paired groups, and in composite bars compare both parts and totals
Pie chart sector angle =frequencytotal×360∘=\dfrac{\text{frequency}}{\text{total}}\times360^\circ compare proportions; convert an angle back with angle360∘×total\dfrac{\text{angle}}{360^\circ}\times\text{total}
Pictogram use one stated key for every symbol, including fractional symbols multiply symbols by the key before comparing frequencies
Stem-and-leaf split each value into a stem and leaf, order every row and give a key recover the original values and read their distribution
Frequency distribution list each value or category once with its frequency check the total frequency and identify common or rare values

For bars, choose a linear scale that reaches the largest frequency, label both axes and plot each height accurately. A dual bar chart needs a key and the same scale for both groups. A composite bar's segment heights add to its total; read a segment from the difference between its two boundaries. Unequal widths or a non-linear unmarked scale are misleading.

For frequencies 9,14,79,14,7, the total is 3030. The sector angles are 9÷30×360∘=108∘9\div30\times360^\circ=108^\circ, 14÷30×360∘=168∘14\div30\times360^\circ=168^\circ and 7÷30×360∘=84∘7\div30\times360^\circ=84^\circ. They sum to 360∘360^\circ, which checks the chart.

For 13,15,21,21,28,3113,15,21,21,28,31, use stems 1,2,31,2,3 and ordered leaves: 1∣3 51\mid3\ 5, 2∣1 1 82\mid1\ 1\ 8, 3∣13\mid1. The key 1∣3=131\mid3=13 fixes the place value. Every original value must appear exactly once.

Interpret only what the display supports: read the scale and key before calculating, and compare like with like. Bars show frequencies, pie sectors show proportions, and a stem-and-leaf diagram retains individual values. This objective uses simple frequency distributions; grouped-data estimates and scatter-diagram correlation belong elsewhere.

C9.5 Scatter diagrams

Syllabus
0580–2028–2029
Topic
C9.5
Level
Core

Plot and interpret a scatter diagram

A scatter diagram plots one point for each paired observation, so the overall relationship between two numerical variables can be seen.

Put the explanatory variable on the horizontal axis and the other variable on the vertical axis when the context makes that choice clear. Label both axes with units, use linear scales, and plot each pair as a small clear cross. For (23,31.2)(23,31.2), move to 2323 on the horizontal scale and 31.231.2 on the vertical scale; the point must satisfy both coordinates.

Read the whole cloud of points before describing it. State the variables and direction, for example: 'as time in the shop increases, the number of items bought tends to increase.' A point far from the overall pattern is an unusual point; identify it from its coordinates.

Do not join consecutive points. A scatter diagram shows paired observations and a general pattern, not a time sequence or proof that one variable causes the other.

Recognise positive, negative and zero correlation

Correlation describes the direction of the overall relationship between two variables in a scatter diagram.

Type Pattern from left to right Meaning
Positive points tend to rise as one variable increases, the other tends to increase
Negative points tend to fall as one variable increases, the other tends to decrease
Zero no clear upward or downward pattern changes in one variable give no consistent direction for the other

If higher temperature is generally paired with more ice creams sold, the correlation is positive. If greater car power is generally paired with a shorter acceleration time, it is negative. Judge the trend of the whole cloud, not the slope between two selected points.

Correlation can be weak or affected by unusual points, and a very small set may not reveal a reliable pattern. Correlation describes association; it does not by itself prove causation.

Draw and use a straight line of best fit

A line of best fit is a single straight ruled line drawn by inspection to represent the central trend of a scatter diagram.

First identify the direction and centre of the point cloud. Draw one ruled line across the full data set, with roughly even numbers of points above and below it over its whole length. The line need not pass through any point or through the origin, and an isolated unusual point should not pull it away from the main pattern.

To estimate yy from a given xx, start at the xx-axis value, move vertically to the line, then move horizontally to the yy-axis and read the scale. Reverse these moves to estimate xx from yy. An estimate is approximate, so report a precision justified by the scale.

If the best-fit line for two test scores crosses near (40,48)(40,48), a test 1 score of 4040 gives an estimated test 2 score of about 4848. The estimate comes from the line, not from choosing the nearest plotted student.

Estimates within the horizontal range of the data are interpolation. Extending beyond that range is extrapolation and is less reliable because the observed pattern may not continue.