9. Statistics
- Syllabus
- 0580–2028–2029
- Section
- 9
- Level
- Extended
Classifying data means placing every observation into exactly one suitable category, then tabulating the counts so they can be checked and compared.
Choose categories that do not overlap and together include every possible observation. Record each observation once: omission and double-counting both change the frequencies.
| Table | Use | Recording rule |
|---|---|---|
| Tally table | one variable has several categories | add one tally per observation; bundle the fifth tally across the previous four, then write the numerical frequency |
| Two-way table | each observation has one category from each of two variables | place it in the single cell where its row and column categories meet |
For shirt colours blue, red, blue, green, red, blue, the frequencies are blue 3, red 2, green 1. Their sum is 6, matching the six observations; tally marks are the counting record, not the final numerical frequencies.
| Part-time | Full-time | Total | |
|---|---|---|---|
| Lecturer | 8 | 12 | 20 |
| Professor | 3 | 7 | 10 |
| Total | 11 | 19 | 30 |
Add across each row and down each column. Both sets of totals must give the same grand total. Find a missing cell by subtracting the other known cells from its correct row or column total, not from an unrelated total.
Reading a statistical display means extracting values accurately; an inference combines those values into a conclusion that the display supports.
Check the title, variables, units, axis intervals and key before reading. Locate the relevant value, quartile, interval or category, then calculate any required total, difference, fraction, percentage or estimate. State the unit and show which values support the conclusion.
| Group | Lower quartile | Median | Upper quartile | Number in group |
|---|---|---|---|---|
| A | 145 | 162 | 178 | 136 |
| B | 138 | 154 | 170 | 144 |
For Group A, half the values are above the median 162. Therefore an estimate of the number above 162 is 0.5×136=68. This is an estimate because a box plot summarises positions rather than listing every value.
Do not infer about a different variable or a wider population than the display describes. An evidence-backed statement names the statistic or pattern used and is qualified as an estimate when the diagram cannot give an exact count.
A complete comparison describes both a typical value and the variation, using the same statistical measures and units for both data sets.
| Data set | Median | Interquartile range | Range |
|---|---|---|---|
| P | 56 | 18 | 42 |
| Q | 49 | 35 | 61 |
Set P has the higher typical value because its median is 56, compared with 49 for Q. Name the statistic and direction; 'P is better' is not a statistical comparison.
Set P is more consistent because its interquartile range is smaller: 18 compared with 35. The IQR compares the spread of the middle half; the range compares the full span and is more affected by extreme values.
Make like-for-like comparisons: mean with mean, median with median, and the same measure of spread for both groups. Two useful comments usually address different features—one centre and one spread—not the same feature twice.
A statistical conclusion is limited by the data collected and by what its summary measures leave hidden.
| Restriction | Why it matters | Safer response |
|---|---|---|
| small or unrepresentative sample | it may not reflect the intended population | limit the claim to the sampled group |
| extreme value | it can pull the mean and enlarge the range | inspect the data and compare median or IQR |
| one average or spread measure | different distributions can share the same summary | compare centre and spread together |
| different units, scales or time periods | the comparison is not like-for-like | standardise before comparing |
| association between variables | another factor may explain the pattern | describe association, not cause |
For 24,25,26,27,98, the mean is 40 but the median is 26. The single extreme value makes the mean look much larger than most observations, so the median better represents the centre of this set.
A lower median but larger IQR means one group is typically lower yet more variable; neither statement cancels the other. Conclusions should identify the exact statistic and avoid turning a summary comparison into a claim about every individual.
A limitation does not make data useless. It defines what can be claimed safely: state the group and period covered, the measures compared and the specific uncertainty that remains.
Mean, median and mode describe centre; quartiles divide ordered data into quarters; range and interquartile range describe spread.
| Measure | Calculation | Main use |
|---|---|---|
| Mean | sum of values ÷ number of values | uses every numerical value; affected by extremes |
| Median | middle of ordered data | resistant to extremes |
| Mode | most frequent value | identifies the most common value or category |
| Quartiles | Q1 and Q3 mark the lower and upper quarter positions | locate the middle half of ordered data |
| Range | maximum − minimum | full spread; sensitive to extremes |
| IQR | Q3−Q1 | spread of the middle half; resistant to extremes |
For ordered data 2,4,5,6,7,8,9,11, the median is (6+7)÷2=6.5. The lower half is 2,4,5,6, so Q1=(4+5)÷2=4.5; the upper half gives Q3=(8+9)÷2=8.5. Thus IQR =8.5−4.5=4, while range =11−2=9.
Use centre and spread together: median with IQR is often more representative when extremes are present; mean with range may expose the influence of the full data. State the measure used rather than saying only 'average' or 'variation'.
Order the data before finding median or quartiles. Apply one consistent quartile convention, do not confuse a quartile value with a quarter of the numerical range, and remember that IQR ignores the outer quarters.
For grouped data, estimate the mean by representing every value in a class by that class midpoint.
estimated mean=∑f∑fm,m=2lower boundary+upper boundary
| Class | Frequency f | Midpoint m | fm |
|---|---|---|---|
| 0<x≤20 | 3 | 10 | 30 |
| 20<x≤40 | 5 | 30 | 150 |
| 40<x≤60 | 2 | 50 | 100 |
| Total | 10 | 280 |
The estimate is 280÷10=28. Use the frequency, not the class width, as the multiplier. The same midpoint method works for grouped discrete or grouped continuous intervals when their frequencies are known.
The result is an estimate because the exact values inside each class are unknown and need not all equal the midpoint. Preserve units and avoid rounding midpoints or intermediate totals unnecessarily.
The modal class is the class interval with the greatest frequency in a grouped frequency distribution.
| Height h (m) | Frequency |
|---|---|
| 1.3<h≤1.4 | 6 |
| 1.4<h≤1.5 | 11 |
| 1.5<h≤1.6 | 18 |
| 1.6<h≤1.7 | 9 |
The greatest frequency is 18, so the modal class is 1.5<h≤1.6 m. Give the complete interval with its inequality signs and unit, not the frequency 18.
Grouped data do not reveal the exact most common value inside the interval, so report a modal class rather than inventing a mode. Do not choose the widest class or the interval with the largest boundary.
A statistical display represents frequencies so that categories, proportions or the shape of individual data can be compared. A correct display preserves every frequency and uses a clear scale or key.
| Display | Construction rule | What to interpret |
|---|---|---|
| Bar chart | equal-width bars, consistent gaps and a labelled linear frequency scale | compare heights; in dual bars compare paired groups, and in composite bars compare parts and totals |
| Pie chart | sector angle =totalfrequency×360∘ | compare proportions; convert angle back with 360∘angle×total |
| Pictogram | use one stated key for every symbol, including fractional symbols | multiply symbols by the key before comparing frequencies |
| Stem-and-leaf | split each value into a stem and leaf, order each row and give a key | recover original values and read their distribution |
| Frequency distribution | list each value or category once with its frequency | check total frequency and identify common or rare values |
For bars, choose a linear scale reaching the largest frequency, label both axes and plot each height accurately. A dual bar chart needs a key and the same scale for both groups. A composite bar's segments add to its total; read a segment from the difference between its boundaries. Unequal widths or a non-linear unmarked scale are misleading.
For frequencies 9,14,7, the total is 30. The sector angles are 9÷30×360∘=108∘, 14÷30×360∘=168∘ and 7÷30×360∘=84∘. They sum to 360∘, checking the chart.
For 13,15,21,21,28,31, use stems 1,2,3 and ordered leaves: 1∣3 5, 2∣1 1 8, 3∣1. The key 1∣3=13 fixes the place value. Every original value must appear exactly once.
Interpret only what the display supports: read the scale and key before calculating, and compare like with like. Bars show frequencies, pie sectors show proportions, and stem-and-leaf retains individual values. Box plots, histograms and scatter diagrams belong to other frozen objectives.
A scatter diagram plots one point for each paired observation, so the overall relationship between two numerical variables can be seen.
Put the explanatory variable on the horizontal axis and the other variable on the vertical axis when the context makes that choice clear. Label both axes with units, use linear scales, and plot each pair as a small clear cross. For (23,31.2), move to 23 on the horizontal scale and 31.2 on the vertical scale; the point must satisfy both coordinates.
Read the whole cloud of points before describing it. State the variables and direction, for example: 'as time in the shop increases, the number of items bought tends to increase.' A point far from the overall pattern is an unusual point; identify it from its coordinates.
Do not join consecutive points. A scatter diagram shows paired observations and a general pattern, not a time sequence or proof that one variable causes the other.
Correlation describes the direction of the overall relationship between two variables in a scatter diagram.
| Type | Pattern from left to right | Meaning |
|---|---|---|
| Positive | points tend to rise | as one variable increases, the other tends to increase |
| Negative | points tend to fall | as one variable increases, the other tends to decrease |
| Zero | no clear upward or downward pattern | changes in one variable give no consistent direction for the other |
If higher temperature is generally paired with more ice creams sold, the correlation is positive. If greater car power is generally paired with a shorter acceleration time, it is negative. Judge the trend of the whole cloud, not the slope between two selected points.
Correlation can be weak or affected by unusual points, and a very small set may not reveal a reliable pattern. Correlation describes association; it does not by itself prove causation.
A line of best fit is a single straight ruled line drawn by inspection to represent the central trend of a scatter diagram.
First identify the direction and centre of the point cloud. Draw one ruled line across the full data set, with roughly even numbers of points above and below it over its whole length. The line need not pass through any point or through the origin, and an isolated unusual point should not pull it away from the main pattern.
To estimate y from a given x, start at the x-axis value, move vertically to the line, then move horizontally to the y-axis and read the scale. Reverse these moves to estimate x from y. An estimate is approximate, so report a precision justified by the scale.
If the best-fit line for two test scores crosses near (40,48), a test 1 score of 40 gives an estimated test 2 score of about 48. The estimate comes from the line, not from choosing the nearest plotted student.
Estimates within the horizontal range of the data are interpolation. Extending beyond that range is extrapolation and is less reliable because the observed pattern may not continue.
Cumulative frequency is a running total: at each boundary it counts every observation up to and including that boundary.
| Class | Frequency | Upper boundary | Cumulative frequency |
|---|---|---|---|
| 0<x≤10 | 3 | 10 | 3 |
| 10<x≤20 | 5 | 20 | 8 |
| 20<x≤30 | 4 | 30 | 12 |
Add frequencies successively, then plot each upper class boundary against its cumulative frequency as a small clear cross. Include the lower boundary with cumulative frequency 0 when it is known. Label both axes and join the points with a smooth increasing curve; cumulative frequency must never decrease.
number in a<x≤b≈CF(b)−CF(a),number greater than a≈N−CF(a)
Read values from the curve using the axis scales, not from the class frequencies in isolation. Plot at upper boundaries, not class midpoints. Graph readings are estimates, and the final cumulative frequency must equal the total N.
A percentile is the data value below which a stated percentage of the observations lie. On a cumulative curve, convert the percentage into a cumulative-frequency position first.
| Measure | Cumulative-frequency position |
|---|---|
| lower quartile Q1 | 0.25N |
| median Q2 | 0.50N |
| upper quartile Q3 | 0.75N |
| pth percentile | 100pN |
Mark the required position on the cumulative-frequency axis, move horizontally to the curve, then move vertically to the data axis. For N=200, the median is read at cumulative frequency 100 and the 80th percentile at 160.
IQR=Q3−Q1
The median divides the ordered observations into two halves. The IQR measures the spread of the middle 50%, so a smaller IQR means the central half is more tightly clustered. Keep graph readings consistent with the scale and report them as estimates.
Do not read the 80th percentile at CF=80 unless N=100. Do not subtract cumulative-frequency positions to obtain the IQR: first read the two data values Q1 and Q3, then subtract them.
A histogram represents grouped continuous data so that each bar's area represents its class frequency. This is why unequal class widths require frequency density on the vertical axis.
| Feature | Histogram rule |
|---|---|
| horizontal axis | mark continuous class boundaries; bar width equals class width |
| vertical axis | label Frequency density and use a linear scale |
| bars | draw touching rectangles over the exact class intervals |
| frequency | compare bar areas, not heights alone |
A tall narrow bar can have less frequency than a shorter wide bar. To compare classes, use area =class width×frequency density. The tallest bar identifies the greatest frequency density, while the greatest area identifies the greatest frequency.
If one class has width 10 and density 4, its area and frequency are 40. A class of width 25 and density 2 has frequency 50: it is shorter but contains more observations.
Do not leave gaps between adjacent continuous intervals, plot class midpoints as bar positions, or label the vertical axis Frequency when widths differ. A bar chart compares categories using heights; a histogram encodes grouped frequency through area.
Frequency density adjusts frequency for class width, allowing unequal-width classes to be compared fairly in a histogram.
frequency density=class widthfrequency,frequency=density×class width
| Class | Width | Frequency | Density |
|---|---|---|---|
| 0<x≤10 | 10 | 20 | 20÷10=2 |
| 10<x≤30 | 20 | 30 | 30÷20=1.5 |
| 30<x≤35 | 5 | 25 | 25÷5=5 |
If a printed histogram gives bar heights in centimetres rather than a labelled density scale, all heights are proportional to density. Find the common multiplier from one known bar, then apply it to every density. To fill a missing class when total frequency is known, subtract the frequencies represented by the existing bar areas first.
To recover a frequency from the diagram, read the density and multiply by the class width. For only part of a class, the corresponding rectangle area gives an estimate based on the distribution within that class.
Class width is upper boundary minus lower boundary, not the midpoint or number of labels. Do not compare raw frequencies to choose bar height, and do not use a different height multiplier for different bars.