6 Statistics and probability
- Syllabus
- 2017
- Section
- 6
- Level
- Higher

Choose a display by matching it to the data and the comparison you need. Pictograms and bar charts show discrete categories; pie charts show parts of one whole; two-way tables classify each item by two categorical variables.
| Display | Best use | Essential check |
|---|---|---|
| pictogram | small category frequencies | read the key, including partial symbols |
| bar chart | compare discrete categories | equal-width separated bars and a labelled scale |
| pie chart | compare proportions of one total | sectors total 360∘ |
| two-way table | cross-classify two categories | row and column totals agree with the grand total |
For a pie chart, sector angle=totalfrequency×360∘. Reverse the calculation with frequency=360∘angle×total.
A taller bar may represent a larger frequency, but a wider bar in an ordinary bar chart does not. Do not use bar area unless the graph is a histogram.
A frequency table organises every observation into one category or class, records tallies in groups of five, and converts each tally into a frequency.
| Step | Action |
|---|---|
| 1 | define categories or non-overlapping class intervals |
| 2 | process the raw list once, adding one tally per observation |
| 3 | convert each tally to a numerical frequency |
| 4 | add the frequencies and compare with the number of observations |
Categories must be exhaustive and mutually exclusive: every value belongs somewhere and no value belongs twice. Interval notation such as 10<x≤20 fixes where a boundary value goes.
Do not recount separately for each category; a single pass with tallies reduces omissions and double-counting.
Before reading a statistical diagram, identify its title, variable, units, category labels or class intervals, and the scale on each axis or in its key.
| Task | Reliable method |
|---|---|
| read a value | trace from the mark or bar to the labelled scale |
| compare categories | read both values, then subtract or form the requested ratio |
| find a total | convert every symbol, bar or sector to frequency before adding |
| interpret a pie sector | use its fraction of 360∘, not its visual width alone |
If one vertical grid step represents 2 games and a bar is 5 steps high, its frequency is 5×2=10 games. The scale must be decoded before the height is used.
A diagram supports statements about the displayed data, not an unexplained causal claim. Also check for truncated axes or unequal scales before comparing visual size.
A histogram displays grouped continuous data. Bars touch, class intervals determine bar widths, and the vertical axis is frequency density so that each bar's area represents frequency.
| Quantity | Relationship |
|---|---|
| class width | upper boundary − lower boundary |
| frequency density | class widthfrequency |
| frequency | class width × frequency density |
Compute every class width and frequency density, mark the continuous class boundaries on the horizontal axis, then draw touching rectangles. To estimate a frequency over part of a class, use the corresponding fraction of that bar's area.
When class widths differ, bar height is not frequency. Compare or add bar areas; treating a histogram as an ordinary bar chart gives the wrong result.
Cumulative frequency is a running total. Each plotted point shows how many observations are at or below an upper class boundary.
| Class frequency | Cumulative frequency | Plot at |
|---|---|---|
| 3 | 3 | first upper boundary |
| 16 | 3+16=19 | second upper boundary |
| 24 | 19+24=43 | third upper boundary |
Add frequencies successively, pair each running total with its upper class boundary, include the lower starting boundary with cumulative frequency 0 when appropriate, plot the points, and join them with a smooth increasing curve or line segments.
Do not plot at class midpoints and do not use the separate class frequencies as vertical coordinates. A cumulative-frequency graph must never decrease.
A cumulative frequency diagram converts between a value and the number at or below that value. Read up from a value to the curve and across to cumulative frequency, or reverse those steps to find a value for a chosen cumulative count.
| Required estimate for total N | Cumulative position |
|---|---|
| lower quartile Q1 | N/4 |
| median | N/2 |
| upper quartile Q3 | 3N/4 |
| interquartile range | Q3−Q1 |
If the graph gives cumulative frequency c at value x, then about c observations are ≤x and about N−c are >x. For an interval a<x≤b, subtract the readings: CF(b)−CF(a).
Graph readings are estimates, so use sensible precision. When comparing two groups, state what the medians show about typical value and what the interquartile ranges show about spread.
An average is a single value used to describe a typical or central feature of a data set. Different averages answer different questions, so the context and data shape determine which is useful.
| Measure | What it represents | Sensitive to extreme values? |
|---|---|---|
| mean | equal-share value using every observation | yes |
| median | middle of ordered data | much less |
| mode | most frequent value or category | no |
For a frequency table, frequencies tell how many times each value occurs; expand conceptually or use cumulative counts to locate the middle.
There is no universally 'best' average. A data set may have no mode or several modes, and a mean can be unrepresentative when a few extreme values pull it away from most observations.
For discrete data, calculate each measure from the actual values and their frequencies, keeping centre and spread distinct.
| Measure | Rule |
|---|---|
| mean | n∑x, or ∑f∑fx for a frequency table |
| median | middle ordered value; average the two middle values when n is even |
| mode | value with greatest frequency |
| range | maximum − minimum |
In reverse-mean problems, first convert a mean to a total: total=mean×number. Combine or subtract totals before dividing by the new number of values.
Order the data before finding the median. Do not divide by the number of table rows: for a frequency table the divisor is ∑f.
Grouped data do not reveal the exact observations, so represent every value in a class by its class midpoint and calculate an estimated mean.
| Step | Calculation |
|---|---|
| midpoint | (lower boundary+upper boundary)/2 |
| estimated class total | midpoint × frequency |
| estimated mean | ∑f∑(midpoint×f) |
For 100<h≤110 with frequency 12, use midpoint 105, contributing 105×12=1260 to the estimated total.
The answer is an estimate because actual values need not equal the midpoint. Never multiply frequency by class width when estimating the mean.
The modal class is the class interval containing the greatest frequency. State the complete interval, including its boundary convention.
| Class | Frequency |
|---|---|
| 0<d≤4 | 9 |
| 4<d≤8 | 15 |
| 8<d≤12 | 7 |
The modal class in the example is 4<d≤8 because 15 is the largest frequency.
The modal class is an interval, not its midpoint. If information is shown only by a histogram with unequal widths, compare frequency density or bar area according to the task rather than assuming the tallest-looking width gives the class frequency.
For total frequency N, the median lies at cumulative frequency N/2. Read horizontally from N/2 to the cumulative-frequency curve, then vertically to the data axis.
| Step | Action |
|---|---|
| 1 | read the final cumulative total N |
| 2 | calculate N/2 |
| 3 | move from N/2 on the vertical axis to the curve |
| 4 | move down to estimate the median value |
To compare typical values for two groups, compare their medians in context: the group with the larger median tends to have the larger observed value.
Use the cumulative-frequency axis position, not half the horizontal-axis range. The result is an estimate and should not claim more precision than the graph supports.
A measure of spread describes how variable or dispersed the data are. Larger spread means values are less tightly clustered; smaller spread means greater consistency.
| Measure | Uses | Sensitivity |
|---|---|---|
| range =max−min | full data width | strongly affected by extremes |
| interquartile range =Q3−Q1 | middle 50% width | resistant to extremes |
Compare groups with both centre and spread: median describes a typical value, while IQR describes consistency of the middle half. Write comparisons in the units and context of the data.
A higher median does not mean greater spread, and a smaller IQR does not mean a smaller typical value; centre and variation answer different questions.
Order the data, locate the lower quartile Q1 and upper quartile Q3 using the course's median-of-halves convention, then calculate IQR=Q3−Q1.
| Ordered data count | Split for quartiles |
|---|---|
| odd | exclude the overall median, then find medians of lower and upper halves |
| even | split into equal lower and upper halves, then find each half's median |
For 11 ordered values, the 6th is the overall median; Q1 is the median of values 1–5 and Q3 is the median of values 7–11.
Quartiles come from ordered positions, not one quarter and three quarters of the numerical range. State and use one consistent convention.
For total frequency N, read the lower quartile at cumulative frequency N/4 and the upper quartile at 3N/4, then subtract.
| Quantity | Cumulative frequency position |
|---|---|
| Q1 | N/4 |
| median | N/2 |
| Q3 | 3N/4 |
| IQR | Q3−Q1 |
A larger estimated IQR means the middle half is more spread out; a smaller estimated IQR means it is more consistent. Support the statement with both estimated IQRs when comparing groups.
Do not subtract cumulative frequencies 3N/4−N/4; those positions locate quartile values on the horizontal axis, and the horizontal readings must be subtracted.
An outcome is one possible result of a trial. An event is a specified outcome or set of outcomes. A random trial has an uncertain result, even though its possible outcomes may be known.
| Term | Meaning |
|---|---|
| impossible | cannot occur |
| unlikely | probability below 21 |
| evens | probability 21 |
| likely | probability above 21 |
| certain | must occur |
Equally likely outcomes have the same chance. Do not assume outcomes are equally likely merely because they can all occur.
Probability language compares likelihood, not frequency already observed in a small sample. 'Random' does not mean every result must appear equally often.
Every probability lies from 0 to 1 inclusive. Probability 0 means impossible, 21 means an even chance, and 1 means certain.
| Probability | Interpretation |
|---|---|
| 0 | impossible |
| between 0 and 21 | unlikely |
| 21 | evens |
| between 21 and 1 | likely |
| 1 | certain |
Fractions, decimals and percentages can name the same position: 43=0.75=75%.
A value below 0 or above 1 cannot be a probability. When marking a scale, use its subdivisions rather than estimating from the page width.
When all elementary outcomes are equally likely, theoretical probability is the number of favourable outcomes divided by the total number of possible outcomes.
| Quantity | Rule |
|---|---|
| probability of event A | P(A)=all equally likely outcomesfavourable outcomes |
| check | 0≤P(A)≤1 |
A fair six-sided die has three even faces, so P(even)=63=21.
Count outcomes, not labels. If sectors, objects or mechanisms are not equally likely, favourable-count over total-count is not justified without weighting them.
A Venn diagram partitions the universal set into disjoint regions. Read the region named by the event, add its frequencies, then divide by the total frequency when one item is chosen at random.
| Event | Regions included |
|---|---|
| A∩B | overlap of A and B |
| A∪B | every region in A or B, overlap once |
| A′ | every region outside A |
| neither A nor B | outside both circles |
For a conditional statement such as 'given B', restrict the denominator to the total inside B before counting the favourable part.
Do not count an overlap twice when finding a union, and do not omit the region outside all circles from the universal total.
A sample space lists every possible outcome of an experiment. An event is a subset of that sample space.
| First coin | Second coin | Ordered outcome |
|---|---|---|
| H | H | HH |
| H | T | HT |
| T | H | TH |
| T | T | TT |
For fair independent coins, the event 'exactly one head' is {HT,TH}, so its probability is 42=21.
Outcomes such as HT and TH are different when order records successive results. A sample space must be exhaustive and must not repeat an outcome.
For two successive choices or trials, use an ordered list, two-way table or tree so every first-stage option is paired with every permitted second-stage option.
| Step | Action |
|---|---|
| 1 | fix one first-stage outcome |
| 2 | pair it with every allowed second-stage outcome |
| 3 | repeat for each first-stage outcome |
| 4 | check restrictions such as no repetition or order ignored |
If there are m first-stage choices and n independent second-stage choices, there are mn ordered outcomes.
Do not use mn when a choice is removed, repeats are forbidden, or AB and BA represent the same unordered pair; adjust the list to the actual rules.
Experimental probability uses observed results: divide the frequency of the event by the total number of trials or observations.
| Quantity | Calculation |
|---|---|
| experimental probability | number of trialsevent frequency |
| estimated event count in N future trials | N×experimental probability |
A larger relevant sample usually gives a more stable estimate, but it does not guarantee the next outcome or make the estimate exact.
Use the total number represented by the data as the denominator. Do not confuse a cumulative frequency, a subgroup total or a graph reading with the whole sample.
An event and its complement cover every outcome without overlap, so their probabilities add to 1.
| Required event | Complement method |
|---|---|
| not A | P(A′)=1−P(A) |
| at least one success | 1−P(no successes) |
| any unlisted category | 1−sum of listed category probabilities |
If P(packed lunch)=0.79, then P(no packed lunch)=1−0.79=0.21.
Subtract from 1 only when the event being subtracted is exactly the complement of the required event. 'At least one' is complemented by 'none', not by 'exactly one'.
Mutually exclusive events cannot happen on the same trial. For such events, the probability of one or the other is the sum of their probabilities.
| Condition | Addition rule |
|---|---|
| A and B mutually exclusive | P(A∪B)=P(A)+P(B) |
| A and B may overlap | P(A∪B)=P(A)+P(B)−P(A∩B) |
If a spinner lands on exactly one colour, red and yellow are mutually exclusive, so P(red or yellow)=P(red)+P(yellow).
The word 'or' does not by itself prove mutual exclusivity. Check whether both events can occur together before adding without an overlap correction.
Expected frequency is the long-run estimate of how often an event occurs: multiply its probability by the number of trials.
| Known information | Calculation |
|---|---|
| probability p, trials n | expected frequency =np |
| event frequency f, probability p | estimated trials =f/p |
| expected frequency f, trials n | probability =f/n |
If P(blue)=0.4 over 280 spins, the expected frequency is 280×0.4=112.
Expected frequency is an estimate, not a guaranteed result. First find any missing probability, and keep trials and event counts in consistent units.
A probability tree shows successive stages. Branches leaving one node list all possible next outcomes and must sum to 1.
| Step | Action |
|---|---|
| 1 | label every stage and branch outcome |
| 2 | complete each sibling pair or group so it sums to 1 |
| 3 | multiply probabilities along one path |
| 4 | add probabilities of the distinct paths that satisfy the event |
Decide whether later branch probabilities stay the same or change according to replacement, dependence and earlier outcomes.
Do not add probabilities along a path or multiply alternative paths. Multiplication means successive events on one path; addition combines disjoint completed paths.
Events are independent when knowing one occurred does not change the probability of the other. Then the probability of both is the product of their probabilities.
| Relationship | Rule |
|---|---|
| independent A and B | P(A∩B)=P(A)P(B) |
| repeated independent event | use the same branch probabilities at each repetition |
| exactly one of two events | add the two orders AB′ and A′B |
For at least one success in independent repetitions, the complement is often shorter: 1−P(all failures).
Repeated-looking trials are not automatically independent. Replacement or a separate mechanism may preserve probabilities; removing an item usually changes them.
Without replacement, the first selection changes both the total number of objects and possibly the favourable number. Later probabilities are conditional on the earlier path.
| After one object is removed | Update |
|---|---|
| denominator | subtract 1 from the total |
| favourable numerator | subtract 1 only if the removed object was favourable |
| one order | multiply its successive conditional probabilities |
| several valid orders | calculate each order and add |
From 7 red and 5 blue counters, P(two red without replacement)=127×116.
Do not reuse the original fraction on the second draw. For 'one of each', include both red-then-blue and blue-then-red unless the order is fixed.
A multi-step probability problem is solved by defining the required event, choosing a representation, applying the correct rule at each stage, and checking that the answer is plausible.
| Feature in the problem | Useful move |
|---|---|
| all outcomes can be listed | sample space or systematic table |
| successive stages | probability tree |
| 'not' or 'at least one' | test a complement |
| disjoint alternatives | add path probabilities |
| without replacement | update conditional branches |
| repeated trials estimate | expected frequency =np |
State assumptions, keep exact fractions until the end when practical, and verify 0≤P≤1 and that sibling branches sum to 1.
Do not select a rule from a keyword alone. Translate the event first, then check for overlap, independence, replacement and whether the question asks for probability or expected count.