5. Probability & Statistics 1
- Syllabus
- 9709–2028–2029
- Section
- 5
- Level
- AS

| Need/data | Suitable representation | Advantage | Limitation |
|---|---|---|---|
| preserve small raw dataset | stem-and-leaf | retains individual values/order | crowded for large range |
| compare centre/spread | box plot | compact five-number comparison | hides shape/details |
| continuous class shape | histogram | area shows frequency/density | loses raw values |
| percentiles/proportions | cumulative-frequency graph | reads cumulative ranks | graph readings are estimates |
Identify variable type, sample size, grouping and comparison question. Choose a display whose encoding answers that question, label scale/units/key and state one relevant advantage and disadvantage.
For raw numerical data, a stem-and-leaf diagram may preserve more information than grouped displays. Grouping improves compactness but loses within-class positions.
Use the same scale/class boundaries where direct comparison is intended; otherwise visual differences may be caused by encoding choices.
No display is universally “best”. Histograms are for continuous intervals, not separated categorical bars; scatter plots are not part of this objective’s named one-variable displays.
| Display | Construction invariant | Main reading |
|---|---|---|
| stem-and-leaf / back-to-back | ordered leaves, common stem, key | raw values, shape, median/mode |
| box-and-whisker | min, Q1, median, Q3, max on scale | centre, IQR/range, skew comparison |
| histogram | contiguous class boundaries; density=f/class width | bar area = frequency |
| cumulative-frequency | cumulative totals at upper class boundaries, monotone curve | quantiles/percentiles/proportions |
For back-to-back stems, use one shared key and order leaves away from the stem consistently on each side so both datasets can be reconstructed.
\text{frequency density}=\frac{\text{frequency}}{\text{class width}},\qquad \text{frequency}=\text{bar area}.
State comparisons using numerical features: median/centre, IQR/range, shape or modal class. Label axes, units and sample identity.
Histogram height is not frequency for unequal class widths. A cumulative-frequency graph uses running totals, not ordinary class frequencies.
Mean, median and mode describe location. Range, interquartile range, variance and standard deviation describe spread; each responds differently to outliers and skew.
Use the mean when all values and squared deviations are meaningful, the median for skewed or ordinal data, and state which spread measure matches the centre.
One extreme value can raise the mean and standard deviation while leaving the median and IQR almost unchanged.
A larger mean does not imply greater variability, and “average” is not automatically the arithmetic mean.
For total frequency $N$: median rank $N/2$, quartiles $N/4$ and $3N/4$, percentile $p$ rank $pN/100$. Read the corresponding data value from the horizontal axis.
To estimate the number/proportion below value x, read cumulative count C(x), then use C(x)/N. Above x: [N−C(x)]/N.
For $a<X\le b$, estimated count isC(b)-C(a),and estimated proportion is $[C(b)-C(a)]/N$.
With N=80, Q3 is the x-value at cumulative frequency 60. If C(50)=62 and C(30)=18, the estimated proportion between 30 and 50 is (62−18)/80=0.55.
Graph readings from grouped data are estimates. Cumulative frequency is a count/rank, not itself the percentile’s data value.
For $n=\sum f$:\bar x=\frac{\sum fx}{n},\qquad \sigma=\sqrt{\frac{\sum fx^2}{n}-\bar x^2}.For raw data take $f=1$; for grouped data use class midpoint $x$ and label results estimates.
Build/use n,∑x,∑x2 (or ∑f,∑fx,∑fx2). Given totals can be inserted directly; do not reconstruct fictional raw values.
If $y=(x-a)/b$, then\bar x=a+b\bar y,\qquad \sigma_x=|b|\sigma_y.Coded totals $\sum(x-a)$ or $\sum(x-a)^2$ can also be expanded algebraically to recover totals.
For two datasets, add n, ∑x and ∑x2, then recompute the combined mean and SD. Do not average the two means/SDs unless sample sizes and formulas justify it.
∑x2 means sum of squared observations, not (∑x)2. Grouped midpoint results are estimates; combined variance must be rebuilt from combined totals.
nPr counts ordered arrangements of r objects from n; nCr counts unordered selections. Use nPr=n!/(n−r)! and nCr=n!/[r!(n−r)!].
Decide whether positions matter before choosing a formula, and avoid counting the same outcome under different descriptions.
Choosing president and secretary from 8 people uses 8P2; choosing a two-person committee uses 8C2.
The same two people can form two ordered roles but only one unordered committee.
| Structure | Counting method |
|---|---|
| repeated identical objects | n!/(a!b!⋯) |
| specified objects together | treat as one block, then multiply internal orders |
| specified objects not together | total arrangements − together arrangements |
| people in two or more rows | choose/allocate people to labelled seats/rows, then arrange within rows |
NEEDLESS has $8$ letters with $E$ repeated $3$ times and $S$ repeated $2$ times, so\frac{8!}{3!2!}distinctlineararrangements.
For 5 distinct books with two specified together: 4!×2!. For them not together: 5!−4!×2!. State whether ends/particular seats create separate cases.
If rows or seats are distinguishable, allocate occupants with combinations or direct seat choices, then arrange each row. Avoid counting the same final seating through different allocation orders.
Circular arrangements are explicitly excluded. Do not identify rotations or use (n−1)!; every task here is a line or labelled-row arrangement.
Forequiprobableelementaryoutcomes,P(A)=\frac{\text{number of outcomes in }A}{\text{total number of outcomes}}.Thetwocountsmustusethesameoutcomeunit.
| Situation | Build the probability |
|---|---|
| small sample space | enumerate every equally likely elementary outcome |
| ordered choices | count with permutations |
| unordered selections | count with combinations |
For two fair dice, use 36 ordered pairs. A total of 8 occurs for (2,6),(3,5),(4,4),(5,3),(6,2), so the probability is 5/36.
From 5 red and 3 blue balls, two chosen together have (28) equiprobable unordered pairs. Exactly one of each has (15)(13) pairs, giving 15/28.
Do not count favourable outcomes as ordered arrangements and total outcomes as unordered selections. First write what one elementary outcome means.
| Event structure | Operation |
|---|---|
| successive stages on one path | multiply branch probabilities |
| mutually exclusive alternative paths | add path probabilities |
| at least one / not / neither | consider 1−P(complement) |
A bag has 3 red and 2 blue balls. Two are drawn without replacement. The path red then blue has probability 3/5×2/4=3/10; blue then red has probability 2/5×3/4=3/10. Therefore one of each has probability 3/10+3/10=3/5.
Without replacement, update both the favourable count and total after the first draw. With replacement, the second-stage probabilities return to their original values.
Do not add probabilities along one path or multiply alternative paths. The syllabus does not require explicit use of the general overlapping-events addition formula here.
Events are mutually exclusive when A∩B=∅. They are independent when P(A∩B)=P(A)P(B), equivalently P(A|B)=P(A) when defined.
Disjoint non-zero events cannot be independent: learning that one occurred makes the other impossible. Test the stated relationship with the correct equation.
A single die roll being even and odd is mutually exclusive; two independent coin tosses are not mutually exclusive across different tosses.
“Independent” does not mean unrelated in everyday language, and mutually exclusive events are not independent unless one has probability zero.
If $P(B)>0$, thenP(A\mid B)=\frac{P(A\cap B)}{P(B)}.Read $A\mid B$ as ‘$A$ given that $B$ has occurred’.
| Representation | What the condition changes |
|---|---|
| equiprobable sample space | keep only outcomes satisfying the given event; this becomes the new denominator |
| tree diagram | start from the branch or branches compatible with the given information |
A fair die is known to show more than 2. The restricted outcomes are {3,4,5,6}. Given this information, the probability of an even result is 2/4=1/2, not 3/6.
A bag contains 5 red and 3 blue balls. If the first ball drawn without replacement is known to be red, then P(second red∣first red)=4/7. Equivalently, a joint probability can be built as P(A∩B)=P(B)P(A∣B).
The condition belongs after the vertical bar and determines the denominator. P(A∣B) and P(B∣A) are usually different.
For a discrete random variable X, list each possible value x with P(X=x). Every probability must be between 0 and 1 and the probabilities must sum to 1; use this total first to find any unknown probability.
E(X)=\sum xP(X=x),\qquad E(X^2)=\sum x^2P(X=x),\operatorname{Var}(X)=E(X^2)-[E(X)]^2.
| x | 0 | 1 | 2 |
|---|---|---|---|
| P(X=x) | 0.2 | 0.5 | 0.3 |
| xP(X=x) | 0 | 0.5 | 0.6 |
| x2P(X=x) | 0 | 0.5 | 1.2 |
Thus E(X)=1.1, E(X2)=1.7 and Var(X)=1.7−1.12=0.49. A probability table may come from enumeration or from combining outcomes that give the same value of X.
Do not average the listed x-values unless they are equally likely, and do not confuse [E(X)]2 with E(X2).
| Model | Random variable | Conditions | Support |
|---|---|---|---|
| X∼B(n,p) | number of successes in fixed n trials | independent trials, two outcomes, constant p | 0,1,…,n |
| X∼Geo(p) | trial number of the first success | independent repeated trials, constant p | 1,2,3,… |
P(X=r)=\binom nr p^r(1-p)^{n-r}\quad\text{for }X\sim B(n,p),P(X=r)=p(1-p)^{r-1}\quad\text{for }X\sim Geo(p).
If 8 independent items are defective with probability 0.1, the probability of exactly 2 defective items is (28)(0.1)2(0.9)6. For ranges such as at least 2, sum the relevant values or use a shorter complement.
If each attempt succeeds with probability 0.2, then the first success on attempt 4 has probability (0.8)3(0.2). ‘After attempt 4’ means four failures, so its probability is (0.8)4.
Geometric r includes the successful trial, whereas binomial r counts successes. Reject either model if independence or constant p is not reasonable.
| Distribution | Expectation | Variance |
|---|---|---|
| X∼B(n,p) | E(X)=np | Var(X)=np(1−p) |
| Y∼Geo(p) | E(Y)=1/p | not required in this syllabus objective |
These are long-run summaries, not necessarily possible single outcomes. A binomial expectation can be non-integer; a geometric expectation is the average trial number on which the first success occurs.
For X∼B(20,0.3), E(X)=20(0.3)=6 and Var(X)=20(0.3)(0.7)=4.2. The standard deviation, if requested, is 4.2.
For Y∼Geo(0.2), E(Y)=1/0.2=5: over many repetitions, the first success occurs on trial 5 on average.
Use the parameters of the stated model. Do not use np for a geometric variable or forget the factor 1−p in binomial variance. Proofs of these formulae are not required.
X∼N(μ,σ2) models a continuous variable whose distribution is approximately bell-shaped, symmetric and unimodal. The centre is μ; σ controls spread; total area under the curve is 1.
| Probability | Sketch/shading instruction |
|---|---|
| P(X<a) | shade left of a |
| P(X>a) | shade right of a |
| P(a<X<b) | shade between a and b |
StandardisewithZ=\frac{X-\mu}{\sigma},\qquad Z\sim N(0,1),thenusethestandardnormaltableandsymmetryorcomplementsasneeded.
For X∼N(10,4), the standard deviation is 2, so P(X<12)=P(Z<(12−10)/2)=P(Z<1)=0.8413.
Curve height is not a probability; area is. Because the model is continuous, P(X=a)=0, so < and ≤ give the same probability. In N(μ,σ2), the second parameter is variance.
| Step | Direct probability | Inverse/relationship problem |
|---|---|---|
| 1 | sketch and identify lower, upper or interval area | convert the stated area to the correct signed z-quantile |
| 2 | display Z=(X−μ)/σ with the numerical boundary | write (x−μ)/σ=z |
| 3 | read the table; complement or subtract if needed | rearrange to x=μ+zσ or the required relationship |
If X∼N(50,62), then P(X>62)=P(Z>662−50)=P(Z>2)=1−0.9772=0.0228. This displays the full standardisation required.
If P(X<x)=0.90, tables give z≈1.282, so (x−μ)/σ=1.282 and x=μ+1.282σ. If x is known, the same equation gives a relationship between μ and σ.
For a central probability, split the excluded area equally between two tails before finding z. For example, central 90% leaves 0.05 per tail and uses bounds μ±1.645σ.
‘At least’ and ‘more than’ are upper-tail events. Do not feed an upper-tail probability directly into a lower-tail table, and do not omit the sign of z for a boundary below μ.
For $X\sim B(n,p)$, let $q=1-p$. Use the normal approximation only whennp>5\quad\text{and}\quad nq>5,thenapproximatewithY\sim N(np,npq).
| Binomial event | Continuous normal event |
|---|---|
| X≤k | Y<k+0.5 |
| X<k (that is X≤k−1) | Y<k−0.5 |
| X≥k | Y>k−0.5 |
| X>k (that is X≥k+1) | Y>k+0.5 |
| a≤X≤b | a−0.5<Y<b+0.5 |
If X∼B(100,0.4), then np=40 and nq=60, so the conditions hold and Y∼N(40,24). To approximate P(X≤45), use P(Y<45.5)=P(Z<2445.5−40).
The correction gives each integer value its full unit-wide bar: the bar for X=45 extends from 44.5 to 45.5. Sketching the included integer bars makes the correct boundary visible.
Check both approximation conditions before calculating. Use variance npq but standard deviation npq when standardising, and apply continuity correction before converting to z.