5. Probability & Statistics 1

Syllabus
9709–2028–2029
Section
5
Level
A2

5.1 Representation of data

Syllabus
9709–2028–2029
Topic
5.1
Level
A2

Choose a display by what must remain visible

Need/data Suitable representation Advantage Limitation
preserve small raw dataset stem-and-leaf retains individual values/order crowded for large range
compare centre/spread box plot compact five-number comparison hides shape/details
continuous class shape histogram area shows frequency/density loses raw values
percentiles/proportions cumulative-frequency graph reads cumulative ranks graph readings are estimates

Identify variable type, sample size, grouping and comparison question. Choose a display whose encoding answers that question, label scale/units/key and state one relevant advantage and disadvantage.

For raw numerical data, a stem-and-leaf diagram may preserve more information than grouped displays. Grouping improves compactness but loses within-class positions.

Use the same scale/class boundaries where direct comparison is intended; otherwise visual differences may be caused by encoding choices.

No display is universally “best”. Histograms are for continuous intervals, not separated categorical bars; scatter plots are not part of this objective’s named one-variable displays.

Construct and read each named statistical display

Display Construction invariant Main reading
stem-and-leaf / back-to-back ordered leaves, common stem, key raw values, shape, median/mode
box-and-whisker min, Q1Q_1, median, Q3Q_3, max on scale centre, IQR/range, skew comparison
histogram contiguous class boundaries; density=f/class width\text{density}=f/\text{class width} bar area = frequency
cumulative-frequency cumulative totals at upper class boundaries, monotone curve quantiles/percentiles/proportions

For back-to-back stems, use one shared key and order leaves away from the stem consistently on each side so both datasets can be reconstructed.

\text{frequency density}=\frac{\text{frequency}}{\text{class width}},\qquad \text{frequency}=\text{bar area}.

State comparisons using numerical features: median/centre, IQR/range, shape or modal class. Label axes, units and sample identity.

Histogram height is not frequency for unequal class widths. A cumulative-frequency graph uses running totals, not ordinary class frequencies.

Averages describe centre while spread describes variation around it

Mean, median and mode describe location. Range, interquartile range, variance and standard deviation describe spread; each responds differently to outliers and skew.

Use the mean when all values and squared deviations are meaningful, the median for skewed or ordinal data, and state which spread measure matches the centre.

One extreme value can raise the mean and standard deviation while leaving the median and IQR almost unchanged.

A larger mean does not imply greater variability, and “average” is not automatically the arithmetic mean.

Convert cumulative ranks, values and intervals in both directions

For total frequency $N$: median rank $N/2$, quartiles $N/4$ and $3N/4$, percentile $p$ rank $pN/100$. Read the corresponding data value from the horizontal axis.

To estimate the number/proportion below value xx, read cumulative count C(x)C(x), then use C(x)/NC(x)/N. Above xx: [N−C(x)]/N[N-C(x)]/N.

For $a<X\le b$, estimated count isC(b)-C(a),and estimated proportion is $[C(b)-C(a)]/N$.

With N=80N=80, Q3Q_3 is the x-value at cumulative frequency 6060. If C(50)=62C(50)=62 and C(30)=18C(30)=18, the estimated proportion between 3030 and 5050 is (62−18)/80=0.55(62-18)/80=0.55.

Graph readings from grouped data are estimates. Cumulative frequency is a count/rank, not itself the percentile’s data value.

Compute mean and standard deviation from totals, coding or combined sets

For $n=\sum f$:\bar x=\frac{\sum fx}{n},\qquad \sigma=\sqrt{\frac{\sum fx^2}{n}-\bar x^2}.For raw data take $f=1$; for grouped data use class midpoint $x$ and label results estimates.

Build/use n,∑x,∑x2n,\sum x,\sum x^2 (or ∑f,∑fx,∑fx2\sum f,\sum fx,\sum fx^2). Given totals can be inserted directly; do not reconstruct fictional raw values.

If $y=(x-a)/b$, then\bar x=a+b\bar y,\qquad \sigma_x=|b|\sigma_y.Coded totals $\sum(x-a)$ or $\sum(x-a)^2$ can also be expanded algebraically to recover totals.

For two datasets, add nn, ∑x\sum x and ∑x2\sum x^2, then recompute the combined mean and SD. Do not average the two means/SDs unless sample sizes and formulas justify it.

∑x2\sum x^2 means sum of squared observations, not (∑x)2(\sum x)^2. Grouped midpoint results are estimates; combined variance must be rebuilt from combined totals.

5.2 Permutations and combinations

Syllabus
9709–2028–2029
Topic
5.2
Level
A2

Permutations count ordered selections while combinations ignore order

nPr counts ordered arrangements of r objects from n; nCr counts unordered selections. Use nPr=n!/(n−r)! and nCr=n!/[r!(n−r)!].

Decide whether positions matter before choosing a formula, and avoid counting the same outcome under different descriptions.

Choosing president and secretary from 8 people uses 8P2; choosing a two-person committee uses 8C2.

The same two people can form two ordered roles but only one unordered committee.

Choose division, blocks, complements or row allocation

Structure Counting method
repeated identical objects n!/(a!b!⋯ )n!/(a!b!\cdots)
specified objects together treat as one block, then multiply internal orders
specified objects not together total arrangements −- together arrangements
people in two or more rows choose/allocate people to labelled seats/rows, then arrange within rows

NEEDLESS has $8$ letters with $E$ repeated $3$ times and $S$ repeated $2$ times, so\frac{8!}{3!2!}distinctlineararrangements.distinct linear arrangements.

For 5 distinct books with two specified together: 4!×2!4!\times2!. For them not together: 5!−4!×2!5!-4!\times2!. State whether ends/particular seats create separate cases.

If rows or seats are distinguishable, allocate occupants with combinations or direct seat choices, then arrange each row. Avoid counting the same final seating through different allocation orders.

Circular arrangements are explicitly excluded. Do not identify rotations or use (n−1)!(n-1)!; every task here is a line or labelled-row arrangement.

5.3 Probability

Syllabus
9709–2028–2029
Topic
5.3
Level
A2

Count the same kind of equally likely outcome

Forequiprobableelementaryoutcomes,For equiprobable elementary outcomes,P(A)=\frac{\text{number of outcomes in }A}{\text{total number of outcomes}}.Thetwocountsmustusethesameoutcomeunit.The two counts must use the same outcome unit.

Situation Build the probability
small sample space enumerate every equally likely elementary outcome
ordered choices count with permutations
unordered selections count with combinations

For two fair dice, use 36 ordered pairs. A total of 8 occurs for (2,6),(3,5),(4,4),(5,3),(6,2)(2,6),(3,5),(4,4),(5,3),(6,2), so the probability is 5/365/36.

From 5 red and 3 blue balls, two chosen together have (82)\binom82 equiprobable unordered pairs. Exactly one of each has (51)(31)\binom51\binom31 pairs, giving 15/2815/28.

Do not count favourable outcomes as ordered arrangements and total outcomes as unordered selections. First write what one elementary outcome means.

Multiply along a path and add alternative paths

Event structure Operation
successive stages on one path multiply branch probabilities
mutually exclusive alternative paths add path probabilities
at least one / not / neither consider 1−P(complement)1-P(\text{complement})

A bag has 3 red and 2 blue balls. Two are drawn without replacement. The path red then blue has probability 3/5×2/4=3/103/5\times2/4=3/10; blue then red has probability 2/5×3/4=3/102/5\times3/4=3/10. Therefore one of each has probability 3/10+3/10=3/53/10+3/10=3/5.

Without replacement, update both the favourable count and total after the first draw. With replacement, the second-stage probabilities return to their original values.

Do not add probabilities along one path or multiply alternative paths. The syllabus does not require explicit use of the general overlapping-events addition formula here.

Mutual exclusivity and independence describe different relationships

Events are mutually exclusive when A∩B=∅. They are independent when P(A∩B)=P(A)P(B), equivalently P(A|B)=P(A) when defined.

Disjoint non-zero events cannot be independent: learning that one occurred makes the other impossible. Test the stated relationship with the correct equation.

A single die roll being even and odd is mutually exclusive; two independent coin tosses are not mutually exclusive across different tosses.

“Independent” does not mean unrelated in everyday language, and mutually exclusive events are not independent unless one has probability zero.

Conditioning restricts the possible outcomes

If $P(B)>0$, thenP(A\mid B)=\frac{P(A\cap B)}{P(B)}.Read $A\mid B$ as ‘$A$ given that $B$ has occurred’.

Representation What the condition changes
equiprobable sample space keep only outcomes satisfying the given event; this becomes the new denominator
tree diagram start from the branch or branches compatible with the given information

A fair die is known to show more than 2. The restricted outcomes are {3,4,5,6}\{3,4,5,6\}. Given this information, the probability of an even result is 2/4=1/22/4=1/2, not 3/63/6.

A bag contains 5 red and 3 blue balls. If the first ball drawn without replacement is known to be red, then P(second red∣first red)=4/7P(\text{second red}\mid\text{first red})=4/7. Equivalently, a joint probability can be built as P(A∩B)=P(B)P(A∣B)P(A\cap B)=P(B)P(A\mid B).

The condition belongs after the vertical bar and determines the denominator. P(A∣B)P(A\mid B) and P(B∣A)P(B\mid A) are usually different.

5.4 Discrete random variables

Syllabus
9709–2028–2029
Topic
5.4
Level
A2

A probability table gives both centre and spread

For a discrete random variable XX, list each possible value xx with P(X=x)P(X=x). Every probability must be between 0 and 1 and the probabilities must sum to 1; use this total first to find any unknown probability.

E(X)=\sum xP(X=x),\qquad E(X^2)=\sum x^2P(X=x),\operatorname{Var}(X)=E(X^2)-[E(X)]^2.

xx 0 1 2
P(X=x)P(X=x) 0.2 0.5 0.3
xP(X=x)xP(X=x) 0 0.5 0.6
x2P(X=x)x^2P(X=x) 0 0.5 1.2

Thus E(X)=1.1E(X)=1.1, E(X2)=1.7E(X^2)=1.7 and Var⁡(X)=1.7−1.12=0.49\operatorname{Var}(X)=1.7-1.1^2=0.49. A probability table may come from enumeration or from combining outcomes that give the same value of XX.

Do not average the listed xx-values unless they are equally likely, and do not confuse [E(X)]2[E(X)]^2 with E(X2)E(X^2).

Choose binomial for a fixed count, geometric for first success

Model Random variable Conditions Support
X∼B(n,p)X\sim B(n,p) number of successes in fixed nn trials independent trials, two outcomes, constant pp 0,1,…,n0,1,\ldots,n
X∼Geo(p)X\sim Geo(p) trial number of the first success independent repeated trials, constant pp 1,2,3,…1,2,3,\ldots

P(X=r)=\binom nr p^r(1-p)^{n-r}\quad\text{for }X\sim B(n,p),P(X=r)=p(1-p)^{r-1}\quad\text{for }X\sim Geo(p).

If 8 independent items are defective with probability 0.1, the probability of exactly 2 defective items is (82)(0.1)2(0.9)6\binom82(0.1)^2(0.9)^6. For ranges such as at least 2, sum the relevant values or use a shorter complement.

If each attempt succeeds with probability 0.2, then the first success on attempt 4 has probability (0.8)3(0.2)(0.8)^3(0.2). ‘After attempt 4’ means four failures, so its probability is (0.8)4(0.8)^4.

Geometric rr includes the successful trial, whereas binomial rr counts successes. Reject either model if independence or constant pp is not reasonable.

Read model moments directly from their parameters

Distribution Expectation Variance
X∼B(n,p)X\sim B(n,p) E(X)=npE(X)=np Var⁡(X)=np(1−p)\operatorname{Var}(X)=np(1-p)
Y∼Geo(p)Y\sim Geo(p) E(Y)=1/pE(Y)=1/p not required in this syllabus objective

These are long-run summaries, not necessarily possible single outcomes. A binomial expectation can be non-integer; a geometric expectation is the average trial number on which the first success occurs.

For X∼B(20,0.3)X\sim B(20,0.3), E(X)=20(0.3)=6E(X)=20(0.3)=6 and Var⁡(X)=20(0.3)(0.7)=4.2\operatorname{Var}(X)=20(0.3)(0.7)=4.2. The standard deviation, if requested, is 4.2\sqrt{4.2}.

For Y∼Geo(0.2)Y\sim Geo(0.2), E(Y)=1/0.2=5E(Y)=1/0.2=5: over many repetitions, the first success occurs on trial 5 on average.

Use the parameters of the stated model. Do not use npnp for a geometric variable or forget the factor 1−p1-p in binomial variance. Proofs of these formulae are not required.

5.5 The normal distribution

Syllabus
9709–2028–2029
Topic
5.5
Level
A2

A normal probability is area under a symmetric continuous curve

X∼N(μ,σ2)X\sim N(\mu,\sigma^2) models a continuous variable whose distribution is approximately bell-shaped, symmetric and unimodal. The centre is μ\mu; σ\sigma controls spread; total area under the curve is 1.

Probability Sketch/shading instruction
P(X<a)P(X<a) shade left of aa
P(X>a)P(X>a) shade right of aa
P(a<X<b)P(a<X<b) shade between aa and bb

StandardisewithStandardise withZ=\frac{X-\mu}{\sigma},\qquad Z\sim N(0,1),thenusethestandardnormaltableandsymmetryorcomplementsasneeded.then use the standard normal table and symmetry or complements as needed.

For X∼N(10,4)X\sim N(10,4), the standard deviation is 2, so P(X<12)=P(Z<(12−10)/2)=P(Z<1)=0.8413P(X<12)=P(Z<(12-10)/2)=P(Z<1)=0.8413.

Curve height is not a probability; area is. Because the model is continuous, P(X=a)=0P(X=a)=0, so << and ≤\le give the same probability. In N(μ,σ2)N(\mu,\sigma^2), the second parameter is variance.

Show the tail, standardise fully, then solve or invert

Step Direct probability Inverse/relationship problem
1 sketch and identify lower, upper or interval area convert the stated area to the correct signed zz-quantile
2 display Z=(X−μ)/σZ=(X-\mu)/\sigma with the numerical boundary write (x−μ)/σ=z(x-\mu)/\sigma=z
3 read the table; complement or subtract if needed rearrange to x=μ+zσx=\mu+z\sigma or the required relationship

If X∼N(50,62)X\sim N(50,6^2), then P(X>62)=P(Z>62−506)=P(Z>2)=1−0.9772=0.0228.P(X>62)=P\left(Z>\frac{62-50}{6}\right)=P(Z>2)=1-0.9772=0.0228. This displays the full standardisation required.

If P(X<x)=0.90P(X<x)=0.90, tables give z≈1.282z\approx1.282, so (x−μ)/σ=1.282(x-\mu)/\sigma=1.282 and x=μ+1.282σx=\mu+1.282\sigma. If xx is known, the same equation gives a relationship between μ\mu and σ\sigma.

For a central probability, split the excluded area equally between two tails before finding zz. For example, central 90% leaves 0.05 per tail and uses bounds μ±1.645σ\mu\pm1.645\sigma.

‘At least’ and ‘more than’ are upper-tail events. Do not feed an upper-tail probability directly into a lower-tail table, and do not omit the sign of zz for a boundary below μ\mu.

Move each binomial boundary by half a unit

For $X\sim B(n,p)$, let $q=1-p$. Use the normal approximation only whennp>5\quad\text{and}\quad nq>5,thenapproximatewiththen approximate withY\sim N(np,npq).

Binomial event Continuous normal event
X≤kX\le k Y<k+0.5Y<k+0.5
X<kX<k (that is X≤k−1X\le k-1) Y<k−0.5Y<k-0.5
X≥kX\ge k Y>k−0.5Y>k-0.5
X>kX>k (that is X≥k+1X\ge k+1) Y>k+0.5Y>k+0.5
a≤X≤ba\le X\le b a−0.5<Y<b+0.5a-0.5<Y<b+0.5

If X∼B(100,0.4)X\sim B(100,0.4), then np=40np=40 and nq=60nq=60, so the conditions hold and Y∼N(40,24)Y\sim N(40,24). To approximate P(X≤45)P(X\le45), use P(Y<45.5)=P(Z<45.5−4024).P(Y<45.5)=P\left(Z<\frac{45.5-40}{\sqrt{24}}\right).

The correction gives each integer value its full unit-wide bar: the bar for X=45X=45 extends from 44.5 to 45.5. Sketching the included integer bars makes the correct boundary visible.

Check both approximation conditions before calculating. Use variance npqnpq but standard deviation npq\sqrt{npq} when standardising, and apply continuity correction before converting to zz.