6. Probability & Statistics 2

Syllabus
9709–2028–2029
Section
6
Level
A2

6.1 The Poisson distribution

Syllabus
9709–2028–2029
Topic
6.1
Level
A2

A Poisson model counts events in a fixed interval at a constant average rate

If events occur independently at constant mean rate λ per interval, X~Po(λ) and P(X=r)=e^{−λ}λ^r/r!.

Match λ to the interval length, check that events are countable and rare enough for the model, and use complements for “at least one” questions.

If a call centre averages 3 calls per minute, P(2 calls in one minute)=e^{−3}3²/2.

Changing the interval changes λ proportionally; it is not a universal parameter for every time window.

A Poisson model’s mean and variance both equal its parameter

For X~Po(λ), E(X)=λ and Var(X)=λ. For a time or area interval scaled by k, the mean becomes kλ under a constant-rate model.

Use the equality as a model check, not as a statement that every sample has equal mean and variance. Estimate λ from appropriate exposure.

If the observed average is 4 events per hour, a two-hour interval has Po(8), not Po(4).

Sample variance need not equal sample mean exactly; the equality describes the distributional parameter.

Use Poisson only for stable independent random-event counts

Check Poisson-model requirement
response a count 0,1,2,…0,1,2,\ldots in fixed time, length, area or volume
rate constant average rate over the exposure
occurrence events occur independently
simultaneity in a very small exposure, more than one event is negligible

If the average is $r$ events per unit exposure, then an exposure of size $t$ usesX\sim Po(rt).

If flaws occur randomly at an average of 0.8 per metre, the number in 5 metres can be modelled by Po(4)Po(4), provided the rate is stable and flaws do not cluster.

Question the model when the rate varies systematically, one event triggers another, observations are capped by a fixed number of trials, or the exposure itself is unclear. State the modelling assumption when context makes it relevant.

A count variable is not automatically Poisson. The theorem that independent Poisson variables add belongs to linear combinations and does not replace the model-suitability checks here.

Approximate rare binomial successes with $Po(np)$

For $X\sim B(n,p)$, the syllabus guide isn>50\quad\text{and}\quad np<5(approximately).Whensuitable,use(approximately). When suitable, useX\approx Y,\qquad Y\sim Po(\lambda),\quad\lambda=np.

Step Action
1 check large nn and small pp using the stated guide
2 calculate λ=np\lambda=np
3 keep the same integer event (=,≤,≥=,\le,\ge)
4 calculate with the Poisson formula/table; no continuity correction

If X∼B(200,0.01)X\sim B(200,0.01), then n>50n>50 and np=2<5np=2<5, so use Y∼Po(2)Y\sim Po(2). Thus P(X≥2)≈P(Y≥2)=1−P(Y=0)−P(Y=1)=1−e−2(1+2).P(X\ge2)\approx P(Y\ge2)=1-P(Y=0)-P(Y=1)=1-e^{-2}(1+2).

The approximation replaces many rare independent success opportunities by a random-event count with the same expected value npnp.

Small pp alone is insufficient. Do not use a half-unit continuity correction: both binomial and Poisson variables are discrete.

Approximate a large-mean Poisson count with $N(\lambda,\lambda)$

For $X\sim Po(\lambda)$, when $\lambda$ is large (the syllabus guide is approximately $\lambda>15$), useY\sim N(\lambda,\lambda).Hence the normal standard deviation is $\sqrt{\lambda}$.

Poisson event Corrected normal event
X≤kX\le k Y<k+0.5Y<k+0.5
X≥kX\ge k Y>k−0.5Y>k-0.5
a≤X≤ba\le X\le b a−0.5<Y<b+0.5a-0.5<Y<b+0.5

If X∼Po(25)X\sim Po(25), the guide condition holds and Y∼N(25,25)Y\sim N(25,25). Therefore P(X≥30)≈P(Y>29.5)=P(Z>29.5−255)=P(Z>0.9).P(X\ge30)\approx P(Y>29.5)=P\left(Z>\frac{29.5-25}{5}\right)=P(Z>0.9).

State the large-mean check, write the approximating normal distribution, correct the integer boundary, then standardise with λ\sqrt{\lambda}.

Do not mix in binomial conditions. For a Poisson variable both normal mean and variance are λ\lambda, but the denominator in the zz-score is λ\sqrt{\lambda}.

6.2 Linear combinations of random variables

Syllabus
9709–2028–2029
Topic
6.2
Level
A2

Means follow signs; independent variances use squared coefficients

Quantity Result Condition
E(aX+b)E(aX+b) aE(X)+baE(X)+b always
Var⁡(aX+b)\operatorname{Var}(aX+b) a2Var⁡(X)a^2\operatorname{Var}(X) always
E(aX+bY)E(aX+bY) aE(X)+bE(Y)aE(X)+bE(Y) always
Var⁡(aX+bY)\operatorname{Var}(aX+bY) a2Var⁡(X)+b2Var⁡(Y)a^2\operatorname{Var}(X)+b^2\operatorname{Var}(Y) X,YX,Y independent
Starting distributions Resulting distribution
XX normal aX+baX+b is normal
independent normal X,YX,Y aX+bYaX+bY is normal
independent X∼Po(λ),Y∼Po(μ)X\sim Po(\lambda),Y\sim Po(\mu) X+Y∼Po(λ+μ)X+Y\sim Po(\lambda+\mu)

If independent X∼N(10,4)X\sim N(10,4) and Y∼N(6,9)Y\sim N(6,9), then X−Y∼N(10−6,  4+9)=N(4,13).X-Y\sim N(10-6,\;4+9)=N(4,13). The mean subtracts, but the variances add because the coefficient of YY is squared: (−1)2=1(-1)^2=1.

If independent counts A∼Po(2)A\sim Po(2) and B∼Po(3.5)B\sim Po(3.5) refer to compatible exposure, then A+B∼Po(5.5)A+B\sim Po(5.5). A difference of Poisson variables is not Poisson by this result.

Define the required total, difference or cost as a linear combination; calculate its mean; calculate its variance with squared coefficients and an independence check; identify the resulting distribution; only then standardise or evaluate its probability.

Constants shift expectation but contribute no variance. Never subtract variances for a difference of independent variables, and do not use the independent-variance or Poisson-sum result when dependence is present.

6.3 Continuous random variables

Syllabus
9709–2028–2029
Topic
6.3
Level
A2

Density height becomes probability only through area

Density property on its support II Meaning
f(x)≥0f(x)\ge0 area cannot create negative probability
∫If(x) dx=1\int_I f(x)\,dx=1 all possible values have total probability 1
P(a<X<b)=∫abf(x) dxP(a<X<b)=\int_a^b f(x)\,dx interval probability is area
P(X=x)=0P(X=x)=0 a single point has zero width

State the single support interval and take f(x)=0f(x)=0 outside it. The support may be finite or infinite. Endpoint inclusion makes no difference to a continuous probability.

For f(x)=3/x4f(x)=3/x^4 on x≥1x\ge1, ∫1∞3x−4 dx=1,\int_1^\infty 3x^{-4}\,dx=1, so it is a valid density. Also P(X>2)=∫2∞3x−4 dx=1/8P(X>2)=\int_2^\infty3x^{-4}\,dx=1/8.

A density can exceed 1 over a narrow interval; only its total area must be 1. Density carries inverse units so that integrated probability is unitless.

f(x)f(x) is not P(X=x)P(X=x). Work directly with density areas; explicit cumulative-distribution-function knowledge is not part of this syllabus objective.

One density gives probabilities, moments and percentiles

Goal Direct density calculation
unknown constant set total support area equal to 1
probability integrate f(x)f(x) over the event interval
mean E(X)=∫xf(x) dxE(X)=\int x f(x)\,dx
variance E(X2)=∫x2f(x) dxE(X^2)=\int x^2f(x)\,dx, then Var(X)=E(X2)−[E(X)]2Var(X)=E(X^2)-[E(X)]^2
ppth percentile cc set area to the left of cc equal to pp

Let f(x)=kxf(x)=kx for 0≤x≤20\le x\le2. Normalisation gives 1=∫02kx dx=2k,1=\int_0^2kx\,dx=2k, so k=1/2k=1/2. Therefore P(X>1)=∫12x/2 dx=3/4P(X>1)=\int_1^2x/2\,dx=3/4.

For this density, E(X)=∫02x(x/2) dx=4/3,E(X)=\int_0^2x(x/2)\,dx=4/3, E(X2)=∫02x2(x/2) dx=2,E(X^2)=\int_0^2x^2(x/2)\,dx=2, so Var(X)=2−(4/3)2=2/9.Var(X)=2-(4/3)^2=2/9.

The median mm divides the total area in half: ∫0mx/2 dx=1/2.\int_0^m x/2\,dx=1/2. Hence m2/4=1/2m^2/4=1/2 and, since mm lies in the support, m=2m=\sqrt2. Other percentiles use the same direct-area equation.

Normalise before every later calculation. A median is found from accumulated area, not from f(m)=0.5f(m)=0.5; explicit CDF notation is neither needed nor included.

6.4 Sampling and estimation

Syllabus
9709–2028–2029
Topic
6.4
Level
A2

A population is the target group while a sample is the observed subset

A population contains all units of interest; a sample is selected to estimate population characteristics. A parameter describes the population, while a statistic is calculated from the sample.

Define the target population before sampling and consider coverage, non-response and selection bias. Larger samples reduce random error but do not automatically remove systematic bias.

A survey of 500 randomly chosen voters is a sample; the proportion supporting a policy in all eligible voters is a population parameter.

A sample statistic is not the parameter itself, and a large biased sample can still mislead.

Random numbers select from a frame; they cannot repair a bad frame

Random-number sampling step Action
frame list and number every covered population unit 1,…,N1,\ldots,N
generate use a random-number generator/table with a fixed reading rule
screen ignore values outside 1,…,N1,\ldots,N and repeated values when sampling without replacement
continue select until the required sample size is reached
Unsatisfactory feature Why results may mislead
convenience/voluntary participation selected people may differ systematically from the target
incomplete frame uncovered groups have no chance of selection
non-response responders may differ from non-responders
leading question or faulty measurement responses are shifted even if selection was random

Choosing random numbers from a school email list gives random selection from that list, but it cannot represent pupils omitted from the list. Explain the missing group and the likely direction of distortion when context supports it.

The syllabus asks for simple criticism and elementary random-number use. Knowledge of named schemes such as quota or stratified sampling is not required.

‘Random’ is not a cure-all: it randomises selection within the sampling frame but does not remove frame, non-response or measurement bias.

The sample mean estimates a population mean and has its own sampling distribution

For independent observations Xᵢ with mean μ and variance σ², the sample mean X̄ has E(X̄)=μ and Var(X̄)=σ²/n. Its standard error is σ/√n (or estimated with s/√n).

Distinguish variability of individual observations from variability of means, and account for finite-population or dependence conditions when relevant.

If σ=12 and n=36, the standard error of X̄ is 2, even though individual values vary with standard deviation 12.

Increasing n reduces standard error by √n, not by n, and does not necessarily reduce measurement bias.

A normal population gives an exactly normal sample mean

IfindependentobservationscomefromIf independent observations come fromX\sim N(\mu,\sigma^2),then for every sample size $n$,\bar X\sim N\left(\mu,\frac{\sigma^2}{n}\right).Thisisexact,notalarge−sampleapproximation.This is exact, not a large-sample approximation.

For a sample-mean boundary $a$, standardise withZ=\frac{\bar X-\mu}{\sigma/\sqrt n}.

If individual values are N(100,152)N(100,15^2) and n=25n=25, then Xˉ∼N(100,9)\bar X\sim N(100,9). Thus P(Xˉ>106)=P(Z>106−1003)=P(Z>2).P(\bar X>106)=P\left(Z>\frac{106-100}{3}\right)=P(Z>2).

Do not require large nn when the population is normal. The mean stays μ\mu, but the variance becomes σ2/n\sigma^2/n; use σ/n\sigma/\sqrt n, not σ\sigma, in the zz-score. CLT belongs to the next objective.

CLT makes a large-sample mean approximately normal

For a large independent random sample from a population with finite mean $\mu$ and variance $\sigma^2$, the Central Limit Theorem gives\bar X\approx N\left(\mu,\frac{\sigma^2}{n}\right).Onlyaninformalunderstandingofthisresultisrequired.Only an informal understanding of this result is required.

Population Distribution of Xˉ\bar X
normal exactly normal for every nn
not normal / shape unspecified approximately normal when nn is sufficiently large

A more skewed or heavy-tailed population generally needs a larger sample before the approximation is convincing. Randomness and independence still matter; CLT does not remove sampling bias or dependence.

A non-normal population has mean 10 and standard deviation 4. For an independent sample of size 64, use Xˉ≈N(10,0.25)\bar X\approx N(10,0.25), so P(Xˉ>11)≈P(Z>11−100.5)=P(Z>2).P(\bar X>11)\approx P\left(Z>\frac{11-10}{0.5}\right)=P(Z>2).

CLT says the distribution of the sample mean becomes approximately normal; it does not say the individual observations or the population become normal.

Use $n-1$ for the unbiased variance estimate

Data form Unbiased mean estimate Unbiased variance estimate
raw x1,…,xnx_1,\ldots,x_n xˉ=∑x/n\bar x=\sum x/n s2={∑x2−(∑x)2/n}/(n−1)s^2=\{\sum x^2-(\sum x)^2/n\}/(n-1)
frequencies ff xˉ=∑fx/∑f\bar x=\sum fx/\sum f s2={∑fx2−(∑fx)2/n}/(n−1)s^2=\{\sum fx^2-(\sum fx)^2/n\}/(n-1), n=∑fn=\sum f

Equivalently,Equivalently,s^2=\frac{\sum (x-\bar x)^2}{n-1}.The denominator $n-1$ corrects the downward bias caused by estimating $\mu$ with the same sample mean.

For data 2,4,62,4,6, xˉ=4\bar x=4 and ∑(x−xˉ)2=4+0+4=8\sum(x-\bar x)^2=4+0+4=8, so the unbiased variance estimate is s2=8/(3−1)=4s^2=8/(3-1)=4.

Unbiased means the estimation procedure gives the true population value on average over repeated random samples. One particular estimate can still be above or below the parameter.

Do not divide by nn when the question asks for an unbiased population-variance estimate, and do not report ss when variance s2s^2 is requested.

A mean confidence interval is estimate plus/minus normal margin

Allowed case Standard error used
normal population, known variance σ2\sigma^2 σ/n\sigma/\sqrt n
large sample s/ns/\sqrt n when σ\sigma is unknown, using the sample estimate

A $100(1-\alpha)\%$ interval is\bar x\pm z_{1-\alpha/2}\times\operatorname{SE}(\bar X).Common two-sided values are $1.645$ for 90%, $1.96$ for 95% and $2.576$ for 99%.

A normal population has known σ=12\sigma=12. From n=36n=36 observations, xˉ=50\bar x=50. A 95% interval is 50±1.96(126)=50±3.92,50\pm1.96\left(\frac{12}{6}\right)=50\pm3.92, giving (46.08,53.92)(46.08,53.92).

Interpret in context: we are 95% confident that the population mean lies between 46.08 and 53.92. Repeated use of this method captures the fixed population mean in about 95% of intervals.

Use the standard error, not the population spread itself. The Cambridge 9709 scope here uses the normal/large-sample cases stated above; do not introduce a t distribution unless another specification explicitly asks for it.

Estimate a population proportion with a large-sample normal interval

From $x$ successes in a large sample of size $n$,\hat p=\frac{x}{n},\qquad \widehat{\operatorname{SE}}(\hat p)=\sqrt{\frac{\hat p(1-\hat p)}{n}}.Checkthattheestimatedsuccessandfailurecountsarebothsufficientlylargeforthenormalapproximation.Check that the estimated success and failure counts are both sufficiently large for the normal approximation.

A $100(1-\alpha)\%$ approximate interval is\hat p\pm z_{1-\alpha/2}\sqrt{\frac{\hat p(1-\hat p)}{n}}.

If 40 of 100 sampled people support a proposal, p^=0.40\hat p=0.40 and SE≈0.4(0.6)/100=0.0490SE\approx\sqrt{0.4(0.6)/100}=0.0490. A 95% interval is 0.40±1.96(0.0490)=0.40±0.096,0.40\pm1.96(0.0490)=0.40\pm0.096, giving approximately (0.304,0.496)(0.304,0.496).

We are 95% confident that the population proportion supporting the proposal lies between about 30.4% and 49.6%, assuming the sample was random and the large-sample approximation is appropriate.

Do not stop after finding the standard error or treat the interval as a range for individual 0/1 responses. The interval estimates one population proportion.

6.5 Hypothesis tests

Syllabus
9709–2028–2029
Topic
6.5
Level
A2

The alternative hypothesis chooses the evidence tail

Term Role in the test
null hypothesis H0H_0 parameter claim used to calculate the reference distribution
alternative H1H_1 upper, lower or two-sided claim that selects the tail(s)
test statistic sample evidence compared with the null distribution
significance level α\alpha maximum/target long-run Type I error rate
rejection (critical) region values sufficiently extreme to reject H0H_0
acceptance region values for which H0H_0 is not rejected
Alternative Test form
parameter >> null value upper-tailed
parameter << null value lower-tailed
parameter ≠\ne null value two-tailed

Calculate a null tail probability (p-value) or compare the statistic with a pre-set rejection region. Reject H0H_0 when the p-value is at most α\alpha or the statistic lies in the rejection region; otherwise do not reject H0H_0.

A two-sided p-value of 0.04 gives sufficient evidence against H0H_0 at 5% but not at 1%. The final statement must name the contextual claim in H1H_1, such as evidence that a mean has changed.

A p-value is not P(H0 is true)P(H_0\text{ is true}). Failure to reject means the evidence is insufficient at the chosen level; it does not prove H0H_0 or justify saying ‘accept H0H_0’ as certainty.

Test a binomial or Poisson count under its null model

Step Discrete-count test
hypotheses state H0:p=p0H_0:p=p_0 or H0:λ=λ0H_0:\lambda=\lambda_0 and a contextual one/two-sided H1H_1
null model use B(n,p0)B(n,p_0) or Po(λ0)Po(\lambda_0), not an observed estimate
extremity include the observed value and outcomes at least as extreme in the H1H_1 direction
calculation evaluate directly, or use a suitable normal approximation with continuity correction
decision compare tail probability/critical region with α\alpha and conclude in context

To test whether a coin has probability greater than 0.50.5 of heads, use H0:p=0.5H_0:p=0.5, H1:p>0.5H_1:p>0.5. If 15 heads occur in 20 tosses, the direct p-value is P(X≥15∣X∼B(20,0.5))=∑r=1520(20r)(0.5)20.P(X\ge15\mid X\sim B(20,0.5))=\sum_{r=15}^{20}\binom{20}{r}(0.5)^{20}. Compare this value with the stated significance level.

For a suitable binomial null model use N(np0,np0(1−p0))N(np_0,np_0(1-p_0)); for a suitable Poisson null model use N(λ0,λ0)N(\lambda_0,\lambda_0). Apply continuity correction to the observed/critical integer boundary before standardising.

For a two-tailed discrete test, construct both extreme regions under H0H_0 so their combined null probability does not exceed the significance level. Discreteness may make the actual significance smaller than the nominal level.

Do not use p^\hat p or the observed rate inside the null distribution. Exact binomial/Poisson calculations need no continuity correction; the half-unit correction belongs only to a normal approximation.

Test a population mean with the null standard error

Syllabus case Null standard error
normal population, known σ2\sigma^2 σ/n\sigma/\sqrt n
large sample, σ\sigma unknown s/ns/\sqrt n using the sample estimate

For $H_0:\mu=\mu_0$, useZ=\frac{\bar X-\mu_0}{\operatorname{SE}(\bar X)}.The alternative $\mu>\mu_0$, $\mu<\mu_0$ or $\mu\ne\mu_0$ selects the upper, lower or two-tailed probability.

A normal population has known σ=10\sigma=10. Test H0:μ=50H_0:\mu=50 against H1:μ>50H_1:\mu>50 using n=100n=100 and xˉ=51.8\bar x=51.8. Then SE=1SE=1 and z=1.8z=1.8, giving upper-tail p-value P(Z≥1.8)≈0.0359P(Z\ge1.8)\approx0.0359. Reject at 5% and conclude there is evidence that the population mean exceeds 50.

The same z=1.8z=1.8 would not reject a two-sided test at 5%, because the two-sided 5% critical values are approximately ±1.96\pm1.96. Always formulate H1H_1 before using the data.

Divide by the standard error, not by σ\sigma itself. This syllabus objective uses the stated normal-known-variance or large-sample normal method; do not introduce a t distribution.

Type I rejects a true null; Type II misses a false null

True state Reject H0H_0 Do not reject H0H_0
H0H_0 true Type I error correct decision
H0H_0 false correct decision Type II error

\alpha=P(\text{reject }H_0\mid H_0\text{ true}),\beta=P(\text{do not reject }H_0\mid\text{specified alternative is true}).

Type I is a false positive: evidence is declared when the null state is actually true. Type II is a false negative: the test fails to detect a specified departure from the null.

For a fixed test rule, the Type I probability is evaluated under the null parameter. The Type II probability depends on which alternative parameter value is actually true; it is not one universal number.

Do not reverse the labels and do not say a Type II error means ‘accepting a true null’. It means not rejecting a null that is false.

Keep the rejection rule fixed, then change the true model

Error probability Keep fixed Evaluate under Probability region
Type I rejection rule null parameter rejection region
Type II at a stated alternative same rejection rule stated alternative parameter acceptance region

Suppose X1,…,X25X_1,\ldots,X_{25} are normal with known σ=10\sigma=10, and a 5% upper-tail test of H0:μ=50H_0:\mu=50 rejects when Xˉ>53.29\bar X>53.29 (since 50+1.645(10/5)=53.2950+1.645(10/5)=53.29). Under H0H_0, P(Type I)=P(Xˉ>53.29∣μ=50)≈0.05.P(\text{Type I})=P(\bar X>53.29\mid\mu=50)\approx0.05.

If the true mean is instead 54, the same fixed rule makes a Type II error when Xˉ≤53.29\bar X\le53.29. Hence P(Type II at μ=54)=P(Z≤53.29−542)=P(Z≤−0.355)≈0.361.P(\text{Type II at }\mu=54)=P\left(Z\le\frac{53.29-54}{2}\right)=P(Z\le-0.355)\approx0.361.

For a binomial rule ‘reject H0:p=p0H_0:p=p_0 when X≥cX\ge c’, Type I is Pp0(X≥c)P_{p_0}(X\ge c) and Type II at p=p1p=p_1 is Pp1(X<c)P_{p_1}(X<c). For a Poisson test, use the same regions with Po(λ0)Po(\lambda_0) and the stated alternative Po(λ1)Po(\lambda_1). Evaluate these discrete probabilities directly when required.

Do not move the critical boundary after switching to the alternative distribution. Type II uses the acceptance region, and its probability must name the alternative parameter value.