6. Probability & Statistics 2
- Syllabus
- 9709–2028–2029
- Section
- 6
- Level
- A2
If events occur independently at constant mean rate λ per interval, X~Po(λ) and P(X=r)=e^{−λ}λ^r/r!.
Match λ to the interval length, check that events are countable and rare enough for the model, and use complements for “at least one” questions.
If a call centre averages 3 calls per minute, P(2 calls in one minute)=e^{−3}3²/2.
Changing the interval changes λ proportionally; it is not a universal parameter for every time window.
For X~Po(λ), E(X)=λ and Var(X)=λ. For a time or area interval scaled by k, the mean becomes kλ under a constant-rate model.
Use the equality as a model check, not as a statement that every sample has equal mean and variance. Estimate λ from appropriate exposure.
If the observed average is 4 events per hour, a two-hour interval has Po(8), not Po(4).
Sample variance need not equal sample mean exactly; the equality describes the distributional parameter.
| Check | Poisson-model requirement |
|---|---|
| response | a count 0,1,2,… in fixed time, length, area or volume |
| rate | constant average rate over the exposure |
| occurrence | events occur independently |
| simultaneity | in a very small exposure, more than one event is negligible |
If the average is $r$ events per unit exposure, then an exposure of size $t$ usesX\sim Po(rt).
If flaws occur randomly at an average of 0.8 per metre, the number in 5 metres can be modelled by Po(4), provided the rate is stable and flaws do not cluster.
Question the model when the rate varies systematically, one event triggers another, observations are capped by a fixed number of trials, or the exposure itself is unclear. State the modelling assumption when context makes it relevant.
A count variable is not automatically Poisson. The theorem that independent Poisson variables add belongs to linear combinations and does not replace the model-suitability checks here.
For $X\sim B(n,p)$, the syllabus guide isn>50\quad\text{and}\quad np<5(approximately).Whensuitable,useX\approx Y,\qquad Y\sim Po(\lambda),\quad\lambda=np.
| Step | Action |
|---|---|
| 1 | check large n and small p using the stated guide |
| 2 | calculate λ=np |
| 3 | keep the same integer event (=,≤,≥) |
| 4 | calculate with the Poisson formula/table; no continuity correction |
If X∼B(200,0.01), then n>50 and np=2<5, so use Y∼Po(2). Thus P(X≥2)≈P(Y≥2)=1−P(Y=0)−P(Y=1)=1−e−2(1+2).
The approximation replaces many rare independent success opportunities by a random-event count with the same expected value np.
Small p alone is insufficient. Do not use a half-unit continuity correction: both binomial and Poisson variables are discrete.
For $X\sim Po(\lambda)$, when $\lambda$ is large (the syllabus guide is approximately $\lambda>15$), useY\sim N(\lambda,\lambda).Hence the normal standard deviation is $\sqrt{\lambda}$.
| Poisson event | Corrected normal event |
|---|---|
| X≤k | Y<k+0.5 |
| X≥k | Y>k−0.5 |
| a≤X≤b | a−0.5<Y<b+0.5 |
If X∼Po(25), the guide condition holds and Y∼N(25,25). Therefore P(X≥30)≈P(Y>29.5)=P(Z>529.5−25)=P(Z>0.9).
State the large-mean check, write the approximating normal distribution, correct the integer boundary, then standardise with λ.
Do not mix in binomial conditions. For a Poisson variable both normal mean and variance are λ, but the denominator in the z-score is λ.
| Quantity | Result | Condition |
|---|---|---|
| E(aX+b) | aE(X)+b | always |
| Var(aX+b) | a2Var(X) | always |
| E(aX+bY) | aE(X)+bE(Y) | always |
| Var(aX+bY) | a2Var(X)+b2Var(Y) | X,Y independent |
| Starting distributions | Resulting distribution |
|---|---|
| X normal | aX+b is normal |
| independent normal X,Y | aX+bY is normal |
| independent X∼Po(λ),Y∼Po(μ) | X+Y∼Po(λ+μ) |
If independent X∼N(10,4) and Y∼N(6,9), then X−Y∼N(10−6,4+9)=N(4,13). The mean subtracts, but the variances add because the coefficient of Y is squared: (−1)2=1.
If independent counts A∼Po(2) and B∼Po(3.5) refer to compatible exposure, then A+B∼Po(5.5). A difference of Poisson variables is not Poisson by this result.
Define the required total, difference or cost as a linear combination; calculate its mean; calculate its variance with squared coefficients and an independence check; identify the resulting distribution; only then standardise or evaluate its probability.
Constants shift expectation but contribute no variance. Never subtract variances for a difference of independent variables, and do not use the independent-variance or Poisson-sum result when dependence is present.
| Density property on its support I | Meaning |
|---|---|
| f(x)≥0 | area cannot create negative probability |
| ∫If(x)dx=1 | all possible values have total probability 1 |
| P(a<X<b)=∫abf(x)dx | interval probability is area |
| P(X=x)=0 | a single point has zero width |
State the single support interval and take f(x)=0 outside it. The support may be finite or infinite. Endpoint inclusion makes no difference to a continuous probability.
For f(x)=3/x4 on x≥1, ∫1∞3x−4dx=1, so it is a valid density. Also P(X>2)=∫2∞3x−4dx=1/8.
A density can exceed 1 over a narrow interval; only its total area must be 1. Density carries inverse units so that integrated probability is unitless.
f(x) is not P(X=x). Work directly with density areas; explicit cumulative-distribution-function knowledge is not part of this syllabus objective.
| Goal | Direct density calculation |
|---|---|
| unknown constant | set total support area equal to 1 |
| probability | integrate f(x) over the event interval |
| mean | E(X)=∫xf(x)dx |
| variance | E(X2)=∫x2f(x)dx, then Var(X)=E(X2)−[E(X)]2 |
| pth percentile c | set area to the left of c equal to p |
Let f(x)=kx for 0≤x≤2. Normalisation gives 1=∫02kxdx=2k, so k=1/2. Therefore P(X>1)=∫12x/2dx=3/4.
For this density, E(X)=∫02x(x/2)dx=4/3, E(X2)=∫02x2(x/2)dx=2, so Var(X)=2−(4/3)2=2/9.
The median m divides the total area in half: ∫0mx/2dx=1/2. Hence m2/4=1/2 and, since m lies in the support, m=2. Other percentiles use the same direct-area equation.
Normalise before every later calculation. A median is found from accumulated area, not from f(m)=0.5; explicit CDF notation is neither needed nor included.
A population contains all units of interest; a sample is selected to estimate population characteristics. A parameter describes the population, while a statistic is calculated from the sample.
Define the target population before sampling and consider coverage, non-response and selection bias. Larger samples reduce random error but do not automatically remove systematic bias.
A survey of 500 randomly chosen voters is a sample; the proportion supporting a policy in all eligible voters is a population parameter.
A sample statistic is not the parameter itself, and a large biased sample can still mislead.
| Random-number sampling step | Action |
|---|---|
| frame | list and number every covered population unit 1,…,N |
| generate | use a random-number generator/table with a fixed reading rule |
| screen | ignore values outside 1,…,N and repeated values when sampling without replacement |
| continue | select until the required sample size is reached |
| Unsatisfactory feature | Why results may mislead |
|---|---|
| convenience/voluntary participation | selected people may differ systematically from the target |
| incomplete frame | uncovered groups have no chance of selection |
| non-response | responders may differ from non-responders |
| leading question or faulty measurement | responses are shifted even if selection was random |
Choosing random numbers from a school email list gives random selection from that list, but it cannot represent pupils omitted from the list. Explain the missing group and the likely direction of distortion when context supports it.
The syllabus asks for simple criticism and elementary random-number use. Knowledge of named schemes such as quota or stratified sampling is not required.
‘Random’ is not a cure-all: it randomises selection within the sampling frame but does not remove frame, non-response or measurement bias.
For independent observations Xᵢ with mean μ and variance σ², the sample mean X̄ has E(X̄)=μ and Var(X̄)=σ²/n. Its standard error is σ/√n (or estimated with s/√n).
Distinguish variability of individual observations from variability of means, and account for finite-population or dependence conditions when relevant.
If σ=12 and n=36, the standard error of X̄ is 2, even though individual values vary with standard deviation 12.
Increasing n reduces standard error by √n, not by n, and does not necessarily reduce measurement bias.
IfindependentobservationscomefromX\sim N(\mu,\sigma^2),then for every sample size $n$,\bar X\sim N\left(\mu,\frac{\sigma^2}{n}\right).Thisisexact,notalarge−sampleapproximation.
For a sample-mean boundary $a$, standardise withZ=\frac{\bar X-\mu}{\sigma/\sqrt n}.
If individual values are N(100,152) and n=25, then Xˉ∼N(100,9). Thus P(Xˉ>106)=P(Z>3106−100)=P(Z>2).
Do not require large n when the population is normal. The mean stays μ, but the variance becomes σ2/n; use σ/n, not σ, in the z-score. CLT belongs to the next objective.
For a large independent random sample from a population with finite mean $\mu$ and variance $\sigma^2$, the Central Limit Theorem gives\bar X\approx N\left(\mu,\frac{\sigma^2}{n}\right).Onlyaninformalunderstandingofthisresultisrequired.
| Population | Distribution of Xˉ |
|---|---|
| normal | exactly normal for every n |
| not normal / shape unspecified | approximately normal when n is sufficiently large |
A more skewed or heavy-tailed population generally needs a larger sample before the approximation is convincing. Randomness and independence still matter; CLT does not remove sampling bias or dependence.
A non-normal population has mean 10 and standard deviation 4. For an independent sample of size 64, use Xˉ≈N(10,0.25), so P(Xˉ>11)≈P(Z>0.511−10)=P(Z>2).
CLT says the distribution of the sample mean becomes approximately normal; it does not say the individual observations or the population become normal.
| Data form | Unbiased mean estimate | Unbiased variance estimate |
|---|---|---|
| raw x1,…,xn | xˉ=∑x/n | s2={∑x2−(∑x)2/n}/(n−1) |
| frequencies f | xˉ=∑fx/∑f | s2={∑fx2−(∑fx)2/n}/(n−1), n=∑f |
Equivalently,s^2=\frac{\sum (x-\bar x)^2}{n-1}.The denominator $n-1$ corrects the downward bias caused by estimating $\mu$ with the same sample mean.
For data 2,4,6, xˉ=4 and ∑(x−xˉ)2=4+0+4=8, so the unbiased variance estimate is s2=8/(3−1)=4.
Unbiased means the estimation procedure gives the true population value on average over repeated random samples. One particular estimate can still be above or below the parameter.
Do not divide by n when the question asks for an unbiased population-variance estimate, and do not report s when variance s2 is requested.
| Allowed case | Standard error used |
|---|---|
| normal population, known variance σ2 | σ/n |
| large sample | s/n when σ is unknown, using the sample estimate |
A $100(1-\alpha)\%$ interval is\bar x\pm z_{1-\alpha/2}\times\operatorname{SE}(\bar X).Common two-sided values are $1.645$ for 90%, $1.96$ for 95% and $2.576$ for 99%.
A normal population has known σ=12. From n=36 observations, xˉ=50. A 95% interval is 50±1.96(612)=50±3.92, giving (46.08,53.92).
Interpret in context: we are 95% confident that the population mean lies between 46.08 and 53.92. Repeated use of this method captures the fixed population mean in about 95% of intervals.
Use the standard error, not the population spread itself. The Cambridge 9709 scope here uses the normal/large-sample cases stated above; do not introduce a t distribution unless another specification explicitly asks for it.
From $x$ successes in a large sample of size $n$,\hat p=\frac{x}{n},\qquad \widehat{\operatorname{SE}}(\hat p)=\sqrt{\frac{\hat p(1-\hat p)}{n}}.Checkthattheestimatedsuccessandfailurecountsarebothsufficientlylargeforthenormalapproximation.
A $100(1-\alpha)\%$ approximate interval is\hat p\pm z_{1-\alpha/2}\sqrt{\frac{\hat p(1-\hat p)}{n}}.
If 40 of 100 sampled people support a proposal, p^=0.40 and SE≈0.4(0.6)/100=0.0490. A 95% interval is 0.40±1.96(0.0490)=0.40±0.096, giving approximately (0.304,0.496).
We are 95% confident that the population proportion supporting the proposal lies between about 30.4% and 49.6%, assuming the sample was random and the large-sample approximation is appropriate.
Do not stop after finding the standard error or treat the interval as a range for individual 0/1 responses. The interval estimates one population proportion.
| Term | Role in the test |
|---|---|
| null hypothesis H0 | parameter claim used to calculate the reference distribution |
| alternative H1 | upper, lower or two-sided claim that selects the tail(s) |
| test statistic | sample evidence compared with the null distribution |
| significance level α | maximum/target long-run Type I error rate |
| rejection (critical) region | values sufficiently extreme to reject H0 |
| acceptance region | values for which H0 is not rejected |
| Alternative | Test form |
|---|---|
| parameter > null value | upper-tailed |
| parameter < null value | lower-tailed |
| parameter = null value | two-tailed |
Calculate a null tail probability (p-value) or compare the statistic with a pre-set rejection region. Reject H0 when the p-value is at most α or the statistic lies in the rejection region; otherwise do not reject H0.
A two-sided p-value of 0.04 gives sufficient evidence against H0 at 5% but not at 1%. The final statement must name the contextual claim in H1, such as evidence that a mean has changed.
A p-value is not P(H0 is true). Failure to reject means the evidence is insufficient at the chosen level; it does not prove H0 or justify saying ‘accept H0’ as certainty.
| Step | Discrete-count test |
|---|---|
| hypotheses | state H0:p=p0 or H0:λ=λ0 and a contextual one/two-sided H1 |
| null model | use B(n,p0) or Po(λ0), not an observed estimate |
| extremity | include the observed value and outcomes at least as extreme in the H1 direction |
| calculation | evaluate directly, or use a suitable normal approximation with continuity correction |
| decision | compare tail probability/critical region with α and conclude in context |
To test whether a coin has probability greater than 0.5 of heads, use H0:p=0.5, H1:p>0.5. If 15 heads occur in 20 tosses, the direct p-value is P(X≥15∣X∼B(20,0.5))=r=15∑20(r20)(0.5)20. Compare this value with the stated significance level.
For a suitable binomial null model use N(np0,np0(1−p0)); for a suitable Poisson null model use N(λ0,λ0). Apply continuity correction to the observed/critical integer boundary before standardising.
For a two-tailed discrete test, construct both extreme regions under H0 so their combined null probability does not exceed the significance level. Discreteness may make the actual significance smaller than the nominal level.
Do not use p^ or the observed rate inside the null distribution. Exact binomial/Poisson calculations need no continuity correction; the half-unit correction belongs only to a normal approximation.
| Syllabus case | Null standard error |
|---|---|
| normal population, known σ2 | σ/n |
| large sample, σ unknown | s/n using the sample estimate |
For $H_0:\mu=\mu_0$, useZ=\frac{\bar X-\mu_0}{\operatorname{SE}(\bar X)}.The alternative $\mu>\mu_0$, $\mu<\mu_0$ or $\mu\ne\mu_0$ selects the upper, lower or two-tailed probability.
A normal population has known σ=10. Test H0:μ=50 against H1:μ>50 using n=100 and xˉ=51.8. Then SE=1 and z=1.8, giving upper-tail p-value P(Z≥1.8)≈0.0359. Reject at 5% and conclude there is evidence that the population mean exceeds 50.
The same z=1.8 would not reject a two-sided test at 5%, because the two-sided 5% critical values are approximately ±1.96. Always formulate H1 before using the data.
Divide by the standard error, not by σ itself. This syllabus objective uses the stated normal-known-variance or large-sample normal method; do not introduce a t distribution.
| True state | Reject H0 | Do not reject H0 |
|---|---|---|
| H0 true | Type I error | correct decision |
| H0 false | correct decision | Type II error |
\alpha=P(\text{reject }H_0\mid H_0\text{ true}),\beta=P(\text{do not reject }H_0\mid\text{specified alternative is true}).
Type I is a false positive: evidence is declared when the null state is actually true. Type II is a false negative: the test fails to detect a specified departure from the null.
For a fixed test rule, the Type I probability is evaluated under the null parameter. The Type II probability depends on which alternative parameter value is actually true; it is not one universal number.
Do not reverse the labels and do not say a Type II error means ‘accepting a true null’. It means not rejecting a null that is false.
| Error probability | Keep fixed | Evaluate under | Probability region |
|---|---|---|---|
| Type I | rejection rule | null parameter | rejection region |
| Type II at a stated alternative | same rejection rule | stated alternative parameter | acceptance region |
Suppose X1,…,X25 are normal with known σ=10, and a 5% upper-tail test of H0:μ=50 rejects when Xˉ>53.29 (since 50+1.645(10/5)=53.29). Under H0, P(Type I)=P(Xˉ>53.29∣μ=50)≈0.05.
If the true mean is instead 54, the same fixed rule makes a Type II error when Xˉ≤53.29. Hence P(Type II at μ=54)=P(Z≤253.29−54)=P(Z≤−0.355)≈0.361.
For a binomial rule ‘reject H0:p=p0 when X≥c’, Type I is Pp0(X≥c) and Type II at p=p1 is Pp1(X<c). For a Poisson test, use the same regions with Po(λ0) and the stated alternative Po(λ1). Evaluate these discrete probabilities directly when required.
Do not move the critical boundary after switching to the alternative distribution. Type II uses the acceptance region, and its probability must name the alternative parameter value.