Unit S2: Statistics 2

Syllabus
2019
Section
—
Level
A2

S2.1 - The Binomial and Poisson distributions

Syllabus
2019
Topic
S2.1
Level
A2

Choose and use Binomial or Poisson

Binomial and Poisson distributions both model counts, but their random processes differ. Choose the model from the conditions before calculating a probability, and keep the count interval and event inequality explicit.

Model Conditions and parameter
X∼B(n,p)X\sim B(n,p) a fixed number nn of independent trials; each trial has success/failure outcomes; success probability pp is constant
X∼Po(λ)X\sim Po(\lambda) events occur independently and singly at a constant average rate in a fixed interval; λ\lambda is the mean count in that interval

PB(X=x)=(nx)px(1−p)n−x,PPo(X=x)=e−λλxx!P_{B}(X=x)=\binom nxp^x(1-p)^{n-x},\qquad P_{Po}(X=x)=e^{-\lambda}\frac{\lambda^x}{x!}

Translate words into integer events before using a formula or cumulative table: 'fewer than 3' is P(X≤2)P(X\le2), 'at least 6' is 1−P(X≤5)1-P(X\le5), and 4<X<104<X<10 is P(X≤9)−P(X≤4)P(X\le9)-P(X\le4). For X∼Po(4.5)X\sim Po(4.5), the last probability is 0.9829−0.5321=0.45080.9829-0.5321=0.4508.

Poisson means scale with exposure. If complaints occur at mean 3 per day, then a 7-day total has distribution Po(21)Po(21). More generally, independent Poisson counts add: if X∼Po(λ1)X\sim Po(\lambda_1) and Y∼Po(λ2)Y\sim Po(\lambda_2) independently, then X+Y∼Po(λ1+λ2)X+Y\sim Po(\lambda_1+\lambda_2).

A calculation is only as sound as the model. Repeated trials with changing success probabilities are not binomial; clustering, dependence or a changing rate undermines Poisson. State assumptions in context—such as failures being independent with constant probability, or events occurring independently at a constant average rate—and comment when the situation contradicts them.

Use Binomial and Poisson moments

The mean gives the long-run expected count and the variance measures its spread. Binomial and Poisson models have characteristic mean–variance relationships that can also identify unknown parameters or test plausibility.

Distribution Mean Variance
X∼B(n,p)X\sim B(n,p) npnp np(1−p)np(1-p)
X∼Po(λ)X\sim Po(\lambda) λ\lambda λ\lambda

If a binomial variable has mean 90 and variance 36, thennp=90,np(1−p)=36.np=90,\qquad np(1-p)=36.Dividing the second equation by the first gives 1−p=36/90=0.41-p=36/90=0.4, so p=0.6p=0.6 and n=90/0.6=150n=90/0.6=150. Check that nn is a positive integer and 0≤p≤10\le p\le1.

For Y∼B(20,0.3)Y\sim B(20,0.3), E(Y)=6E(Y)=6 and Var⁡(Y)=4.2\operatorname{Var}(Y)=4.2. For W∼Po(4)W\sim Po(4), both the mean and variance are 4. An observed count dataset whose sample mean is far from its sample variance may therefore cast doubt on a Poisson model, although closeness alone does not prove all Poisson assumptions.

Do not interchange variance and standard deviation, or assume every count with equal mean and variance must be Poisson. The specification requires use of these moment formulae but not their derivations.

Approximate a Binomial count with Poisson

When X∼B(n,p)X\sim B(n,p) has many trials and a small success probability, the rare-success count can be approximated by a Poisson variable with the same mean.

X∼B(n,p) ≈ Y∼Po(λ),λ=npX\sim B(n,p)\ \approx\ Y\sim Po(\lambda),\qquad \lambda=np

First check that the binomial conditions hold and that the context really describes many independent opportunities for a rare event. Then calculate λ=np\lambda=np, preserve the original integer event exactly, and evaluate it with the Poisson formula or table.

If X∼B(200,0.012)X\sim B(200,0.012), use Y∼Po(2.4)Y\sim Po(2.4). ThenP(X≤1)≈P(Y≤1)=e−2.4(1+2.4)=0.3084.P(X\le1)\approx P(Y\le1)=e^{-2.4}(1+2.4)=0.3084.For 'at least 2', use the complementary Poisson event 1−P(Y≤1)1-P(Y\le1).

This is a discrete-to-discrete approximation, so no continuity correction is used. It becomes less convincing when the success probability is not small or trials are dependent; use the exact binomial distribution when the approximation is not justified. Normal approximations and continuity correction belong to S2.3, not this objective.

S2.2 - Continuous random variables

Syllabus
2019
Topic
S2.2
Level
A2

Understand a continuous random variable

A continuous random variable can take any real value in an interval, so it models a measurement rather than a count. Waiting time, mass and length are typical examples; the number of calls is discrete even when calls arrive through time.

Probability is assigned to intervals, not isolated values. For every particular number aa, P(X=a)=0P(X=a)=0. Consequently, endpoint choices do not change a continuous probability:P(a<X<b)=P(a≤X<b)=P(a<X≤b)=P(a≤X≤b).P(a<X<b)=P(a\le X<b)=P(a<X\le b)=P(a\le X\le b).

This does not say that observing a particular rounded value is impossible. A recorded value such as 2.3 seconds represents a measurement interval determined by the instrument's precision; the exact mathematical point still has probability zero.

Do not list a probability for each possible value as for a discrete distribution, and do not infer that every variable written with decimals is continuous. The random mechanism—measurement across a continuum versus counting separate outcomes—decides the model.

Use density and cumulative distribution functions

A probability density function ff describes probability through area. It must satisfy f(x)≥0f(x)\ge0 on its support and have total area 1. For a continuous random variable,P(a<X≤b)=∫abf(x) dx.P(a<X\le b)=\int_a^b f(x)\,dx.The height f(x)f(x) is not itself the probability P(X=x)P(X=x), which is zero.

F(x0)=P(X≤x0)=∫−∞x0f(t) dtF(x_0)=P(X\le x_0)=\int_{-\infty}^{x_0} f(t)\,dt

A cumulative distribution function is non-decreasing, remains between 0 and 1, and runs from 0 below the support to 1 above it. Once FF is known, interval probabilities are differences:P(a<X≤b)=F(b)−F(a),P(X>b)=1−F(b).P(a<X\le b)=F(b)-F(a),\qquad P(X>b)=1-F(b).

For a simple piecewise polynomial densityf(x)={x,0≤x≤1,2−x,1<x≤2,0,otherwise,f(x)=\begin{cases}x,&0\le x\le1,\\2-x,&1<x\le2,\\0,&\text{otherwise},\end{cases}both pieces are non-negative and their two triangular areas total 1. Accumulating area from the left givesF(x)={0,x<0,x2/2,0≤x≤1,2x−x2/2−1,1<x≤2,1,x>2.F(x)=\begin{cases}0,&x<0,\\x^2/2,&0\le x\le1,\\2x-x^2/2-1,&1<x\le2,\\1,&x>2.\end{cases}The second middle expression starts with the area already accumulated up to 1; integrating that piece as though the accumulation restarted at zero would make the CDF jump or finish at the wrong value.

For every piecewise model, check non-negativity, total area 1, matching CDF values at internal boundaries, and final value 1. Conditional probabilities still use the ordinary ratio of events, with each probability found from density areas or CDF differences.

Move between a CDF and a density

Where the cumulative distribution function is differentiable, its rate of increase is the density:f(x)=dF(x)dx.f(x)=\frac{dF(x)}{dx}.In the other direction, integrating the density from the lower end of the support reconstructs FF.

SupposeF(x)={0,x<0,x2/4,0≤x<2,1,x≥2.F(x)=\begin{cases}0,&x<0,\\x^2/4,&0\le x<2,\\1,&x\ge2.\end{cases}Differentiating each interval givesf(x)={x/2,0<x<2,0,otherwise.f(x)=\begin{cases}x/2,&0<x<2,\\0,&\text{otherwise}.\end{cases}Its integral over 0<x<20<x<2 is 1, so the result is a valid density.

Because a CDF cannot decrease, a candidate derivative must not be negative. Corners at piece boundaries may make the derivative undefined at isolated points, but changing a density at finitely many individual points changes no continuous probability.

Differentiate the CDF piece by piece and include zero outside the support. Do not differentiate the constants 0 and 1 into extra support, and do not treat a jump as acceptable here: a continuous random variable has a continuous CDF.

Calculate continuous means and variances

For a continuous random variable, moments are density-weighted integrals across the whole support. Split the integral wherever the density formula changes.

E(X)=∫−∞∞xf(x) dx,E(X2)=∫−∞∞x2f(x) dx,Var⁡(X)=E(X2)−[E(X)]2E(X)=\int_{-\infty}^{\infty}x f(x)\,dx,\qquad E(X^2)=\int_{-\infty}^{\infty}x^2 f(x)\,dx,\qquad \operatorname{Var}(X)=E(X^2)-[E(X)]^2

If f(x)=x/2f(x)=x/2 for 0<x<20<x<2 and is zero otherwise, thenE(X)=∫02x22 dx=43,E(X)=\int_0^2\frac{x^2}{2}\,dx=\frac43,E(X2)=∫02x32 dx=2,E(X^2)=\int_0^2\frac{x^3}{2}\,dx=2,soVar⁡(X)=2−(43)2=29.\operatorname{Var}(X)=2-\left(\frac43\right)^2=\frac29.The variance is non-negative and has squared units.

For constants aa and bb, use linear-transformation rules rather than reintegrating:E(aX+b)=aE(X)+b,Var⁡(aX+b)=a2Var⁡(X).E(aX+b)=aE(X)+b,\qquad \operatorname{Var}(aX+b)=a^2\operatorname{Var}(X).A shift changes the mean but not the variance.

Do not calculate variance as ∫(x−E(X))f(x) dx\int(x-E(X))f(x)\,dx; that integral is zero. Use either ∫(x−E(X))2f(x) dx\int(x-E(X))^2f(x)\,dx or the stated E(X2)−[E(X)]2E(X^2)-[E(X)]^2 identity, and distinguish variance from standard deviation.

Find mode, median and quartiles

Mode, median and quartiles describe different features of a continuous distribution. The mode is located from density height; the median and quartiles are located from cumulative probability.

Measure Continuous-distribution condition
mode value where f(x)f(x) is greatest; compare stationary points and support endpoints
lower quartile Q1Q_1 F(Q1)=0.25F(Q_1)=0.25
median mm F(m)=0.50F(m)=0.50
upper quartile Q3Q_3 F(Q3)=0.75F(Q_3)=0.75
interquartile range Q3−Q1Q_3-Q_1

For f(x)=x/2f(x)=x/2 on 0<x<20<x<2, F(x)=x2/4F(x)=x^2/4. HenceQ1=1,m=2,Q3=3,Q_1=1,\qquad m=\sqrt2,\qquad Q_3=\sqrt3,and the interquartile range is 3−1\sqrt3-1. Since the density increases throughout its support, its greatest height is at the upper endpoint, so the modal location is 2.

First decide which piece of a piecewise CDF contains 0.25, 0.5 or 0.75, then solve only within that piece and check the answer lies in its interval. For a mode, differentiating the density can locate interior candidates, but endpoints and any corners must also be compared.

Do not set the density equal to 0.5 to find the median, and do not assume the mean, median and mode coincide. That only happens for particular distribution shapes.

S2.3 - Continuous distributions

Syllabus
2019
Topic
S2.3
Level
A2

Derive and use the continuous uniform distribution

A continuous uniform variable gives equal probability to intervals of equal length. If X∼U(a,b)X\sim U(a,b) with a<ba<b, its density is a rectangle of width b−ab-a. Unit total area fixes its height:

f(x)={1b−a,a≤x≤b,0,otherwise.f(x)=\begin{cases}\dfrac{1}{b-a},&a\le x\le b,\\0,&\text{otherwise}.\end{cases}

Probability is therefore a length ratio. Intersect the requested event with the support first: for a≤c<d≤ba\le c<d\le b,P(c<X<d)=∫cd1b−a dx=d−cb−a.P(c<X<d)=\int_c^d\frac{1}{b-a}\,dx=\frac{d-c}{b-a}.For X∼U(−5,19)X\sim U(-5,19), P(∣X∣>3.5)=[1.5+15.5]/24=17/24P(|X|>3.5)=[1.5+15.5]/24=17/24. Endpoints make no difference for a continuous variable.

Accumulating rectangular area from the left derives the cumulative distribution function:F(x)={0,x<a,x−ab−a,a≤x≤b,1,x>b.F(x)=\begin{cases}0,&x<a,\\\dfrac{x-a}{b-a},&a\le x\le b,\\1,&x>b.\end{cases}The middle piece is linear because each extra unit of xx adds the same area 1/(b−a)1/(b-a).

The mean follows fromE(X)=∫abxb−a dx=b2−a22(b−a)=a+b2.E(X)=\int_a^b\frac{x}{b-a}\,dx=\frac{b^2-a^2}{2(b-a)}=\frac{a+b}{2}.Likewise,E(X2)=b3−a33(b−a)=a2+ab+b23,E(X^2)=\frac{b^3-a^3}{3(b-a)}=\frac{a^2+ab+b^2}{3},soVar⁡(X)=E(X2)−[E(X)]2=(b−a)212.\operatorname{Var}(X)=E(X^2)-[E(X)]^2=\frac{(b-a)^2}{12}.Thus the mean is the rectangle's midpoint, while variance depends only on its width.

Uniform means constant density over a stated interval, not that every exact value has a positive equal probability. Include the zero-density outer pieces, clip transformed event intervals to the support, and use the width b-a rather than b as the denominator.

Apply a Normal approximation with continuity correction

A Normal distribution can approximate a Binomial or Poisson count when the count distribution is sufficiently spread out and not strongly skewed. Match the discrete mean and variance before converting the integer event to a continuous boundary.

Discrete model Approximating model
X∼B(n,p)X\sim B(n,p) Y∼N(np, np(1−p))Y\sim N(np,\,np(1-p))
X∼Po(λ)X\sim Po(\lambda) Y∼N(λ, λ)Y\sim N(\lambda,\,\lambda)
Discrete event Continuity-corrected Normal event
X≤kX\le k Y<k+0.5Y<k+0.5
X<kX<k Y<k−0.5Y<k-0.5
X≥kX\ge k Y>k−0.5Y>k-0.5
X>kX>k Y>k+0.5Y>k+0.5
a≤X≤ba\le X\le b a−0.5<Y<b+0.5a-0.5<Y<b+0.5

Each integer count represents a unit-width bar extending 0.5 on either side of its centre. Moving the Normal boundary to the outer edge of the last included bar preserves approximately the same area. First rewrite words such as 'fewer than 32' as the integer event X ≤ 31; then its corrected boundary is 31.5.

Suppose X∼Po(36)X\sim Po(36) and we need P(X<32)P(X<32). Use Y∼N(36,36)Y\sim N(36,36), so its standard deviation is 6. With the correction,P(X<32)≈P(Y<31.5)=P(Z<31.5−366)=P(Z<−0.75)=0.2266.P(X<32)\approx P(Y<31.5)=P\left(Z<\frac{31.5-36}{6}\right)=P(Z<-0.75)=0.2266.The value 32 is not the boundary because the count 32 is excluded.

For a Binomial example, X∼B(200,0.4)X\sim B(200,0.4) gives Y∼N(80,48)Y\sim N(80,48). ThenP(X≥90)≈P(Y>89.5)=P(Z>89.5−8048).P(X\ge90)\approx P(Y>89.5)=P\left(Z>\frac{89.5-80}{\sqrt{48}}\right).Keep the second Normal parameter as the variance, but divide by its square root when standardising.

Do not apply a continuity correction to an already continuous event, omit it for a discrete-to-Normal approximation, or use a Normal model when a small mean or extreme Binomial probability leaves the count distribution strongly skewed. This objective concerns probability approximation; hypothesis-test decisions belong to the following Topic.

S2.4 - Hypothesis tests

Syllabus
2019
Topic
S2.4
Level
A2

Distinguish a population, census and sample

A statistical investigation begins by defining exactly what is being studied. The population is the complete set of units of interest; a census seeks data from every population unit, while a sample survey collects data from only a subset.

Term Meaning
population every unit about which the investigation aims to draw conclusions
sampling unit one individual member or item that can be selected
sampling frame the operational list or other representation from which units are selected
census investigation intended to include every population unit
sample selected subset used to learn about the population
Approach Main advantages Main limitations
census no sampling variation; detailed information on small subgroups costly and slow; may become outdated; non-response and measurement errors can remain; unsuitable for destructive testing
sample survey quicker and cheaper; can be repeated; permits destructive testing of a limited number conclusions vary from sample to sample; selection bias or an incomplete frame can make it unrepresentative

To study the fill mass of all cans produced during one shift, the population is every can from that shift, one can is a sampling unit, and the production list or numbered stream may provide the frame. Measuring 100 selected cans is a sample survey; measuring every can is a census.

A large sample is not automatically representative, and a census is not automatically error-free. The target population and sampling frame may differ: units missing from the frame cannot be selected, while duplicated entries may be overrepresented.

Understand statistics and sampling distributions

A statistic is a quantity calculated entirely from sample observations. It may estimate or test a population feature, but its formula cannot contain an unknown population parameter. Thus Xˉ\bar X, the sample range and the number of successes are statistics; (X1−μ)/σ(X_1-\mu)/\sigma is not when μ\mu and σ\sigma are unknown.

If the same random-sampling procedure were repeated with a fixed sample size, the statistic would usually change. Its sampling distribution lists every possible value of the statistic together with its probability under the stated population model.

Suppose two independent Bernoulli observations are drawn from a population with success probability 0.50.5, and let TT be the sample proportion of successes. The four equally likely ordered samples giveT=0, 12, 1T=0,\ \frac12,\ 1with probabilitiesP(T=0)=14,P(T=12)=12,P(T=1)=14.P(T=0)=\frac14,\quad P\left(T=\frac12\right)=\frac12,\quad P(T=1)=\frac14.This probability model is the sampling distribution of TT, not the distribution of a single observation.

A test statistic is a chosen statistic whose sampling distribution is known under the null hypothesis. That distribution lets an observed sample value be judged as ordinary or unusually extreme.

Do not confuse a sample distribution—the observed data values in one sample—with a sampling distribution—the distribution of a statistic over all possible samples generated by the same design.

Form and interpret a hypothesis test

A hypothesis test asks whether sample evidence is sufficiently unusual under a stated model to justify rejecting that model in favour of a specific alternative. It can therefore refine a mathematical model, but it cannot prove either hypothesis.

Component Role
null hypothesis H0H_0 precise baseline parameter value used to calculate probabilities, such as p=0.35p=0.35 or λ=8\lambda=8
alternative hypothesis H1H_1 claim supported by departure in the stated direction: <<, >> or ≠\ne
significance level α\alpha maximum chosen probability of rejecting H0H_0 through the critical region when H0H_0 is true

State both hypotheses before using the data. Assume H0, identify the sampling distribution of a suitable test statistic, and calculate the probability of the observed value or values at least as supportive of H1. If this p-value is no greater than alpha—or the statistic lies in the precomputed critical region—reject H0; otherwise do not reject H0.

A historical breakdown rate is 88 per week and a refurbishment is claimed to have changed it. UseH0:λ=8,H1:λ≠8.H_0:\lambda=8,\qquad H_1:\lambda\ne8.If the observed count is outside the two-tailed critical region, do not reject H0H_0: there is insufficient evidence that the mean breakdown rate changed.

Write conclusions in the context and with evidential language. 'Do not reject H0' means the sample was not sufficiently inconsistent with H0; it does not mean H0 has been accepted or shown true.

Construct and use a critical region

A critical region is the set of test-statistic values that cause rejection of the null hypothesis. It is chosen from the sampling distribution under H0 so that its total probability—the actual significance level—does not exceed the stated level and is as close as the discrete distribution permits.

actual significance=P(test statistic lies in the critical region∣H0)\text{actual significance}=P(\text{test statistic lies in the critical region}\mid H_0)

For an upper-tailed test with J∼Po(9)J\sim Po(9) under H0H_0, tables giveP(J≤13)=0.9261,P(J≤14)=0.9585.P(J\le13)=0.9261,\qquad P(J\le14)=0.9585.HenceP(J≥14)=0.0739>0.05,P(J≥15)=0.0415≤0.05.P(J\ge14)=0.0739>0.05,\qquad P(J\ge15)=0.0415\le0.05.The 5% critical region is therefore J≥15J\ge15. An observed value of 15 or more is significant at the 5% level.

Because a discrete tail changes in jumps, the actual significance is often below the nominal level. Moving the boundary one integer further inward here would make the rejection probability too large; moving it outward would be valid but unnecessarily conservative.

State a critical region as values such as J ≥ 15, not as the probability statement P(J ≥ 15). The region contains outcomes; its probability under H0 measures the chance of entering it.

Choose a one-tailed or two-tailed test

The wording of the alternative hypothesis fixes the tail structure before the sample is inspected. A directional claim uses one tail; a claim of any change uses both tails.

Claim Alternative Extreme evidence
parameter has decreased H1:θ<θ0H_1:\theta<\theta_0 small test-statistic values: lower tail
parameter has increased H1:θ>θ0H_1:\theta>\theta_0 large test-statistic values: upper tail
parameter has changed H1:θ≠θ0H_1:\theta\ne\theta_0 unusually small or large values: two tails

For a one-tailed test, the critical-region probability is placed in the stated tail. For a two-tailed test at nominal level alpha, choose lower and upper regions with probabilities as close as practical to alpha/2 while keeping their combined probability no greater than alpha. Discreteness may prevent equal tails or an exact total.

Testing whether a historical proportion 0.350.35 has changed requires H0:p=0.35H_0:p=0.35 and H1:p≠0.35H_1:p\ne0.35, so both unusually few and unusually many successes oppose H0H_0. Testing whether it has fallen uses H1:p<0.35H_1:p<0.35 and only the lower tail.

Do not choose one or two tails after seeing whether the sample result is high or low. That changes the testing rule to favour the observed data and invalidates the stated significance level.

Test Binomial proportions and Poisson means

For a Binomial test, the parameter is a population success probability p; for a Poisson test, it is the mean event rate lambda for a stated exposure. Under H0, that exact parameter determines the sampling distribution used for the test.

Test Distribution under H0H_0
nn independent fixed-probability trials, testing p=p0p=p_0 X∼B(n,p0)X\sim B(n,p_0)
Poisson count over tt base exposure units, testing rate λ=λ0\lambda=\lambda_0 per unit X∼Po(tλ0)X\sim Po(t\lambda_0)

Use this order: (1) define the parameter and state H0 and H1; (2) translate the observed count and direction into the correct tail or tails; (3) calculate an exact tail probability from tables/formulae or construct a critical region; (4) compare with alpha; (5) reject or do not reject H0 and give a contextual conclusion.

Suppose 8 of 40 items fail and the historical failure proportion is 0.350.35. To test whether it is lower, useH0:p=0.35,H1:p<0.35,X∼B(40,0.35).H_0:p=0.35,\qquad H_1:p<0.35,\qquad X\sim B(40,0.35).The lower-tail probability isP(X≤8)=0.0303<0.05.P(X\le8)=0.0303<0.05.Reject H0H_0: there is sufficient evidence at the 5% level that the failure proportion is lower than 0.350.35.

Use an approximation only when suitable, preserving the null model: a rare-event Binomial may use Po(np0)Po(np_0), while a sufficiently spread Binomial or Poisson count may use a matched Normal distribution with continuity correction. For example, six weeks at a null Poisson rate of 6 per week gives total mean 36 and may be approximated by N(36,36)N(36,36).

For a discrete two-tailed test, construct both rejection tails and add their null probabilities to obtain the actual significance. Do not automatically double one observed-tail probability unless that rule is justified by the chosen test; discrete tails are often unequal.

Keep the parameter in the hypotheses, not the observed statistic: write H0:p=p0 or H0:lambda=lambda0, not H0:X=x. A non-significant result is insufficient evidence for H1, not evidence that the null parameter is exactly correct.