Unit S2: Statistics 2
- Syllabus
- 2019
- Section
- —
- Level
- A2

Binomial and Poisson distributions both model counts, but their random processes differ. Choose the model from the conditions before calculating a probability, and keep the count interval and event inequality explicit.
| Model | Conditions and parameter |
|---|---|
| X∼B(n,p) | a fixed number n of independent trials; each trial has success/failure outcomes; success probability p is constant |
| X∼Po(λ) | events occur independently and singly at a constant average rate in a fixed interval; λ is the mean count in that interval |
PB(X=x)=(xn)px(1−p)n−x,PPo(X=x)=e−λx!λx
Translate words into integer events before using a formula or cumulative table: 'fewer than 3' is P(X≤2), 'at least 6' is 1−P(X≤5), and 4<X<10 is P(X≤9)−P(X≤4). For X∼Po(4.5), the last probability is 0.9829−0.5321=0.4508.
Poisson means scale with exposure. If complaints occur at mean 3 per day, then a 7-day total has distribution Po(21). More generally, independent Poisson counts add: if X∼Po(λ1) and Y∼Po(λ2) independently, then X+Y∼Po(λ1+λ2).
A calculation is only as sound as the model. Repeated trials with changing success probabilities are not binomial; clustering, dependence or a changing rate undermines Poisson. State assumptions in context—such as failures being independent with constant probability, or events occurring independently at a constant average rate—and comment when the situation contradicts them.
The mean gives the long-run expected count and the variance measures its spread. Binomial and Poisson models have characteristic mean–variance relationships that can also identify unknown parameters or test plausibility.
| Distribution | Mean | Variance |
|---|---|---|
| X∼B(n,p) | np | np(1−p) |
| X∼Po(λ) | λ | λ |
If a binomial variable has mean 90 and variance 36, thennp=90,np(1−p)=36.Dividing the second equation by the first gives 1−p=36/90=0.4, so p=0.6 and n=90/0.6=150. Check that n is a positive integer and 0≤p≤1.
For Y∼B(20,0.3), E(Y)=6 and Var(Y)=4.2. For W∼Po(4), both the mean and variance are 4. An observed count dataset whose sample mean is far from its sample variance may therefore cast doubt on a Poisson model, although closeness alone does not prove all Poisson assumptions.
Do not interchange variance and standard deviation, or assume every count with equal mean and variance must be Poisson. The specification requires use of these moment formulae but not their derivations.
When X∼B(n,p) has many trials and a small success probability, the rare-success count can be approximated by a Poisson variable with the same mean.
X∼B(n,p) ≈ Y∼Po(λ),λ=np
First check that the binomial conditions hold and that the context really describes many independent opportunities for a rare event. Then calculate λ=np, preserve the original integer event exactly, and evaluate it with the Poisson formula or table.
If X∼B(200,0.012), use Y∼Po(2.4). ThenP(X≤1)≈P(Y≤1)=e−2.4(1+2.4)=0.3084.For 'at least 2', use the complementary Poisson event 1−P(Y≤1).
This is a discrete-to-discrete approximation, so no continuity correction is used. It becomes less convincing when the success probability is not small or trials are dependent; use the exact binomial distribution when the approximation is not justified. Normal approximations and continuity correction belong to S2.3, not this objective.
A continuous random variable can take any real value in an interval, so it models a measurement rather than a count. Waiting time, mass and length are typical examples; the number of calls is discrete even when calls arrive through time.
Probability is assigned to intervals, not isolated values. For every particular number a, P(X=a)=0. Consequently, endpoint choices do not change a continuous probability:P(a<X<b)=P(a≤X<b)=P(a<X≤b)=P(a≤X≤b).
This does not say that observing a particular rounded value is impossible. A recorded value such as 2.3 seconds represents a measurement interval determined by the instrument's precision; the exact mathematical point still has probability zero.
Do not list a probability for each possible value as for a discrete distribution, and do not infer that every variable written with decimals is continuous. The random mechanism—measurement across a continuum versus counting separate outcomes—decides the model.
A probability density function f describes probability through area. It must satisfy f(x)≥0 on its support and have total area 1. For a continuous random variable,P(a<X≤b)=∫abf(x)dx.The height f(x) is not itself the probability P(X=x), which is zero.
F(x0)=P(X≤x0)=∫−∞x0f(t)dt
A cumulative distribution function is non-decreasing, remains between 0 and 1, and runs from 0 below the support to 1 above it. Once F is known, interval probabilities are differences:P(a<X≤b)=F(b)−F(a),P(X>b)=1−F(b).
For a simple piecewise polynomial densityf(x)=⎩⎨⎧x,2−x,0,0≤x≤1,1<x≤2,otherwise,both pieces are non-negative and their two triangular areas total 1. Accumulating area from the left givesF(x)=⎩⎨⎧0,x2/2,2x−x2/2−1,1,x<0,0≤x≤1,1<x≤2,x>2.The second middle expression starts with the area already accumulated up to 1; integrating that piece as though the accumulation restarted at zero would make the CDF jump or finish at the wrong value.
For every piecewise model, check non-negativity, total area 1, matching CDF values at internal boundaries, and final value 1. Conditional probabilities still use the ordinary ratio of events, with each probability found from density areas or CDF differences.
Where the cumulative distribution function is differentiable, its rate of increase is the density:f(x)=dxdF(x).In the other direction, integrating the density from the lower end of the support reconstructs F.
SupposeF(x)=⎩⎨⎧0,x2/4,1,x<0,0≤x<2,x≥2.Differentiating each interval givesf(x)={x/2,0,0<x<2,otherwise.Its integral over 0<x<2 is 1, so the result is a valid density.
Because a CDF cannot decrease, a candidate derivative must not be negative. Corners at piece boundaries may make the derivative undefined at isolated points, but changing a density at finitely many individual points changes no continuous probability.
Differentiate the CDF piece by piece and include zero outside the support. Do not differentiate the constants 0 and 1 into extra support, and do not treat a jump as acceptable here: a continuous random variable has a continuous CDF.
For a continuous random variable, moments are density-weighted integrals across the whole support. Split the integral wherever the density formula changes.
E(X)=∫−∞∞xf(x)dx,E(X2)=∫−∞∞x2f(x)dx,Var(X)=E(X2)−[E(X)]2
If f(x)=x/2 for 0<x<2 and is zero otherwise, thenE(X)=∫022x2dx=34,E(X2)=∫022x3dx=2,soVar(X)=2−(34)2=92.The variance is non-negative and has squared units.
For constants a and b, use linear-transformation rules rather than reintegrating:E(aX+b)=aE(X)+b,Var(aX+b)=a2Var(X).A shift changes the mean but not the variance.
Do not calculate variance as ∫(x−E(X))f(x)dx; that integral is zero. Use either ∫(x−E(X))2f(x)dx or the stated E(X2)−[E(X)]2 identity, and distinguish variance from standard deviation.
Mode, median and quartiles describe different features of a continuous distribution. The mode is located from density height; the median and quartiles are located from cumulative probability.
| Measure | Continuous-distribution condition |
|---|---|
| mode | value where f(x) is greatest; compare stationary points and support endpoints |
| lower quartile Q1 | F(Q1)=0.25 |
| median m | F(m)=0.50 |
| upper quartile Q3 | F(Q3)=0.75 |
| interquartile range | Q3−Q1 |
For f(x)=x/2 on 0<x<2, F(x)=x2/4. HenceQ1=1,m=2,Q3=3,and the interquartile range is 3−1. Since the density increases throughout its support, its greatest height is at the upper endpoint, so the modal location is 2.
First decide which piece of a piecewise CDF contains 0.25, 0.5 or 0.75, then solve only within that piece and check the answer lies in its interval. For a mode, differentiating the density can locate interior candidates, but endpoints and any corners must also be compared.
Do not set the density equal to 0.5 to find the median, and do not assume the mean, median and mode coincide. That only happens for particular distribution shapes.
A continuous uniform variable gives equal probability to intervals of equal length. If X∼U(a,b) with a<b, its density is a rectangle of width b−a. Unit total area fixes its height:
f(x)=⎩⎨⎧b−a1,0,a≤x≤b,otherwise.
Probability is therefore a length ratio. Intersect the requested event with the support first: for a≤c<d≤b,P(c<X<d)=∫cdb−a1dx=b−ad−c.For X∼U(−5,19), P(∣X∣>3.5)=[1.5+15.5]/24=17/24. Endpoints make no difference for a continuous variable.
Accumulating rectangular area from the left derives the cumulative distribution function:F(x)=⎩⎨⎧0,b−ax−a,1,x<a,a≤x≤b,x>b.The middle piece is linear because each extra unit of x adds the same area 1/(b−a).
The mean follows fromE(X)=∫abb−axdx=2(b−a)b2−a2=2a+b.Likewise,E(X2)=3(b−a)b3−a3=3a2+ab+b2,soVar(X)=E(X2)−[E(X)]2=12(b−a)2.Thus the mean is the rectangle's midpoint, while variance depends only on its width.
Uniform means constant density over a stated interval, not that every exact value has a positive equal probability. Include the zero-density outer pieces, clip transformed event intervals to the support, and use the width b-a rather than b as the denominator.
A Normal distribution can approximate a Binomial or Poisson count when the count distribution is sufficiently spread out and not strongly skewed. Match the discrete mean and variance before converting the integer event to a continuous boundary.
| Discrete model | Approximating model |
|---|---|
| X∼B(n,p) | Y∼N(np,np(1−p)) |
| X∼Po(λ) | Y∼N(λ,λ) |
| Discrete event | Continuity-corrected Normal event |
|---|---|
| X≤k | Y<k+0.5 |
| X<k | Y<k−0.5 |
| X≥k | Y>k−0.5 |
| X>k | Y>k+0.5 |
| a≤X≤b | a−0.5<Y<b+0.5 |
Each integer count represents a unit-width bar extending 0.5 on either side of its centre. Moving the Normal boundary to the outer edge of the last included bar preserves approximately the same area. First rewrite words such as 'fewer than 32' as the integer event X ≤ 31; then its corrected boundary is 31.5.
Suppose X∼Po(36) and we need P(X<32). Use Y∼N(36,36), so its standard deviation is 6. With the correction,P(X<32)≈P(Y<31.5)=P(Z<631.5−36)=P(Z<−0.75)=0.2266.The value 32 is not the boundary because the count 32 is excluded.
For a Binomial example, X∼B(200,0.4) gives Y∼N(80,48). ThenP(X≥90)≈P(Y>89.5)=P(Z>4889.5−80).Keep the second Normal parameter as the variance, but divide by its square root when standardising.
Do not apply a continuity correction to an already continuous event, omit it for a discrete-to-Normal approximation, or use a Normal model when a small mean or extreme Binomial probability leaves the count distribution strongly skewed. This objective concerns probability approximation; hypothesis-test decisions belong to the following Topic.
A statistical investigation begins by defining exactly what is being studied. The population is the complete set of units of interest; a census seeks data from every population unit, while a sample survey collects data from only a subset.
| Term | Meaning |
|---|---|
| population | every unit about which the investigation aims to draw conclusions |
| sampling unit | one individual member or item that can be selected |
| sampling frame | the operational list or other representation from which units are selected |
| census | investigation intended to include every population unit |
| sample | selected subset used to learn about the population |
| Approach | Main advantages | Main limitations |
|---|---|---|
| census | no sampling variation; detailed information on small subgroups | costly and slow; may become outdated; non-response and measurement errors can remain; unsuitable for destructive testing |
| sample survey | quicker and cheaper; can be repeated; permits destructive testing of a limited number | conclusions vary from sample to sample; selection bias or an incomplete frame can make it unrepresentative |
To study the fill mass of all cans produced during one shift, the population is every can from that shift, one can is a sampling unit, and the production list or numbered stream may provide the frame. Measuring 100 selected cans is a sample survey; measuring every can is a census.
A large sample is not automatically representative, and a census is not automatically error-free. The target population and sampling frame may differ: units missing from the frame cannot be selected, while duplicated entries may be overrepresented.
A statistic is a quantity calculated entirely from sample observations. It may estimate or test a population feature, but its formula cannot contain an unknown population parameter. Thus Xˉ, the sample range and the number of successes are statistics; (X1−μ)/σ is not when μ and σ are unknown.
If the same random-sampling procedure were repeated with a fixed sample size, the statistic would usually change. Its sampling distribution lists every possible value of the statistic together with its probability under the stated population model.
Suppose two independent Bernoulli observations are drawn from a population with success probability 0.5, and let T be the sample proportion of successes. The four equally likely ordered samples giveT=0, 21, 1with probabilitiesP(T=0)=41,P(T=21)=21,P(T=1)=41.This probability model is the sampling distribution of T, not the distribution of a single observation.
A test statistic is a chosen statistic whose sampling distribution is known under the null hypothesis. That distribution lets an observed sample value be judged as ordinary or unusually extreme.
Do not confuse a sample distribution—the observed data values in one sample—with a sampling distribution—the distribution of a statistic over all possible samples generated by the same design.
A hypothesis test asks whether sample evidence is sufficiently unusual under a stated model to justify rejecting that model in favour of a specific alternative. It can therefore refine a mathematical model, but it cannot prove either hypothesis.
| Component | Role |
|---|---|
| null hypothesis H0 | precise baseline parameter value used to calculate probabilities, such as p=0.35 or λ=8 |
| alternative hypothesis H1 | claim supported by departure in the stated direction: <, > or = |
| significance level α | maximum chosen probability of rejecting H0 through the critical region when H0 is true |
State both hypotheses before using the data. Assume H0, identify the sampling distribution of a suitable test statistic, and calculate the probability of the observed value or values at least as supportive of H1. If this p-value is no greater than alpha—or the statistic lies in the precomputed critical region—reject H0; otherwise do not reject H0.
A historical breakdown rate is 8 per week and a refurbishment is claimed to have changed it. UseH0:λ=8,H1:λ=8.If the observed count is outside the two-tailed critical region, do not reject H0: there is insufficient evidence that the mean breakdown rate changed.
Write conclusions in the context and with evidential language. 'Do not reject H0' means the sample was not sufficiently inconsistent with H0; it does not mean H0 has been accepted or shown true.
A critical region is the set of test-statistic values that cause rejection of the null hypothesis. It is chosen from the sampling distribution under H0 so that its total probability—the actual significance level—does not exceed the stated level and is as close as the discrete distribution permits.
actual significance=P(test statistic lies in the critical region∣H0)
For an upper-tailed test with J∼Po(9) under H0, tables giveP(J≤13)=0.9261,P(J≤14)=0.9585.HenceP(J≥14)=0.0739>0.05,P(J≥15)=0.0415≤0.05.The 5% critical region is therefore J≥15. An observed value of 15 or more is significant at the 5% level.
Because a discrete tail changes in jumps, the actual significance is often below the nominal level. Moving the boundary one integer further inward here would make the rejection probability too large; moving it outward would be valid but unnecessarily conservative.
State a critical region as values such as J ≥ 15, not as the probability statement P(J ≥ 15). The region contains outcomes; its probability under H0 measures the chance of entering it.
The wording of the alternative hypothesis fixes the tail structure before the sample is inspected. A directional claim uses one tail; a claim of any change uses both tails.
| Claim | Alternative | Extreme evidence |
|---|---|---|
| parameter has decreased | H1:θ<θ0 | small test-statistic values: lower tail |
| parameter has increased | H1:θ>θ0 | large test-statistic values: upper tail |
| parameter has changed | H1:θ=θ0 | unusually small or large values: two tails |
For a one-tailed test, the critical-region probability is placed in the stated tail. For a two-tailed test at nominal level alpha, choose lower and upper regions with probabilities as close as practical to alpha/2 while keeping their combined probability no greater than alpha. Discreteness may prevent equal tails or an exact total.
Testing whether a historical proportion 0.35 has changed requires H0:p=0.35 and H1:p=0.35, so both unusually few and unusually many successes oppose H0. Testing whether it has fallen uses H1:p<0.35 and only the lower tail.
Do not choose one or two tails after seeing whether the sample result is high or low. That changes the testing rule to favour the observed data and invalidates the stated significance level.
For a Binomial test, the parameter is a population success probability p; for a Poisson test, it is the mean event rate lambda for a stated exposure. Under H0, that exact parameter determines the sampling distribution used for the test.
| Test | Distribution under H0 |
|---|---|
| n independent fixed-probability trials, testing p=p0 | X∼B(n,p0) |
| Poisson count over t base exposure units, testing rate λ=λ0 per unit | X∼Po(tλ0) |
Use this order: (1) define the parameter and state H0 and H1; (2) translate the observed count and direction into the correct tail or tails; (3) calculate an exact tail probability from tables/formulae or construct a critical region; (4) compare with alpha; (5) reject or do not reject H0 and give a contextual conclusion.
Suppose 8 of 40 items fail and the historical failure proportion is 0.35. To test whether it is lower, useH0:p=0.35,H1:p<0.35,X∼B(40,0.35).The lower-tail probability isP(X≤8)=0.0303<0.05.Reject H0: there is sufficient evidence at the 5% level that the failure proportion is lower than 0.35.
Use an approximation only when suitable, preserving the null model: a rare-event Binomial may use Po(np0), while a sufficiently spread Binomial or Poisson count may use a matched Normal distribution with continuity correction. For example, six weeks at a null Poisson rate of 6 per week gives total mean 36 and may be approximated by N(36,36).
For a discrete two-tailed test, construct both rejection tails and add their null probabilities to obtain the actual significance. Do not automatically double one observed-tail probability unless that rule is justified by the chosen test; discrete tails are often unequal.
Keep the parameter in the hypotheses, not the observed statistic: write H0:p=p0 or H0:lambda=lambda0, not H0:X=x. A non-significant result is insufficient evidence for H1, not evidence that the null parameter is exactly correct.