Unit S3: Statistics 3

Syllabus
2019
Section
—
Level
A2

S3.1 - Combinations of random variables

Syllabus
2019
Topic
S3.1
Level
A2

Combine independent Normal random variables

A linear combination of independent Normal random variables is also Normal. IfX∼N(μx,σx2),Y∼N(μy,σy2)X\sim N(\mu_x,\sigma_x^2),\qquad Y\sim N(\mu_y,\sigma_y^2)independently, then its mean follows the signs in the combination, while each variance contribution uses the square of its coefficient.

aX±bY∼N(aμx±bμy, a2σx2+b2σy2)aX\pm bY\sim N\left(a\mu_x\pm b\mu_y,\ a^2\sigma_x^2+b^2\sigma_y^2\right)

Combination Mean Variance
aX+bYaX+bY aμx+bμya\mu_x+b\mu_y a2σx2+b2σy2a^2\sigma_x^2+b^2\sigma_y^2
aX−bYaX-bY aμx−bμya\mu_x-b\mu_y a2σx2+b2σy2a^2\sigma_x^2+b^2\sigma_y^2

Subtraction changes the centre because it reverses Y's contribution to the value. It does not subtract uncertainty: deviations in either variable create spread in the combination, so independent variance contributions add after scaling by squared coefficients.

For independent copies, add one mean and one variance contribution for each copy. If male load M∼N(80,100)M\sim N(80,100) and female load W∼N(69,25)W\sim N(69,25), then the load of six men and three women isT=M1+⋯+M6+W1+⋯+W3∼N(687,675).T=M_1+\cdots+M_6+W_1+\cdots+W_3\sim N(687,675).ThereforeP(T>700)=P(Z>700−687675)=P(Z>0.500…)≈0.3085.P(T>700)=P\left(Z>\frac{700-687}{\sqrt{675}}\right)=P(Z>0.500\ldots)\approx0.3085.

Convert comparisons into one variable before standardising. If A∼N(20,9)A\sim N(20,9) and B∼N(8,4)B\sim N(8,4) independently, thenD=A−2B∼N(4,25),D=A-2B\sim N(4,25),soP(A>2B)=P(D>0)=P(Z>0−45)=P(Z>−0.8)=0.7881.P(A>2B)=P(D>0)=P\left(Z>\frac{0-4}{5}\right)=P(Z>-0.8)=0.7881.

Check both Normality and independence before using the rule. Keep the second parameter of N(mean, variance) as a variance, take its square root only when standardising, and never subtract variances for a difference. No proof of the combination rule is required.

S3.2 - Sampling

Syllabus
2019
Topic
S3.2
Level
A2

Take a simple random sample with random numbers

In a simple random sample of size n, every possible set of n distinct population units has the same chance of selection. A complete sampling frame is therefore needed so that every unit can be identified before random selection.

Number the N units consistently, for example 00–69 for 70 students. Read random digits in groups with enough places to represent the largest label. Accept a number only when it is in range and has not already been selected; ignore out-of-range values and repeats until n distinct units have been obtained.

For 70 students labelled 00–69, two-digit random numbers are required. If the stream begins 47, 83, 05, 47, 62, then select 47, ignore 83 because it is outside 00–69, select 05, ignore the repeated 47, and select 62.

The labels identify units but must not influence selection; the random-number source supplies the chance mechanism. Rejecting invalid and repeated labels preserves a sample without replacement while treating eligible labels symmetrically.

A table of individually random digits is not automatically suitable if its displayed structure does not generate the required multi-digit numbers randomly. Do not choose convenient replacements for rejected values, and do not confuse random selection with a haphazard sample.

Choose stratified, systematic or quota sampling

Choose a sampling method from the available frame, the population structure and the practical constraints. Stratified, systematic and quota sampling can all spread a sample across a population, but only the first two use a specified probability-based selection rule when carried out correctly.

Method How it is taken Useful when Main advantage Main limitation
stratified split into non-overlapping strata; allocate proportionally; randomly sample within each known subgroups should be represented reflects population structure and supports subgroup summaries needs stratum information and a frame; more organization
systematic order and number the frame; choose a random start; then take every kkth unit a complete ordered frame is available and an evenly spread sample is convenient quick and simple after the start periodicity or ordering can bias the result
quota set target numbers for categories; interviewers select available units until each quota is full no complete frame is available and speed/cost dominate quick, inexpensive and enforces category totals non-random within quotas, so interviewer or availability bias remains

For population size NN and stratum size NhN_h, allocatenh≈nNhN,n_h\approx n\frac{N_h}{N},then adjust rounding so all allocations total nn. With 800 employees split 430, 250 and 120 across three cities and sample size 100, proportional allocations are approximately 54, 31 and 15.

For an ordered list of N=280N=280 students and sample size n=40n=40, the interval is k=N/n=7k=N/n=7. Choose a random start from 1 to 7, then select that unit and every seventh unit afterwards until 40 are chosen.

Stratification does not remove bias if selection inside a stratum is not random. A systematic sample needs a random start and can fail when the list has a pattern related to the variable. Quota sampling matches selected category counts but is not the same as stratified random sampling.

S3.3 - Estimation, confidence intervals and tests

Syllabus
2019
Topic
S3.3
Level
A2

Estimate parameters without systematic bias

An estimator is a rule calculated from a random sample to estimate a population parameter. Its observed value is an estimate. An estimator θ^\hat\theta is unbiased when E(θ^)=θE(\hat\theta)=\theta; its bias is E(θ^)−θE(\hat\theta)-\theta.

The standard error is the standard deviation of an estimator's sampling distribution. It measures how much estimates would vary across repeated samples, rather than how much individual observations vary.

\bar x=\frac1n\sum_{i=1}^{n}x_i,\qquad s^2=\frac1{n-1}\sum_{i=1}^{n}(x_i-\bar x)^2

The sample mean xˉ\bar x is an unbiased estimate of μ\mu. Dividing the squared deviations by n−1n-1, not nn, makes s2s^2 an unbiased estimate of σ2\sigma^2. When σ\sigma is unknown, s/ns/\sqrt n estimates the standard error of the sample mean.

For the sample 2, 4, 6, xˉ=4\bar x=4 and ∑(xi−xˉ)2=8\sum(x_i-\bar x)^2=8, so s2=8/(3−1)=4s^2=8/(3-1)=4. The estimated standard error of xˉ\bar x is 2/32/\sqrt3.

Unbiased does not mean that one estimate equals the true parameter, and it does not by itself mean most precise. Among unbiased estimators, a smaller variance means greater efficiency.

Use the sampling distribution of the mean

For independent observations from a population with mean μ\mu and variance σ2\sigma^2, the sample mean Xˉ\bar X has the same mean μ\mu but the smaller variance σ2/n\sigma^2/n. Averaging therefore preserves the centre while reducing sampling variability.

E(\bar X)=\mu,\qquad \operatorname{Var}(\bar X)=\frac{\sigma^2}{n},\qquad \operatorname{se}(\bar X)=\frac{\sigma}{\sqrt n}

If each observation is Normal, X∼N(μ,σ2)X\sim N(\mu,\sigma^2), then the result is exact: Xˉ∼N(μ,σ2/n)\bar X\sim N(\mu,\sigma^2/n). No proof is required.

If bag weights are N(50,62)N(50,6^2) and n=36n=36, then Xˉ∼N(50,12)\bar X\sim N(50,1^2). Thus P(Xˉ<48.5)=P(Z<−1.5)≈0.0668P(\bar X<48.5)=P(Z<-1.5)\approx0.0668.

Use σ/n\sigma/\sqrt n, not σ/n\sigma/n, for the standard deviation of Xˉ\bar X. Exact Normality here comes from a Normal population; the large-sample Central Limit theorem for non-Normal populations belongs to the later objective.

Interpret a confidence interval and link it to a test

A confidence interval is a range produced by a sampling procedure for an unknown fixed parameter. A 95% procedure is designed so that, over many independent samples, about 95% of the intervals constructed in the same way contain the true parameter.

After one interval has been calculated, the parameter and its endpoints are fixed: the interval either contains the parameter or it does not. The confidence level describes the long-run success rate of the method, not a 95% probability assigned to this particular fixed interval.

For a two-sided test of H0:μ=μ0H_0:\mu=\mu_0 at significance level α\alpha, use the matching 100(1−α)%100(1-\alpha)\% confidence interval. Reject H0H_0 exactly when μ0\mu_0 lies outside the interval; if it lies inside, do not reject H0H_0.

If a 95% confidence interval for μ\mu is (22.24,22.56)(22.24,22.56), the value 22.50 is compatible with the data at the 5% two-sided level, whereas 23.00 would be rejected.

A value inside the interval is not proved true, and failing to reject it is not the same as accepting it. Match the confidence level to a two-sided test; one-sided tests require the corresponding one-sided procedure.

Find a confidence interval for a Normal mean

For a random sample from a Normal population whose variance σ2\sigma^2 is known, a two-sided 100(1−α)%100(1-\alpha)\% confidence interval for μ\mu is centred on the sample mean and extends by a critical value times its standard error.

\bar x;\pm;z_{1-\alpha/2}\frac{\sigma}{\sqrt n}

Find xˉ\bar x, calculate σ/n\sigma/\sqrt n, select the standard Normal critical value for the required confidence level, then subtract and add the margin of error. State both limits with appropriate accuracy.

For xˉ=22.4\bar x=22.4, known σ=0.4\sigma=0.4 and n=36n=36, a 98% interval uses z0.99=2.3263z_{0.99}=2.3263. The margin is 2.3263(0.4/6)=0.15512.3263(0.4/6)=0.1551, giving (22.245,22.555)(22.245,22.555), or approximately (22.25,22.56)(22.25,22.56).

This formula assumes the population variance is known and the sample is random; with a small sample, the population itself must be Normal. Do not use ss in place of σ\sigma under this objective, and no theoretical derivation is required.

Test a Normal mean when the variance is known

To test a Normal population mean when σ2\sigma^2 is known, compare the observed sample mean with the null value μ0\mu_0 in units of its standard error.

Z=\frac{\bar X-\mu_0}{\sigma/\sqrt n}\sim N(0,1)\quad\text{under }H_0

State H0H_0 and a one- or two-sided H1H_1 in terms of μ\mu. Calculate zz, compare it with the appropriate standard Normal critical value (or use its pp-value), then give a conclusion about H0H_0 and a contextual conclusion.

Suppose H0:μ=1010H_0:\mu=1010 against H1:μ≠1010H_1:\mu\ne1010, with known σ=8\sigma=8, n=100n=100 and xˉ=1008.47\bar x=1008.47. Then z=(1008.47−1010)/(8/10)=−1.9125z=(1008.47-1010)/(8/10)=-1.9125. Since ∣z∣<1.96|z|<1.96, do not reject H0H_0 at 5%.

Choose the tail before seeing the result. A non-significant result means insufficient evidence against H0H_0; it does not prove H0H_0. This exact test requires a Normal population and known variance.

Use large-sample inference for one mean

The Central Limit theorem makes the sampling distribution of Xˉ\bar X approximately Normal for a sufficiently large independent random sample, even when the population distribution is not Normal. This extends Normal confidence intervals and hypothesis tests beyond Normal populations.

\frac{\bar X-\mu}{S/\sqrt n};\text{is treated approximately as }N(0,1)\text{ when }n\text{ is large.}

When σ2\sigma^2 is unknown, use the unbiased sample variance s2s^2 and estimated standard error s/ns/\sqrt n. Then use standard Normal critical values for the required large-sample confidence interval or test.

For a large sample with n=100n=100, xˉ=12.4\bar x=12.4 and s=3.0s=3.0, an approximate 95% interval is 12.4±1.96(3/10)12.4\pm1.96(3/10), giving (11.812,12.988)(11.812,12.988). The same standard error is used to standardise a hypothesised mean.

Large sample size supports both the Normal approximation and replacing σ\sigma by ss; it does not remove the need for random, independent observations or protect against a badly biased sample. Knowledge of the tt-distribution is not required.

Test the difference between two Normal means

For two independent Normal populations with known variances, the difference of sample means is Normal. Test a claimed population-mean difference δ0\delta_0 by comparing Xˉ−Yˉ\bar X-\bar Y with δ0\delta_0 using the combined standard error.

Z=\frac{(\bar X-\bar Y)-\delta_0}{\sqrt{\sigma_x^2/n_x+\sigma_y^2/n_y}}\sim N(0,1)\quad\text{under }H_0:\mu_x-\mu_y=\delta_0

Independence makes the variances of the two sample means add. The null difference is often 0, but use the value actually claimed and keep the subtraction order consistent in the hypotheses, numerator and conclusion.

With xˉ=52\bar x=52, yˉ=48\bar y=48, σx=6\sigma_x=6, nx=36n_x=36, σy=8\sigma_y=8, ny=64n_y=64 and δ0=0\delta_0=0, the standard error is 1+1=2\sqrt{1+1}=\sqrt2. Hence z=4/2=2.83z=4/\sqrt2=2.83, so reject equality at the 5% two-sided level.

This result requires independent samples from Normal populations and known variances. Do not subtract standard errors; variances add. Paired observations require a different analysis of within-pair differences.

Compare two means with large samples and unknown variances

For two independent large samples, the Central Limit theorem allows inference for μx−μy\mu_x-\mu_y even when the populations are not Normal. If the population variances are unknown, estimate each one from its own sample.

Z=\frac{(\bar X-\bar Y)-\delta_0}{\sqrt{S_x^2/n_x+S_y^2/n_y}};\text{is treated approximately as }N(0,1).

State hypotheses in terms of μx−μy\mu_x-\mu_y, calculate the two-sample standard error from sx2s_x^2 and sy2s_y^2, standardise, then compare with the correct one- or two-tailed Normal critical value. Interpret the decision in context.

For nx=64n_x=64, xˉ=75.2\bar x=75.2, sx=8s_x=8 and ny=81n_y=81, yˉ=71.8\bar y=71.8, sy=9s_y=9, testing H0:μx−μy=0H_0:\mu_x-\mu_y=0 gives standard error 82/64+92/81=2\sqrt{8^2/64+9^2/81}=\sqrt2 and z=3.4/2=2.40z=3.4/\sqrt2=2.40. This exceeds 1.6449 for a 5% upper-tail test.

The samples must be independently and randomly obtained, and both must be large enough for the approximations. Do not pool the two sample variances unless a separate model justifies it. Knowledge of the tt-distribution is not required.

S3.4 - Goodness of fit and contingency tables

Syllabus
2019
Topic
S3.4
Level
A2

Use a chi-squared test for fit or association

A chi-squared test compares observed frequencies OiO_i with frequencies EiE_i expected under a null model. Large discrepancies produce a large statistic and evidence against H0H_0; the alternative says the stated model or independence claim does not hold.

For goodness of fit, write H0H_0: the named distribution is a suitable model, and H1H_1: it is not suitable. Required models include discrete uniform, binomial, Normal, Poisson and continuous uniform (rectangular). For a contingency table, write H0H_0: the two named variables are independent (no association), and H1H_1: they are associated.

\chi^2=\sum_i\frac{(O_i-E_i)^2}{E_i}

For goodness of fit, calculate Ei=NpiE_i=Np_i from the null distribution. For an r×cr\times c contingency table, Eij=(row total)(column total)/NE_{ij}=(\text{row total})(\text{column total})/N. Compare the statistic with the upper-tail critical value for the chosen significance level and degrees of freedom.

For 50 observations in five equally likely categories, suppose O=(8,12,9,11,10)O=(8,12,9,11,10) and E=(10,10,10,10,10)E=(10,10,10,10,10). Then χ2=(4+4+1+1+0)/10=1.0\chi^2=(4+4+1+1+0)/10=1.0. With 4 degrees of freedom, 1.0<9.4881.0<9.488 at 5%, so do not reject the discrete-uniform model.

A small statistic does not prove the model or independence; it means the data do not provide sufficient evidence against H0H_0. Use frequencies, not raw measurements or percentages, and do not describe association as correlation.

Choose chi-squared degrees of freedom correctly

Degrees of freedom select the correct chi-squared reference distribution. Count categories after any required combining, then subtract constraints created by the total frequency and by parameters estimated from the same data.

\nu=k-1-p\quad\text{for goodness of fit}

Here kk is the final number of cells and pp is the number of distribution parameters estimated from the sample. If every parameter is specified in advance, p=0p=0. For example, six final Poisson cells with the mean estimated from the data give ν=6−1−1=4\nu=6-1-1=4. If combining reduces the table to five cells, use ν=3\nu=3.

\nu=(r-1)(c-1)\quad\text{for an }r\times c\text{ contingency table}

The chi-squared approximation requires adequate expected frequencies. When Ei<5E_i<5, combine suitable neighbouring or tail categories before calculating the final statistic; recompute each combined OO and EE, and base kk and the degrees of freedom on the combined table.

A 3×43\times4 contingency table has ν=(3−1)(4−1)=6\nu=(3-1)(4-1)=6. Its expected frequencies still come from row total ×\times column total /N/N; the observed counts do not determine ν\nu.

Subtract only parameters estimated from these data, not parameters supplied by the null model. Do not use the original number of cells after combining, and do not apply Yates' correction; it is not required in this specification.

S3.5 - Regression and correlation

Syllabus
2019
Topic
S3.5
Level
A2

Measure monotonic association with Spearman's rank

Spearman's rank correlation coefficient rsr_s measures the strength and direction of a monotonic association between two variables by comparing their ranks. It is useful for ordinal data or when the relationship is monotonic but a linear model for the original measurements is unsuitable.

Rank both variables in a consistent direction, calculate did_i as one rank minus the other for each pair, square the differences, and sum them. Reversing both ranking directions changes no differences; reversing only one changes the sign of the coefficient.

r_s=1-\frac{6\sum d_i^2}{n(n^2-1)}\qquad\text{when there are no tied ranks}

For five pairs with ranks (1,2,3,4,5)(1,2,3,4,5) and (1,3,2,5,4)(1,3,2,5,4), ∑d2=4\sum d^2=4. Hence rs=1−6(4)/[5(25−1)]=0.8r_s=1-6(4)/[5(25-1)]=0.8, showing a fairly strong positive monotonic association.

When values are tied, assign average ranks: each tied value receives the average of the rank positions it occupies. Numerical questions involving ties will not be set, but the correct approach is to calculate the product moment correlation coefficient of the two average-rank lists rather than use the no-ties shortcut unchanged.

rsr_s lies between −1-1 and 11: its sign gives direction and its magnitude gives strength of monotonic association. It uses order, not the sizes of gaps, so it loses quantitative information; it does not prove causation and need not describe a non-monotonic relationship well.

Test whether a population correlation is zero

A correlation test asks whether a sample coefficient is extreme enough to provide evidence that the corresponding population correlation is not zero. Select the table by coefficient, sample size nn, significance level and test direction.

For Spearman use H0:ρs=0H_0:\rho_s=0; for product moment correlation use H0:ρ=0H_0:\rho=0. Choose a >0>0 or <0<0 alternative for a predicted direction, or ≠0\ne0 when either direction counts. A two-sided claim requires the two-tailed table value.

Calculate rsr_s or rr and compare it with the matching critical value. For an upper-tail test reject H0H_0 above the positive critical value; for a two-tailed test reject when the coefficient's absolute value exceeds the critical magnitude.

Suppose n=7n=7, rs=0.679r_s=0.679 and the 5% upper-tail Spearman critical value is 0.7143. Since 0.679<0.71430.679<0.7143, do not reject H0H_0; there is insufficient evidence of a positive population rank correlation at the 5% level.

Product moment correlation targets linear association in the original values; Spearman targets monotonic association through ranks. Do not interchange their tables.

Choose the direction before inspecting the coefficient. Non-significance does not establish zero correlation, and significance does not establish causation. Conclude using the named variables and tested direction.