4. Further Probability & Statistics

Syllabus
9231–2028–2029
Section
4
Level
A2

4.1 Continuous random variables

Syllabus
9231–2028–2029
Topic
4.1
Level
A2

A piecewise PDF must be non-negative and have total area one

AprobabilitydensityfunctionsatisfiesA probability density function satisfiesf(x)\ge0,\qquad \int_{-\infty}^{\infty}f(x),dx=1.Probabilitiesareareas:Probabilities are areas:P(a<X<b)=\int_a^b f(x),dx.ForapiecewisePDF,spliteveryintegralateachbranchboundary.For a piecewise PDF, split every integral at each branch boundary.

SupposeSupposef(x)=\begin{cases}kx,&0\le x<1,\k(2-x),&1\le x\le2,\0,&\text{otherwise}.\end{cases}NormalizationgivesNormalization gives1=k\int_0^1x,dx+k\int_1^2(2-x),dx=\frac{k}{2}+\frac{k}{2},so $k=1$.

Usingthesecondbranch,Using the second branch,P(X>1.5)=\int_{1.5}^{2}(2-x),dx=0.125.If an interval crossed $x=1$, its probability would be the sum of two integrals, one from each branch.

A density value f(x) is not P(X=x); for a continuous variable P(X=x)=0. The formulas need not be continuous at a branch join unless other conditions impose continuity, but the PDF must remain non-negative and normalized.

Expected functions are integrated against the continuous density

If continuous $X$ has PDF $f$, then for a function $g$,E[g(X)]=\int_{-\infty}^{\infty}g(x)f(x),dx,integratedovertheactualsupport.Importantcasesareintegrated over the actual support. Important cases areE[X]=\int xf(x),dx,\qquad E[X^2]=\int x^2f(x),dx,and $\operatorname{Var}(X)=E[X^2]-E[X]^2$.

For $f(x)=2x$ on $0\le x\le1$,E[X]=\int_0^1 2x^2,dx=\frac23,andandE[X^2]=\int_0^1 2x^3,dx=\frac12.ThereforeTherefore\operatorname{Var}(X)=\frac12-\left(\frac23\right)^2=\frac1{18}.

For any requested transformation, replace g(x) directly inside the integral. For example, E[(3X-1) squared] uses g(x)=(3x-1) squared; it does not require first finding the distribution of 3X-1.

This is a continuous integral, not a discrete sum. Do not omit f(x), and do not assume E[g(X)]=g(E[X]); that equality fails for most nonlinear g.

The CDF accumulates the PDF and percentiles invert the CDF

For continuous $X$,F(x)=P(X\le x)=\int_{-\infty}^{x}f(t),dt.Hence, where differentiable, $f(x)=F'(x)$, andP(a<X\le b)=F(b)-F(a).A percentile $q_p$ satisfies $F(q_p)=p$.

A CDF is non-decreasing, right-continuous, tends to 0 as x tends to negative infinity and tends to 1 as x tends to positive infinity. For a bounded continuous PDF, write the outside-support CDF branches explicitly as 0 and 1.

If $f(x)=2x$ on $0\le x\le1$, then $F(x)=x^2$ on that interval. ThusP(0.2<X\le0.8)=0.8^2-0.2^2=0.60.The 75th percentile solves $q^2=0.75$, soq_{0.75}=\sqrt{0.75}.

A PDF height is not a probability and a percentile solves F(q)=p, not f(q)=p. For continuous X, endpoint choices do not change an interval probability because P(X=x)=0.

Transform a continuous variable by translating its cumulative event

For $Y=h(X)$, begin withF_Y(y)=P(Y\le y)=P(h(X)\le y).Translate this inequality into an event for $X$, using monotonicity and the support. Substitute into $F_X$, state the transformed support, then differentiate to obtain $f_Y(y)=F_Y'(y)$.

Let $f_X(x)=2x$ for $0\le x\le1$, so $F_X(x)=x^2$ there, and set $Y=X^3$. For $0\le y\le1$,F_Y(y)=P(X^3\le y)=P(X\le y^{1/3})=F_X(y^{1/3})=y^{2/3}.

ThereforeThereforeF_Y(y)=\begin{cases}0,&y<0,\y^{2/3},&0\le y\le1,\1,&y>1,\end{cases}and on $0<y<1$,f_Y(y)=\frac{d}{dy}y^{2/3}=\frac23y^{-1/3}.Its improper integral over $(0,1)$ is one.

Do not substitute y cubed into F_X: the event requires the inverse transformation, here the cube root. If h is decreasing, the inequality reverses; if it is not one-to-one on the support, split the event into all contributing branches.

4.2 Inference using normal and t-distributions

Syllabus
9231–2028–2029
Topic
4.2
Level
A2

A small-sample unknown-variance mean test uses t with n minus 1 degrees of freedom

For a random sample of size $n$ from a normal population with unknown variance, test $H_0:\mu=\mu_0$ usingT=\frac{\bar X-\mu_0}{S/\sqrt n}\sim t_{n-1}\quad\text{under }H_0.The alternative $H_1$ determines whether one or both tails are critical.

A sample has $n=10$, $\bar x=52.0$ and $s=3.0$. Test $H_0:\mu=50$ against $H_1:\mu>50$ at 5%. The observed statistic ist=\frac{52-50}{3/\sqrt{10}}=2.108.With 9 degrees of freedom, the one-tail critical value is $1.833$, so reject $H_0$.

Conclude in context: there is sufficient evidence at the 5% level to suggest that the population mean exceeds 50. State the assumptions: a random sample and a normally distributed underlying population.

Hypotheses concern the population mean mu, not the sample mean. Failing to reject H0 means insufficient evidence for H1, not proof that H0 is true; statistical significance also does not establish practical importance.

Pool within-sample squared deviations when two variances estimate one common variance

For independent samples of sizes $n_1,n_2$ from populations assumed to share variance $\sigma^2$, the pooled estimate iss_p^2=\frac{(n_1-1)s_1^2+(n_2-1)s_2^2}{n_1+n_2-2}.Equivalently, add the two within-sample sums $\sum(x-\bar x)^2$ and divide by the combined degrees of freedom.

Fromrawdata,eachwithin−samplesumcanbecalculatedasFrom raw data, each within-sample sum can be calculated as\sum(x-\bar x)^2=\sum x^2-\frac{(\sum x)^2}{n}.Dothisseparatelyforeachsample;differencesbetweenthetwosamplemeansdonotenterthepooledvariancenumerator.Do this separately for each sample; differences between the two sample means do not enter the pooled variance numerator.

If $n_1=8,s_1^2=9$ and $n_2=12,s_2^2=16$, thens_p^2=\frac{7(9)+11(16)}{18}=\frac{239}{18}=13.28,so $s_p=3.64$ approximately.

Do not average sample standard deviations or sample means. Pooling is justified only under the common population-variance model; the weights are degrees of freedom, not simply equal weights.

Choose the difference-of-means test from the sampling design and variance model

design/information analyse statistic denominator reference
paired observations differences did_i sd/ns_d/\sqrt n tn−1t_{n-1}
independent normal samples, unknown equal variances Xˉ1−Xˉ2\bar X_1-\bar X_2 sp1/n1+1/n2s_p\sqrt{1/n_1+1/n_2} tn1+n2−2t_{n_1+n_2-2}
independent samples with population variances known or a justified normal model Xˉ1−Xˉ2\bar X_1-\bar X_2 σ12/n1+σ22/n2\sqrt{\sigma_1^2/n_1+\sigma_2^2/n_2} standard normal

For paired data define one signed difference consistently. Test $H_0:\mu_d=0$ usingT=\frac{\bar d}{s_d/\sqrt n}.If $n=9$, $\bar d=2.0$ and $s_d=2.4$, then $t=2.5$ with 8 df. Against $H_1:\mu_d>0$ at 5%, $2.5>1.860$, so reject $H_0$.

For a paired t-test, the population distribution of differences must be normal; normality of each marginal sample is not the relevant statement. For the pooled two-sample t-test, the two underlying populations are normal with equal variances and the samples are independent.

Pairing is not a cosmetic label: it removes between-pair variation and changes the sample to the differences. Always define the subtraction order so the alternative, statistic sign and contextual conclusion agree.

A small-sample mean interval uses a t margin with n minus 1 degrees of freedom

For a random sample from a normal population with unknown variance, a $100(1-\alpha)\%$ confidence interval for $\mu$ is\bar x\pm t_{n-1,,1-\alpha/2}\frac{s}{\sqrt n},where $s^2$ is the unbiased sample variance.

Suppose $n=13$, $\bar x=55.6$ and $s=13.07$. For a 90% interval, $t_{12,0.95}=1.782$. The margin is1.782\frac{13.07}{\sqrt{13}}=6.46,givingapproximatelygiving approximately(49.1,62.1).Reportendpointswithsuitableaccuracyandunits.Report endpoints with suitable accuracy and units.

The method has 90% long-run coverage under its assumptions. The population mean is fixed; the interval is random before sampling. A necessary small-sample assumption is that the underlying population, not merely the sample mean, is normally distributed.

Do not use a standard normal critical value when the small-sample population variance is unknown. The interval estimates the population mean; it does not contain 90% of individual observations.

A difference interval must match the design and preserve subtraction order

model interval centre standard error and critical reference
independent normal, unknown equal variances xˉ1−xˉ2\bar x_1-\bar x_2 sp1/n1+1/n2s_p\sqrt{1/n_1+1/n_2} with tn1+n2−2t_{n_1+n_2-2}
independent, known variances/normal method xˉ1−xˉ2\bar x_1-\bar x_2 σ12/n1+σ22/n2\sqrt{\sigma_1^2/n_1+\sigma_2^2/n_2} with zz
paired dˉ\bar d sd/ns_d/\sqrt n with tn−1t_{n-1}

For independent equal-variance samples, let $n_1=10,n_2=8$, $\bar x_1-\bar x_2=5.7$ and $s_p^2=341.725$. A 95% interval uses $t_{16,0.975}=2.120$:5.7\pm2.120\sqrt{341.725\left(\frac1{10}+\frac18\right)},givingapproximatelygiving approximately(-12.9,24.3).

The interval is for mu1 minus mu2 in that order. Because zero lies inside this 95% interval, the corresponding two-sided 5% test would not reject equality of the population means. This is insufficient evidence of a difference, not proof of equality.

Do not construct separate intervals for the two means and subtract their endpoints: the standard error of a difference must be derived from the joint design. Reversing the subtraction order negates both endpoints but does not change whether zero is included.

4.3 Chi-squared tests

Syllabus
9231–2028–2029
Topic
4.3
Level
A2

Fit a prescribed distribution by turning its probabilities into expected counts

Under the stated hypothesis, identify the theoretical distribution and estimate only the parameters the question leaves unknown. Calculate each class probability from that fitted model, then multiply by the total frequency N: expected count E=Np. Preserve exhaustive classes, including a final tail such as X at least 4.

ForafittedPoissonmodel,estimateFor a fitted Poisson model, estimate\hat\lambda=\bar x=\frac{\sum xf}{N},thenthenP(X=r)=e^{-\hat\lambda}\frac{\hat\lambda^r}{r!},\qquad E_r=N P(X=r).AtailclasshasA tail class hasE_{\ge k}=N\left(1-\sum_{r=0}^{k-1}P(X=r)\right).

If $N=100$ and the fitted value is $\hat\lambda=1.2$, then the expected frequency for $X=0$ is100e^{-1.2}=30.12,and for $X=1$ it is100e^{-1.2}(1.2)=36.14.Continuewithunroundedprobabilitiesbeforecombininganysmalltailclasses.Continue with unrounded probabilities before combining any small tail classes.

Fitting does not mean copying observed proportions into expected counts. Matching an estimated mean does not establish good fit; it only supplies the model probabilities that the chi-squared analysis will assess.

Goodness of fit uses final combined classes and parameter-adjusted degrees of freedom

Set H0 to the prescribed distribution and H1 to not that distribution. Before calculating the statistic, combine adjacent classes so every final expected frequency is at least 5; combine their observed and expected counts together.

For the final $k$ classes,X^2=\sum\frac{(O-E)^2}{E}.If $p$ distribution parameters were estimated from these data,\text{df}=k-1-p.Compareintheuppertail:alargestatisticisevidenceagainstthefittedmodel.Compare in the upper tail: a large statistic is evidence against the fitted model.

With final observed counts $(20,30,50)$ and expected counts $(25,25,50)$,X^2=\frac{25}{25}+\frac{25}{25}+0=2.00.If no parameter was estimated, df $=3-1=2$; since $2.00<5.991$ at 5%, do not reject H0.

Count classes after combining and subtract every parameter estimated from the same data. A non-significant result means the data are compatible with the model at that level; it does not prove the model true.

Independence predicts each cell from its row and column margins

For an $r\times c$ contingency table, $H_0$ states that the two categorical variables are independent. The expected count in cell $(i,j)$ isE_{ij}=\frac{(\text{row }i\text{ total})(\text{column }j\text{ total})}{\text{grand total}},and df $=(r-1)(c-1)$.

ForobservedcountsFor observed counts\begin{pmatrix}30&20\10&40\end{pmatrix},rowtotalsare50,50andcolumntotals40,60,soexpectedcountsarerow totals are 50,50 and column totals 40,60, so expected counts are\begin{pmatrix}20&30\20&30\end{pmatrix}.ThusThusX^2=\frac{100}{20}+\frac{100}{30}+\frac{100}{20}+\frac{100}{30}=16.67.

Here df=1 and 16.67 exceeds the 5% critical value 3.841, so reject independence and conclude that there is evidence of an association. Every final expected cell must be at least 5; where needed, combine meaningful rows or columns before recalculating margins and df.

Use counts, not percentages, and do not apply Yates' correction in this syllabus. Association is not causation, and independence is not the same as mutually exclusive categories.

4.4 Non-parametric tests

Syllabus
9231–2028–2029
Topic
4.4
Level
A2

Non-parametric tests use signs or ranks when a normal model is not credible

feature parametric mean test non-parametric sign/rank test
main numerical information original magnitudes signs or order ranks
typical target mean/model parameter median, paired shift or identity of distributions
distributional demand often normality and variance conditions fewer shape conditions, but still design/test-specific assumptions
useful when model assumptions are credible normality is doubtful, data are skewed/outlier-prone or only ordering is reliable

Choose from the sampling design first: one sample, matched pairs or two independent samples. Then identify whether only direction is defensible or whether ranks of magnitudes can be used. State population hypotheses before calculating a statistic.

Signs and ranks reduce sensitivity to extreme magnitudes and avoid a normal-mean model, but they discard some metric information. When a valid parametric model holds, that discarded information can make a non-parametric test less powerful.

Non-parametric does not mean assumption-free. Randomness and independence or valid pairing remain essential, and the Wilcoxon tests in this syllabus require symmetrical distributions.

Sign, signed-rank and rank-sum tests retain different information

test construct under H0H_0 statistic basis key condition
sign signs of observations-minus-median or paired differences positive count B∼Bin⁡(n,1/2)B\sim\operatorname{Bin}(n,1/2) independent signs/valid pairs
Wilcoxon signed-rank rank absolute differences, restore signs positive and negative rank sums symmetric difference distribution
Wilcoxon rank-sum pool two independent samples and rank all values rank sum for a named sample independent samples; symmetry condition in this syllabus

For $n$ signed ranks with no zero differences or ties, the positive-rank sum $W^+$ has, under $H_0$,E(W^+)=\frac{n(n+1)}4,\qquad \operatorname{Var}(W^+)=\frac{n(n+1)(2n+1)}{24}.Thesesupportanormalapproximationwhenappropriate.These support a normal approximation when appropriate.

For independent samples of sizes $n_1,n_2$ and $N=n_1+n_2$, the rank sum $W_1$ hasE(W_1)=\frac{n_1(N+1)}2,\qquad \operatorname{Var}(W_1)=\frac{n_1n_2(N+1)}{12}.

Signed-rank and rank-sum are different tests: the first ranks within-pair or one-sample absolute differences, while the second ranks two independent samples together. The syllabus excludes tied ranks and zero differences in application questions.

A single-sample median test chooses between signs and signed ranks

method process for testing median m0m_0 information used
sign test record signs of xi−m0x_i-m_0 and use B∼Bin⁡(n,1/2)B\sim\operatorname{Bin}(n,1/2) direction only
signed-rank rank ∣xi−m0∣|x_i-m_0|, restore signs and sum ranks direction plus ordered magnitude; requires symmetry

Test $H_0:m= m_0$ against $H_1:m>m_0$. If 10 of 12 observations exceed $m_0$, then under $H_0$P(B\ge10)=\frac{\binom{12}{10}+\binom{12}{11}+\binom{12}{12}}{2^{12}}=\frac{79}{4096}=0.0193.At5At 5%, reject $H_0$.

For a large sign test, $B$ may be approximated byN\left(\frac n2,\frac n4\right),using a continuity correction. For signed-rank, use the null mean and variance from the previous card and correct the discrete rank-sum boundary by $0.5$ when using a normal approximation.

The alternative determines the tail before counting. These are tests about a population median or symmetric location, not a mean. Syllabus application questions contain no observations equal to the tested median and no tied ranks.

Matched designs test differences; independent designs test pooled ranks

data/design appropriate test statistic
paired, direction only paired sign number of positive differences
paired, symmetric differences Wilcoxon matched-pairs signed-rank signed rank sum of within-pair differences
two independent samples, syllabus symmetry condition met Wilcoxon rank-sum rank sum for a named sample after pooling

For independent samples $n_1=20,n_2=25$, let $W_1=560$ and $N=45$. Under identical populations,E(W_1)=20(46)/2=460,\operatorname{Var}(W_1)=20(25)(46)/12=1916.67.Foranupper−tailnormalapproximation,continuitycorrectiongivesFor an upper-tail normal approximation, continuity correction givesz=\frac{559.5-460}{\sqrt{1916.67}}=2.27.

Since 2.27 exceeds the 5% one-tail critical value 1.645, reject the null in the direction attached to sample 1's high ranks. For paired tests, define every difference in one order before applying the sign or signed-rank procedure.

Rank-sum is not a test for paired data, and matched-pairs signed-rank is not formed by ranking the two columns separately. Questions in this syllabus avoid tied ranks and zero-difference pairs; use exact tables instead of a normal approximation when the sample sizes make that appropriate.

4.5 Probability generating functions

Syllabus
9231–2028–2029
Topic
4.5
Level
A2

A PGF stores each probability as a power-series coefficient

For a non-negative integer-valued random variable $X$,G_X(s)=E[s^X]=\sum_{r=0}^{\infty}P(X=r)s^r.Thus $[s^r]G_X(s)=P(X=r)$, $G_X(0)=P(X=0)$ and $G_X(1)=1$.

distribution and support PGF
discrete uniform on 1,…,n1,\ldots,n s(1−sn)n(1−s)\dfrac{s(1-s^n)}{n(1-s)} for s≠1s\ne1, with G(1)=1G(1)=1
Bin⁡(n,p)\operatorname{Bin}(n,p), q=1−pq=1-p (q+ps)n(q+ps)^n
geometric P(X=r)=pqr−1P(X=r)=pq^{r-1}, r=1,2,…r=1,2,\ldots ps1−qs\dfrac{ps}{1-qs}
Po⁡(λ)\operatorname{Po}(\lambda) exp⁡(λ(s−1))\exp(\lambda(s-1))

For example, if $G(s)=0.2+0.5s+0.3s^2$, thenP(X=0)=0.2,\quad P(X=1)=0.5,\quad P(X=2)=0.3.Thecoefficientrulealsoletsaclosedformbeexpandedtorecoverprobabilities.The coefficient rule also lets a closed form be expanded to recover probabilities.

State the geometric support convention: starting at one produces the numerator ps, while a failures-before-success convention starts at zero. A PGF is not an MGF, and G(1) must equal one for a valid probability distribution.

PGF derivatives at one give factorial moments, then mean and variance

For $G(s)=E[s^X]$,E[X]=G'(1),\qquad E[X(X-1)]=G''(1).Since $X^2=X(X-1)+X$,\operatorname{Var}(X)=G''(1)+G'(1)-[G'(1)]^2.Differentiate before setting $s=1$.

For $X\sim\operatorname{Po}(\lambda)$,G(s)=e^{\lambda(s-1)},\quad G'(s)=\lambda e^{\lambda(s-1)},\quad G''(s)=\lambda^2e^{\lambda(s-1)}.Hence $G'(1)=\lambda$ and $G''(1)=\lambda^2$.

ThereforeThereforeE[X]=\lambda,andand\operatorname{Var}(X)=\lambda^2+\lambda-\lambda^2=\lambda.ThecancellationchecksthefamiliarPoissonequalityofmeanandvariance.The cancellation checks the familiar Poisson equality of mean and variance.

G double-prime at one is not E[X squared]; it omits one copy of E[X]. If an alleged variance is negative, recheck differentiation, substitution at one and the subtraction of the squared mean.

Independence turns the PGF of a sum into a product

If $X_1,\ldots,X_k$ are independent non-negative integer-valued variables and $S=\sum X_i$, thenG_S(s)=E\left[s^{\sum X_i}\right]=E\left[\prod s^{X_i}\right]=\prod G_{X_i}(s).Independencejustifiesthefactorisationoftheexpectation.Independence justifies the factorisation of the expectation.

If $X\sim\operatorname{Po}(\lambda)$ and $Y\sim\operatorname{Po}(\mu)$ independently,G_{X+Y}(s)=e^{\lambda(s-1)}e^{\mu(s-1)}=e^{(\lambda+\mu)(s-1)},sosoX+Y\sim\operatorname{Po}(\lambda+\mu).

If $X\sim\operatorname{Bin}(n_1,p)$ and $Y\sim\operatorname{Bin}(n_2,p)$ independently,G_{X+Y}(s)=(1-p+ps)^{n_1+n_2},hence $X+Y\sim\operatorname{Bin}(n_1+n_2,p)$. If the success probabilities differ, multiply the PGFs but do not label the result binomial without further justification.

Multiplying PGFs requires independence. Means still add for dependent variables with finite expectations, but the PGF factorisation and the standard-family conclusions above can fail; coefficients of the actual product may be read when no named family results.