4. Further Probability & Statistics
- Syllabus
- 9231–2028–2029
- Section
- 4
- Level
- A2
Aprobabilitydensityfunctionsatisfiesf(x)\ge0,\qquad \int_{-\infty}^{\infty}f(x),dx=1.Probabilitiesareareas:P(a<X<b)=\int_a^b f(x),dx.ForapiecewisePDF,spliteveryintegralateachbranchboundary.
Supposef(x)=\begin{cases}kx,&0\le x<1,\k(2-x),&1\le x\le2,\0,&\text{otherwise}.\end{cases}Normalizationgives1=k\int_0^1x,dx+k\int_1^2(2-x),dx=\frac{k}{2}+\frac{k}{2},so $k=1$.
Usingthesecondbranch,P(X>1.5)=\int_{1.5}^{2}(2-x),dx=0.125.If an interval crossed $x=1$, its probability would be the sum of two integrals, one from each branch.
A density value f(x) is not P(X=x); for a continuous variable P(X=x)=0. The formulas need not be continuous at a branch join unless other conditions impose continuity, but the PDF must remain non-negative and normalized.
If continuous $X$ has PDF $f$, then for a function $g$,E[g(X)]=\int_{-\infty}^{\infty}g(x)f(x),dx,integratedovertheactualsupport.ImportantcasesareE[X]=\int xf(x),dx,\qquad E[X^2]=\int x^2f(x),dx,and $\operatorname{Var}(X)=E[X^2]-E[X]^2$.
For $f(x)=2x$ on $0\le x\le1$,E[X]=\int_0^1 2x^2,dx=\frac23,andE[X^2]=\int_0^1 2x^3,dx=\frac12.Therefore\operatorname{Var}(X)=\frac12-\left(\frac23\right)^2=\frac1{18}.
For any requested transformation, replace g(x) directly inside the integral. For example, E[(3X-1) squared] uses g(x)=(3x-1) squared; it does not require first finding the distribution of 3X-1.
This is a continuous integral, not a discrete sum. Do not omit f(x), and do not assume E[g(X)]=g(E[X]); that equality fails for most nonlinear g.
For continuous $X$,F(x)=P(X\le x)=\int_{-\infty}^{x}f(t),dt.Hence, where differentiable, $f(x)=F'(x)$, andP(a<X\le b)=F(b)-F(a).A percentile $q_p$ satisfies $F(q_p)=p$.
A CDF is non-decreasing, right-continuous, tends to 0 as x tends to negative infinity and tends to 1 as x tends to positive infinity. For a bounded continuous PDF, write the outside-support CDF branches explicitly as 0 and 1.
If $f(x)=2x$ on $0\le x\le1$, then $F(x)=x^2$ on that interval. ThusP(0.2<X\le0.8)=0.8^2-0.2^2=0.60.The 75th percentile solves $q^2=0.75$, soq_{0.75}=\sqrt{0.75}.
A PDF height is not a probability and a percentile solves F(q)=p, not f(q)=p. For continuous X, endpoint choices do not change an interval probability because P(X=x)=0.
For $Y=h(X)$, begin withF_Y(y)=P(Y\le y)=P(h(X)\le y).Translate this inequality into an event for $X$, using monotonicity and the support. Substitute into $F_X$, state the transformed support, then differentiate to obtain $f_Y(y)=F_Y'(y)$.
Let $f_X(x)=2x$ for $0\le x\le1$, so $F_X(x)=x^2$ there, and set $Y=X^3$. For $0\le y\le1$,F_Y(y)=P(X^3\le y)=P(X\le y^{1/3})=F_X(y^{1/3})=y^{2/3}.
ThereforeF_Y(y)=\begin{cases}0,&y<0,\y^{2/3},&0\le y\le1,\1,&y>1,\end{cases}and on $0<y<1$,f_Y(y)=\frac{d}{dy}y^{2/3}=\frac23y^{-1/3}.Its improper integral over $(0,1)$ is one.
Do not substitute y cubed into F_X: the event requires the inverse transformation, here the cube root. If h is decreasing, the inequality reverses; if it is not one-to-one on the support, split the event into all contributing branches.
For a random sample of size $n$ from a normal population with unknown variance, test $H_0:\mu=\mu_0$ usingT=\frac{\bar X-\mu_0}{S/\sqrt n}\sim t_{n-1}\quad\text{under }H_0.The alternative $H_1$ determines whether one or both tails are critical.
A sample has $n=10$, $\bar x=52.0$ and $s=3.0$. Test $H_0:\mu=50$ against $H_1:\mu>50$ at 5%. The observed statistic ist=\frac{52-50}{3/\sqrt{10}}=2.108.With 9 degrees of freedom, the one-tail critical value is $1.833$, so reject $H_0$.
Conclude in context: there is sufficient evidence at the 5% level to suggest that the population mean exceeds 50. State the assumptions: a random sample and a normally distributed underlying population.
Hypotheses concern the population mean mu, not the sample mean. Failing to reject H0 means insufficient evidence for H1, not proof that H0 is true; statistical significance also does not establish practical importance.
For independent samples of sizes $n_1,n_2$ from populations assumed to share variance $\sigma^2$, the pooled estimate iss_p^2=\frac{(n_1-1)s_1^2+(n_2-1)s_2^2}{n_1+n_2-2}.Equivalently, add the two within-sample sums $\sum(x-\bar x)^2$ and divide by the combined degrees of freedom.
Fromrawdata,eachwithin−samplesumcanbecalculatedas\sum(x-\bar x)^2=\sum x^2-\frac{(\sum x)^2}{n}.Dothisseparatelyforeachsample;differencesbetweenthetwosamplemeansdonotenterthepooledvariancenumerator.
If $n_1=8,s_1^2=9$ and $n_2=12,s_2^2=16$, thens_p^2=\frac{7(9)+11(16)}{18}=\frac{239}{18}=13.28,so $s_p=3.64$ approximately.
Do not average sample standard deviations or sample means. Pooling is justified only under the common population-variance model; the weights are degrees of freedom, not simply equal weights.
| design/information | analyse | statistic denominator | reference |
|---|---|---|---|
| paired observations | differences di | sd/n | tn−1 |
| independent normal samples, unknown equal variances | Xˉ1−Xˉ2 | sp1/n1+1/n2 | tn1+n2−2 |
| independent samples with population variances known or a justified normal model | Xˉ1−Xˉ2 | σ12/n1+σ22/n2 | standard normal |
For paired data define one signed difference consistently. Test $H_0:\mu_d=0$ usingT=\frac{\bar d}{s_d/\sqrt n}.If $n=9$, $\bar d=2.0$ and $s_d=2.4$, then $t=2.5$ with 8 df. Against $H_1:\mu_d>0$ at 5%, $2.5>1.860$, so reject $H_0$.
For a paired t-test, the population distribution of differences must be normal; normality of each marginal sample is not the relevant statement. For the pooled two-sample t-test, the two underlying populations are normal with equal variances and the samples are independent.
Pairing is not a cosmetic label: it removes between-pair variation and changes the sample to the differences. Always define the subtraction order so the alternative, statistic sign and contextual conclusion agree.
For a random sample from a normal population with unknown variance, a $100(1-\alpha)\%$ confidence interval for $\mu$ is\bar x\pm t_{n-1,,1-\alpha/2}\frac{s}{\sqrt n},where $s^2$ is the unbiased sample variance.
Suppose $n=13$, $\bar x=55.6$ and $s=13.07$. For a 90% interval, $t_{12,0.95}=1.782$. The margin is1.782\frac{13.07}{\sqrt{13}}=6.46,givingapproximately(49.1,62.1).Reportendpointswithsuitableaccuracyandunits.
The method has 90% long-run coverage under its assumptions. The population mean is fixed; the interval is random before sampling. A necessary small-sample assumption is that the underlying population, not merely the sample mean, is normally distributed.
Do not use a standard normal critical value when the small-sample population variance is unknown. The interval estimates the population mean; it does not contain 90% of individual observations.
| model | interval centre | standard error and critical reference |
|---|---|---|
| independent normal, unknown equal variances | xˉ1−xˉ2 | sp1/n1+1/n2 with tn1+n2−2 |
| independent, known variances/normal method | xˉ1−xˉ2 | σ12/n1+σ22/n2 with z |
| paired | dˉ | sd/n with tn−1 |
For independent equal-variance samples, let $n_1=10,n_2=8$, $\bar x_1-\bar x_2=5.7$ and $s_p^2=341.725$. A 95% interval uses $t_{16,0.975}=2.120$:5.7\pm2.120\sqrt{341.725\left(\frac1{10}+\frac18\right)},givingapproximately(-12.9,24.3).
The interval is for mu1 minus mu2 in that order. Because zero lies inside this 95% interval, the corresponding two-sided 5% test would not reject equality of the population means. This is insufficient evidence of a difference, not proof of equality.
Do not construct separate intervals for the two means and subtract their endpoints: the standard error of a difference must be derived from the joint design. Reversing the subtraction order negates both endpoints but does not change whether zero is included.
Under the stated hypothesis, identify the theoretical distribution and estimate only the parameters the question leaves unknown. Calculate each class probability from that fitted model, then multiply by the total frequency N: expected count E=Np. Preserve exhaustive classes, including a final tail such as X at least 4.
ForafittedPoissonmodel,estimate\hat\lambda=\bar x=\frac{\sum xf}{N},thenP(X=r)=e^{-\hat\lambda}\frac{\hat\lambda^r}{r!},\qquad E_r=N P(X=r).AtailclasshasE_{\ge k}=N\left(1-\sum_{r=0}^{k-1}P(X=r)\right).
If $N=100$ and the fitted value is $\hat\lambda=1.2$, then the expected frequency for $X=0$ is100e^{-1.2}=30.12,and for $X=1$ it is100e^{-1.2}(1.2)=36.14.Continuewithunroundedprobabilitiesbeforecombininganysmalltailclasses.
Fitting does not mean copying observed proportions into expected counts. Matching an estimated mean does not establish good fit; it only supplies the model probabilities that the chi-squared analysis will assess.
Set H0 to the prescribed distribution and H1 to not that distribution. Before calculating the statistic, combine adjacent classes so every final expected frequency is at least 5; combine their observed and expected counts together.
For the final $k$ classes,X^2=\sum\frac{(O-E)^2}{E}.If $p$ distribution parameters were estimated from these data,\text{df}=k-1-p.Compareintheuppertail:alargestatisticisevidenceagainstthefittedmodel.
With final observed counts $(20,30,50)$ and expected counts $(25,25,50)$,X^2=\frac{25}{25}+\frac{25}{25}+0=2.00.If no parameter was estimated, df $=3-1=2$; since $2.00<5.991$ at 5%, do not reject H0.
Count classes after combining and subtract every parameter estimated from the same data. A non-significant result means the data are compatible with the model at that level; it does not prove the model true.
For an $r\times c$ contingency table, $H_0$ states that the two categorical variables are independent. The expected count in cell $(i,j)$ isE_{ij}=\frac{(\text{row }i\text{ total})(\text{column }j\text{ total})}{\text{grand total}},and df $=(r-1)(c-1)$.
Forobservedcounts\begin{pmatrix}30&20\10&40\end{pmatrix},rowtotalsare50,50andcolumntotals40,60,soexpectedcountsare\begin{pmatrix}20&30\20&30\end{pmatrix}.ThusX^2=\frac{100}{20}+\frac{100}{30}+\frac{100}{20}+\frac{100}{30}=16.67.
Here df=1 and 16.67 exceeds the 5% critical value 3.841, so reject independence and conclude that there is evidence of an association. Every final expected cell must be at least 5; where needed, combine meaningful rows or columns before recalculating margins and df.
Use counts, not percentages, and do not apply Yates' correction in this syllabus. Association is not causation, and independence is not the same as mutually exclusive categories.
| feature | parametric mean test | non-parametric sign/rank test |
|---|---|---|
| main numerical information | original magnitudes | signs or order ranks |
| typical target | mean/model parameter | median, paired shift or identity of distributions |
| distributional demand | often normality and variance conditions | fewer shape conditions, but still design/test-specific assumptions |
| useful when | model assumptions are credible | normality is doubtful, data are skewed/outlier-prone or only ordering is reliable |
Choose from the sampling design first: one sample, matched pairs or two independent samples. Then identify whether only direction is defensible or whether ranks of magnitudes can be used. State population hypotheses before calculating a statistic.
Signs and ranks reduce sensitivity to extreme magnitudes and avoid a normal-mean model, but they discard some metric information. When a valid parametric model holds, that discarded information can make a non-parametric test less powerful.
Non-parametric does not mean assumption-free. Randomness and independence or valid pairing remain essential, and the Wilcoxon tests in this syllabus require symmetrical distributions.
| test | construct under H0 | statistic basis | key condition |
|---|---|---|---|
| sign | signs of observations-minus-median or paired differences | positive count B∼Bin(n,1/2) | independent signs/valid pairs |
| Wilcoxon signed-rank | rank absolute differences, restore signs | positive and negative rank sums | symmetric difference distribution |
| Wilcoxon rank-sum | pool two independent samples and rank all values | rank sum for a named sample | independent samples; symmetry condition in this syllabus |
For $n$ signed ranks with no zero differences or ties, the positive-rank sum $W^+$ has, under $H_0$,E(W^+)=\frac{n(n+1)}4,\qquad \operatorname{Var}(W^+)=\frac{n(n+1)(2n+1)}{24}.Thesesupportanormalapproximationwhenappropriate.
For independent samples of sizes $n_1,n_2$ and $N=n_1+n_2$, the rank sum $W_1$ hasE(W_1)=\frac{n_1(N+1)}2,\qquad \operatorname{Var}(W_1)=\frac{n_1n_2(N+1)}{12}.
Signed-rank and rank-sum are different tests: the first ranks within-pair or one-sample absolute differences, while the second ranks two independent samples together. The syllabus excludes tied ranks and zero differences in application questions.
| method | process for testing median m0 | information used |
|---|---|---|
| sign test | record signs of xi−m0 and use B∼Bin(n,1/2) | direction only |
| signed-rank | rank ∣xi−m0∣, restore signs and sum ranks | direction plus ordered magnitude; requires symmetry |
Test $H_0:m= m_0$ against $H_1:m>m_0$. If 10 of 12 observations exceed $m_0$, then under $H_0$P(B\ge10)=\frac{\binom{12}{10}+\binom{12}{11}+\binom{12}{12}}{2^{12}}=\frac{79}{4096}=0.0193.At5
For a large sign test, $B$ may be approximated byN\left(\frac n2,\frac n4\right),using a continuity correction. For signed-rank, use the null mean and variance from the previous card and correct the discrete rank-sum boundary by $0.5$ when using a normal approximation.
The alternative determines the tail before counting. These are tests about a population median or symmetric location, not a mean. Syllabus application questions contain no observations equal to the tested median and no tied ranks.
| data/design | appropriate test | statistic |
|---|---|---|
| paired, direction only | paired sign | number of positive differences |
| paired, symmetric differences | Wilcoxon matched-pairs signed-rank | signed rank sum of within-pair differences |
| two independent samples, syllabus symmetry condition met | Wilcoxon rank-sum | rank sum for a named sample after pooling |
For independent samples $n_1=20,n_2=25$, let $W_1=560$ and $N=45$. Under identical populations,E(W_1)=20(46)/2=460,\operatorname{Var}(W_1)=20(25)(46)/12=1916.67.Foranupper−tailnormalapproximation,continuitycorrectiongivesz=\frac{559.5-460}{\sqrt{1916.67}}=2.27.
Since 2.27 exceeds the 5% one-tail critical value 1.645, reject the null in the direction attached to sample 1's high ranks. For paired tests, define every difference in one order before applying the sign or signed-rank procedure.
Rank-sum is not a test for paired data, and matched-pairs signed-rank is not formed by ranking the two columns separately. Questions in this syllabus avoid tied ranks and zero-difference pairs; use exact tables instead of a normal approximation when the sample sizes make that appropriate.
For a non-negative integer-valued random variable $X$,G_X(s)=E[s^X]=\sum_{r=0}^{\infty}P(X=r)s^r.Thus $[s^r]G_X(s)=P(X=r)$, $G_X(0)=P(X=0)$ and $G_X(1)=1$.
| distribution and support | PGF |
|---|---|
| discrete uniform on 1,…,n | n(1−s)s(1−sn) for s=1, with G(1)=1 |
| Bin(n,p), q=1−p | (q+ps)n |
| geometric P(X=r)=pqr−1, r=1,2,… | 1−qsps |
| Po(λ) | exp(λ(s−1)) |
For example, if $G(s)=0.2+0.5s+0.3s^2$, thenP(X=0)=0.2,\quad P(X=1)=0.5,\quad P(X=2)=0.3.Thecoefficientrulealsoletsaclosedformbeexpandedtorecoverprobabilities.
State the geometric support convention: starting at one produces the numerator ps, while a failures-before-success convention starts at zero. A PGF is not an MGF, and G(1) must equal one for a valid probability distribution.
For $G(s)=E[s^X]$,E[X]=G'(1),\qquad E[X(X-1)]=G''(1).Since $X^2=X(X-1)+X$,\operatorname{Var}(X)=G''(1)+G'(1)-[G'(1)]^2.Differentiate before setting $s=1$.
For $X\sim\operatorname{Po}(\lambda)$,G(s)=e^{\lambda(s-1)},\quad G'(s)=\lambda e^{\lambda(s-1)},\quad G''(s)=\lambda^2e^{\lambda(s-1)}.Hence $G'(1)=\lambda$ and $G''(1)=\lambda^2$.
ThereforeE[X]=\lambda,and\operatorname{Var}(X)=\lambda^2+\lambda-\lambda^2=\lambda.ThecancellationchecksthefamiliarPoissonequalityofmeanandvariance.
G double-prime at one is not E[X squared]; it omits one copy of E[X]. If an alleged variance is negative, recheck differentiation, substitution at one and the subtraction of the squared mean.
If $X_1,\ldots,X_k$ are independent non-negative integer-valued variables and $S=\sum X_i$, thenG_S(s)=E\left[s^{\sum X_i}\right]=E\left[\prod s^{X_i}\right]=\prod G_{X_i}(s).Independencejustifiesthefactorisationoftheexpectation.
If $X\sim\operatorname{Po}(\lambda)$ and $Y\sim\operatorname{Po}(\mu)$ independently,G_{X+Y}(s)=e^{\lambda(s-1)}e^{\mu(s-1)}=e^{(\lambda+\mu)(s-1)},soX+Y\sim\operatorname{Po}(\lambda+\mu).
If $X\sim\operatorname{Bin}(n_1,p)$ and $Y\sim\operatorname{Bin}(n_2,p)$ independently,G_{X+Y}(s)=(1-p+ps)^{n_1+n_2},hence $X+Y\sim\operatorname{Bin}(n_1+n_2,p)$. If the success probabilities differ, multiply the PGFs but do not label the result binomial without further justification.
Multiplying PGFs requires independence. Means still add for dependent variables with finite expectations, but the PGF factorisation and the standard-family conclusions above can fail; coefficients of the actual product may be read when no named family results.