10 Statistics and probability

Syllabus
2016
Topic
10
Level

Learning objectives

Choose and read data diagrams

Data / purpose Representation Essential scale
separate categories bar chart equal-width separated bars; height shows frequency
parts of a whole pie chart sector angle =360°×frequencytotal=360°\times\frac{frequency}{total}
continuous grouped data histogram touching bars; area shows frequency

\text{frequency density}=\frac{\text{frequency}}{\text{class width}},\qquad \text{frequency}=\text{density}\times\text{width}

Label axes and units, use a consistent scale, and interpret what bar height or area represents before extracting a value. Unequal histogram class widths require frequency density, not raw frequency, on the vertical scale.

A histogram is not a bar chart: its bars touch and their areas, not merely heights, represent frequencies. Cumulative frequency graphs are excluded from this syllabus.

Compare mean, median and mode

Measure Calculation / identification What it describes
mean xn\frac{\sum x}{n} or fxf\frac{\sum fx}{\sum f} balance point using every value
median middle ordered value; average the two middle values if needed central position
mode most frequent value most common outcome

Order raw data before finding median or mode. For a discrete frequency table, use cumulative running totals only to locate positions, multiply each value by its frequency for the mean, and divide by total frequency.

The mean uses all values but is affected by extremes; the median is resistant to extremes; the mode can describe the most common value but may be absent or multiple.

Do not divide fx\sum fx by the number of different values; divide by total frequency. These are exact measures for a discrete data set, not grouped-data estimates.

Estimate a mean from grouped data

m=\frac{\text{lower boundary}+\text{upper boundary}}2

\text{estimated mean}=\frac{\sum fm}{\sum f}

Find each class midpoint, multiply it by the class frequency, total the products, then divide by total frequency. Keep enough working precision and state that the result is an estimate.

The original values inside each interval are unknown, so the calculation assumes every observation in a class is represented by its midpoint.

Use class boundaries, not frequencies, to find midpoints. The answer is not exact unless the underlying values all equal their midpoints. Weighted and moving means will not be set.

Locate modal and median classes

The modal class is the class interval with the greatest frequency. If evidence is a histogram with unequal widths, recover frequency from bar area before comparing classes.

\text{median position lies around }\frac{N}{2}\text{ in the ordered grouped data}

Add frequencies cumulatively until the running total first reaches or passes the middle position. The interval containing that observation is the median class.

A class is an interval, so report the whole interval, not a midpoint. The tallest histogram bar is the modal class only when equal widths make height proportional to frequency.

Use probability language and complements

Idea Meaning
sample space every possible outcome
event a set of outcomes of interest
probability scale 00 impossible, 11 certain
relative frequency observed event count divided by trial count
complement event does not occur

P(A)=\frac{\text{favourable equally likely outcomes}}{\text{total equally likely outcomes}},\qquad P(\text{not }A)=1-P(A)

Relative frequency estimates probability from experiment. With more trials it often stabilises, but it need not equal the theoretical probability exactly.

Probabilities must lie from 0 to 1 and exhaustive outcome probabilities total 1. Counting favourable outcomes over total outcomes is valid only when outcomes are equally likely.

Add mutually exclusive probabilities

Mutually exclusive events cannot occur on the same trial, so they have no overlapping outcomes.

P(A\text{ or }B)=P(A)+P(B)\quad\text{when }A\text{ and }B\text{ are mutually exclusive}

For two or more disjoint events, add their probabilities. Check the events do not share an outcome and that the total does not exceed 1.

Do not use simple addition when events overlap: a shared outcome would be counted twice. Mutually exclusive describes whether events can happen together; it does not mean independent.

Multiply independent probabilities

Independent events do not change one another's probabilities. Repeating with replacement often produces independence; without replacement usually does not.

P(A\text{ and }B)=P(A)\times P(B)\quad\text{for independent }A,B

For a specified sequence of independent events, multiply the probability at every stage. Include complements when a stage says an event does not occur.

The product rule here requires independence. Do not assume events are independent merely because they occur at different times; check whether the first outcome changes the second probability.

Calculate independent events with trees

Draw one set of branches per stage, label every outcome, and put its probability on the branch. At each split, outgoing probabilities total 1.

Multiply along a path to find the probability of that complete outcome. Add the probabilities of mutually exclusive paths that satisfy the requested combined event.

P(AA)=P(A)P(A),\qquad P(\text{exactly one }A)=P(A)P(A')+P(A')P(A)

For independent repetitions, branch probabilities repeat at later stages. Do not add down a path or multiply different successful paths together.

Update probabilities along combined-event paths

In a combined experiment, the first outcome can change what remains available and therefore change later probabilities. A tree records the probability appropriate after each history.

Write first-stage probabilities, update counts or conditions separately on every second-stage branch, multiply along each complete path, then add the mutually exclusive paths matching the event.

P(\text{first red, then blue})=\frac{r}{n}\times\frac{b}{n-1}\quad\text{without replacement}

At every node, outgoing probabilities sum to 1; all terminal path probabilities together also sum to 1.

Do not reuse the original denominator after an item is removed. Conditional information changes the relevant sample space; independence must not be assumed.

Find probability within a restricted group

When told that an outcome lies in a particular group, discard outcomes outside that group and calculate within the reduced sample space.

\text{required probability}=\frac{\text{outcomes satisfying both facts}}{\text{outcomes satisfying the given fact}}

Circle or count the given group first, then count the favourable members inside it. A table, Venn diagram or short list can make the new denominator visible.

The denominator is the size or probability of the given group, not the original total. Use plain language such as ‘among the selected group’; the notation P(AB)P(A\mid B) will not be used.

Convert probability to expected frequency

\text{expected frequency}=\text{number of trials}\times\text{probability}

Expected frequency is the long-run average number of occurrences predicted over that many trials. It is a model-based expectation, not a guaranteed observed count.

If expected frequency and number of trials are known, divide to recover the probability. Check probability lies from 0 to 1 and expected frequency from 0 to the trial count.

An expected frequency may be non-integer even though an actual count must be whole. Do not round early or claim the event must occur exactly that many times.