Q BankQuestion BankDocsDocuments

5.1 Representation of data

Syllabus
9709–2028–2029
Topic
5.1
Level
AS

Data representation should preserve scale, units and the variable type

Choose a display suited to the data: bar charts for categories, histograms for continuous intervals, and scatter plots for paired numerical variables. Axes need labels, units and honest scales.

Keep class widths visible in histograms and avoid implying continuity for categorical bars. A graph is a model of the data, not decoration.

A histogram with unequal class widths uses frequency density so each bar area represents frequency.

A histogram’s bar height is not always frequency; unequal widths require density.

Charts and diagrams communicate comparisons through a truthful visual scale

A chart should encode the intended quantity with a consistent scale, labelled axes and a legend when needed. Pie charts show parts of a whole; box plots show distribution summaries.

Check that categories do not overlap, totals match the denominator and truncated axes are clearly marked. Use the chart to compare evidence, not to infer causation.

A box plot’s median line and quartiles compare centre and spread without displaying every observation.

A visually larger sector or bar is meaningful only when the scale and category totals are comparable.

Averages describe centre while spread describes variation around it

Mean, median and mode describe location. Range, interquartile range, variance and standard deviation describe spread; each responds differently to outliers and skew.

Use the mean when all values and squared deviations are meaningful, the median for skewed or ordinal data, and state which spread measure matches the centre.

One extreme value can raise the mean and standard deviation while leaving the median and IQR almost unchanged.

A larger mean does not imply greater variability, and “average” is not automatically the arithmetic mean.

A cumulative-frequency graph estimates medians, quartiles and percentiles

Cumulative frequency totals observations up to each boundary. On a cumulative-frequency graph, read the median at N/2, quartiles at N/4 and 3N/4, and use differences to estimate an interquartile range.

Use class boundaries, draw a smooth monotone curve only as an approximation, and read values against the horizontal axis carefully.

For N=80, the upper quartile is the x-value at cumulative frequency 60; IQR is Q3−Q1.

Cumulative frequency itself is not a percentile value, and a graph gives an estimate rather than an exact raw-data quartile.

Grouped-data mean and standard deviation use class midpoints as estimates

For grouped data, estimate the mean with Σfx/Σf using class midpoints x, and use Σfx² to estimate variance. The result depends on treating each class as concentrated at its midpoint.

Keep a table of f, x, fx and fx², use the same units, and describe the answer as an estimate because within-class positions are unknown.

A class 10≤x<20 contributes frequency×15 to the estimated total, not frequency×10 or ×20.

Grouped-data statistics are not exact raw-data statistics; changing class widths or boundaries changes the estimate.

Objective notes

5 learning objectives
ConceptA-Level CAIE Mathematics AS