EuraStudy
This topic covers the summary of univariate data by measures of location (mean, median, mode) and spread (range, interquartile range, variance and standard deviation), including grouped data and the effect of linear coding. It then treats skewness, the detection of outliers, and the choice, construction and interpretation of the standard statistical diagrams: histograms with frequency density, box plots and cumulative frequency curves.
5 sections~14 min reading time3 competenciesLevel Foundation 1 · Standard 3 · Advanced 1
basic level
The AS foundation expects the measures of location and spread, standard deviation from the summation formulae, and reading the standard diagrams.
higher level
The full A-Level expects fluent use of coding, both coefficients of skewness, outlier rules, and the critical comparison of two distributions.
Reading depth: In depth
Text size: Standard
Mean of raw and frequency data
The second form weights each value (or class midpoint) by its frequency.
Journey times (minutes) are recorded: 0-10 (8 people), 10-20 (15), 20-40 (24), 40-70 (9). Estimate the mean journey time.
The class midpoints are 5, 15, 30 and 55 minutes.
.
The total frequency is , so the estimated mean is minutes (3 s.f.).
It is an estimate because we assumed each person's time equals their class midpoint.
Result: The estimated mean journey time is 26.4 minutes (3 s.f.), an estimate because grouped data hide the exact values.
Typical mistakes
Active revision
The times (minutes) taken by 40 pupils to complete a task are grouped. Estimate the mean using the class midpoints and explain why your answer is only an estimate.
Active recall
Recall the key points — then reveal.
Sources: Pearson Edexcel Level 3 Advanced GCE in Statistics (9ST0) Specification (Pearson Edexcel)
Variance and standard deviation
is the corrected sum of squares; the standard deviation has the units of the data.
Effect of linear coding
A shift leaves the spread unchanged; a scale multiplies both centre and spread.
Five values have and . Find the mean and the (population) standard deviation.
.
.
.
(3 s.f.).
Result: The mean is 12 and the standard deviation is (3 s.f.).
Typical mistakes
Active revision
For a data set with , and , find the mean and standard deviation.
Active recall
Recall the key points — then reveal.
Sources: Pearson Edexcel Level 3 Advanced GCE in Statistics (9ST0) Specification (Pearson Edexcel)
Box-and-whisker plot with an outlier
Quartile (fence) rule for outliers
With ; an alternative rule flags values more than 2 or 3 standard deviations from the mean.
For a data set and . Determine whether the value 72 is an outlier by the quartile rule.
.
.
Since , the value lies beyond the upper fence.
72 is an outlier by the quartile rule and should be plotted separately and investigated.
Result: The upper fence is 70, and because 72 exceeds it, 72 is an outlier.
Typical mistakes
Active revision
A data set has , and includes the value 72. Determine, using the rule, whether 72 is an outlier.
Active recall
Recall the key points — then reveal.
Sources: Pearson Edexcel Level 3 Advanced GCE in Statistics (9ST0) Specification (Pearson Edexcel)
Pearson coefficient of skewness
Positive for right skew, negative for left skew, zero for symmetry.
Quartile coefficient of skewness
Uses only the quartiles, so it is robust to outliers.
A distribution has , and . Find the quartile coefficient of skewness and interpret it.
.
.
.
The quartiles are symmetric about the median, so the distribution shows no skew by this measure.
Result: The quartile coefficient of skewness is 0, indicating a symmetric middle half of the distribution.
Typical mistakes
Active revision
A distribution has mean 52, median 55 and standard deviation 12. Calculate the Pearson coefficient of skewness and describe the skew.
Active recall
Recall the key points — then reveal.
Sources: GCE AS and A level subject content (Statistics) (Department for Education / Ofqual)
Histogram with unequal class widths
Frequency density (histogram)
The bar height for a histogram, so that area = frequency even with unequal class widths.
Cumulative frequency curve
For the journey-time data with total frequency 56 and cumulative frequencies 8, 23, 47, 56 at 10, 20, 40, 70 minutes, estimate the median.
The median is at cumulative frequency .
Cumulative frequency reaches 23 at 20 minutes and 47 at 40 minutes, so the median lies in the 20-40 class.
minutes.
About half of the journeys take under 24 minutes; it is an estimate because grouping hides the exact times.
Result: The estimated median journey time is about 24.2 minutes (from linear interpolation in the 20-40 class).
Typical mistakes
Active revision
Using the journey-time data (classes 0-10, 10-20, 20-40, 40-70 with frequencies 8, 15, 24, 9), estimate the median from the cumulative frequency curve.
Active recall
Recall the key points — then reveal.
Sources: Pearson Edexcel Level 3 Advanced GCE in Statistics (9ST0) Specification (Pearson Edexcel)
References & sources
Department for Education / Ofqual