EuraStudy
Notes/Statistics/Normal distribution
Notes · StatisticsUK · A-Levels

Normal distribution

This topic introduces the normal distribution N(mu, sigma-squared) as a continuous model for data that cluster symmetrically about a mean. It covers the properties of the bell curve, the standard normal Z following N(0,1) and standardisation, finding probabilities as areas and quantiles by the inverse normal, recovering unknown parameters from given probabilities, and judging the suitability of the model.

5 sections·~11 min reading time·3 competencies·Level Foundation 1 · Standard 2 · Advanced 2

T·0888 / 18
Exam profile
AO1 · Standardise values and find normal probabilities and quantilesAO2 · Interpret areas under the curve and the meaning of standardisationAO3 · Find unknown parameters and judge whether the normal model fits a context
Operators:findcalculatestandardisedetermineshow thatassess

basic level

The AS foundation expects standardising to the standard normal and finding probabilities using tables or a calculator.

higher level

The full A-Level expects inverse-normal work, finding unknown mu or sigma (including two-equation problems), and judging the model.

Depth

Reading depth: In depth

Text

Text size: Standard

Contents · 5 sections▾
  1. Normal distribution
    • 01The normal model and its properties○
    • 02The standard normal and standardisation◐
    • 03Finding probabilities under the curve◐
    • 04Inverse normal and finding unknown parameters●
    • 05Judging the suitability of the normal model●
§ 01

The normal model and its properties#

●○○FoundationLPPearson Edexcel 9ST0, Topic 6 (Paper 1)

The normal curve with the central 95% region

Function graph, N(0,1) = (1/sqrt(2*pi))*exp(-x^2/2)Graph of N(0,1), maximum at (0, 0.399), y-intercept at y = 0.399, on the interval x from -4 to 4−4−3−2−112340.10.20.30.4−1.961.96N(0,1)densityz (standard deviations from m…
Fig. 1The central 95% of a normal distribution lies within 1.96 standard deviations of the mean.

Key points

The normal distribution is a continuous model whose density is the symmetric bell curve, written X∼N(μ,σ2)X \sim N(\mu, \sigma^2)X∼N(μ,σ2) where μ\muμ is the mean and σ2\sigma^2σ2 the variance. It is the natural model for quantities that arise as the sum of many small independent effects — heights, measurement errors, examination marks — which is why it appears throughout statistics and underlies the sampling theory of later topics.
The curve is symmetric about μ\muμ, so the mean, median and mode all coincide there, and the total area beneath it is 1. The parameter μ\muμ locates the centre and σ\sigmaσ controls the width: a larger σ\sigmaσ spreads the curve wider and flatter, a smaller σ\sigmaσ makes it tall and narrow, while the area always remains 1. Points of inflection occur one standard deviation either side of the mean, at μ±σ\mu \pm \sigmaμ±σ.
Probabilities are areas under the curve, so for a continuous variable P(X=a)=0P(X = a) = 0P(X=a)=0 and there is no distinction between strict and non-strict inequalities: P(X<a)=P(X≤a)P(X < a) = P(X \leq a)P(X<a)=P(X≤a). This is a genuine difference from the discrete distributions and removes the endpoint fussiness that dogs binomial and Poisson calculations.
The empirical (68-95-99.7) rule captures the spread: about 68% of the distribution lies within one standard deviation of the mean, about 95% within two, and about 99.7% within three. More precisely, the central 95% lies within 1.96σ1.96\sigma1.96σ of the mean — the value that generates the standard 95% confidence interval later. These landmarks let one sanity-check any normal probability at a glance.
X∼N(μ,σ2):P(μ−σ<X<μ+σ)≈0.68,P(μ−1.96σ<X<μ+1.96σ)=0.95X \sim N(\mu, \sigma^2): \quad P(\mu - \sigma < X < \mu + \sigma) \approx 0.68, \quad P(\mu - 1.96\sigma < X < \mu + 1.96\sigma) = 0.95X∼N(μ,σ2):P(μ−σ<X<μ+σ)≈0.68,P(μ−1.96σ<X<μ+1.96σ)=0.95

The normal model and the empirical rule

Symmetric about μ\muμ; 68% within one standard deviation, 95% within 1.96.

Worked example

Using the empirical rule

Adult heights are modelled by N(170,82)N(170, 8^2)N(170,82) cm. Use the empirical rule to describe the interval containing about 95% of heights.

  1. 01Identify the parameters

    μ=170\mu = 170μ=170 cm and σ=8\sigma = 8σ=8 cm.

  2. 02Apply the rule

    About 95% lie within 1.96σ=1.96×8=15.71.96\sigma = 1.96 \times 8 = 15.71.96σ=1.96×8=15.7 cm of the mean.

  3. 03State the interval

    170±15.7170 \pm 15.7170±15.7, i.e. roughly 154 cm to 186 cm.

    170±1.96(8)=(154.3,185.7)170 \pm 1.96(8) = (154.3, 185.7)170±1.96(8)=(154.3,185.7)

Result: About 95% of heights lie between roughly 154 cm and 186 cm.

Exam focus

  • State the properties of the normal curve (symmetry, area 1, inflection at μ±σ\mu \pm \sigmaμ±σ).
  • Use the empirical rule as a check on a computed probability.

Typical mistakes

  • Distinguishing P(X<a)P(X < a)P(X<a) from P(X≤a)P(X \leq a)P(X≤a) for a continuous variable; they are equal.
  • Confusing the variance σ2\sigma^2σ2 with the standard deviation σ\sigmaσ in N(μ,σ2)N(\mu, \sigma^2)N(μ,σ2).

Active revision

The masses of apples are modelled by N(150,202)N(150, 20^2)N(150,202) grams. Using the empirical rule, state an interval containing about 95% of the apples.

Active recall

Recall the key points — then reveal.

Sources: Pearson Edexcel Level 3 Advanced GCE in Statistics (9ST0) Specification (Pearson Edexcel)

§ 02

The standard normal and standardisation#

●●○StandardLPPearson Edexcel 9ST0, Topic 6 (Paper 1)

A standard normal tail probability

Function graph, (1/sqrt(2*pi))*exp(-x^2/2)Graph, maximum at (0, 0.399), y-intercept at y = 0.399, on the interval x from -4 to 4−4−3−2−112340.10.20.30.4z = 1.25densityz
Fig. 2P(Z < 1.25) = 0.8944 is the shaded area to the left of z = 1.25; the upper tail is 0.1056.

Key points

Every normal distribution can be converted to the standard normal Z∼N(0,1)Z \sim N(0, 1)Z∼N(0,1), which has mean 0 and standard deviation 1, by standardising: z=x−μσz = \dfrac{x - \mu}{\sigma}z=σx−μ​. The standardised value zzz counts how many standard deviations xxx lies above (positive) or below (negative) the mean. This single transformation reduces every normal probability to a question about ZZZ, for which tables and calculators are built.
Standardisation works because a linear transformation of a normal variable is still normal: subtracting μ\muμ shifts the mean to 0, and dividing by σ\sigmaσ scales the standard deviation to 1, without changing the shape. The area to the left of xxx under N(μ,σ2)N(\mu, \sigma^2)N(μ,σ2) equals the area to the left of zzz under N(0,1)N(0,1)N(0,1), so the two probabilities are identical.
The distribution function of the standard normal is written Φ(z)=P(Z≤z)\Phi(z) = P(Z \leq z)Φ(z)=P(Z≤z), and by symmetry Φ(−z)=1−Φ(z)\Phi(-z) = 1 - \Phi(z)Φ(−z)=1−Φ(z). The figure shades P(Z<1.25)=Φ(1.25)=0.8944P(Z < 1.25) = \Phi(1.25) = 0.8944P(Z<1.25)=Φ(1.25)=0.8944; the corresponding upper tail is P(Z>1.25)=1−0.8944=0.1056P(Z > 1.25) = 1 - 0.8944 = 0.1056P(Z>1.25)=1−0.8944=0.1056. A quick sketch of the required area, before reading any table, prevents the frequent tail-direction error.
In practice a calculator gives normal probabilities directly for any μ\muμ and σ\sigmaσ, but standardisation remains essential: it is the step required to find an unknown μ\muμ or σ\sigmaσ, to combine normal variables, and to connect to the zzz-based confidence intervals and tests. Understanding the zzz-value as a number of standard deviations is the key idea.
z=x−μσ,Φ(z)=P(Z≤z),Φ(−z)=1−Φ(z)z = \frac{x - \mu}{\sigma}, \qquad \Phi(z) = P(Z \leq z), \qquad \Phi(-z) = 1 - \Phi(z)z=σx−μ​,Φ(z)=P(Z≤z),Φ(−z)=1−Φ(z)

Standardisation and the standard normal cdf

The zzz-value is the number of standard deviations from the mean; Φ\PhiΦ is symmetric.

Worked example

Standardising to find a probability

Heights follow X∼N(170,82)X \sim N(170, 8^2)X∼N(170,82) cm. Find the probability that a randomly chosen adult is shorter than 180 cm.

  1. 01Standardise

    z=180−1708=1.25z = \dfrac{180 - 170}{8} = 1.25z=8180−170​=1.25.

  2. 02Read the probability

    P(X<180)=P(Z<1.25)=Φ(1.25)=0.8944P(X < 180) = P(Z < 1.25) = \Phi(1.25) = 0.8944P(X<180)=P(Z<1.25)=Φ(1.25)=0.8944.

    P(X<180)=Φ(180−1708)=Φ(1.25)=0.8944P(X < 180) = \Phi\left(\frac{180 - 170}{8}\right) = \Phi(1.25) = 0.8944P(X<180)=Φ(8180−170​)=Φ(1.25)=0.8944
  3. 03Interpret

    About 89.4% of adults are shorter than 180 cm under this model.

Result: P(X<180)=0.8944P(X < 180) = 0.8944P(X<180)=0.8944, so about 89.4% are below 180 cm.

Exam focus

  • Standardise a value with z=x−μσz = \frac{x - \mu}{\sigma}z=σx−μ​ and use Φ(−z)=1−Φ(z)\Phi(-z) = 1 - \Phi(z)Φ(−z)=1−Φ(z).
  • Sketch the required area to fix the tail direction before reading a probability.

Typical mistakes

  • Dividing by the variance σ2\sigma^2σ2 instead of the standard deviation σ\sigmaσ when standardising.
  • Reading the wrong tail, e.g. quoting Φ(z)\Phi(z)Φ(z) when 1−Φ(z)1 - \Phi(z)1−Φ(z) is required.

Active revision

For X∼N(170,82)X \sim N(170, 8^2)X∼N(170,82), standardise X=180X = 180X=180 and find P(X<180)P(X < 180)P(X<180).

Active recall

Recall the key points — then reveal.

Sources: Pearson Edexcel Level 3 Advanced GCE in Statistics (9ST0) Specification (Pearson Edexcel)

§ 03

Finding probabilities under the curve#

●●○StandardLPPearson Edexcel 9ST0, Topic 6 (Paper 1)

Key points

A probability over an interval is the area between two standardised bounds: P(a<X<b)=Φ ⁣(b−μσ)−Φ ⁣(a−μσ)P(a < X < b) = \Phi\!\left(\dfrac{b - \mu}{\sigma}\right) - \Phi\!\left(\dfrac{a - \mu}{\sigma}\right)P(a<X<b)=Φ(σb−μ​)−Φ(σa−μ​). Standardise both ends, read each cumulative probability, and subtract. A sketch of the shaded strip makes the subtraction unambiguous and guards against sign errors.
Upper-tail and 'between' probabilities are handled by the symmetry of the curve. P(X>b)=1−Φ ⁣(b−μσ)P(X > b) = 1 - \Phi\!\left(\dfrac{b - \mu}{\sigma}\right)P(X>b)=1−Φ(σb−μ​), and a central interval symmetric about the mean, such as the shaded 95% region in the previous section, uses Φ(z)−Φ(−z)=2Φ(z)−1\Phi(z) - \Phi(-z) = 2\Phi(z) - 1Φ(z)−Φ(−z)=2Φ(z)−1. Recognising symmetry often halves the work.
For continuous data the endpoints carry no probability, so 'more than', 'at least' and 'over' all standardise identically — a welcome simplification after the discrete distributions. The only care needed is to standardise consistently and to keep track of which side of the mean each bound lies.
Calculators return these areas directly for any μ\muμ and σ\sigmaσ, and in the examination that is the expected method; the standardised working is shown to communicate the reasoning and to make partial credit available. A well-labelled sketch plus the standardised expression is the model answer, even when the final number comes from technology.
P(a<X<b)=Φ ⁣(b−μσ)−Φ ⁣(a−μσ)P(a < X < b) = \Phi\!\left(\frac{b - \mu}{\sigma}\right) - \Phi\!\left(\frac{a - \mu}{\sigma}\right)P(a<X<b)=Φ(σb−μ​)−Φ(σa−μ​)

Interval probability

Standardise both bounds and subtract the cumulative probabilities.

Worked example

A 'between' probability

The contents of cartons follow X∼N(50,62)X \sim N(50, 6^2)X∼N(50,62) ml. Find P(44<X<59)P(44 < X < 59)P(44<X<59).

  1. 01Standardise both bounds

    z1=44−506=−1z_1 = \dfrac{44 - 50}{6} = -1z1​=644−50​=−1 and z2=59−506=1.5z_2 = \dfrac{59 - 50}{6} = 1.5z2​=659−50​=1.5.

  2. 02Read cumulative values

    Φ(1.5)=0.9332\Phi(1.5) = 0.9332Φ(1.5)=0.9332 and Φ(−1)=1−0.8413=0.1587\Phi(-1) = 1 - 0.8413 = 0.1587Φ(−1)=1−0.8413=0.1587.

  3. 03Subtract

    P(44<X<59)=0.9332−0.1587=0.7745P(44 < X < 59) = 0.9332 - 0.1587 = 0.7745P(44<X<59)=0.9332−0.1587=0.7745.

    P(44<X<59)=Φ(1.5)−Φ(−1)=0.7745P(44 < X < 59) = \Phi(1.5) - \Phi(-1) = 0.7745P(44<X<59)=Φ(1.5)−Φ(−1)=0.7745

Result: P(44<X<59)=0.7745P(44 < X < 59) = 0.7745P(44<X<59)=0.7745 (4 d.p.).

Exam focus

  • Find a 'between' probability by subtracting two standardised cumulative values.
  • Use symmetry for central and upper-tail probabilities.

Typical mistakes

  • Subtracting the cumulative values in the wrong order, giving a negative probability.
  • Adding tail areas that should be subtracted, or vice versa, from a mis-drawn sketch.

Active revision

For X∼N(50,62)X \sim N(50, 6^2)X∼N(50,62), find P(44<X<59)P(44 < X < 59)P(44<X<59).

Active recall

Recall the key points — then reveal.

Sources: Pearson Edexcel Level 3 Advanced GCE in Statistics (9ST0) Specification (Pearson Edexcel)

§ 04

Inverse normal and finding unknown parameters#

●●●AdvancedLPPearson Edexcel 9ST0, Topic 6 (Paper 1)

Key points

The inverse problem asks for the value xxx with a given probability below it. First find the zzz-value with Φ(z)\Phi(z)Φ(z) equal to the required probability — the inverse normal — then unstandardise with x=μ+zσx = \mu + z\sigmax=μ+zσ. For example, the 90th percentile of the standard normal is z=1.2816z = 1.2816z=1.2816, so the 90th percentile of N(μ,σ2)N(\mu, \sigma^2)N(μ,σ2) is μ+1.2816σ\mu + 1.2816\sigmaμ+1.2816σ.
The same zzz-values recur and are worth knowing: the central 90% uses ±1.6449\pm 1.6449±1.6449, the central 95% uses ±1.96\pm 1.96±1.96, and the central 99% uses ±2.5758\pm 2.5758±2.5758. These are exactly the critical values of the zzz-tests and confidence intervals in the inference topics, so fluency here transfers directly.
Finding an unknown μ\muμ or σ\sigmaσ reverses standardisation. If a single probability is given, one equation x−μσ=z\dfrac{x - \mu}{\sigma} = zσx−μ​=z determines the unknown parameter. If both μ\muμ and σ\sigmaσ are unknown, two given probabilities yield two simultaneous equations, which are solved by subtracting to eliminate μ\muμ and find σ\sigmaσ, then back-substituting.
Careful reading of the tail is essential: 'the heaviest 2%' fixes P(X>x)=0.02P(X > x) = 0.02P(X>x)=0.02, so Φ(z)=0.98\Phi(z) = 0.98Φ(z)=0.98 and z=2.0537z = 2.0537z=2.0537, whereas 'the lightest 2%' fixes P(X<x)=0.02P(X < x) = 0.02P(X<x)=0.02 and a negative zzz. A quick sketch showing which tail is meant prevents the sign errors that most often cost marks in these questions.
x=μ+zσwhereΦ(z)=P(X≤x)x = \mu + z\sigma \quad\text{where}\quad \Phi(z) = P(X \leq x)x=μ+zσwhereΦ(z)=P(X≤x)

Inverse normal (unstandardising)

Find zzz from the probability, then convert back to the xxx-scale.

Worked example

Finding an unknown mean and standard deviation

Bags of flour have normally distributed mass with 5% below 1000 g and 10% above 1050 g. Find the mean and standard deviation.

  1. 01Standardise both facts

    P(X<1000)=0.05P(X < 1000) = 0.05P(X<1000)=0.05 gives 1000−μσ=−1.6449\dfrac{1000 - \mu}{\sigma} = -1.6449σ1000−μ​=−1.6449; P(X>1050)=0.10P(X > 1050) = 0.10P(X>1050)=0.10 gives 1050−μσ=1.2816\dfrac{1050 - \mu}{\sigma} = 1.2816σ1050−μ​=1.2816.

  2. 02Eliminate the mean

    Subtracting the equations: 1050−1000σ=1.2816−(−1.6449)=2.9265\dfrac{1050 - 1000}{\sigma} = 1.2816 - (-1.6449) = 2.9265σ1050−1000​=1.2816−(−1.6449)=2.9265, so σ=502.9265=17.1\sigma = \dfrac{50}{2.9265} = 17.1σ=2.926550​=17.1 g.

    σ=502.9265=17.1 g\sigma = \frac{50}{2.9265} = 17.1\text{ g}σ=2.926550​=17.1 g
  3. 03Find the mean

    μ=1000+1.6449σ=1000+1.6449(17.08)=1028\mu = 1000 + 1.6449\sigma = 1000 + 1.6449(17.08) = 1028μ=1000+1.6449σ=1000+1.6449(17.08)=1028 g (3 s.f.).

  4. 04Check

    1050−102817.08=1.28\dfrac{1050 - 1028}{17.08} = 1.2817.081050−1028​=1.28, matching the required upper zzz.

Result: The mean is about 1028 g and the standard deviation about 17.1 g.

Exam focus

  • Use the inverse normal to find a percentile, then unstandardise to the xxx-scale.
  • Solve two simultaneous equations to find an unknown μ\muμ and σ\sigmaσ together.

Typical mistakes

  • Taking the wrong sign for zzz by misreading which tail the probability refers to.
  • Rearranging z=x−μσz = \frac{x - \mu}{\sigma}z=σx−μ​ incorrectly when solving for μ\muμ or σ\sigmaσ.

Active revision

For X∼N(500,252)X \sim N(500, 25^2)X∼N(500,252), find the value exceeded by only 5% of the distribution.

Active recall

Recall the key points — then reveal.

Sources: Pearson Edexcel Level 3 Advanced GCE in Statistics (9ST0) Specification (Pearson Edexcel)

§ 05

Judging the suitability of the normal model#

●●●AdvancedLPPearson Edexcel 9ST0, Topic 6 (Paper 1)

Key points

The normal model is appropriate when data are roughly symmetric and unimodal with no strong skew or heavy tails, and when values well away from the mean are genuinely plausible. Physical measurements, biological dimensions and aggregated errors often qualify. Skewed quantities such as incomes or waiting times usually do not, and forcing a normal model on them gives misleading probabilities.
Evidence for or against normality comes from the data's shape: a roughly symmetric histogram, mean and median close together, and a small skewness coefficient support the model, whereas marked skew, multimodality or outliers argue against it. A formal check, the chi-squared goodness-of-fit test against a normal distribution, is developed later in the course.
A practical caution is that the normal distribution has infinite tails, so it always assigns a small positive probability to impossible values — a normal model for heights technically allows a negative height. This is acceptable when the mean is many standard deviations from the impossible region, but it is a reason to treat the model as an approximation and to check that the tail in question is not spuriously large.
As always, the conclusion belongs in context: state whether the normal model is reasonable, cite the specific evidence, and note the limitation. The examination rewards a judgement backed by the shape of the data far more than an unexamined assumption of normality, and this critical stance is the bridge to the inference that assumes normality in the topics that follow.
Worked example

Assessing normality from the data

A sample of 200 reaction times has mean 0.29 s, median 0.26 s and a long right tail. Discuss whether a normal model is appropriate.

  1. 01Compare mean and median

    The mean (0.29) exceeds the median (0.26), indicating right skew.

  2. 02Consider the tail

    A long right tail of slow responses is asymmetric, unlike the symmetric normal curve.

  3. 03Conclude

    The normal model is not appropriate; the data are positively skewed, so a right-skewed model (or a transformation) would fit better.

Result: The right skew (mean above median, long upper tail) makes the normal model inappropriate for these reaction times.

Exam focus

  • Judge whether the normal model suits a data set, citing symmetry, skewness and the plausibility of tail values.
  • Recognise the infinite-tail limitation and when it matters.

Typical mistakes

  • Assuming normality for clearly skewed data such as incomes or waiting times.
  • Ignoring that the model gives non-zero probability to impossible values in the far tail.

Active revision

A data set of household incomes is strongly right-skewed. Explain why a normal model is inappropriate and suggest what the shape implies for the mean and median.

Active recall

Recall the key points — then reveal.

Sources: GCE AS and A level subject content (Statistics) (Department for Education / Ofqual)

Contents

Section -- / 05

    • 01The normal model and its properties○
    • 02The standard normal and standardisation◐
    • 03Finding probabilities under the curve◐
    • 04Inverse normal and finding unknown parameters●
    • 05Judging the suitability of the normal model●

0/5 Read

From notes into training

Normal distribution

Reinforce this topic with matching tasks from the question bank.

~11
min
3
Competencies
Practise

References & sources

Sources

Pearson Edexcel

  • Pearson Edexcel Level 3 Advanced GCE in Statistics (9ST0) Specification

Department for Education / Ofqual

  • GCE AS and A level subject content (Statistics)

Previous topic

Exponential and Poisson distributions

Next topic

Probability distributions

EuraStudy·Notes T·08·MMXXVI

Carry on to the next topic — your learning path is kept.