EuraStudy
Notes/Statistics/Population and samples
Notes · StatisticsUK · A-Levels

Population and samples

This topic sets out the vocabulary of data collection: the population and its sampling frame, the difference between a census and a sample, and the standard random and non-random sampling methods. It stresses how a method is chosen for a real enquiry, and how bias and non-response threaten the validity of any conclusions drawn from a sample.

5 sections·~16 min reading time·3 competencies·Level Foundation 1 · Standard 3 · Advanced 1

T·0111 / 18
Exam profile
AO1 · Define and distinguish census, sample, sampling frame and the standard sampling methodsAO2 · Evaluate the suitability of a sampling method and identify sources of biasAO3 · Design and justify a sampling approach for a real statistical enquiry
Operators:definedescribeexplainidentifyassessjustifycomment

basic level

The AS foundation expects the definitions of population, census and sample and the three random sampling methods, with straightforward advantages and disadvantages.

higher level

The full A-Level expects fluent critique of a method in an unfamiliar context, recognition of subtle bias and non-response effects, and the design of a defensible sampling scheme within the statistical enquiry cycle.

Depth

Reading depth: In depth

Text

Text size: Standard

Contents · 5 sections▾
  1. Population and samples
    • 01Populations, censuses and samples○
    • 02Random sampling methods◐
    • 03Non-random sampling methods◐
    • 04Bias and non-response◐
    • 05Choosing a sampling method in context●
§ 01

Populations, censuses and samples#

●○○FoundationLPPearson Edexcel 9ST0, Topic 3 (Paper 1)

Key points

A population is the entire collection of individuals or items about which we want information — every voter in a constituency, every light bulb produced in a shift, every tree in a forest. A characteristic that varies across the population, such as height or lifetime, is a variable, and the number that summarises the whole population (its true mean, say) is a parameter. Almost all of statistics exists because we rarely have the resources to measure a whole population, so we must reason about the parameter from a smaller, manageable set of observations.
A census observes every member of the population. Its great strength is that it is complete and, in principle, gives the parameter exactly with no sampling error. Its weaknesses are practical and decisive: a census is expensive, slow, and often impossible — you cannot test every light bulb for its lifetime, because testing destroys it. Where the population is large, dynamic, or the test is destructive, a census is the wrong tool.
A sample is a subset of the population that is actually observed. The individual people or items selected are the sampling units, and the summary computed from the sample (the sample mean, for instance) is a statistic that we use to estimate the unknown parameter. A sampling frame is the list of all sampling units from which the sample is drawn — an electoral register, a list of employee numbers, a stock database. If the frame omits part of the population or is out of date, every sample drawn from it inherits that flaw, however careful the later selection.
The central trade-off is therefore between the completeness of a census and the economy of a sample. A sample is cheaper and faster and makes destructive or infinite-population testing possible, at the price of sampling error — the natural variation between one sample and another — and the risk of bias if the sample is not representative. The whole of sampling theory is about controlling those two risks.
Worked example

Census or sample?

A supermarket chain with 240 stores wants to know the average weekly waste of fresh produce per store. Discuss whether a census or a sample is appropriate, naming the population and the sampling units.

  1. 01Identify the population

    The population is the collection of all 240 stores (for a given week), and each store is a sampling unit. The variable is weekly produce waste.

  2. 02Weigh a census

    With only 240 stores and waste figures already recorded internally, a census is actually feasible and would give the true mean with no sampling error.

  3. 03Weigh a sample

    If the figures were costly to collect (for example, requiring a manual audit at each store), a sample of stores would be cheaper and faster, at the cost of sampling error.

  4. 04Conclude

    Because the population is small and the data are already held centrally, a census is reasonable here; a sample would be preferred only if measurement were expensive.

Result: A census is appropriate because the population is small (240 stores) and the data are already recorded; a sample would be justified only if collecting each figure were costly.

Exam focus

  • State clearly what the population, sampling units and sampling frame are in a described situation.
  • Give a specific, context-linked reason for preferring a sample to a census (for example, that testing is destructive).

Typical mistakes

  • Confusing the population with the sampling frame — the frame is the list you sample from, and it may not cover the whole population.
  • Claiming a census has no disadvantages; in most real contexts cost, time or destructive testing rule it out.

Active revision

A manufacturer wants to know the mean lifetime of the batteries it produces. Explain why a census is not sensible here, and identify the population, a sampling unit and a possible sampling frame.

Active recall

Recall the key points — then reveal.

Sources: Pearson Edexcel Level 3 Advanced GCE in Statistics (9ST0) Specification (Pearson Edexcel)

§ 02

Random sampling methods#

●●○StandardLPPearson Edexcel 9ST0, Topic 3 (Paper 1)

Classification of sampling methods

Sampling methodsProbability tree, 6 paths, Data: Random → Simple random; Random → Systematic; Random → Stratified; Non-random → Quota; Non-random → Opportunity; Non-random → Self-selectionRandomNon-randomSampling meth…Simple randomSystematicStratifiedQuotaOpportunitySelf-selection
Fig. 1Sampling methods divide into random (probability) methods and non-random methods.

Key points

In a random (probability) sampling method every unit has a known, non-zero probability of selection, which is what allows the theory of sampling distributions to be applied later. There are three standard random methods, distinguished by how that probability is realised. Understanding when each is appropriate — and its cost — is a recurring assessment demand.
Simple random sampling gives every possible sample of the required size an equal chance of being chosen; in practice each unit in the frame is numbered and a random number generator or random number table selects the sample without replacement. It is unbiased and simple to justify, but it needs a complete numbered frame and can, by chance, miss important subgroups — a simple random sample of a school might contain very few Year 7 pupils.
Systematic sampling orders the frame and then selects every kkkth unit after a random start, where k=Nnk = \frac{N}{n}k=nN​ is the sampling interval for a population of size NNN and a sample of size nnn. It is quick and spreads the sample evenly through the list, which is convenient at a production line or a queue. Its danger is periodicity: if the list has a repeating pattern with the same period as kkk, the sample can be badly unrepresentative.
Stratified sampling divides the population into non-overlapping strata (year groups, regions, age bands) and takes a simple random sample from each, with the number drawn from each stratum proportional to its size: the number sampled from a stratum is n×stratum sizeNn \times \frac{\text{stratum size}}{N}n×Nstratum size​. It guarantees that every subgroup is represented in proportion to its share of the population, which usually reduces sampling error, but it requires the strata to be known in advance and is more work to organise.
k=Nnk = \frac{N}{n}k=nN​

Systematic sampling interval

For a population of size NNN and a sample of size nnn; select every kkkth unit after a random start.

nstratum=n×stratum sizeNn_{\text{stratum}} = n \times \frac{\text{stratum size}}{N}nstratum​=n×Nstratum size​

Proportional stratified allocation

The number taken from each stratum is proportional to that stratum's share of the population.

Worked example

A stratified sample from two year groups

A college has 600 Year 12 and 400 Year 13 students (1000 in total). A stratified sample of 50 is required. Find the number from each year group and describe the selection.

  1. 01Find the sampling fraction

    The overall fraction is 501000=120\frac{50}{1000} = \frac{1}{20}100050​=201​, so each stratum is sampled at the same rate.

  2. 02Allocate Year 12

    50×6001000=3050 \times \frac{600}{1000} = 3050×1000600​=30 students.

    n12=50×6001000=30n_{12} = 50 \times \frac{600}{1000} = 30n12​=50×1000600​=30
  3. 03Allocate Year 13

    50×4001000=2050 \times \frac{400}{1000} = 2050×1000400​=20 students; the two parts sum to 50 as required.

  4. 04Describe selection

    Number the students in each year group and use random numbers to draw 30 from Year 12 and 20 from Year 13 without replacement.

Result: Take 30 from Year 12 and 20 from Year 13 (a total of 50), each by simple random sampling within the year group.

Exam focus

  • Describe precisely how to carry out each random method, including numbering the frame and using random numbers.
  • Compute a systematic interval kkk or a proportional stratified allocation and state the random start.

Typical mistakes

  • Describing systematic sampling without a random starting point — the start must be chosen at random within the first interval.
  • Rounding stratified allocations so the parts no longer sum to the intended sample size; adjust the rounding so the total is preserved.

Active revision

A college has 600 students in Year 12 and 400 in Year 13. Describe how to take a stratified sample of 50 students, stating how many come from each year group.

Active recall

Recall the key points — then reveal.

Sources: Pearson Edexcel Level 3 Advanced GCE in Statistics (9ST0) Specification (Pearson Edexcel)

§ 03

Non-random sampling methods#

●●○StandardLPPearson Edexcel 9ST0, Topic 3 (Paper 1)

Key points

Non-random methods select units without a known probability of inclusion. They are cheaper and quicker and are used when there is no sampling frame, but because selection is not random the standard theory of sampling distributions does not strictly apply, and the risk of bias is higher. The three named methods are quota, opportunity and self-selection sampling.
Quota sampling instructs the interviewer to fill fixed quotas that mirror the population's structure — say, 20 men and 20 women in given age bands — but leaves the choice of who fills each quota to the interviewer. It reproduces the population's proportions cheaply and without a frame, but the interviewer's discretion introduces bias (approachable people are over-selected) and the quotas themselves may be based on out-of-date figures.
Opportunity (or convenience) sampling simply takes whoever is available — the first thirty shoppers who pass, the students in one's own class. It is the easiest and cheapest method, but it is the most exposed to bias, because the people who happen to be available are rarely representative of the whole population.
Self-selection sampling relies on individuals volunteering, for example by returning a questionnaire or responding to an online poll. Volunteers tend to hold stronger opinions than non-volunteers, so self-selected samples are typically biased towards the views of those who care most, and the response rate is often low. In every non-random method, the key examined skill is to name the specific way bias enters and to say in which direction it is likely to distort the result.
Worked example

Diagnosing the bias in a phone-in poll

An online news site runs a poll asking readers whether a new bypass should be built, and 78% of the 4000 respondents say yes. Explain why this figure may overstate support among all residents.

  1. 01Name the method

    Respondents choose to take part, so this is self-selection sampling.

  2. 02Coverage bias

    Only readers of that site who saw the poll could respond, so residents without internet access or who read other sources are excluded.

  3. 03Volunteer bias

    People with strong feelings — often those campaigning for the bypass — are more likely to respond, inflating the apparent level of support.

  4. 04Conclude

    The 78% describes vocal respondents, not the resident population; it cannot be generalised without a properly designed random sample.

Result: The poll is a self-selected sample subject to coverage and volunteer bias, so 78% probably overstates true support among all residents.

Exam focus

  • Match a described procedure to the correct named method (quota, opportunity or self-selection).
  • Explain, in context, a specific mechanism by which the method introduces bias.

Typical mistakes

  • Confusing quota sampling with stratified sampling; stratified sampling selects randomly within strata, whereas quota sampling does not.
  • Saying only that a method is biased without naming who is over- or under-represented and why.

Active revision

A radio station invites listeners to phone in to vote on a local issue and reports the percentage in favour. Identify the sampling method and explain two reasons why the result may not represent the views of all residents.

Active recall

Recall the key points — then reveal.

Sources: Pearson Edexcel Level 3 Advanced GCE in Statistics (9ST0) Specification (Pearson Edexcel)

§ 04

Bias and non-response#

●●○StandardLPPearson Edexcel 9ST0, Topic 3 (Paper 1)

Key points

Bias is any systematic tendency for a sample statistic to differ from the population parameter — a distortion that does not average out as the sample grows. It must be distinguished sharply from sampling error, which is the ordinary random variation between samples and does shrink as nnn increases. Increasing the sample size reduces sampling error but never removes bias; a large biased sample is still biased.
Sampling (selection) bias arises when the selection mechanism systematically favours some units — an out-of-date frame that omits recent movers, or convenience sampling that reaches only the readily available. Measurement (response) bias arises when the way data are collected distorts the answers, for example a leading question, a poorly calibrated instrument, or respondents giving socially desirable rather than truthful answers.
Non-response bias occurs when those who do not respond differ systematically from those who do. If a survey on working hours has a low response rate, and the busiest people are precisely those least likely to reply, the sample will understate typical working hours regardless of its size. Reporting the response rate, and comparing respondents with known population figures, are the standard defences.
Good practice therefore attends to the whole chain: a complete and current sampling frame, a random selection method, neutral and well-tested measurement, and active follow-up to raise the response rate. In the statistical enquiry cycle, anticipating these threats at the planning stage is far more effective than trying to correct for them after the data are in.
Worked example

Why a bigger biased sample does not help

A researcher estimates the mean commuting time of a city's workers by surveying people leaving a city-centre railway station at 8:45 am. They plan to survey 5000 people to be 'accurate'. Explain why this will not give a reliable estimate.

  1. 01Spot the selection bias

    Only rail commuters arriving at that time are sampled; those who drive, cycle, walk or work different hours are systematically excluded.

  2. 02Effect of increasing n

    Surveying 5000 rather than 500 shrinks the random sampling error but leaves the systematic exclusion untouched.

  3. 03Direction of bias

    Rail commuters often travel longer distances, so the estimate is likely to be systematically too high.

  4. 04Conclude

    The estimate is biased; a representative frame of all workers and a random method are needed, not merely a larger sample.

Result: The design has selection bias that a larger sample cannot remove; the mean commuting time is likely to be systematically overestimated.

Exam focus

  • Distinguish clearly between bias (systematic) and sampling error (random), and explain that a larger sample cures only the latter.
  • Identify the type of bias (selection, measurement or non-response) present in a described study.

Typical mistakes

  • Believing a very large sample must be representative; size does not remove bias.
  • Treating non-response as harmless; it biases the result whenever non-responders differ from responders.

Active revision

A survey about satisfaction with a bus service is handed out on board the buses one weekday morning. Explain two distinct sources of bias in this design.

Active recall

Recall the key points — then reveal.

Sources: GCE AS and A level subject content (Statistics) (Department for Education / Ofqual)

§ 05

Choosing a sampling method in context#

●●●AdvancedLPPearson Edexcel 9ST0, Topic 3 (Paper 1)

Proportional stratified allocation

Stratified sample of 50Column chart: number sampled by year group, Data: Students sampled · Year 12 (600): 30; Students sampled · Year 13 (400): 20051015202530Year 12 (60…Year 13 (40…3020number sampledyear group
Fig. 2A stratified sample of 50 allocated in proportion to year-group sizes of 600 and 400.

Key points

Choosing a method is a judgement that balances representativeness against cost and feasibility, and the reason must always be tied to the specific context. The first questions are: is there a usable sampling frame; are there important subgroups that must be represented; is the population ordered in a way that makes systematic selection convenient or dangerous; and what resources are available? The answers point to a method.
If a complete frame exists and no subgroup is of special concern, simple random sampling is the natural, defensible choice. If the population contains distinct subgroups whose proportions matter — an examiner comparing schools, a pollster balancing regions — stratified sampling protects representation and usually lowers sampling error. If there is no frame at all, only a non-random method is possible, and the task becomes minimising and disclosing the resulting bias.
A stratified sample allocates the total sample across strata in proportion to their sizes, so the illustration below shows a sample of 50 split between two year groups of 600 and 400. Making that allocation explicit, and preserving the total after rounding, is exactly the kind of calculation the exam rewards.
Finally, whatever the method, the enquiry cycle demands honesty about limitations: state the frame used, the method and its known weaknesses, the response rate achieved, and any direction in which the estimate may be biased. A well-argued acknowledgement of a design's limits is worth more marks than a false claim that the sample is perfectly representative.
nstratum=n×stratum sizeNn_{\text{stratum}} = n \times \frac{\text{stratum size}}{N}nstratum​=n×Nstratum size​

Proportional allocation (recap)

Used to split the total sample nnn across strata in proportion to their sizes.

Worked example

Selecting a method for a multi-site survey

A charity operates 5 large centres and 20 small centres and wishes to sample 100 service users to estimate mean satisfaction, ensuring both centre types are represented. Recommend a method and outline the allocation.

  1. 01Identify the structure

    There are two clear strata (large and small centres) whose users may differ, so representation across both matters.

  2. 02Choose stratified sampling

    Stratify by centre type and allocate the 100-user sample in proportion to the number of users in each stratum.

  3. 03Illustrate allocation

    If large centres hold 60% of users and small centres 40%, take 100×0.6=60100 \times 0.6 = 60100×0.6=60 from large and 100×0.4=40100 \times 0.4 = 40100×0.4=40 from small, by simple random sampling within each.

  4. 04State a limitation

    It requires an up-to-date list of users per centre; without it, allocation and random selection are not possible.

Result: Use stratified sampling by centre type (here 60 from large, 40 from small), with the limitation that a current frame of users per centre is required.

Exam focus

  • Recommend a specific method for a described enquiry and justify it against the alternatives.
  • State the limitations of the chosen design, including any residual bias and how it might be reduced.

Typical mistakes

  • Recommending simple random sampling when no sampling frame exists — random selection needs a frame.
  • Giving a generic advantage instead of one tied to the context of the question.

Active revision

A supermarket wants to survey customer satisfaction across its three store formats (large, medium and small), which serve different numbers of customers. Recommend and justify a sampling method, and state one limitation.

Active recall

Recall the key points — then reveal.

Sources: Pearson Edexcel Level 3 Advanced GCE in Statistics (9ST0) Specification (Pearson Edexcel)

Contents

Section -- / 05

    • 01Populations, censuses and samples○
    • 02Random sampling methods◐
    • 03Non-random sampling methods◐
    • 04Bias and non-response◐
    • 05Choosing a sampling method in context●

0/5 Read

From notes into training

Population and samples

Reinforce this topic with matching tasks from the question bank.

~16
min
3
Competencies
Practise

References & sources

Sources

Pearson Edexcel

  • Pearson Edexcel Level 3 Advanced GCE in Statistics (9ST0) Specification

Department for Education / Ofqual

  • GCE AS and A level subject content (Statistics)

Next topic

Numerical measures, graphs and diagrams

EuraStudy·Notes T·01·MMXXVI

Carry on to the next topic — your learning path is kept.