EuraStudy
Notes/Psychology/Research methods
Notes · PsychologyUK · A-Levels

Research methods

Research methods are the scientific tools psychologists use to investigate mind and behaviour, and they underpin every study in the specification. This chapter covers experimental and non-experimental methods, sampling, the control of variables, reliability, validity and ethics, the features of science, and the descriptive and inferential statistics - including how to select and compute the sign test, chi-squared and Spearman's rho - that carry the mathematical requirement of the A-Level.

7 sections·~28 min reading time·3 competencies·Level Foundation 1 · Standard 2 · Advanced 4

T·0777 / 17
Exam profile
AO1 · Describe experimental and non-experimental methods, sampling, statistics and the features of scienceAO2 · Design a study and select and compute a statistical test for given data (the mathematical requirement)AO3 · Evaluate designs and methods for reliability, validity and ethics and interpret significance
Operators:describeexplainapplycalculatedesignevaluatejustify

basic level

AS-Level covers experimental design, non-experimental methods, sampling, control, reliability, validity, ethics and descriptive statistics.

higher level

The full A-Level adds the features of science and peer review, the normal distribution, and inferential testing including choosing and computing a statistical test and understanding Type I and Type II errors.

Depth

Reading depth: In depth

Text

Text size: Standard

Contents · 7 sections▾
  1. Research methods
    • 01Experimental design: variables, hypotheses and designs○
    • 02Non-experimental methods, sampling and pilot studies◐
    • 03Control, reliability, validity and ethics◐
    • 04Features of science and peer review●
    • 05Descriptive statistics and the normal distribution●
    • 06Inferential statistics: probability, significance and choosing a test●
    • 07Computing inferential tests: the sign test, chi-squared and Spearman's rho●
§ 01

Experimental design: variables, hypotheses and designs#

●○○FoundationLPAQA 7182 3.2.3LPDfE GCE Psychology - the experimental method

Key points

An experiment manipulates one variable to measure its effect on another. The independent variable (IV) is the variable the researcher deliberately changes; the dependent variable (DV) is the variable that is measured. Both must be operationalised - defined precisely in a way that can be measured (for example, 'memory' operationalised as 'the number of words recalled from a list of 20'). A hypothesis is a testable, precise statement of the expected outcome; a directional (one-tailed) hypothesis predicts the direction of the difference ('participants recall more words in the quiet condition'), whereas a non-directional (two-tailed) hypothesis predicts only that there will be a difference ('recall differs between the quiet and noisy conditions'). The null hypothesis states there is no effect, and is what the statistical test actually tests.
Anything other than the IV that could affect the DV must be considered. An extraneous variable is any variable, other than the IV, that could affect the DV and that we try to control (for example, the time of day). A confounding variable is an extraneous variable that has not been controlled and that varies systematically with the IV, so that we cannot tell whether the IV or the confound caused the change in the DV - a confound wrecks the experiment's validity. Demand characteristics (cues that reveal the aim, so participants change their behaviour) and investigator effects (the researcher unintentionally influencing the result) are important sources of extraneous influence.
There are three experimental designs. In an independent groups design, different participants do each condition; it avoids order effects but introduces participant variables (the groups may differ). In a repeated measures design, the same participants do every condition; it controls participant variables but introduces order effects (practice or fatigue), which are dealt with by counterbalancing. In a matched pairs design, different but matched participants (paired on relevant characteristics) do each condition, combining some advantages of both but being time-consuming and never a perfect match.
There are also types of experiment defined by where and how the IV is manipulated. A laboratory experiment manipulates the IV in a controlled setting (high control, high internal validity, but possibly artificial). A field experiment manipulates the IV in a natural setting (more realistic, less control). A natural experiment uses a naturally occurring IV that the researcher does not control (e.g. a natural disaster). A quasi-experiment uses an IV based on an existing difference between people (such as age or gender) that cannot be manipulated. In the last two the researcher does not manipulate the IV, which limits causal claims.
Worked example

Building a study from a research question

A psychologist wants to know whether caffeine improves reaction time. Operationalise the variables, write a directional hypothesis, and identify one confound to control.

  1. 01Operationalise the variables

    IV = whether the participant is given a caffeinated (200 mg) or a decaffeinated drink; DV = mean reaction time in milliseconds on a computer task.

  2. 02Write a directional hypothesis

    'Participants who drink the caffeinated drink will have a faster (lower) mean reaction time than those who drink the decaffeinated drink.'

  3. 03Identify and control a confound

    Time of day affects alertness and could confound the result, so all participants should be tested at the same time of day (standardisation).

Result: A well-operationalised, directional hypothesis with a named confound (time of day) controlled by standardisation.

Exam focus

  • Identify and operationalise the IV and DV in a study and write a suitable directional or non-directional hypothesis.
  • Choose an experimental design for a scenario and justify it against order effects and participant variables.

Typical mistakes

  • Writing a vague hypothesis - it must be operationalised and testable, not 'memory will be affected'.
  • Confusing an extraneous variable (uncontrolled but not systematic) with a confounding variable (varies with the IV).

Active revision

A researcher tests whether background music affects concentration. Identify the IV and DV, operationalise the DV, write a non-directional hypothesis, and choose a design with justification.

Active recall

Recall the key points — then reveal.

Sources: GCE AS and A level subject content for psychology (Department for Education) · AQA A-level Psychology 7182 specification (AQA)

§ 02

Non-experimental methods, sampling and pilot studies#

●●○StandardLPAQA 7182 3.2.3LPDfE GCE Psychology - non-experimental methods and sampling

Sampling methods compared

Five sampling methodsTable with 4 columns and 5 rows, Data: Method · How it works · Strength · Weakness; Random · equal chance for all (lottery) · unbiased selection · needs a full list; may be unrepresentative; Systematic · every nth member · objective · not truly random; Stratified · proportional subgroups · most representative · time-consuming; Opportunity · whoever is available · quick and easy · biased sample; Volunteer · self-selected via advert · reaches willing participants · volunteer biasMETHODHOW IT WORKSSTRENGTHWEAKNESSRandomequal chance for all(lottery)unbiased selectionneeds a full list; may beunrepresentativeSystematicevery nth memberobjectivenot truly randomStratifiedproportional subgroupsmost representativetime-consumingOpportunitywhoever is availablequick and easybiased sampleVolunteerself-selected via advertreaches willing participantsvolunteer bias
Fig. 1Sampling methods trade representativeness against practicality; stratified is the most representative, opportunity the most convenient.

Key points

Not all research is experimental. Observational techniques record behaviour as it occurs and may be naturalistic or controlled, covert or overt, participant or non-participant; they have high ecological validity but observer bias and no manipulation of an IV. Self-report methods gather data directly from participants through questionnaires (efficient, but open to social desirability and only as good as the questions) and interviews (structured, unstructured or semi-structured; richer, but time-consuming and open to interviewer effects). A correlational analysis measures the relationship between two co-variables without manipulating anything, expressed as a correlation coefficient - crucially, a correlation cannot establish cause and effect. Case studies are detailed investigations of a single person or small group (rich and useful for rare cases such as brain damage, but not generalisable), and content analysis turns qualitative material (such as television adverts) into quantitative data by coding it into categories.
A sample is the group of participants actually studied, drawn from the target population the researcher wants to generalise to. The goal is a representative sample so that findings generalise, and the specification names five sampling methods. Random sampling gives everyone in the population an equal chance of selection (e.g. drawing names from a hat or using a random number generator); it is unbiased but can still, by chance, be unrepresentative and requires a full list of the population. Systematic sampling selects every nth member from a list (e.g. every 10th name); it is objective but not truly random.
Stratified sampling divides the population into strata (subgroups, such as age bands) and samples from each in proportion to its size in the population; it produces the most representative sample but is time-consuming and requires knowing the strata. Opportunity sampling uses whoever is available and willing; it is quick and convenient but highly biased (unrepresentative of the wider population). Volunteer (self-selected) sampling relies on participants responding to an advert; it is easy and reaches motivated participants but attracts a particular type of person (volunteer bias).
Before the main study, a pilot study - a small-scale trial run - is carried out to check that the procedures, materials and measures work, so that problems (ambiguous questions, flawed timings) can be corrected before investing in the full study. This saves time and money and improves the quality of the design. Choosing a sampling method involves a trade-off between representativeness and practicality, and the exam often asks students to identify a method, state one strength and one weakness, and suggest how a sample was, or should be, obtained.
Worked example

Selecting a stratified sample

A school of 1,000 students has 600 girls and 400 boys. A researcher wants a stratified sample of 50. How many girls and boys should be selected?

  1. 01Find each stratum's proportion

    Girls are 600/1000 = 0.6 of the school; boys are 400/1000 = 0.4.

  2. 02Apply the proportions to the sample

    Girls: 0.6 x 50 = 30; boys: 0.4 x 50 = 20.

    girls=6001000×50=30,boys=4001000×50=20\text{girls} = \frac{600}{1000} \times 50 = 30, \quad \text{boys} = \frac{400}{1000} \times 50 = 20girls=1000600​×50=30,boys=1000400​×50=20
  3. 03Check

    30 + 20 = 50, matching the required sample size, with each subgroup represented in proportion to the population.

Result: Select 30 girls and 20 boys - a stratified sample proportional to the school population.

Exam focus

  • Distinguish the five sampling methods and give a strength and weakness of each.
  • Explain why a correlation cannot show cause and effect and what a pilot study is for.

Typical mistakes

  • Calling opportunity sampling 'random' - random sampling requires every member to have an equal chance.
  • Claiming a strong correlation proves one variable causes the other.

Active revision

A researcher stops shoppers in one town centre to complete a questionnaire. Identify the sampling method, give one weakness, and suggest a more representative method.

Active recall

Recall the key points — then reveal.

Sources: GCE AS and A level subject content for psychology (Department for Education) · AQA A-level Psychology 7182 specification (AQA)

§ 03

Control, reliability, validity and ethics#

●●○StandardLPAQA 7182 3.2.3LPDfE GCE Psychology - reliability, validity and ethics

Key points

Control keeps everything except the IV constant so that any change in the DV can be attributed to the IV. Randomisation uses chance to decide the order of conditions or materials (reducing investigator bias), and standardisation keeps the procedure identical for every participant (using the same instructions and conditions). Single-blind (participants do not know the condition) and double-blind (neither participants nor the researcher administering the study know) procedures reduce demand characteristics and investigator effects.
Reliability means consistency: a measure is reliable if it gives the same result when repeated. Internal reliability (consistency within a measure) can be checked by the split-half method (comparing two halves of a test); external reliability (consistency over time) by test-retest (giving the same test again later); and inter-observer reliability (consistency between observers) by correlating two observers' records - a correlation of about +0.8 or above is usually taken as acceptable. Reliability is improved by standardising procedures, training observers and operationalising behavioural categories clearly.
Validity means accuracy or truth: whether a study measures what it claims to. Internal validity is whether the effect was really due to the IV (threatened by confounds and demand characteristics); external validity is whether the findings generalise, including ecological validity (to other settings and everyday life), population validity (to other people) and temporal validity (to other times). Validity is assessed by face validity and concurrent validity (comparing with an established measure) and improved by controlling confounds, using a control group and lifelike tasks.
Ethical issues arise when the interests of participants conflict with the aims of the research, and psychologists follow the British Psychological Society (BPS) code of ethics. The main issues and their remedies are: informed consent (tell participants enough to decide, and gain agreement); deception (avoid it unless justified, and debrief afterwards); protection from harm (participants should leave in the state they arrived, with the right to withdraw); privacy and confidentiality (protect personal data and anonymise it). Where deception is necessary, cost-benefit analysis, ethics committees, presumptive consent and full debriefing are used to manage it. Reliability, validity and ethics are the three lenses through which almost every study in the specification is evaluated.
Worked example

Assessing inter-observer reliability

Two observers independently record the number of aggressive acts in ten short film clips, and their scores correlate at +0.65. State whether this is acceptable and suggest how to improve it.

  1. 01Apply the criterion

    Inter-observer reliability is usually judged acceptable at a correlation of about +0.8 or above.

  2. 02Evaluate

    A correlation of +0.65 is below this threshold, so the observers are not recording consistently and the measure is unreliable.

  3. 03Improve

    Operationalise 'aggression' into clearer behavioural categories and train the observers together beforehand, then re-check the correlation.

Result: +0.65 is below the +0.8 guideline, so reliability is inadequate; clearer categories and observer training should raise it.

Exam focus

  • Distinguish reliability (consistency) from validity (accuracy) and name a way to assess and improve each.
  • Identify an ethical issue in a study and state the appropriate way of dealing with it from the BPS code.

Typical mistakes

  • Confusing reliability (consistency) with validity (accuracy) - a measure can be reliable but not valid.
  • Confusing ecological validity (settings/tasks) with population validity (the people sampled).

Active revision

Two observers rating aggression in a playground agree on only 60% of instances. Identify the problem, and suggest two ways to improve it.

Active recall

Recall the key points — then reveal.

Sources: GCE AS and A level subject content for psychology (Department for Education) · AQA A-level Psychology 7182 specification (AQA)

§ 04

Features of science and peer review#

●●●AdvancedLPAQA 7182 3.2.3LPDfE GCE Psychology - features of science

Key points

For Psychology to count as a science it must share the features of the natural sciences. Empiricism means knowledge comes from direct observation and measurement rather than opinion. Objectivity means observations are free from the researcher's bias, which is why controlled methods and standardisation matter. Replicability means a study can be repeated to check its findings, which depends on detailed, standardised reporting. Together these make the empirical method the basis of scientific psychology.
Two ideas from the philosophy of science are named. Falsifiability (Popper) is the principle that a genuine scientific theory must be capable of being proved wrong: a theory that can explain any possible outcome (Popper's criticism of some psychodynamic ideas) is not scientific. Scientists therefore try to test and refute theories, and a theory survives by resisting attempts to falsify it. Theory construction and hypothesis testing describe the cycle by which science advances: observations lead to a theory, from which testable hypotheses are deduced, tested and used to revise the theory.
Kuhn added the idea of a paradigm - a set of shared assumptions and methods that defines a science at a given time. Kuhn argued that a mature science has a single dominant paradigm, and that scientific progress can involve a paradigm shift, when accumulating anomalies overturn the old paradigm and a new one replaces it. Kuhn suggested Psychology is 'pre-paradigmatic' because it has several competing approaches rather than one agreed paradigm, which is used both to question and to defend its scientific status.
Peer review is the process by which research is evaluated by other experts in the field before it is published. Its purposes are to check the quality and validity of the research, to allocate research funding, and to help maintain the integrity of published science. It is evaluated as essential but imperfect: reviewers may be biased (against unconventional findings or towards established researchers), anonymity can be misused, and it tends to favour positive results and established theories (publication bias), which can slow the acceptance of genuinely new ideas. The economic implications of research are also on the specification: psychological research contributes to the economy, for example by improving the treatment of mental illness (reducing absence from work) and informing effective childcare and education.
Worked example

Applying falsifiability

A theory claims that all behaviour is driven by unconscious wishes, and interprets any behaviour - and its opposite - as evidence for this. Explain why this is a problem for the theory's scientific status.

  1. 01State the principle

    Popper argued a scientific theory must be falsifiable - it must make predictions that could, in principle, be shown to be wrong.

  2. 02Apply it

    If the theory can explain a behaviour and its exact opposite equally well, then no observation could ever contradict it.

  3. 03Draw the conclusion

    Because nothing could falsify it, the theory fails the test of falsifiability and cannot be considered scientific in Popper's sense, however insightful it may be.

Result: A theory that cannot be falsified is not scientific by Popper's criterion, regardless of its explanatory appeal.

Exam focus

  • Define empiricism, objectivity, replicability and falsifiability and apply them to whether a theory is scientific.
  • Explain the purpose of peer review and evaluate its limitations.

Typical mistakes

  • Saying a theory is scientific because it explains everything - a scientific theory must be falsifiable (capable of being disproved).
  • Confusing a paradigm (shared assumptions) with a single theory.

Active revision

Using the features of science, discuss whether the psychodynamic approach can be considered scientific.

Active recall

Recall the key points — then reveal.

Sources: GCE AS and A level subject content for psychology (Department for Education) · AQA A-level Psychology 7182 specification (AQA)

§ 05

Descriptive statistics and the normal distribution#

●●●AdvancedLPAQA 7182 3.2.3LPDfE GCE Psychology - descriptive statistics

The normal distribution and the empirical rule

Function graph, distribution = exp(-(x*x)/2)Graph of distribution, maximum at (0, 1), y-intercept at y = 1, on the interval x from -4 to 4−4−3−2−112340.20.40.60.81−1 SD+1 SD−2 SD+2 SDdistributionfrequencystandard deviations from the …
Fig. 2About 68% of scores lie within one standard deviation of the mean and about 95% within two - the empirical rule.

Key points

Descriptive statistics summarise data. Measures of central tendency give a typical value: the mean (the arithmetic average, xˉ=∑xn\bar{x} = \frac{\sum x}{n}xˉ=n∑x​) uses all the data but is distorted by extreme values; the median (the middle value when ordered) is unaffected by extremes but ignores their size; the mode (the most frequent value) is the only measure that can be used with categorical data. The choice depends on the level of measurement and the shape of the data.
Measures of dispersion describe the spread. The range (highest minus lowest, usually plus one) is quick but affected by extremes and ignores the middle values. The standard deviation is more informative: it is the average amount by which scores deviate from the mean, so a larger standard deviation means more spread-out (more variable) data. It uses every score, and a small standard deviation indicates the scores cluster tightly around the mean - useful when comparing the consistency of two conditions.
Levels of measurement classify data and determine which statistics and tests are appropriate. Nominal data are in named categories (e.g. counts of people preferring each of three brands). Ordinal data can be ranked or ordered but the intervals are not equal (e.g. positions in a race, or ratings on a subjective scale). Interval data are measured on a scale with equal intervals (e.g. temperature, reaction time in milliseconds). This hierarchy - nominal, ordinal, interval - is central to choosing a statistical test in the next sections.
Many psychological variables (such as IQ) follow a normal distribution - a symmetrical, bell-shaped curve in which the mean, median and mode all fall at the centre. Its defining feature is captured by the empirical rule: about 68% of scores lie within one standard deviation of the mean, about 95% within two standard deviations, and about 99.7% within three. This is the basis of the statistical-infrequency definition of abnormality and of understanding how unusual a particular score is. A skewed distribution is asymmetrical: in a positive skew most scores are low with a long tail of high scores (mode < median < mean), and in a negative skew the reverse.
xˉ=∑xn\bar{x} = \frac{\sum x}{n}xˉ=n∑x​

Mean

Add all the scores and divide by the number of scores.

s=∑(x−xˉ)2n−1s = \sqrt{\dfrac{\sum (x - \bar{x})^2}{n - 1}}s=n−1∑(x−xˉ)2​​

Standard deviation

The square root of the mean squared deviation from the mean; dividing by n-1 gives the (unbiased) sample standard deviation.

Comparing the spread of two conditions

Recall scores by conditionBox plot: words recalled by condition, Data: Quiet: min 11, Q1 14, Md 16, Q3 18, max 20; Noisy: min 6, Q1 9, Md 11, Q3 14, max 1968101214161820QuietNoisywords recalledcondition
Fig. 3A boxplot displays the median, quartiles and range; the quiet condition here has a higher median and a smaller spread (illustrative data).
Worked example

Calculating the mean and standard deviation

Five participants score 4, 6, 8, 5 and 7 on a memory test. Calculate the mean and the (sample) standard deviation.

  1. 01Calculate the mean

    Sum = 4 + 6 + 8 + 5 + 7 = 30; mean = 30 / 5 = 6.

    xˉ=305=6\bar{x} = \frac{30}{5} = 6xˉ=530​=6
  2. 02Find the squared deviations

    Deviations from 6 are -2, 0, 2, -1, 1; squared they are 4, 0, 4, 1, 1, which sum to 10.

  3. 03Divide by n-1 and take the square root

    10 / (5 - 1) = 2.5; the square root of 2.5 is about 1.58.

    s=104=2.5≈1.58s = \sqrt{\dfrac{10}{4}} = \sqrt{2.5} \approx 1.58s=410​​=2.5​≈1.58

Result: The mean is 6 and the sample standard deviation is about 1.58.

Exam focus

  • Select and justify a measure of central tendency and dispersion for a given data set and level of measurement.
  • Use the empirical rule (68-95-99.7) to judge how unusual a score is, and identify skew.

Typical mistakes

  • Using the mean with skewed data or with ordinal data - the median is usually more appropriate there.
  • Confusing the direction of skew - in a positive skew the long tail is on the high (right) side.

Active revision

For a set of reaction times with one very slow outlier, state which measure of central tendency and which measure of dispersion you would report, and why.

Active recall

Recall the key points — then reveal.

Sources: GCE AS and A level subject content for psychology (Department for Education) · AQA A-level Psychology 7182 specification (AQA)

§ 06

Inferential statistics: probability, significance and choosing a test#

●●●AdvancedLPAQA 7182 3.2.3LPDfE GCE Psychology - inferential testing

Choosing a statistical test

Statistical test decision treeProbability tree, 8 paths, Data: Difference → Nominal → Unrelated: chi-squared; Difference → Nominal → Related: sign test; Difference → Ordinal → Unrelated: Mann-Whitney; Difference → Ordinal → Related: Wilcoxon; Difference → Interval → Unrelated: t-test; Difference → Interval → Related: t-test; Correlation → Ordinal: Spearman's rho; Correlation → Interval: Pearson's rNominalOrdinalIntervalDifferenceCorrelationResearch aim?Unrelated: chi-squaredRelated: sign testUnrelated: Mann-WhitneyRelated: WilcoxonUnrelated: t-testRelated: t-testOrdinal: Spearman's rhoInterval: Pearson's r
Fig. 4Answer three questions - difference or correlation, level of measurement, related or unrelated - to reach the correct test.

Key points

Inferential statistics let us decide whether a result is likely to be genuine or just due to chance. Because we cannot be 100% certain, we work with probability: the significance level is the probability that the result occurred by chance, and psychology conventionally uses p≤0.05p \leq 0.05p≤0.05 (a 5% or less probability that the result is due to chance). If the test shows the result would occur by chance 5% of the time or less, we reject the null hypothesis and accept that the effect is 'significant'; otherwise we retain the null hypothesis.
Two kinds of error can occur. A Type I error (a false positive) is rejecting the null hypothesis when it is actually true - claiming an effect that is not really there; it is more likely if the significance level is too lenient (e.g. p≤0.10p \leq 0.10p≤0.10). A Type II error (a false negative) is retaining the null hypothesis when it is actually false - missing a real effect; it is more likely if the significance level is too stringent (e.g. p≤0.01p \leq 0.01p≤0.01). The 0.05 level is a compromise that balances the two risks.
Every test yields a calculated (observed) value, which is compared with a critical value from a statistical table. Finding the right critical value requires three things: the significance level, whether the hypothesis is one- or two-tailed, and the degrees of freedom or the sample size N. Depending on the test, the result is significant either when the calculated value is greater than or equal to the critical value (e.g. chi-squared, Spearman's rho) or less than or equal to it (e.g. the sign test, Mann-Whitney, Wilcoxon) - so it is essential to know the rule for the particular test.
Choosing the correct test depends on three questions: (1) Is the study looking for a difference or a correlation (association)? (2) What is the level of measurement of the data - nominal, ordinal or interval? (3) For a test of difference, is the design related (repeated measures or matched pairs) or unrelated (independent groups)? Answering these three questions leads to a single appropriate test, as the decision tree below shows. Learning to justify a test choice from these three features is a guaranteed exam skill.
p≤0.05p \leq 0.05p≤0.05

Significance level

The conventional 5% threshold: a result is significant if the probability it arose by chance is 0.05 or less.

Worked example

Selecting a test from a scenario

Researchers compare the number of words recalled (a difference) by two separate groups of participants (independent groups), where the data are interval. Which test should they use?

  1. 01Difference or correlation?

    They are comparing two conditions for a difference, not measuring an association - so a test of difference is needed.

  2. 02Level of measurement?

    Words recalled counted on an equal-interval scale is interval data.

  3. 03Related or unrelated?

    Two separate groups of participants means an unrelated (independent groups) design; a test of difference on interval, unrelated data is the unrelated t-test.

Result: The unrelated (independent) t-test, chosen because it is a test of difference on interval data from an unrelated design.

Exam focus

  • Select the correct statistical test from the three deciding features and justify the choice.
  • Explain Type I and Type II errors and how the significance level affects the risk of each.

Typical mistakes

  • Assuming all tests are significant when the calculated value is larger - some (sign test, Mann-Whitney, Wilcoxon) require it to be smaller than or equal to the critical value.
  • Confusing a Type I error (false positive) with a Type II error (false negative).

Active revision

A study using a repeated measures design measures anxiety on an ordinal rating scale before and after therapy. Name the test that should be used and justify it from the three deciding features.

Active recall

Recall the key points — then reveal.

Sources: GCE AS and A level subject content for psychology (Department for Education) · AQA A-level Psychology 7182 specification (AQA)

§ 07

Computing inferential tests: the sign test, chi-squared and Spearman's rho#

●●●AdvancedLPAQA 7182 3.2.3LPDfE GCE Psychology - inferential testing

A positive correlation (Spearman's rho)

Hours studied against test scoreScatter plot: test score by hours studied, Data: (5, 40); (8, 50); (3, 35); (9, 60); (6, 55); (2, 30)3035404550556023456789test scorehours studied
Fig. 5The scatter for the Spearman worked example: as hours studied rise, test score tends to rise - a strong positive correlation.

Key points

The sign test is the one test AQA requires students to be able to compute in full. It is used for a test of difference, with nominal data, from a related design (the same participants measured twice). The procedure is: work out the sign (+ or -) of the change for each participant, discard any participants who show no change, count how many of each sign there are, and take the calculated value S as the number of the less frequent sign. N is the number of participants left after discarding ties. S is significant (the result is genuine) if it is less than or equal to the critical value from the sign-test table for that N, significance level and number of tails.
Chi-squared (χ2\chi^2χ2) is a test of difference (or association) for nominal data from an unrelated design, applied to a contingency table of frequencies. The observed frequencies (O) are compared with the frequencies expected if the null hypothesis were true (E), where each cell's E=row total×column totalgrand totalE = \frac{\text{row total} \times \text{column total}}{\text{grand total}}E=grand totalrow total×column total​, using χ2=∑(O−E)2E\chi^2 = \sum \frac{(O-E)^2}{E}χ2=∑E(O−E)2​. The degrees of freedom are (rows−1)(columns−1)(\text{rows}-1)(\text{columns}-1)(rows−1)(columns−1); for a 2x2 table df=1df = 1df=1 and the critical value at p≤0.05p \leq 0.05p≤0.05 is 3.84. The result is significant when the calculated χ2\chi^2χ2 is greater than or equal to the critical value.
Spearman's rho (rsr_srs​) is a test of correlation for ordinal data (or interval data treated as ordinal). Each of the two co-variables is ranked, the difference d between the two ranks is found for each pair, and rs=1−6∑d2n(n2−1)r_s = 1 - \frac{6\sum d^2}{n(n^2-1)}rs​=1−n(n2−1)6∑d2​ is calculated, where n is the number of pairs. The coefficient ranges from -1 (perfect negative) through 0 (no correlation) to +1 (perfect positive); it is significant when the calculated value is greater than or equal to the critical value for that n and significance level.
Interpreting the outcome is as important as the calculation. Whichever test is used, the final step is always to compare the calculated value with the correct critical value and to state, in the context of the study, whether the null hypothesis is rejected or retained. A significant result lets the researcher reject the null hypothesis and accept the alternative (there is a real effect or relationship); a non-significant result means the null hypothesis is retained. Reporting should also acknowledge the possibility of Type I or Type II errors.
χ2=∑(O−E)2E\chi^2 = \sum \frac{(O - E)^2}{E}χ2=∑E(O−E)2​

Chi-squared

Sum, over every cell, the squared difference between observed and expected frequencies divided by the expected frequency.

rs=1−6∑d2n(n2−1)r_s = 1 - \dfrac{6 \sum d^2}{n(n^2 - 1)}rs​=1−n(n2−1)6∑d2​

Spearman's rho

d is the difference between the two ranks for each pair; n is the number of pairs.

Worked example

A full sign-test calculation

Fifteen participants rate their mood before and after a therapy. Twelve improve (+), three get worse (-) and none stay the same. Test at p <= 0.05 two-tailed (critical S for N = 15 is 3).

  1. 01Assign signs and discard ties

    Improvements are + (12), deteriorations are - (3); no participant showed no change, so none is discarded and N = 15.

  2. 02Find S

    S is the number of the less frequent sign; here the minus signs are fewer, so S = 3.

  3. 03Compare with the critical value

    The critical value of S for N = 15, two-tailed, p <= 0.05 is 3; the result is significant if S is less than or equal to this. Since 3 <= 3, the result is significant.

  4. 04State the conclusion

    We reject the null hypothesis: the therapy produced a significant change in mood (p <= 0.05).

Result: S = 3, which is less than or equal to the critical value of 3, so the result is significant and the null hypothesis is rejected.

Worked example

A chi-squared calculation

In a study of two teaching methods, of 40 students taught by Method A, 30 passed and 10 failed; of 40 taught by Method B, 20 passed and 20 failed. Test for an association at p <= 0.05 (df = 1, critical value 3.84).

  1. 01Find the expected frequencies

    Each cell's E = (row total x column total) / grand total. With row totals 40 and 40, column totals 50 (pass) and 30 (fail), grand total 80: each pass cell E = 40 x 50 / 80 = 25; each fail cell E = 40 x 30 / 80 = 15.

  2. 02Apply the formula

    Chi-squared = (30-25)^2/25 + (10-15)^2/15 + (20-25)^2/25 + (20-15)^2/15 = 1 + 1.67 + 1 + 1.67 = 5.33.

    χ2=2525+2515+2525+2515=5.33\chi^2 = \frac{25}{25} + \frac{25}{15} + \frac{25}{25} + \frac{25}{15} = 5.33χ2=2525​+1525​+2525​+1525​=5.33
  3. 03Compare with the critical value

    For df = (2-1)(2-1) = 1 and p <= 0.05 the critical value is 3.84; the result is significant if chi-squared is greater than or equal to this. Since 5.33 >= 3.84, it is significant.

Result: Chi-squared = 5.33 > 3.84, so there is a significant association between teaching method and pass rate.

Worked example

A Spearman's rho calculation

Six participants' hours studied (2, 3, 5, 6, 8, 9) and test scores (30, 35, 40, 55, 50, 60) are recorded. Calculate Spearman's rho and decide significance at p <= 0.05 two-tailed (critical value 0.886).

  1. 01Rank each variable

    Ranking from lowest, hours give ranks 1,2,3,4,5,6 for 2,3,5,6,8,9; the matching scores 30,35,40,55,50,60 rank 1,2,3,5,4,6.

  2. 02Find d and d-squared

    The rank differences d are 0, 0, 0, -1, 1, 0, so d-squared are 0, 0, 0, 1, 1, 0, giving the sum of d-squared = 2.

  3. 03Apply the formula

    rho = 1 - (6 x 2) / (6 x (36 - 1)) = 1 - 12/210 = 0.94.

    rs=1−6×26(62−1)=1−12210≈0.94r_s = 1 - \frac{6 \times 2}{6(6^2 - 1)} = 1 - \frac{12}{210} \approx 0.94rs​=1−6(62−1)6×2​=1−21012​≈0.94
  4. 04Compare with the critical value

    For n = 6, two-tailed, p <= 0.05 the critical value is 0.886; the result is significant if rho is greater than or equal to this. Since 0.94 >= 0.886, it is significant.

Result: Spearman's rho = 0.94, a strong positive correlation that exceeds the critical value 0.886, so it is significant.

Exam focus

  • Compute the sign test in full: signs, discarding ties, N, S as the less frequent sign, and comparison with the critical value.
  • Compute chi-squared (with expected frequencies and df) and Spearman's rho, and state the significance decision in context.

Typical mistakes

  • Taking S as the larger, not the smaller, number of signs, or forgetting to discard participants who show no change.
  • Comparing the calculated value against the wrong critical value (wrong N/df, wrong number of tails, or the wrong direction of the inequality).

Active revision

In a related study, 15 participants are measured before and after an intervention: 12 improve, 3 get worse, and none stay the same. Carry out a sign test at p <= 0.05 (two-tailed; critical value of S for N = 15 is 3) and state the conclusion.

Active recall

Recall the key points — then reveal.

Sources: GCE AS and A level subject content for psychology (Department for Education) · AQA A-level Psychology 7182 specification (AQA)

Contents

Section -- / 07

    • 01Experimental design: variables, hypotheses and designs○
    • 02Non-experimental methods, sampling and pilot studies◐
    • 03Control, reliability, validity and ethics◐
    • 04Features of science and peer review●
    • 05Descriptive statistics and the normal distribution●
    • 06Inferential statistics: probability, significance and choosing a test●
    • 07Computing inferential tests: the sign test, chi-squared and Spearman's rho●

0/7 Read

From notes into training

Research methods

Reinforce this topic with matching tasks from the question bank.

~28
min
3
Competencies
Practise

References & sources

Sources

Department for Education

  • GCE AS and A level subject content for psychology

AQA

  • AQA A-level Psychology 7182 specification

Previous topic

Biopsychology

Next topic

Issues and debates in Psychology

EuraStudy·Notes T·07·MMXXVI

Carry on to the next topic — your learning path is kept.