Methodological concepts
- Syllabus
- 9990–2028–2029
- Topic
- —
- Level
- AS
An aim states what a study intends to investigate. An alternative hypothesis predicts a difference/relationship; a null predicts no difference/relationship beyond chance. Every hypothesis must identify the population and operational variables.
| Statement | Experiment form | Correlation form |
|---|---|---|
| Aim | Investigate whether condition X affects measured Y in population P | Investigate the relationship between measured X and Y in P |
| Directional alternative | P in X1 will score higher/lower on Y than P in X2 | As X increases, Y will increase/decrease in P |
| Non-directional alternative | There will be a difference in Y between X1 and X2 for P | There will be a relationship between X and Y in P |
| Null | There will be no difference in Y between X1 and X2 for P | There will be no relationship between X and Y in P |
Use a directional hypothesis only when prior theory/evidence justifies the direction before data collection. Use non-directional when an effect is expected but its direction is uncertain. Write scores/behaviour, not vague 'better' or causal wording for a correlation.
Audit: population named; IV levels or both co-variables operationalised; DV measure named; comparison/relationship word present; direction only when justified; null is exact logical counterpart.
One-tailed means one predicted direction, not one condition. The null is not 'nothing happened'; it is a population statement of no difference/association. Results support or fail to support a hypothesis—they do not rewrite it after seeing data.
The IV is deliberately changed in an experiment; the DV is the measured outcome. An operational definition states exactly how each variable is created or scored so another researcher could reproduce it.
| Concept | Weak label | Operational version |
|---|---|---|
| IV: background sound | Music vs silence | 70 dB instrumental track through headphones versus identical headphones with no audio during a 10-minute task |
| DV: memory | Memory score | Number of 20 nouns correctly freely recalled in 2 minutes, duplicates/intrusions excluded |
| Correlational co-variable | Stress | Total score on named 10-item scale completed after school |
A DV can be frequency, duration, latency, accuracy/error, test/scale score, choice, physiological value or coded category. State unit, observation window, scoring rules and direction. Pilot floor/ceiling effects and ambiguity.
| Operational gain | Possible cost |
|---|---|
| Precision and replicability | Narrow measure may omit the construct |
| Quantitative comparison | Score may reward speed/strategy rather than target ability |
| Standardisation | Artificial task may reduce ecological validity |
Predictors/co-variables are not IVs unless manipulated. 'Aggression', 'happiness' and 'learning' are constructs, not complete DVs. Operational clarity improves replicability but does not automatically establish validity.
A control holds, removes, measures or balances a variable so conditions differ mainly in the IV. Standardisation gives every participant the same procedure. An uncontrolled variable becomes a confound when it systematically co-varies with the IV and could change the DV.
| Source | Example | Feasible control |
|---|---|---|
| Participant | Baseline memory, age, sleep | Repeated measures, matching, random allocation, baseline measurement |
| Situational | Room noise, time, device, experimenter tone | Same setting/time/material/script; randomise/balance sessions |
| Order | Practice, fatigue, carry-over | Counterbalance, rest/washout, alternate forms |
| Demand/experimenter | Guessing aim, cueing responses | Blind/double-blind, cover story where ethical, standard script |
For each threat: name it; explain how it differs between IV levels; explain how it could alter the operational DV; specify an exact control. A generic 'keep everything the same' earns less than a mechanism-linked control.
Controls strengthen internal validity/reliability but may make tasks artificial, restrict natural variation, increase cost or create ethical issues. Measure rather than eliminate important real-world factors when ecological validity is central.
Not every uncontrolled variable is a confound; it must vary systematically with the IV and explain the DV. Standardisation cannot remove participant differences by itself. Random allocation balances conditions; random sampling addresses population recruitment.
Quantitative versus qualitative describes data form; objective versus subjective describes dependence on personal judgement/experience. The two axes can combine in four ways.
| Objective/externally verifiable | Subjective/judgement-based | |
|---|---|---|
| Quantitative | Reaction time, heart rate, correct-answer count | Participant's 0-8 distress rating; observer category score requiring judgement |
| Qualitative | Verbatim audio transcript/recorded words as raw event | Participant's interpretation or researcher's thematic account |
| Type | Main value | Main limit |
|---|---|---|
| Quantitative | Compact comparison, graphs, replicable scoring | Reduction may omit meaning/context |
| Qualitative | Rich explanations and unexpected themes | Slow coding, interpretation and lower inter-rater reliability |
| Objective | Less affected by self-presentation/interpretation | May measure proxy rather than lived construct |
| Subjective | Direct access to private experience | Social desirability, memory and perspective bias |
Choose data to fit the question, then triangulate: e.g. combine sleep minutes from actigraphy (quantitative/objective), rating of sleep quality (quantitative/subjective) and diary account (qualitative/subjective). Agreement strengthens confidence; disagreement is evidence to explain, not delete.
Numbers are not automatically objective: a Likert score quantifies a subjective judgement. Words are not automatically subjective: an exact recorded utterance is an observable event, though coding its meaning may be subjective. No data type is universally best.
The population is the full group to which a researcher wants to generalise; the sample is the participating subset. Sampling technique controls who gets an invitation, but non-response and eligibility still shape the final sample.
| Technique | Exact procedure | Strength | Bias/limit |
|---|---|---|---|
| Opportunity | Recruit available eligible people at chosen place/time | Fast, cheap, practical | Place/time/researcher-access bias; often unrepresentative |
| Volunteer/self-selecting | Advertise and eligible people opt in | Consent/interest, reaches dispersed people | Volunteer traits: motivation, time, topic interest, payment response |
| Random | Build complete sampling frame; use random generator/lottery so each has equal selection chance | Reduces researcher selection bias; potentially representative | Frame may omit people; selected non-response; costly |
Novel scenario: define target population with inclusion/exclusion; identify/construct sampling frame if random; state recruitment location/channel/time; describe exact selection—not just name; anticipate who is missed/refuses; compare achieved sample demographics with population; bound generalisation accordingly.
| Evidence | Generalisation consequence |
|---|---|
| Narrow age/sex/culture/occupation | Findings may not transfer where relevant mechanism differs |
| Large but biased online volunteer sample | Precision can increase while representativeness remains poor |
| Small random sample | Less selection bias but greater sampling fluctuation |
| Attrition/non-response | Final sample may differ from those initially selected |
Random sampling recruits from a population; random allocation assigns a recruited sample to conditions. Large does not equal representative. Opportunity sampling is based on availability, not deliberate quota matching. Generalisability also depends on task/setting/time, not sample alone.
| Human guideline | Requirement and safeguard |
|---|---|
| Minimise harm/maximise benefit | Risk assess, monitor distress, stop/refer/support; use least harmful effective procedure |
| Valid informed consent | Capacity, understandable purpose/procedure/risks/data use; guardian consent plus assent where relevant |
| Right to withdraw | Leave and remove data without penalty; make route/reminder practical |
| Lack of deception | Disclose truth unless justified/minimal and impossible otherwise; never deceive about material risk |
| Confidentiality | Limit access, pseudonymise, secure storage/reporting |
| Privacy | Observe/collect only where people reasonably expect and consent; minimise intrusion |
| Debriefing | Reveal aim/deception, restore understanding, answer questions, offer data withdrawal/support |
| Animal guideline | Applied question |
|---|---|
| Minimise harm/maximise benefit | Is scientific/welfare value proportionate to pain, distress and lasting effects? |
| Replacement | Can non-animal, simulation, existing data or less sentient model answer it? |
| Species | Is species scientifically appropriate and welfare expertise available? |
| Numbers | Use minimum needed for valid evidence—too few also wastes animals |
| Procedures | Refine handling/anaesthesia/endpoints; suitable housing/social needs; justified reward/deprivation; avoid aversive stimuli |
Exam chain: name guideline → cite exact procedure/sample/data feature → explain likely harm/autonomy/welfare or benefit → judge severity/probability/reversibility → propose feasible safeguard and its methodological trade-off. For animals, use the syllabus animal guideline rather than importing human consent/right-to-withdraw labels.
| Method need | Ethical tension | Design response |
|---|---|---|
| Avoid demand characteristics | Deception/incomplete disclosure | Minimal deception, no risk deception, prior/retrospective consent, prompt debrief |
| Natural public behaviour | Consent/privacy | Public-expectation audit, anonymised low-risk recording, gatekeeper/debrief where possible |
| Stress/emergency simulation | Harm and withdrawal | Lower intensity, screening, stop rule, support and alternative task |
| Animal reinforcement | Deprivation/aversive welfare | Preferred reward, minimal restriction, voluntary participation and welfare endpoints |
Signed consent is not automatically valid if information, capacity or freedom is missing. Debriefing mitigates deception but cannot undo severe harm. Confidentiality concerns data identity; privacy concerns access to the person/behaviour. Ethical acceptability is a reasoned balance, not a checklist score.
Validity is the extent to which a measure/study supports the intended interpretation. Ask separately whether the construct was measured, the IV caused the DV, the behaviour represents real life and the finding transfers to the target population/context.
| Validity question | Main threats | Contextual improvement |
|---|---|---|
| Construct/measurement | Proxy score, subjective coding, social desirability | Operational/pilot/validated measure, blind coding, triangulation |
| Internal/causal | Confounds, allocation/order/experimenter effects | Control/standardise, random allocation, counterbalance, blind |
| Ecological | Artificial task/setting, low mundane realism | More representative task/context, field evidence—while retaining controls |
| Population/generalisation | Narrow/biased sample and cultural/time differences | Broader/stratified/random recruitment, replication across groups |
Demand characteristics arise when participants infer the aim and alter behaviour; reduce through credible neutral instructions, unobtrusive/indirect measures, blinding or justified deception/debrief. Subjectivity can add insight but requires transparent coding/checks. Objectivity reduces judgement, not necessarily proxy invalidity.
Write: identify exact evidence feature → name the interpretation threatened/supported → explain mechanism → judge consequence for this conclusion → propose improvement and trade-off. 'Laboratory means low validity' is incomplete without showing why this task differs from the target behaviour.
A field setting can contain an invalid measure; a laboratory task can validly test a narrow mechanism. Reliability is necessary for many valid measurements but consistent bias remains invalid. Generalisability extends beyond sample to setting, task, culture and time.
Reliability is consistency of measurement/procedure. Replicability is whether documentation/materials are sufficient for another researcher to repeat the study and test whether the pattern recurs.
| Check | Same/different element | Use | Improve when low |
|---|---|---|---|
| Inter-rater | Two raters score same response/product | Interviews, tests, qualitative coding | Rubric, examples, training, blind double-code |
| Inter-observer | Two observers code same behaviour/time | Structured observation | Operational categories, observer training, video recode |
| Test-retest | Same measure to same people at two suitable times | Stable traits/questionnaires/tests | Clarify items, standardise conditions; avoid interval/practice extremes |
| Procedural replication | Independent study repeats documented method | Tests result robustness | Full script/materials/operational rules, sample and analysis transparency |
Plan: standardise instructions, timing, apparatus and scoring; define categories/units; pilot ambiguity; preserve versioned materials; report recruitment/allocation/exclusions; have independent coders; choose a reliability check that targets the likely source of inconsistency.
High reliability narrows random/observer inconsistency and makes differences interpretable. Low reliability can hide or create effects. Yet a consistently wrong scale, biased item or invalid category can be highly reliable.
Two researchers repeating a whole study is replication, not inter-rater reliability. Test-retest uses the same measure/construct at two times, not two experimental conditions. Agreement does not prove validity, objectivity or truth.
Descriptive analysis organises and summarises observed data. Central tendency gives a typical/central value; spread shows variability. Cambridge requires recognition/finding/interpretation, not calculations or unspecified statistical tests.
| Measure | How found | Best feature | Main weakness |
|---|---|---|---|
| Mode | Most frequent value/category | Works with nominal categories; shows common response | May be multiple/none; ignores rest |
| Median | Middle ordered score (average two middles if even) | Resistant to extreme/skew; ordinal suitable | Ignores distances and much data |
| Mean | Sum divided by number | Uses every interval/ratio score; sensitive comparison | Distorted by extremes/skew; unsuitable for categories |
| Measure | Meaning | Interpretation |
|---|---|---|
| Range | Highest minus lowest (state convention if endpoints included) | Overall span; one extreme can dominate |
| Standard deviation | Typical dispersion around mean | Low SD = scores clustered/consistent; high SD = dispersed—not a high/low mean |
| Display | Data structure | Construction/reading rule |
|---|---|---|
| Table | Any organised categories/conditions | Clear title, labelled rows/columns, units, no ambiguous totals |
| Bar chart | Separate categories/conditions | Equal-width separated bars; axes/units; height = frequency/summary |
| Histogram | Continuous scores grouped into adjacent intervals | Bars touch; numerical ordered x-axis; frequency on y |
| Scatter graph | Paired co-variable scores | One dot per pair; both axes operational variables; inspect direction/strength/outliers |
Compare centre and spread together: equal means can hide different consistency; lower mean may accompany wider overlap. Describe exact pattern with units, largest/smallest, difference and variability. Do not infer cause or significance from descriptive displays.
Bar charts display discrete categories with gaps; histograms display continuous intervals with touching bars; scatter graphs do not display group frequencies. Low SD means low variation, not low scores. A mean difference alone does not prove a reliable population effect.