3.1.2 AS research methodology

Syllabus
9990–2028–2029
Section
3.1.2
Level
AS

Research methods

Syllabus
9990–2028–2029
Topic
—
Level
AS

Experiments test causal effects by manipulating an IV and controlling comparison

An experiment deliberately manipulates an independent variable (IV), measures an operational dependent variable (DV), and compares conditions while controlling alternatives. Setting (laboratory/field) and participant design (independent, matched or repeated) are separate choices.

Type Defining feature Typical strength Typical cost
Laboratory experiment IV manipulated in an artificial/high-control setting Standardisation, replication and internal validity Lower mundane/ecological validity; demand characteristics; imposed-task ethics
Field experiment IV manipulated in a natural everyday setting Natural behaviour and lower awareness/demand Less control, harder replication/allocation; consent/debrief/withdrawal problems
Not an experiment No manipulated IV/control comparison Can describe/associate rich real behaviour Cannot claim the measured factor caused the outcome
Design Allocation Main advantage Main threat and remedy
Independent measures Different participants in each condition; random allocation if possible No order effects Participant differences; use random allocation, sufficient sample/control
Matched pairs Different people matched on relevant variable(s), one per condition Reduces selected participant differences without order effects Matching is slow/imperfect; loss of one may lose pair
Repeated measures Same participants complete every condition Controls participant differences; fewer participants Practice/fatigue/carry-over; counterbalance condition order

Novel scenario workflow: write an operational IV with at least two levels; write a measurable DV with units/scoring window; name design and allocation; hold instructions/materials/time/setting constant; add control condition/group; counterbalance if repeated; state consent, withdrawal, harm and debrief safeguards.

Evaluate in context: reliability asks whether the standardised procedure/measure could reproduce; internal validity asks whether IV rather than confounds changed DV; ecological validity asks whether task/setting represents the target behaviour; ethics asks what manipulation, deception or allocation does to participants.

Random sampling selects people from a population; random allocation assigns recruited participants to conditions. A control group receives no/alternative manipulation; a control condition can be completed by the same repeated-measures participants. Laboratory location alone does not make a study experimental.

Self-reports turn private experience into answers through designed questions

A self-report asks participants to report their own behaviour, cognition, emotion or experience. Questionnaires and interviews can both use open and closed questions; interview structure and delivery technique must be specified separately.

Choice Defining feature Gain Cost
Paper/online questionnaire Participant completes written items Large/cheap/anonymous; standardised Misunderstanding, low return, no probing
Structured interview Same questions/order, usually fixed prompts Comparable and reliable Limited depth; interviewer presence/bias
Unstructured interview Flexible conversation guided by participant Rich, unexpected detail Low standardisation, hard replication/analysis
Semi-structured interview Core standard questions plus planned/contingent probes Comparability plus depth Requires interviewer skill; probing can vary
Telephone/face-to-face Remote voice versus co-present delivery Reach/lower visual pressure versus rapport/non-verbal cues Identity/rapport versus interviewer/social desirability effects
Item type Example construction Data/evaluation
Closed 'In the last 7 days, on how many days did you sleep ≥8 hours? 0-7' Quantitative, fast/reliable comparison; options may force/omit answers
Likert/rating 'I felt anxious before the test: 1 strongly disagree … 5 strongly agree' Operational score; midpoint/response-set meaning must be clear
Open 'Describe one situation in which the test changed how you felt.' Qualitative depth/validity; coding is slower and less reliable

Write one idea per neutral item; define time/context; make response options exhaustive and mutually exclusive; avoid leading, loaded, double-barrelled, ambiguous, jargon and double-negative wording; pilot for interpretation; standardise interviewer prompts; plan coding before collection.

Quality question Contextual check Improvement
Reliability Would wording/order/interviewer/coding yield similar scores? Standard script, pilot, coding scheme, inter-rater/test-retest check
Validity Does answer reflect target construct rather than memory/demand? Anonymous/private response, multiple items, open probe/triangulation
Bias Social desirability, acquiescence, recall, interviewer expectations? Neutral wording, balanced items, confidentiality, trained interviewer
Ethics Sensitive disclosure, privacy, storage and right to skip? Informed consent, skip/withdraw option, safeguarding and secure anonymisation

Interview does not mean qualitative: closed structured questions produce quantitative data, while questionnaire open questions produce qualitative data. An anonymous response may reduce social desirability but cannot guarantee honesty, accurate memory or construct validity.

Case studies build intensive, triangulated evidence about one bounded unit

A case study investigates one bounded participant, family, group, institution or event in depth and context. It is a research strategy, not a single data-collection technique: interviews, questionnaires, observations, tests, records, biological measures and follow-up can be combined.

Evidence layer Possible technique Contribution
History/context Records, informant interviews, timeline Explains onset and environmental meaning
Current experience Interview/questionnaire/diary First-person cognition/emotion
Observable behaviour Naturalistic/structured observation or task Behavioural evidence beyond self-report
Objective/standard measure Test, diagnostic scale, physiological/brain measure Repeatable comparison or mechanism evidence
Change over time Treatment phases and follow-up Trajectory, maintenance and alternative explanations

Novel scenario: define the bounded case and why it is theoretically/clinically informative; collect several independent evidence sources; create a dated sequence; operationalise repeated measures; record contradictory evidence; obtain consent from the participant and relevant informants; anonymise identifying detail; avoid treating intervention response as controlled causal proof.

Strength Limitation
Rich holistic/contextual detail gives high ecological/construct insight Unique history and tiny unit prevent statistical population generalisation
Triangulation can test convergence across methods/data types Researcher interpretation and retrospective memory may bias the narrative
Rare/unethical-to-create cases generate hypotheses and practical learning No control group, random allocation or isolated IV means weak causal inference
Repeated follow-up captures change and unusual processes Time/cost, attrition and changing measures reduce reliability
Participant voice can preserve meaning Privacy, stigma, consent capacity and deductive identification are acute ethical risks

Case evidence can show that a phenomenon is possible, reveal a mechanism candidate and generate/refine theory. Transfer analytically by asking whether the new case shares relevant mechanisms/context—not by claiming one case represents everyone.

One participant in a laboratory experiment is not automatically a case study; intensive contextual study is required. A case can contain quantitative data and standardised tests. Depth improves understanding, not automatic validity, causality or generalisability.

Observations turn behaviour into evidence through explicit setting, role and coding choices

Observation systematically records behaviour. Classify it on four independent axes, then specify behavioural categories, sampling and observers. Observation is a technique; it becomes part of an experiment only when an IV is manipulated.

Axis Option A Option B Main trade-off
Awareness Overt: participants know Covert: do not know Consent/transparency versus reactivity/demand
Observer role Participant: joins group Non-participant: remains separate Insider access versus detachment/role bias
Coding Structured: predefined categories/checklist Unstructured: open narrative Reliability/comparison versus depth/unexpected behaviour
Setting/control Naturalistic: normal setting Controlled: arranged setting/task Ecological validity versus standardisation/control
Design step Required operational decision
Define behaviour Mutually exclusive, observable categories—not inferred motives (e.g. 'raises voice above conversational level for ≥2 s')
Sample behaviour Event sampling counts every target event; time sampling records at fixed intervals
Record Frequency, duration, latency, sequence and/or contextual field notes
Train observers Examples/non-examples, blind coding where possible, pilot ambiguous cases
Check reliability Two observers code same segment; compare agreement/correlation and revise categories
Protect participants Consent where feasible, public/private expectation, debrief deception, anonymise and stop if harm occurs
Strength Limitation and contextual remedy
Direct behaviour avoids memory/self-report bias Cannot directly access thought/emotion; triangulate with self-report/measure
Naturalistic/covert designs can capture spontaneous behaviour Low control and ethical problems; standardise window/context and debrief where possible
Structured categories yield quantitative comparison and replication Category reduction may miss meaning; pilot and retain field notes
Unstructured records discover unexpected patterns Observer bias/low reliability; reflexive notes, second coding and clear later scheme

Scenario answer sequence: state all four axes; operationalise at least two categories; choose event/time sampling; say exactly who observes, where and for how long; describe inter-observer reliability; identify reactivity/confounds; add consent/privacy/debrief safeguards.

Naturalistic does not imply covert, unstructured or participant observation. Controlled observation does not automatically manipulate an IV. High observer agreement means consistent coding, not that categories validly represent aggression, empathy or another inferred construct.

Correlations describe the direction and strength of association between measured co-variables

A correlation measures how two co-variables vary together; neither is manipulated. Each participant/unit contributes a paired score plotted as one point. Direction and strength describe association, not cause.

Pattern Meaning What it does not mean
Positive Higher X tends to accompany higher Y (and lower with lower) X is beneficial or causes Y
Negative Higher X tends to accompany lower Y Weak, harmful or no relationship
Strong Points cluster closely around a monotonic trend; coefficient nearer ±1 Valid measurement or causality
Weak/zero Points are dispersed/no monotonic trend; coefficient nearer 0 No non-linear relation or no subgroup pattern

Name both co-variables and give units/scoring window: for example, 'minutes of phone use from device log between 20:00-24:00 averaged over seven days' and 'total sleep minutes from actigraphy on the same nights'. 'Phone use and sleep' is not operational.

Step Decision
Sample paired scores Same units/people measured on both variables
Inspect scatterplot Direction, form, outliers and possible subgroups
Select coefficient Match scale/distribution/monotonic assumptions
Describe State direction and strength, preferably with coefficient
Evaluate Reliability/validity of both measures, range restriction, sampling and outliers
Infer Predict/identify association; do not assign causal direction
Causal threat Example for phone use ↔ sleep
Directionality Phone use may reduce sleep, or inability to sleep may increase use
Third variable Stress, workload or caffeine may increase use and reduce sleep
Selection/measurement A narrow student sample or inaccurate self-report can create/distort association

Strengths: studies naturally occurring variables that cannot ethically/practically be manipulated, quantifies prediction and generates hypotheses. Limits: no causal conclusion, vulnerable to third variables/directionality, and the correlation cannot be more valid or reliable than its two operational measures.

Do not call co-variables IV and DV. A coefficient sign gives direction, while absolute size gives strength. Even a perfect association cannot by itself show which variable causes which or rule out a common cause.

Longitudinal studies follow the same units to explain change and continuity over time

A longitudinal study repeatedly measures the same participant(s), case(s) or cohort across a meaningful period to investigate change, continuity or delayed effects. It can be observational/correlational or experimental if an IV/control comparison is added.

Form Structure Claim boundary
Descriptive longitudinal Same measures at several time points Describes within-unit trajectory, not its cause
Correlational longitudinal Earlier co-variable predicts later outcome Establishes time order but still has third-variable/confounding threats
Longitudinal experiment Manipulated condition/control followed over time Stronger causal and maintenance inference if allocation/control/attrition remain sound
Pre-post only Same unit measured before and after Change is visible, but history, maturation, testing and regression remain alternatives

Specify target interval and measurement schedule; keep operational measures equivalent; record baseline and relevant confounds; preserve participant IDs securely; standardise contact; plan retention and missing-data rules; add comparison/control where causal inference is intended; predefine follow-up outcome and stopping/safeguarding procedures.

Strength Why it matters
Within-person change Separates individual trajectory from one-time age-group differences
Temporal order Shows predictor preceded outcome, narrowing but not eliminating causal explanations
Delayed/maintenance effects Tests whether learning, treatment or brain/behaviour change persists
Rich repeated data Reveals turning points and individual differences hidden by group averages
Threat Consequence Mitigation
Attrition Smaller sample and systematic survivor bias Retention plan, compare dropouts, transparent missing-data analysis
Practice/testing Repetition itself changes scores Alternate forms, spacing, appropriate control
Historical/maturation change Time-related events/development mimic effect Comparison group, repeated baseline/context measures
Measure drift New tools/raters change apparent score Calibrate/equate methods and document changes
Cost/privacy Long commitment and sensitive linked records Proportionate schedule, renewed consent, secure pseudonymous linkage

Novel scenario answer: name the same participants, at least two dated waves, an unchanged operational outcome, expected change, retention method, attrition/practice/history control and ethical re-consent/confidentiality. If treatment is manipulated, also specify allocation and control condition.

Repeated measures compares conditions using the same participants; longitudinal follows the same units over a meaningful time course. A study can be both, but they are not synonyms. Time order improves causal reasoning yet does not remove third variables, history or maturation.

Methodological concepts

Syllabus
9990–2028–2029
Topic
—
Level
AS

Aims ask the question; hypotheses state the testable outcome pattern

An aim states what a study intends to investigate. An alternative hypothesis predicts a difference/relationship; a null predicts no difference/relationship beyond chance. Every hypothesis must identify the population and operational variables.

Statement Experiment form Correlation form
Aim Investigate whether condition X affects measured Y in population P Investigate the relationship between measured X and Y in P
Directional alternative P in X1 will score higher/lower on Y than P in X2 As X increases, Y will increase/decrease in P
Non-directional alternative There will be a difference in Y between X1 and X2 for P There will be a relationship between X and Y in P
Null There will be no difference in Y between X1 and X2 for P There will be no relationship between X and Y in P

Use a directional hypothesis only when prior theory/evidence justifies the direction before data collection. Use non-directional when an effect is expected but its direction is uncertain. Write scores/behaviour, not vague 'better' or causal wording for a correlation.

Audit: population named; IV levels or both co-variables operationalised; DV measure named; comparison/relationship word present; direction only when justified; null is exact logical counterpart.

One-tailed means one predicted direction, not one condition. The null is not 'nothing happened'; it is a population statement of no difference/association. Results support or fail to support a hypothesis—they do not rewrite it after seeing data.

Operational variables make manipulation and measurement exact and replicable

The IV is deliberately changed in an experiment; the DV is the measured outcome. An operational definition states exactly how each variable is created or scored so another researcher could reproduce it.

Concept Weak label Operational version
IV: background sound Music vs silence 70 dB instrumental track through headphones versus identical headphones with no audio during a 10-minute task
DV: memory Memory score Number of 20 nouns correctly freely recalled in 2 minutes, duplicates/intrusions excluded
Correlational co-variable Stress Total score on named 10-item scale completed after school

A DV can be frequency, duration, latency, accuracy/error, test/scale score, choice, physiological value or coded category. State unit, observation window, scoring rules and direction. Pilot floor/ceiling effects and ambiguity.

Operational gain Possible cost
Precision and replicability Narrow measure may omit the construct
Quantitative comparison Score may reward speed/strategy rather than target ability
Standardisation Artificial task may reduce ecological validity

Predictors/co-variables are not IVs unless manipulated. 'Aggression', 'happiness' and 'learning' are constructs, not complete DVs. Operational clarity improves replicability but does not automatically establish validity.

Controls reduce alternative explanations by standardising or balancing relevant variation

A control holds, removes, measures or balances a variable so conditions differ mainly in the IV. Standardisation gives every participant the same procedure. An uncontrolled variable becomes a confound when it systematically co-varies with the IV and could change the DV.

Source Example Feasible control
Participant Baseline memory, age, sleep Repeated measures, matching, random allocation, baseline measurement
Situational Room noise, time, device, experimenter tone Same setting/time/material/script; randomise/balance sessions
Order Practice, fatigue, carry-over Counterbalance, rest/washout, alternate forms
Demand/experimenter Guessing aim, cueing responses Blind/double-blind, cover story where ethical, standard script

For each threat: name it; explain how it differs between IV levels; explain how it could alter the operational DV; specify an exact control. A generic 'keep everything the same' earns less than a mechanism-linked control.

Controls strengthen internal validity/reliability but may make tasks artificial, restrict natural variation, increase cost or create ethical issues. Measure rather than eliminate important real-world factors when ecological validity is central.

Not every uncontrolled variable is a confound; it must vary systematically with the IV and explain the DV. Standardisation cannot remove participant differences by itself. Random allocation balances conditions; random sampling addresses population recruitment.

Data form and data source are separate: quantitative/qualitative and objective/subjective

Quantitative versus qualitative describes data form; objective versus subjective describes dependence on personal judgement/experience. The two axes can combine in four ways.

Objective/externally verifiable Subjective/judgement-based
Quantitative Reaction time, heart rate, correct-answer count Participant's 0-8 distress rating; observer category score requiring judgement
Qualitative Verbatim audio transcript/recorded words as raw event Participant's interpretation or researcher's thematic account
Type Main value Main limit
Quantitative Compact comparison, graphs, replicable scoring Reduction may omit meaning/context
Qualitative Rich explanations and unexpected themes Slow coding, interpretation and lower inter-rater reliability
Objective Less affected by self-presentation/interpretation May measure proxy rather than lived construct
Subjective Direct access to private experience Social desirability, memory and perspective bias

Choose data to fit the question, then triangulate: e.g. combine sleep minutes from actigraphy (quantitative/objective), rating of sleep quality (quantitative/subjective) and diary account (qualitative/subjective). Agreement strengthens confidence; disagreement is evidence to explain, not delete.

Numbers are not automatically objective: a Likert score quantifies a subjective judgement. Words are not automatically subjective: an exact recorded utterance is an observable event, though coding its meaning may be subjective. No data type is universally best.

Sampling links a target population to the people who actually provide data

The population is the full group to which a researcher wants to generalise; the sample is the participating subset. Sampling technique controls who gets an invitation, but non-response and eligibility still shape the final sample.

Technique Exact procedure Strength Bias/limit
Opportunity Recruit available eligible people at chosen place/time Fast, cheap, practical Place/time/researcher-access bias; often unrepresentative
Volunteer/self-selecting Advertise and eligible people opt in Consent/interest, reaches dispersed people Volunteer traits: motivation, time, topic interest, payment response
Random Build complete sampling frame; use random generator/lottery so each has equal selection chance Reduces researcher selection bias; potentially representative Frame may omit people; selected non-response; costly

Novel scenario: define target population with inclusion/exclusion; identify/construct sampling frame if random; state recruitment location/channel/time; describe exact selection—not just name; anticipate who is missed/refuses; compare achieved sample demographics with population; bound generalisation accordingly.

Evidence Generalisation consequence
Narrow age/sex/culture/occupation Findings may not transfer where relevant mechanism differs
Large but biased online volunteer sample Precision can increase while representativeness remains poor
Small random sample Less selection bias but greater sampling fluctuation
Attrition/non-response Final sample may differ from those initially selected

Random sampling recruits from a population; random allocation assigns a recruited sample to conditions. Large does not equal representative. Opportunity sampling is based on availability, not deliberate quota matching. Generalisability also depends on task/setting/time, not sample alone.

Ethical research converts risk-benefit principles into human and animal safeguards

Human guideline Requirement and safeguard
Minimise harm/maximise benefit Risk assess, monitor distress, stop/refer/support; use least harmful effective procedure
Valid informed consent Capacity, understandable purpose/procedure/risks/data use; guardian consent plus assent where relevant
Right to withdraw Leave and remove data without penalty; make route/reminder practical
Lack of deception Disclose truth unless justified/minimal and impossible otherwise; never deceive about material risk
Confidentiality Limit access, pseudonymise, secure storage/reporting
Privacy Observe/collect only where people reasonably expect and consent; minimise intrusion
Debriefing Reveal aim/deception, restore understanding, answer questions, offer data withdrawal/support
Animal guideline Applied question
Minimise harm/maximise benefit Is scientific/welfare value proportionate to pain, distress and lasting effects?
Replacement Can non-animal, simulation, existing data or less sentient model answer it?
Species Is species scientifically appropriate and welfare expertise available?
Numbers Use minimum needed for valid evidence—too few also wastes animals
Procedures Refine handling/anaesthesia/endpoints; suitable housing/social needs; justified reward/deprivation; avoid aversive stimuli

Exam chain: name guideline → cite exact procedure/sample/data feature → explain likely harm/autonomy/welfare or benefit → judge severity/probability/reversibility → propose feasible safeguard and its methodological trade-off. For animals, use the syllabus animal guideline rather than importing human consent/right-to-withdraw labels.

Method need Ethical tension Design response
Avoid demand characteristics Deception/incomplete disclosure Minimal deception, no risk deception, prior/retrospective consent, prompt debrief
Natural public behaviour Consent/privacy Public-expectation audit, anonymised low-risk recording, gatekeeper/debrief where possible
Stress/emergency simulation Harm and withdrawal Lower intensity, screening, stop rule, support and alternative task
Animal reinforcement Deprivation/aversive welfare Preferred reward, minimal restriction, voluntary participation and welfare endpoints

Signed consent is not automatically valid if information, capacity or freedom is missing. Debriefing mitigates deception but cannot undo severe harm. Confidentiality concerns data identity; privacy concerns access to the person/behaviour. Ethical acceptability is a reasoned balance, not a checklist score.

Validity asks whether the evidence supports the intended measure, cause and transfer

Validity is the extent to which a measure/study supports the intended interpretation. Ask separately whether the construct was measured, the IV caused the DV, the behaviour represents real life and the finding transfers to the target population/context.

Validity question Main threats Contextual improvement
Construct/measurement Proxy score, subjective coding, social desirability Operational/pilot/validated measure, blind coding, triangulation
Internal/causal Confounds, allocation/order/experimenter effects Control/standardise, random allocation, counterbalance, blind
Ecological Artificial task/setting, low mundane realism More representative task/context, field evidence—while retaining controls
Population/generalisation Narrow/biased sample and cultural/time differences Broader/stratified/random recruitment, replication across groups

Demand characteristics arise when participants infer the aim and alter behaviour; reduce through credible neutral instructions, unobtrusive/indirect measures, blinding or justified deception/debrief. Subjectivity can add insight but requires transparent coding/checks. Objectivity reduces judgement, not necessarily proxy invalidity.

Write: identify exact evidence feature → name the interpretation threatened/supported → explain mechanism → judge consequence for this conclusion → propose improvement and trade-off. 'Laboratory means low validity' is incomplete without showing why this task differs from the target behaviour.

A field setting can contain an invalid measure; a laboratory task can validly test a narrow mechanism. Reliability is necessary for many valid measurements but consistent bias remains invalid. Generalisability extends beyond sample to setting, task, culture and time.

Reliability checks consistency; replicability makes an independent repetition possible

Reliability is consistency of measurement/procedure. Replicability is whether documentation/materials are sufficient for another researcher to repeat the study and test whether the pattern recurs.

Check Same/different element Use Improve when low
Inter-rater Two raters score same response/product Interviews, tests, qualitative coding Rubric, examples, training, blind double-code
Inter-observer Two observers code same behaviour/time Structured observation Operational categories, observer training, video recode
Test-retest Same measure to same people at two suitable times Stable traits/questionnaires/tests Clarify items, standardise conditions; avoid interval/practice extremes
Procedural replication Independent study repeats documented method Tests result robustness Full script/materials/operational rules, sample and analysis transparency

Plan: standardise instructions, timing, apparatus and scoring; define categories/units; pilot ambiguity; preserve versioned materials; report recruitment/allocation/exclusions; have independent coders; choose a reliability check that targets the likely source of inconsistency.

High reliability narrows random/observer inconsistency and makes differences interpretable. Low reliability can hide or create effects. Yet a consistently wrong scale, biased item or invalid category can be highly reliable.

Two researchers repeating a whole study is replication, not inter-rater reliability. Test-retest uses the same measure/construct at two times, not two experimental conditions. Agreement does not prove validity, objectivity or truth.

Descriptive analysis matches centre, spread and display to the structure of data

Descriptive analysis organises and summarises observed data. Central tendency gives a typical/central value; spread shows variability. Cambridge requires recognition/finding/interpretation, not calculations or unspecified statistical tests.

Measure How found Best feature Main weakness
Mode Most frequent value/category Works with nominal categories; shows common response May be multiple/none; ignores rest
Median Middle ordered score (average two middles if even) Resistant to extreme/skew; ordinal suitable Ignores distances and much data
Mean Sum divided by number Uses every interval/ratio score; sensitive comparison Distorted by extremes/skew; unsuitable for categories
Measure Meaning Interpretation
Range Highest minus lowest (state convention if endpoints included) Overall span; one extreme can dominate
Standard deviation Typical dispersion around mean Low SD = scores clustered/consistent; high SD = dispersed—not a high/low mean
Display Data structure Construction/reading rule
Table Any organised categories/conditions Clear title, labelled rows/columns, units, no ambiguous totals
Bar chart Separate categories/conditions Equal-width separated bars; axes/units; height = frequency/summary
Histogram Continuous scores grouped into adjacent intervals Bars touch; numerical ordered x-axis; frequency on y
Scatter graph Paired co-variable scores One dot per pair; both axes operational variables; inspect direction/strength/outliers

Compare centre and spread together: equal means can hide different consistency; lower mean may accompany wider overlap. Describe exact pattern with units, largest/smallest, difference and variability. Do not infer cause or significance from descriptive displays.

Bar charts display discrete categories with gaps; histograms display continuous intervals with touching bars; scatter graphs do not display group frequencies. Low SD means low variation, not low scores. A mean difference alone does not prove a reliable population effect.