3.1.2 AS research methodology
- Syllabus
- 9990–2028–2029
- Section
- 3.1.2
- Level
- AS

An experiment deliberately manipulates an independent variable (IV), measures an operational dependent variable (DV), and compares conditions while controlling alternatives. Setting (laboratory/field) and participant design (independent, matched or repeated) are separate choices.
| Type | Defining feature | Typical strength | Typical cost |
|---|---|---|---|
| Laboratory experiment | IV manipulated in an artificial/high-control setting | Standardisation, replication and internal validity | Lower mundane/ecological validity; demand characteristics; imposed-task ethics |
| Field experiment | IV manipulated in a natural everyday setting | Natural behaviour and lower awareness/demand | Less control, harder replication/allocation; consent/debrief/withdrawal problems |
| Not an experiment | No manipulated IV/control comparison | Can describe/associate rich real behaviour | Cannot claim the measured factor caused the outcome |
| Design | Allocation | Main advantage | Main threat and remedy |
|---|---|---|---|
| Independent measures | Different participants in each condition; random allocation if possible | No order effects | Participant differences; use random allocation, sufficient sample/control |
| Matched pairs | Different people matched on relevant variable(s), one per condition | Reduces selected participant differences without order effects | Matching is slow/imperfect; loss of one may lose pair |
| Repeated measures | Same participants complete every condition | Controls participant differences; fewer participants | Practice/fatigue/carry-over; counterbalance condition order |
Novel scenario workflow: write an operational IV with at least two levels; write a measurable DV with units/scoring window; name design and allocation; hold instructions/materials/time/setting constant; add control condition/group; counterbalance if repeated; state consent, withdrawal, harm and debrief safeguards.
Evaluate in context: reliability asks whether the standardised procedure/measure could reproduce; internal validity asks whether IV rather than confounds changed DV; ecological validity asks whether task/setting represents the target behaviour; ethics asks what manipulation, deception or allocation does to participants.
Random sampling selects people from a population; random allocation assigns recruited participants to conditions. A control group receives no/alternative manipulation; a control condition can be completed by the same repeated-measures participants. Laboratory location alone does not make a study experimental.
A self-report asks participants to report their own behaviour, cognition, emotion or experience. Questionnaires and interviews can both use open and closed questions; interview structure and delivery technique must be specified separately.
| Choice | Defining feature | Gain | Cost |
|---|---|---|---|
| Paper/online questionnaire | Participant completes written items | Large/cheap/anonymous; standardised | Misunderstanding, low return, no probing |
| Structured interview | Same questions/order, usually fixed prompts | Comparable and reliable | Limited depth; interviewer presence/bias |
| Unstructured interview | Flexible conversation guided by participant | Rich, unexpected detail | Low standardisation, hard replication/analysis |
| Semi-structured interview | Core standard questions plus planned/contingent probes | Comparability plus depth | Requires interviewer skill; probing can vary |
| Telephone/face-to-face | Remote voice versus co-present delivery | Reach/lower visual pressure versus rapport/non-verbal cues | Identity/rapport versus interviewer/social desirability effects |
| Item type | Example construction | Data/evaluation |
|---|---|---|
| Closed | 'In the last 7 days, on how many days did you sleep ≥8 hours? 0-7' | Quantitative, fast/reliable comparison; options may force/omit answers |
| Likert/rating | 'I felt anxious before the test: 1 strongly disagree … 5 strongly agree' | Operational score; midpoint/response-set meaning must be clear |
| Open | 'Describe one situation in which the test changed how you felt.' | Qualitative depth/validity; coding is slower and less reliable |
Write one idea per neutral item; define time/context; make response options exhaustive and mutually exclusive; avoid leading, loaded, double-barrelled, ambiguous, jargon and double-negative wording; pilot for interpretation; standardise interviewer prompts; plan coding before collection.
| Quality question | Contextual check | Improvement |
|---|---|---|
| Reliability | Would wording/order/interviewer/coding yield similar scores? | Standard script, pilot, coding scheme, inter-rater/test-retest check |
| Validity | Does answer reflect target construct rather than memory/demand? | Anonymous/private response, multiple items, open probe/triangulation |
| Bias | Social desirability, acquiescence, recall, interviewer expectations? | Neutral wording, balanced items, confidentiality, trained interviewer |
| Ethics | Sensitive disclosure, privacy, storage and right to skip? | Informed consent, skip/withdraw option, safeguarding and secure anonymisation |
Interview does not mean qualitative: closed structured questions produce quantitative data, while questionnaire open questions produce qualitative data. An anonymous response may reduce social desirability but cannot guarantee honesty, accurate memory or construct validity.
A case study investigates one bounded participant, family, group, institution or event in depth and context. It is a research strategy, not a single data-collection technique: interviews, questionnaires, observations, tests, records, biological measures and follow-up can be combined.
| Evidence layer | Possible technique | Contribution |
|---|---|---|
| History/context | Records, informant interviews, timeline | Explains onset and environmental meaning |
| Current experience | Interview/questionnaire/diary | First-person cognition/emotion |
| Observable behaviour | Naturalistic/structured observation or task | Behavioural evidence beyond self-report |
| Objective/standard measure | Test, diagnostic scale, physiological/brain measure | Repeatable comparison or mechanism evidence |
| Change over time | Treatment phases and follow-up | Trajectory, maintenance and alternative explanations |
Novel scenario: define the bounded case and why it is theoretically/clinically informative; collect several independent evidence sources; create a dated sequence; operationalise repeated measures; record contradictory evidence; obtain consent from the participant and relevant informants; anonymise identifying detail; avoid treating intervention response as controlled causal proof.
| Strength | Limitation |
|---|---|
| Rich holistic/contextual detail gives high ecological/construct insight | Unique history and tiny unit prevent statistical population generalisation |
| Triangulation can test convergence across methods/data types | Researcher interpretation and retrospective memory may bias the narrative |
| Rare/unethical-to-create cases generate hypotheses and practical learning | No control group, random allocation or isolated IV means weak causal inference |
| Repeated follow-up captures change and unusual processes | Time/cost, attrition and changing measures reduce reliability |
| Participant voice can preserve meaning | Privacy, stigma, consent capacity and deductive identification are acute ethical risks |
Case evidence can show that a phenomenon is possible, reveal a mechanism candidate and generate/refine theory. Transfer analytically by asking whether the new case shares relevant mechanisms/context—not by claiming one case represents everyone.
One participant in a laboratory experiment is not automatically a case study; intensive contextual study is required. A case can contain quantitative data and standardised tests. Depth improves understanding, not automatic validity, causality or generalisability.
Observation systematically records behaviour. Classify it on four independent axes, then specify behavioural categories, sampling and observers. Observation is a technique; it becomes part of an experiment only when an IV is manipulated.
| Axis | Option A | Option B | Main trade-off |
|---|---|---|---|
| Awareness | Overt: participants know | Covert: do not know | Consent/transparency versus reactivity/demand |
| Observer role | Participant: joins group | Non-participant: remains separate | Insider access versus detachment/role bias |
| Coding | Structured: predefined categories/checklist | Unstructured: open narrative | Reliability/comparison versus depth/unexpected behaviour |
| Setting/control | Naturalistic: normal setting | Controlled: arranged setting/task | Ecological validity versus standardisation/control |
| Design step | Required operational decision |
|---|---|
| Define behaviour | Mutually exclusive, observable categories—not inferred motives (e.g. 'raises voice above conversational level for ≥2 s') |
| Sample behaviour | Event sampling counts every target event; time sampling records at fixed intervals |
| Record | Frequency, duration, latency, sequence and/or contextual field notes |
| Train observers | Examples/non-examples, blind coding where possible, pilot ambiguous cases |
| Check reliability | Two observers code same segment; compare agreement/correlation and revise categories |
| Protect participants | Consent where feasible, public/private expectation, debrief deception, anonymise and stop if harm occurs |
| Strength | Limitation and contextual remedy |
|---|---|
| Direct behaviour avoids memory/self-report bias | Cannot directly access thought/emotion; triangulate with self-report/measure |
| Naturalistic/covert designs can capture spontaneous behaviour | Low control and ethical problems; standardise window/context and debrief where possible |
| Structured categories yield quantitative comparison and replication | Category reduction may miss meaning; pilot and retain field notes |
| Unstructured records discover unexpected patterns | Observer bias/low reliability; reflexive notes, second coding and clear later scheme |
Scenario answer sequence: state all four axes; operationalise at least two categories; choose event/time sampling; say exactly who observes, where and for how long; describe inter-observer reliability; identify reactivity/confounds; add consent/privacy/debrief safeguards.
Naturalistic does not imply covert, unstructured or participant observation. Controlled observation does not automatically manipulate an IV. High observer agreement means consistent coding, not that categories validly represent aggression, empathy or another inferred construct.
A correlation measures how two co-variables vary together; neither is manipulated. Each participant/unit contributes a paired score plotted as one point. Direction and strength describe association, not cause.
| Pattern | Meaning | What it does not mean |
|---|---|---|
| Positive | Higher X tends to accompany higher Y (and lower with lower) | X is beneficial or causes Y |
| Negative | Higher X tends to accompany lower Y | Weak, harmful or no relationship |
| Strong | Points cluster closely around a monotonic trend; coefficient nearer ±1 | Valid measurement or causality |
| Weak/zero | Points are dispersed/no monotonic trend; coefficient nearer 0 | No non-linear relation or no subgroup pattern |
Name both co-variables and give units/scoring window: for example, 'minutes of phone use from device log between 20:00-24:00 averaged over seven days' and 'total sleep minutes from actigraphy on the same nights'. 'Phone use and sleep' is not operational.
| Step | Decision |
|---|---|
| Sample paired scores | Same units/people measured on both variables |
| Inspect scatterplot | Direction, form, outliers and possible subgroups |
| Select coefficient | Match scale/distribution/monotonic assumptions |
| Describe | State direction and strength, preferably with coefficient |
| Evaluate | Reliability/validity of both measures, range restriction, sampling and outliers |
| Infer | Predict/identify association; do not assign causal direction |
| Causal threat | Example for phone use ↔ sleep |
|---|---|
| Directionality | Phone use may reduce sleep, or inability to sleep may increase use |
| Third variable | Stress, workload or caffeine may increase use and reduce sleep |
| Selection/measurement | A narrow student sample or inaccurate self-report can create/distort association |
Strengths: studies naturally occurring variables that cannot ethically/practically be manipulated, quantifies prediction and generates hypotheses. Limits: no causal conclusion, vulnerable to third variables/directionality, and the correlation cannot be more valid or reliable than its two operational measures.
Do not call co-variables IV and DV. A coefficient sign gives direction, while absolute size gives strength. Even a perfect association cannot by itself show which variable causes which or rule out a common cause.
A longitudinal study repeatedly measures the same participant(s), case(s) or cohort across a meaningful period to investigate change, continuity or delayed effects. It can be observational/correlational or experimental if an IV/control comparison is added.
| Form | Structure | Claim boundary |
|---|---|---|
| Descriptive longitudinal | Same measures at several time points | Describes within-unit trajectory, not its cause |
| Correlational longitudinal | Earlier co-variable predicts later outcome | Establishes time order but still has third-variable/confounding threats |
| Longitudinal experiment | Manipulated condition/control followed over time | Stronger causal and maintenance inference if allocation/control/attrition remain sound |
| Pre-post only | Same unit measured before and after | Change is visible, but history, maturation, testing and regression remain alternatives |
Specify target interval and measurement schedule; keep operational measures equivalent; record baseline and relevant confounds; preserve participant IDs securely; standardise contact; plan retention and missing-data rules; add comparison/control where causal inference is intended; predefine follow-up outcome and stopping/safeguarding procedures.
| Strength | Why it matters |
|---|---|
| Within-person change | Separates individual trajectory from one-time age-group differences |
| Temporal order | Shows predictor preceded outcome, narrowing but not eliminating causal explanations |
| Delayed/maintenance effects | Tests whether learning, treatment or brain/behaviour change persists |
| Rich repeated data | Reveals turning points and individual differences hidden by group averages |
| Threat | Consequence | Mitigation |
|---|---|---|
| Attrition | Smaller sample and systematic survivor bias | Retention plan, compare dropouts, transparent missing-data analysis |
| Practice/testing | Repetition itself changes scores | Alternate forms, spacing, appropriate control |
| Historical/maturation change | Time-related events/development mimic effect | Comparison group, repeated baseline/context measures |
| Measure drift | New tools/raters change apparent score | Calibrate/equate methods and document changes |
| Cost/privacy | Long commitment and sensitive linked records | Proportionate schedule, renewed consent, secure pseudonymous linkage |
Novel scenario answer: name the same participants, at least two dated waves, an unchanged operational outcome, expected change, retention method, attrition/practice/history control and ethical re-consent/confidentiality. If treatment is manipulated, also specify allocation and control condition.
Repeated measures compares conditions using the same participants; longitudinal follows the same units over a meaningful time course. A study can be both, but they are not synonyms. Time order improves causal reasoning yet does not remove third variables, history or maturation.
An aim states what a study intends to investigate. An alternative hypothesis predicts a difference/relationship; a null predicts no difference/relationship beyond chance. Every hypothesis must identify the population and operational variables.
| Statement | Experiment form | Correlation form |
|---|---|---|
| Aim | Investigate whether condition X affects measured Y in population P | Investigate the relationship between measured X and Y in P |
| Directional alternative | P in X1 will score higher/lower on Y than P in X2 | As X increases, Y will increase/decrease in P |
| Non-directional alternative | There will be a difference in Y between X1 and X2 for P | There will be a relationship between X and Y in P |
| Null | There will be no difference in Y between X1 and X2 for P | There will be no relationship between X and Y in P |
Use a directional hypothesis only when prior theory/evidence justifies the direction before data collection. Use non-directional when an effect is expected but its direction is uncertain. Write scores/behaviour, not vague 'better' or causal wording for a correlation.
Audit: population named; IV levels or both co-variables operationalised; DV measure named; comparison/relationship word present; direction only when justified; null is exact logical counterpart.
One-tailed means one predicted direction, not one condition. The null is not 'nothing happened'; it is a population statement of no difference/association. Results support or fail to support a hypothesis—they do not rewrite it after seeing data.
The IV is deliberately changed in an experiment; the DV is the measured outcome. An operational definition states exactly how each variable is created or scored so another researcher could reproduce it.
| Concept | Weak label | Operational version |
|---|---|---|
| IV: background sound | Music vs silence | 70 dB instrumental track through headphones versus identical headphones with no audio during a 10-minute task |
| DV: memory | Memory score | Number of 20 nouns correctly freely recalled in 2 minutes, duplicates/intrusions excluded |
| Correlational co-variable | Stress | Total score on named 10-item scale completed after school |
A DV can be frequency, duration, latency, accuracy/error, test/scale score, choice, physiological value or coded category. State unit, observation window, scoring rules and direction. Pilot floor/ceiling effects and ambiguity.
| Operational gain | Possible cost |
|---|---|
| Precision and replicability | Narrow measure may omit the construct |
| Quantitative comparison | Score may reward speed/strategy rather than target ability |
| Standardisation | Artificial task may reduce ecological validity |
Predictors/co-variables are not IVs unless manipulated. 'Aggression', 'happiness' and 'learning' are constructs, not complete DVs. Operational clarity improves replicability but does not automatically establish validity.
A control holds, removes, measures or balances a variable so conditions differ mainly in the IV. Standardisation gives every participant the same procedure. An uncontrolled variable becomes a confound when it systematically co-varies with the IV and could change the DV.
| Source | Example | Feasible control |
|---|---|---|
| Participant | Baseline memory, age, sleep | Repeated measures, matching, random allocation, baseline measurement |
| Situational | Room noise, time, device, experimenter tone | Same setting/time/material/script; randomise/balance sessions |
| Order | Practice, fatigue, carry-over | Counterbalance, rest/washout, alternate forms |
| Demand/experimenter | Guessing aim, cueing responses | Blind/double-blind, cover story where ethical, standard script |
For each threat: name it; explain how it differs between IV levels; explain how it could alter the operational DV; specify an exact control. A generic 'keep everything the same' earns less than a mechanism-linked control.
Controls strengthen internal validity/reliability but may make tasks artificial, restrict natural variation, increase cost or create ethical issues. Measure rather than eliminate important real-world factors when ecological validity is central.
Not every uncontrolled variable is a confound; it must vary systematically with the IV and explain the DV. Standardisation cannot remove participant differences by itself. Random allocation balances conditions; random sampling addresses population recruitment.
Quantitative versus qualitative describes data form; objective versus subjective describes dependence on personal judgement/experience. The two axes can combine in four ways.
| Objective/externally verifiable | Subjective/judgement-based | |
|---|---|---|
| Quantitative | Reaction time, heart rate, correct-answer count | Participant's 0-8 distress rating; observer category score requiring judgement |
| Qualitative | Verbatim audio transcript/recorded words as raw event | Participant's interpretation or researcher's thematic account |
| Type | Main value | Main limit |
|---|---|---|
| Quantitative | Compact comparison, graphs, replicable scoring | Reduction may omit meaning/context |
| Qualitative | Rich explanations and unexpected themes | Slow coding, interpretation and lower inter-rater reliability |
| Objective | Less affected by self-presentation/interpretation | May measure proxy rather than lived construct |
| Subjective | Direct access to private experience | Social desirability, memory and perspective bias |
Choose data to fit the question, then triangulate: e.g. combine sleep minutes from actigraphy (quantitative/objective), rating of sleep quality (quantitative/subjective) and diary account (qualitative/subjective). Agreement strengthens confidence; disagreement is evidence to explain, not delete.
Numbers are not automatically objective: a Likert score quantifies a subjective judgement. Words are not automatically subjective: an exact recorded utterance is an observable event, though coding its meaning may be subjective. No data type is universally best.
The population is the full group to which a researcher wants to generalise; the sample is the participating subset. Sampling technique controls who gets an invitation, but non-response and eligibility still shape the final sample.
| Technique | Exact procedure | Strength | Bias/limit |
|---|---|---|---|
| Opportunity | Recruit available eligible people at chosen place/time | Fast, cheap, practical | Place/time/researcher-access bias; often unrepresentative |
| Volunteer/self-selecting | Advertise and eligible people opt in | Consent/interest, reaches dispersed people | Volunteer traits: motivation, time, topic interest, payment response |
| Random | Build complete sampling frame; use random generator/lottery so each has equal selection chance | Reduces researcher selection bias; potentially representative | Frame may omit people; selected non-response; costly |
Novel scenario: define target population with inclusion/exclusion; identify/construct sampling frame if random; state recruitment location/channel/time; describe exact selection—not just name; anticipate who is missed/refuses; compare achieved sample demographics with population; bound generalisation accordingly.
| Evidence | Generalisation consequence |
|---|---|
| Narrow age/sex/culture/occupation | Findings may not transfer where relevant mechanism differs |
| Large but biased online volunteer sample | Precision can increase while representativeness remains poor |
| Small random sample | Less selection bias but greater sampling fluctuation |
| Attrition/non-response | Final sample may differ from those initially selected |
Random sampling recruits from a population; random allocation assigns a recruited sample to conditions. Large does not equal representative. Opportunity sampling is based on availability, not deliberate quota matching. Generalisability also depends on task/setting/time, not sample alone.
| Human guideline | Requirement and safeguard |
|---|---|
| Minimise harm/maximise benefit | Risk assess, monitor distress, stop/refer/support; use least harmful effective procedure |
| Valid informed consent | Capacity, understandable purpose/procedure/risks/data use; guardian consent plus assent where relevant |
| Right to withdraw | Leave and remove data without penalty; make route/reminder practical |
| Lack of deception | Disclose truth unless justified/minimal and impossible otherwise; never deceive about material risk |
| Confidentiality | Limit access, pseudonymise, secure storage/reporting |
| Privacy | Observe/collect only where people reasonably expect and consent; minimise intrusion |
| Debriefing | Reveal aim/deception, restore understanding, answer questions, offer data withdrawal/support |
| Animal guideline | Applied question |
|---|---|
| Minimise harm/maximise benefit | Is scientific/welfare value proportionate to pain, distress and lasting effects? |
| Replacement | Can non-animal, simulation, existing data or less sentient model answer it? |
| Species | Is species scientifically appropriate and welfare expertise available? |
| Numbers | Use minimum needed for valid evidence—too few also wastes animals |
| Procedures | Refine handling/anaesthesia/endpoints; suitable housing/social needs; justified reward/deprivation; avoid aversive stimuli |
Exam chain: name guideline → cite exact procedure/sample/data feature → explain likely harm/autonomy/welfare or benefit → judge severity/probability/reversibility → propose feasible safeguard and its methodological trade-off. For animals, use the syllabus animal guideline rather than importing human consent/right-to-withdraw labels.
| Method need | Ethical tension | Design response |
|---|---|---|
| Avoid demand characteristics | Deception/incomplete disclosure | Minimal deception, no risk deception, prior/retrospective consent, prompt debrief |
| Natural public behaviour | Consent/privacy | Public-expectation audit, anonymised low-risk recording, gatekeeper/debrief where possible |
| Stress/emergency simulation | Harm and withdrawal | Lower intensity, screening, stop rule, support and alternative task |
| Animal reinforcement | Deprivation/aversive welfare | Preferred reward, minimal restriction, voluntary participation and welfare endpoints |
Signed consent is not automatically valid if information, capacity or freedom is missing. Debriefing mitigates deception but cannot undo severe harm. Confidentiality concerns data identity; privacy concerns access to the person/behaviour. Ethical acceptability is a reasoned balance, not a checklist score.
Validity is the extent to which a measure/study supports the intended interpretation. Ask separately whether the construct was measured, the IV caused the DV, the behaviour represents real life and the finding transfers to the target population/context.
| Validity question | Main threats | Contextual improvement |
|---|---|---|
| Construct/measurement | Proxy score, subjective coding, social desirability | Operational/pilot/validated measure, blind coding, triangulation |
| Internal/causal | Confounds, allocation/order/experimenter effects | Control/standardise, random allocation, counterbalance, blind |
| Ecological | Artificial task/setting, low mundane realism | More representative task/context, field evidence—while retaining controls |
| Population/generalisation | Narrow/biased sample and cultural/time differences | Broader/stratified/random recruitment, replication across groups |
Demand characteristics arise when participants infer the aim and alter behaviour; reduce through credible neutral instructions, unobtrusive/indirect measures, blinding or justified deception/debrief. Subjectivity can add insight but requires transparent coding/checks. Objectivity reduces judgement, not necessarily proxy invalidity.
Write: identify exact evidence feature → name the interpretation threatened/supported → explain mechanism → judge consequence for this conclusion → propose improvement and trade-off. 'Laboratory means low validity' is incomplete without showing why this task differs from the target behaviour.
A field setting can contain an invalid measure; a laboratory task can validly test a narrow mechanism. Reliability is necessary for many valid measurements but consistent bias remains invalid. Generalisability extends beyond sample to setting, task, culture and time.
Reliability is consistency of measurement/procedure. Replicability is whether documentation/materials are sufficient for another researcher to repeat the study and test whether the pattern recurs.
| Check | Same/different element | Use | Improve when low |
|---|---|---|---|
| Inter-rater | Two raters score same response/product | Interviews, tests, qualitative coding | Rubric, examples, training, blind double-code |
| Inter-observer | Two observers code same behaviour/time | Structured observation | Operational categories, observer training, video recode |
| Test-retest | Same measure to same people at two suitable times | Stable traits/questionnaires/tests | Clarify items, standardise conditions; avoid interval/practice extremes |
| Procedural replication | Independent study repeats documented method | Tests result robustness | Full script/materials/operational rules, sample and analysis transparency |
Plan: standardise instructions, timing, apparatus and scoring; define categories/units; pilot ambiguity; preserve versioned materials; report recruitment/allocation/exclusions; have independent coders; choose a reliability check that targets the likely source of inconsistency.
High reliability narrows random/observer inconsistency and makes differences interpretable. Low reliability can hide or create effects. Yet a consistently wrong scale, biased item or invalid category can be highly reliable.
Two researchers repeating a whole study is replication, not inter-rater reliability. Test-retest uses the same measure/construct at two times, not two experimental conditions. Agreement does not prove validity, objectivity or truth.
Descriptive analysis organises and summarises observed data. Central tendency gives a typical/central value; spread shows variability. Cambridge requires recognition/finding/interpretation, not calculations or unspecified statistical tests.
| Measure | How found | Best feature | Main weakness |
|---|---|---|---|
| Mode | Most frequent value/category | Works with nominal categories; shows common response | May be multiple/none; ignores rest |
| Median | Middle ordered score (average two middles if even) | Resistant to extreme/skew; ordinal suitable | Ignores distances and much data |
| Mean | Sum divided by number | Uses every interval/ratio score; sensitive comparison | Distorted by extremes/skew; unsuitable for categories |
| Measure | Meaning | Interpretation |
|---|---|---|
| Range | Highest minus lowest (state convention if endpoints included) | Overall span; one extreme can dominate |
| Standard deviation | Typical dispersion around mean | Low SD = scores clustered/consistent; high SD = dispersed—not a high/low mean |
| Display | Data structure | Construction/reading rule |
|---|---|---|
| Table | Any organised categories/conditions | Clear title, labelled rows/columns, units, no ambiguous totals |
| Bar chart | Separate categories/conditions | Equal-width separated bars; axes/units; height = frequency/summary |
| Histogram | Continuous scores grouped into adjacent intervals | Bars touch; numerical ordered x-axis; frequency on y |
| Scatter graph | Paired co-variable scores | One dot per pair; both axes operational variables; inspect direction/strength/outliers |
Compare centre and spread together: equal means can hide different consistency; lower mean may accompany wider overlap. Describe exact pattern with units, largest/smallest, difference and variability. Do not infer cause or significance from descriptive displays.
Bar charts display discrete categories with gaps; histograms display continuous intervals with touching bars; scatter graphs do not display group frequencies. Low SD means low variation, not low scores. A mean difference alone does not prove a reliable population effect.