Q BankQuestion BankDocsDocuments

2 Probability, Random Variables, and Probability Distributions

Syllabus
2026
Section
2
Level

Exam analysis

No tagged past-paper evidence yet

Published Concept pages under this syllabus area do not have tagged past-paper appearances in the selected level yet.

Recent 5 years

In this section

Topic 2.1

2.1 Tabular and Graphical Representations for the Distributions of Two Categorical Variables

Objectives in this topic

2.1.A—Compare tabular and graphical representations for the relationship between two categorical variables

Compare tabular and graphical representations for the relationship between two categorical variables.

  • A two-way table, also called a contingency table, can be used to summarize and compare data for two categorical variables. The entries in the cells of the table can be frequencies (i.e., counts) or relative frequencies (i.e., proportions).
  • Side-by-side bar charts, segmented bar charts, and mosaic plots are examples of graphs used to display the relationship between two categorical variables. In these graphs, the frequency or relative frequency of each category, or level, of one of the categorical variables is displayed for each category of the other categorical variable.
  • Graphical representations of two categorical variables can be used to compare the relationship of one categorical variable across the levels of the other categorical variable and determine whether the two variables are associated.

2.1.B—Justify a claim using tabular and graphical representations for the distributions of two categorical variables

Justify a claim using tabular and graphical representations for the distributions of two categorical variables.

  • Tabular and graphical representations for the distributions of two categorical variables may reveal information that can be used to justify claims about the variable in context. Probability, Random Variables, and Probability Distributions UNIT 2

Topic 2.2

2.2 Summary Statistics for Two Categorical Variables

Objectives in this topic

2.2.A—Calculate summary statistics from two-way tables

Calculate summary statistics from two-way tables.

  • A joint relative frequency in a two-way table is a cell frequency divided by the total for the entire table.
  • A marginal relative frequency in a two-way table is a row total divided by the total for the entire table or a column total divided by the total for the entire table.
  • A conditional relative frequency is a relative frequency computed by restricting to a particular level, or category of interest. A conditional relative frequency can be a cell frequency in a row divided by the total for that row or it can be a cell frequency in a column divided by the total for that column.

2.2.B—Compare summary statistics for two categorical variables

Compare summary statistics for two categorical variables.

  • Summary statistics for two categorical variables can be used to compare distributions for evidence of association between the two variables.

2.2.C—Justify a claim using summary statistics for two categorical variables

Justify a claim using summary statistics for two categorical variables.

  • Summary statistics for two categorical variables may reveal information that can be used to justify claims about the variables in context.

Topic 2.3

2.3 Estimating Probabilities Using Simulation

Objectives in this topic

2.3.A—Estimate probabilities using simulations

Estimate probabilities using simulations.

  • A random process generates results that are determined by chance.
  • An outcome is the result of one trial of a random process.
  • An event is a collection of outcomes.
  • Simulation is a way to model random events such that simulated outcomes closely match real-world outcomes. All possible outcomes are associated with a value to be determined by chance. Record the counts of simulated outcomes and the count total.
  • The probability of an outcome or event is its long-run relative frequency—that is, its relative frequency over a large number of trials.
  • The relative frequency of an outcome or event determined from empirical data can be used to estimate the actual, or true, probability of that outcome or event.
  • The law of large numbers states that for independent trials, as the number of trials increases, the long-run relative frequency of the outcome or event gets closer and closer to a single value.

Topic 2.4

2.4 Introduction to Probability

Objectives in this topic

2.4.A—Calculate probabilities for events and their complements

Calculate probabilities for events and their complements.

  • The sample space of a random process is the set of all possible nonoverlapping outcomes. The probability of the sample space is 1.
  • If all outcomes in the sample space are equally likely, then the theoretical probability an event E will occur is Enumber of outcomes in event total number of outcomes in the sample space . The probability of event E occurring is written as P ()E .
  • The probability of an event is a number between 0 and 1, inclusive.
  • The probability of the complement of an event E, which can be written as E′, E, or EC (i.e., the probability of “not E”) is equal to 1- PE() .

Topic 2.5

2.5 Mutually Exclusive Events

Objectives in this topic

2.5.A—Justify why two events are mutually exclusive (or disjoint) using joint probability

Justify why two events are mutually exclusive (or disjoint) using joint probability.

  • The probability that events A and B both will occur, sometimes called the joint probability, is the probability of the intersection of A and B. Joint probability is defined as P ()ABintersect or P (AB∩ .)
  • Two events are mutually exclusive, or disjoint, if they cannot occur at the same time. This means that if two events are mutually exclusive, then P ()AB∩ =0 .

Topic 2.6

2.6 Conditional Probability

Objectives in this topic

2.6.A—Calculate conditional probabilities

Calculate conditional probabilities.

  • The probability that event A will occur given that event B has occurred is called a conditional probability and is written as P(|AB ). Conditional probability is defined as P AB PB PB(| ) = ∩() () A .
  • The general multiplication rule states that the probability that events A and B both will occur is equal to the probability that event A will occur multiplied by the conditional probability that event B will occur given that event A has occurred. The multiplication rule is defined as P AB PP BAA∩() = () . (| ) .

Topic 2.7

2.7 Independent Events and Unions of Events

Objectives in this topic

2.7.A—Calculate probabilities for independent events and for the union of two events

Calculate probabilities for independent events and for the union of two events.

  • Events A and B are independent if and only if knowing whether event A has occurred (or will occur) does not change the probability that event B will occur. When events A and B are independent, then P AB PA(|)=() , P BA PB(|)=() , and P AB PA PB∩() = () . () .
  • The probability that event A or event B (or both) will occur is the probability of A union B. The probability of the union is defined as P (AB∪ .)
  • P ABunion () = P PP AB AB() + () - () intersect , or P AB PPPA B AB∪() = () + () -∩() .

Topic 2.8

2.8 Introduction to Random Variables and Probability Distributions

Objectives in this topic

2.8.A—Construct a probability distribution for a discrete random variable

Construct a probability distribution for a discrete random variable.

  • A random variable is a variable whose values have numerical outcomes that result from a random phenomenon.
  • A probability distribution for a discrete random variable shows the probability associated with every possible value of the random variable. The sum of the probabilities over all possible values of a discrete random variable is 1.
  • A discrete probability distribution can be determined using the rules of probability or estimated with a simulation.
  • A discrete probability distribution can be represented as a graph, table, or function showing the probabilities associated with values of a random variable.
  • A cumulative probability distribution can be represented as a table or function and shows the probability of being less than or equal to each value of the discrete random variable.

Topic 2.9

2.9 Parameters of Random Variables

Objectives in this topic

2.9.A—Calculate the mean and standard deviation for a discrete random variable

Calculate the mean and standard deviation for a discrete random variable.

  • A numerical value measuring a characteristic of a probability distribution of a random variable, or a population, is a parameter. The value of a parameter is a single, fixed value.
  • The expected value (or mean) of a probability distribution is a parameter and is denoted by EX( ) or µX. For a discrete random variable X, the expected value is calculated as μXi =∑xP. ()x ,i where xi is the possible value of the random variable and P()xi is the probability of the possible value of the random variable. The expected value can be interpreted as the long-run average outcome of the random variable. The discrete random variable can only take on values that are countable or finite.
  • The standard deviation of a probability distribution is a parameter represented by SD X() or σX. For a discrete random variable X, the standard deviation is calculated as σ μXi i X xP x=∑ -() . () 2 , where xi is the possible value of the random variable, µX is the mean, and P()xi is the probability of the possible value of the random variable. The standard deviation can be interpreted as the typical deviation of the values of the random variable from the mean value (or expected value) of the random variable over the long run. The square of the standard deviation of a random variable is called the variance of the random variable and is denoted as VX() or σ2 X. Probability, Random Variables, and Probability Distributions UNIT 2

2.9.B—Interpret the mean and standard deviation for a discrete random variable

Interpret the mean and standard deviation for a discrete random variable.

  • The mean and standard deviation for the probability distribution of a discrete random variable should be interpreted in the context of a specific population. Probability, Random Variables, and Probability Distributions UNIT 2 72

Topic 2.10

2.10 The Binomial Distribution

Objectives in this topic

2.10.A—Justify why a random variable is or is not a binomial random variable

Justify why a random variable is or is not a binomial random variable.

  • A binomial random variable, X, is a discrete random variable that counts the number of successes in repeated independent trials, n, that have only two possible outcomes (success or failure), with the probability of success p and the probability of failure 1−p .

2.10.B—Calculate the mean and standard deviation for a binomial distribution

Calculate the mean and standard deviation for a binomial distribution.

  • If a random variable is binomial, its mean, µX, is np and its standard deviation, σX, is np(1-p .)

2.10.C—Interpret the mean, standard deviation, and probabilities for a binomial distribution

Interpret the mean, standard deviation, and probabilities for a binomial distribution.

  • The mean, standard deviation, and probabilities for a binomial distribution should be interpreted in context.

2.10.D—Estimate probabilities of binomial random variables using data from a simulation

Estimate probabilities of binomial random variables using data from a simulation.

  • A probability distribution can be constructed using the rules of probability or estimated with a simulation. Probability, Random Variables, and Probability Distributions UNIT 2

2.10.E—Calculate probabilities for a binomial distribution

Calculate probabilities for a binomial distribution.

  • The probability that a binomial random variable, X, has exactly x successes for n independent trials, when the probability of success is p, is calculated as ( n) PX() | |= x = ppx - nx | x| ()1 - ( ) , x = 012,, ,, ... n. This is called the binomial probability function. Probability, Random Variables, and Probability Distributions UNIT 2 74

Topic 2.11

2.11 The Normal Distribution

Objectives in this topic

2.11.A—Describe a normal distribution

Describe a normal distribution.

  • A continuous random variable is a variable that can take on any value within a specified domain. Every interval within the domain has a probability associated with it.
  • Many continuous random variables are wellmodeled by a normal distribution.
  • A normal distribution can be described as a continuous, unimodal, bell-shaped, and symmetric curve.
  • A normal curve can be used to model a distribution of data and a continuous random variable.
  • The normal distribution, or the normal curve, is identified by two parameters, the mean, µ, and the standard deviation, σ. The smaller the standard deviation, the taller and more concentrated the normal curve is around its mean. The larger the standard deviation, the shorter and less concentrated the normal curve is around its mean.

2.11.B—Calculate the mean and standard deviation for a normal distribution

Calculate the mean and standard deviation for a normal distribution.

  • A standard normal distribution is a normal distribution with mean and standard deviation μ =0 σ =1. Probability, Random Variables, and Probability Distributions UNIT 2

2.11.C—Calculate percentages from a normal distribution using the empirical rule

Calculate percentages from a normal distribution using the empirical rule.

  • The empirical rule can be used to estimate the area of a region under the graph of the normal distribution curve. For a normal distribution, approximately 68% of observations are within 1 standard deviation of the mean, approximately 95% of observations are within 2 standard deviations of the mean, and approximately 99.7% of observations are within 3 standard deviations of the mean. This is called the empirical rule, or the 68–95–99.7 rule.

2.11.D—Calculate the probability that a particular value lies within a given interval of a normal distribution

Calculate the probability that a particular value lies within a given interval of a normal distribution.

  • If the distribution of a random variable is approximately normal, the probability that the random variable takes on values within a particular interval of the random variable is determined by the area under the normal curve within that interval. The total probability or area under the normal curve is 1.

2.11.E—Calculate the associated intervals and areas of a normal distribution

Calculate the associated intervals and areas of a normal distribution.

  • The boundaries of an interval associated with a given area in a normal distribution can be determined using technology or using z-scores and a standard normal table.
  • Intervals associated with a given area in a normal distribution can be determined by assigning appropriate inequalities to the boundaries of the intervals. To determine the intervals, p is defined as a number between 0 and 100, xa is the lower bound, and xb is the upper bound on a normal distribution.
    • i. P Xx p a() <= 100 means that the lowest p% of the values lie to the left of xa . Probability, Random Variables, and Probability Distributions UNIT 2 76
    • ii. P xX x p ab()<< = 100 means that p% of the values lie between xa and xb.
    • iii. pP()Xx>=b 100 means that the highest p% of the values lie to the right of xb.
    • iv. To determine the most extreme p% of values on both sides requires dividing the area associated with p% into two equal areas on either extreme of the distribution: P Xx p a()<= ( (| ) )| 1 2 100 and P Xx p b()>= ( (| ) )| 1 2 100 mean that half of the p% most extreme values lie to the left of xa and half of the p% most extreme values lie to the right of xb.

2.11.F—Compare measures of relative position for distributions

Compare measures of relative position for distributions.

  • Percentiles and proportions may be used to compare relative positions of individual values within a normal distribution or between normal distributions. Probability, Random Variables, and Probability Distributions UNIT 2

Topic 2.12

2.12 Sampling Distributions and the Central Limit Theorem

Objectives in this topic

2.12.A—Describe sampling distributions with simulations

Describe sampling distributions with simulations.

  • A sampling distribution of a statistic is the distribution of values of the statistic for all possible samples of a given size from a given population.
  • The sampling distribution of a statistic can be simulated by repeatedly generating a large number of random samples from the population assuming known value(s) for the parameter(s). The value of the statistic is determined and recorded for each sample. The resulting distribution of the sample statistic values approximates the sampling distribution of the statistic.
  • A randomization distribution is the distribution of a statistic generated by simulation from repeatedly randomly reallocating, or reassigning, the response values to treatment groups. The value of the statistic is determined and recorded for each reallocation, or reassignment. The resulting distribution of the statistic values approximates the sampling distribution of the statistic.
  • The central limit theorem (CL T) states that the sampling distribution of a mean of a random sample has a shape that can be approximated by a normal distribution. The larger the sample is, the better the approximation will be.
ConceptAP Statistics