Q BankQuestion BankDocsDocuments

3 Inference for Categorical Data: Proportions

Syllabus
2026
Section
3
Level

Exam analysis

No tagged past-paper evidence yet

Published Concept pages under this syllabus area do not have tagged past-paper appearances in the selected level yet.

Recent 5 years

In this section

Topic 3.1

3.1 Estimators

Objectives in this topic

3.1.A—Justify why an estimator is or is not unbiased

Justify why an estimator is or is not unbiased.

  • When estimating a population parameter, an estimator is unbiased if, on average, the value of the estimator does not underestimate or overestimate the population parameter.

3.1.B—Calculate estimates for a population parameter

Calculate estimates for a population parameter.

  • A sample statistic is a point estimator of the corresponding population parameter and can be thought of as the estimate of the population parameter. For example, the sample proportion p is a point estimator for the population proportion p.

Topic 3.2

3.2 Sampling Distributions for Sample Proportions

Objectives in this topic

3.2.A—Calculate the mean and standard deviation of a sampling distribution for a sample proportion

Calculate the mean and standard deviation of a sampling distribution for a sample proportion.

  • For a population with population proportion p, when the sampled values are independent, the sampling distribution of a sample proportion p has a mean μp = p and a standard deviation p pp n  1 .

3.2.B—Justify the appropriateness of conditions for the sampling distribution of a sample proportion

Justify the appropriateness of conditions for the sampling distribution of a sample proportion.

  • Sampling without replacement requires that two conditions be met:
    • i. The randomization condition—the data should be collected using a random sample.
    • ii. The 10% condition—the population size must be at least 10 times larger than the sample size nN≤10 , where N is the size of the population and n is the sample size.
  • The sampling distribution of the sample proportion p is approximately normal provided the sample size is large enough. To ensure the sample size is lar ge enough, the following condition must be met: np 10 and n 1 p 10, where np is the expected number of successes and n 1 p is the expected number of failures.

3.2.C—Interpret the mean, standard deviation, and probabilities for a sampling distribution of a sample proportion

Interpret the mean, standard deviation, and probabilities for a sampling distribution of a sample proportion.

  • The mean, standard deviation, and probabilities for a sampling distribution of a sample proportion should be interpreted in the context of a specific population.

Topic 3.3

3.3 Constructing a Confidence Interval for a Population Proportion

Objectives in this topic

3.3.A—Identify an appropriate confidence interval procedure including the parameter for a population proportion

Identify an appropriate confidence interval procedure including the parameter for a population proportion.

  • A confidence interval is an interval estimate for a population parameter. Based on the sample proportion, a confidence interval can be calculated to estimate the value of a single population proportion. The appropriate confidence interval procedure is a one-sample z-interval for a population proportion.
  • The parameter for a confidence interval for a population proportion should reference the proportion, the response variable, and the population in context.

3.3.B—Justify the appropriateness of constructing a confidence interval for a population proportion by verifying…

Justify the appropriateness of constructing a confidence interval for a population proportion by verifying conditions.

  • A one-sample z-interval for a population proportion requires that three conditions be met:
    • i. The randomization condition—the data should be collected using a random sample.
    • ii. The 10% condition—when sampling without replacement, the population size must be at least 10 times larger than the sample size nN10% , where N is the size of the population and n is the sample size.
    • iii. The normality condition—the observed number of successes, , and the observed number of failures, 1 , should be at least 10.

3.3.C—Calculate an appropriate confidence interval for a population proportion

Calculate an appropriate confidence interval for a population proportion.

  • z denotes a critical value, such that z and z represent the boundaries enclosing the middle C% of the standard normal distribution, in which C% is an approximate confidence level with which the population proportion is estimated.
  • An interval estimate can be constructed as point estimate (margin of error). For a population proportion, the one-sam p ple z-interval estimate is p 1 z n .

3.3.D—Calculate the standard error and margin of error of a sample statistic for a confidence interval for a…

Calculate the standard error and margin of error of a sample statistic for a confidence interval for a population proportion, and estimate a given sample size from the margin of error.

  • The standard error (SE) of a statistic is an estimate of the standard deviation of the sampling distribution of the statistic. The standard error of the sample proportion p is SE pp np 1 .
  • The standard error quantifies the typical amount that a statistic will vary from the value of the corresponding population parameter.
  • The margin of error of p is half the width of the confidence interval and is calculated as the critical value z times the standard error (SE) of p, which equals p 1 z n . p
  • The formula for the margin of error (MOE) can be rearranged to solve for n, n zp p MOE 2  1 2 , the minimum sample needed to achieve a given margin of error. For this purpose, if p is not defined or unable to be calculated, use p=0.5 in order to find the upper bound for the sample size that will result in a given mar gin of error. Inference for Categorical Data: Proportions UNIT 3 90

Topic 3.4

3.4 Justifying a Claim Based on a Confidence Interval for a Population Proportion

Objectives in this topic

3.4.A—Interpret a confidence interval in context for a population proportion

Interpret a confidence interval in context for a population proportion.

  • Because the confidence interval for a population proportion is calculated based on a sample from a population, the computed interval may or may not contain the value of the population proportion.
  • The interpretation of the confidence level is that in repeated random sampling with the same sample size, approximately C% of confidence intervals calculated will capture the population proportion, with C representing the numerical value of the confidence level used.
  • When interpreting a C% confidence interval for a population proportion, we say we are C% confident that the interval (,ab ) contains the true value of the parameter for the population, where a represents the lower limit and b represents the upper limit. An interpretation of a confidence interval for a population proportion includes a reference to the parameter with details about the population it represents in the context of the study.

3.4.B—Justify a claim based on a confidence interval for a population proportion

Justify a claim based on a confidence interval for a population proportion.

  • A confidence interval for a population proportion provides a range of plausible values that may serve as convincing evidence to support a particular claim about the population proportion. Inference for Categorical Data: Proportions UNIT 3

3.4.C—Identify the relationships among sample size, confidence interval width, confidence level, and margin of error…

Identify the relationships among sample size, confidence interval width, confidence level, and margin of error for a population proportion.

  • For a given sample, increasing the confidence level will result in the following:
    • i. The critical value will increase.
    • ii. The margin of error will increase.
    • iii. The width of the confidence interval will increase.
  • Increasing the sample size decreases the standard error. Thus, when all other things remain the same, the width of the confidence interval for a population proportion tends to decrease as the sample size increases. For a confidence interval for a population proportion with a given confidence level, the width of the interval is approximately proportional to 1 n . Inference for Categorical Data: Proportions UNIT 3 92

Topic 3.5

3.5 Setting Up a Test for a Population Proportion

Objectives in this topic

3.5.A—Identify an appropriate testing method for a population proportion including the parameter for the population…

Identify an appropriate testing method for a population proportion including the parameter for the population proportion.

  • A hypothesis test is a statistical inference procedure that is used to make a decision about the value of a population parameter. The appropriate hypothesis testing procedure is a one-sample z-test for a population proportion.
  • The parameter for a hypothesis test for a population proportion should reference the population parameter, the response variable, and the population in context.

3.5.B—Identify the null and alternative hypotheses for a population proportion

Identify the null and alternative hypotheses for a population proportion.

  • In the hypothesis testing procedure, the null hypothesis, H0, is the statement about a parameter that is assumed to be correct unless there is convincing statistical evidence suggesting otherwise. It is the status quo condition. The alternative hypothesis, Ha, is the claim or belief about a parameter for which evidence is being collected. A researcher’s claim or belief about the population parameter is represented by the alternative hypothesis.
  • The null hypothesis contains an equality reference ,, or . Although the null hypothesis for a one-sided test may include an inequality symbol, in
  • The null hypothesis for a one-sample z-test for a population proportion is as follows: H:0 p ,p 0 where p0 is the null hypothesized value for the population proportion. A one-sided alternative hypothesis for a one-sample z-test for a population proportion is either H:a0 pp or H: pp 0a . A two-sided alternative hypothesis is H:a pp 0.

3.5.C—Justify the appropriateness of a hypothesis test for a population proportion by verifying conditions

Justify the appropriateness of a hypothesis test for a population proportion by verifying conditions.

  • A one-sample z-test for a population proportion requires that three conditions be met:
    • i. The randomization condition—the data should be collected using a random sample.
    • ii. The 10% condition—when sampling without replacement, the population size must be at least 10 times larger than the sample size nN10 , where N is the size of the population and n is the sample size.
    • iii. The normality condition—the expected number of successes, np0 0, and the expected number of failures, n 1 p0 , should be at least 10. Inference for Categorical Data: Proportions UNIT 3 94

Topic 3.6

3.6 p-Values

Objectives in this topic

3.6.A—Interpret the p-value of a hypothesis test for a population proportion

Interpret the p-value of a hypothesis test for a population proportion.

  • Given the null hypothesis is true, there is a probability distribution of the test statistic called the null distribution. Using the null distribution, the p-value is the probability of obtaining a test statistic as extreme or more extreme (i.e., in the direction of the alternative hypothesis) than the test statistic that is observed given that the null hypothesis is true. That is, when x is the test statistic, the p-value is determined by finding the following:
    • i. The probability at or above the observed value of the test statistic Pz x , if the alternative is >
    • ii. The probability at or below the observed value of the test statistic Pz x , if the alternative is <
    • iii. The probability less than or equal to the negative of the absolute value of the test statistic plus the probability greater than or equal to the absolute value of the test Pz xP z x ,statistic, if the alternative is ≠
  • If the distribution of the test statistic has been simulated, the p-value is the proportion of values in the null distribution that are as extreme or more extreme than the observed value of the test statistic. This is as follows:
    • i. The proportion at or above the observed value of the test statistic, if the alternative is >
    • ii. The proportion at or below the observed value of the test statistic, if the alternative is <
    • iii. The proportion less than or equal to the negative of the absolute value of the test statistic plus the proportion greater than or equal to the absolute value of the test statistic, if the alternative is
  • An interpretation of the p-value of a hypothesis test for a population proportion should include a statement that the p-value is computed by assuming the null hypothesis is true (i.e., by assuming the true population proportion is equal to the particular value stated in the null hypothesis in context).
  • Small p-values indicate that the observed value of the test statistic would be unusual if the null hypothesis were true and therefore provide evidence for the alternative hypothesis. The lower the p-value, the more convincing the statistical evidence for the alternative hypothesis.
  • p-values that are not small indicate that the observed value of the test statistic would not be unusual if the null hypothesis were true and therefore do not provide convincing statistical evidence for the alternative hypothesis, nor do they provide evidence that the null hypothesis is true. Inference for Categorical Data: Proportions UNIT 3 96

Topic 3.7

3.7 Carrying Out a Test for a Population Proportion

Objectives in this topic

3.7.A—Calculate an appropriate test statistic and p-value for testing a hypothesis about a population proportion

Calculate an appropriate test statistic and p-value for testing a hypothesis about a population proportion.

  • The test statistic for testing a population proportion is pp n  0 00 1 . The z-statistic has a standard normal distribution when the null hypothesis is true.
  • The distribution of the test statistic assuming the null hypothesis is true (null distribution) can be approximated by a probability model (e.g., a theoretical distribution such as the standard normal distribution).
  • The p-value of a one-sample z-test for a population proportion is found from the standard normal distribution using a table or technology.

3.7.B—Justify a claim about the population based on the results of a hypothesis test for a population proportion

Justify a claim about the population based on the results of a hypothesis test for a population proportion.

  • The significance level of a hypothesis test, denoted by α, is the predetermined probability of rejecting the null hypothesis given that it is true. The significance level may be given or determined by the researcher. The relationship between a p-value and the significance level of a hypothesis test determines whether a result is statistically significant.
  • A formal decision in a hypothesis test explicitly compares the p-value to the significance level, α. If the p-value , then reject the null hypothesis, 0 pp .0 If the p-value , then fail to reject the null hypothesis.
  • Rejecting the null hypothesis means there is convincing statistical evidence to support the alternative hypothesis. Failing to reject the null hypothesis means there is not convincing statistical evidence to support the alternative hypothesis.
  • A hypothesis test can lead to rejecting or not rejecting the null hypothesis but can never lead to concluding or proving that the null hypothesis is true. Lack of statistical evidence for the alternative hypothesis is not the same as evidence for the null hypothesis.
  • The results of a hypothesis test for a population proportion can serve as the statistical reasoning to support the answer to an investigative question about the population that was sampled.
  • A conclusion for the hypothesis test for a population proportion is stated in context consistent with, and in terms of, the alternative hypothesis using non-definitive language. The conclusion should contain a reference to the parameter and the population. Inference for Categorical Data: Proportions UNIT 3 98

Topic 3.8

3.8 Potential Errors When Performing Tests

Objectives in this topic

3.8.A—Identify Type I and Type II errors

Identify Type I and Type II errors.

  • A Type I error occurs when there is convincing statistical evidence that the alternative hypothesis is true (due to the small p-value), but it is not.
  • A Type II error occurs when there is not convincing statistical evidence that the alternative hypothesis is true (due to the large p-value), but it is.
  • The power of a hypothesis test is the probability that a hypothesis test will correctly reject the false null hypothesis.

3.8.B—Calculate the probability of Type I and Type II errors

Calculate the probability of Type I and Type II errors.

  • The probability of making a Type I error is defined as the significance level, α. For a given study and hypothesis test, the probability of making a Type I error is typically set to a small value (e.g., 0.01, 0.05, 0.10) prior to collecting the data.
  • The probability of making a Type II error is 1 power.

3.8.C—Identify the factors that affect the probability of errors in hypothesis testing

Identify the factors that affect the probability of errors in hypothesis testing.

  • For a given study and hypothesis test, the probability of a Type II error should ideally be small, and thus, the power will be large (e.g., P Type II error 02. 0 and power0 .80). The probability of a Type II error decreases and the power increases when any one of the following occurs, provided the others do not change:
    • i. Sample size(s) increases.
    • ii. Standard error decreases.
    • iii. True parameter value is farther from the null hypothesis.
    • iv. Significance level of a test increases.

3.8.D—Interpret Type I and Type II errors

Interpret Type I and Type II errors.

  • In some studies, making a Type I error may have more serious consequences than making a Type II error. In other studies, making a Type II error may have more serious consequences than making a Type I error. The consequences of each error should be considered prior to conducting the study.
  • Because the significance level, α, is the probability of making a Type I error, the consequences of a Type I error influence decisions about a significance level.
  • Because sample size influences the probability of making a Type II error, the consequences of a Type II error influence decisions about how large the sample size should be. Inference for Categorical Data: Proportions UNIT 3 100

Topic 3.9

3.9 Sampling Distributions for the Difference Between Two Population Proportions

Objectives in this topic

3.9.A—Calculate the mean and standard deviation of the sampling distribution for the difference between two sample…

Calculate the mean and standard deviation of the sampling distribution for the difference between two sample proportions.

  • For two independent populations, with population proportions p1 and p2, when the sampled values are independent, the s ampling distribution for the difference in sample proportions, p p , has a mean, pp pp12 12 and standar 12 d deviation, pp pp n pp n12 11 1 22 2 11 .

3.9.B—Justify the appropriateness of conditions for the sampling distribution for the difference between two sample…

Justify the appropriateness of conditions for the sampling distribution for the difference between two sample proportions.

  • When sampling without replacement, two conditions must be met:
    • i. The randomization condition—the data should be collected using two independent random samples.
    • ii. The 10% condition—the size of each sample should be less than or equal to 10% of the respective population size: n11 10%N and n22 10%N , where N is the size of population 1 and 1 N2 is the size of population 2. The sample sizes are represented as n1 and n2.
  • If the data come from an experiment, the data only need to meet the randomization condition. The treatments must be randomly assigned to the experimental units to meet the randomization condition.
  • The sampling distribution for the difference between sample proportions, p 12 p , will have an approximately normal distribution provided both s ample sizes are large enough. To ensure that both samples are large enough, the data must meet the following conditions: n11p 10, n22 1 p 10, n11 11 p 0, n22p 10, and where n11p and n22p are the expected numb n2 1 p2 er of successes and n11 1 p and are the expected number of failures.

3.9.C—Interpret the mean, standard deviation, and probabilities for the sampling distribution for the difference…

Interpret the mean, standard deviation, and probabilities for the sampling distribution for the difference between two sample proportions.

  • The mean, standard deviation, and probabilities for the sampling distribution for the difference between two sample proportions should be interpreted within the context of two specific populations. Inference for Categorical Data: Proportions UNIT 3 102

Topic 3.10

3.10 Constructing a Confidence Interval for the Difference Between Two Population Proportions

Objectives in this topic

3.10.A—Identify an appropriate confidence interval procedure including the parameters for the difference between two…

Identify an appropriate confidence interval procedure including the parameters for the difference between two population proportions.

  • Based on the sample data, a confidence interval can be calculated to estimate the difference between two population proportions. The appropriate confidence interval procedure is a two-sample z-interval for a difference between population proportions.
  • The parameters of a confidence interval for the difference between two population proportions should refer to the difference in the proportions, the response variable, and the populations in context.

3.10.B—Justify the appropriateness of constructing a confidence interval for the difference between two population…

Justify the appropriateness of constructing a confidence interval for the difference between two population proportions by verifying conditions.

  • A two-sample z-interval for a difference between two population proportions requires that three conditions be met:
    • i. The randomization condition—the data should be collected using two independent random samples or a randomized experiment.
    • ii. The 10% condition—when sampling without replacement, the size of each sample should be less than or equal to 10% of the respective population size: n t N11 10≤ % and n ula N22 10≤ % , where N1 is he size of pop tion 1 and N2 is the size of population 2. The sample sizes are represented as n1 and n2. (Note: This condition is unnecessary when the data are from a randomized experiment.)
    • iii. The normality condition—the number of observed successes, n p1 1  and n p2 2  , and number of observed failures, n p1 11  and n p2 21  , for both samples are all at least 10.

3.10.C—Calculate an appropriate confidence interval for the difference between two population proportions

Calculate an appropriate confidence interval for the difference between two population proportions.

  • The point estimate for the difference between two population proportions is pp 12 .
  • For the difference between two population proportions, the interval estimate can be constructed as point estimate (margin of error). The interval estimate for the difference between two population proportions is  ( )     112 12 2 12 1 1p pp p pp z nn .

3.10.D—Calculate the standard error and margin of error for estimating the difference between two population…

Calculate the standard error and margin of error for estimating the difference between two population proportions. Inference for Categorical Data: Proportions UNIT 3 104

  • The standard error (SE ) for the difference between two population proportions is SE pp n pp npp   12 1 1 2 2 1211 .
  • For the difference between two population proportions, the margin of error is the critical value z times the standard error (SE) of the difference between the two proportions, which equals z pp n pp n   1 1 2 2 1211 .

Topic 3.11

3.11 Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Proportions

Objectives in this topic

3.11.A—Interpret a confidence interval in context for the difference between two population proportions

Interpret a confidence interval in context for the difference between two population proportions.

  • Because the confidence interval for the difference between two population proportions is calculated based on samples from two populations, the computed interval may or may not contain the value for the difference between those two population proportions.
  • The interpretation of the confidence level is as follows: In repeated random sampling with the same sample sizes from the same populations, approximately C% of confidence intervals created will capture the difference between the two population proportions, where C represents the numerical value of the confidence level used.
  • When interpreting a C% confidence interval for the difference between two population proportions, we say we are C% confident that the interval ab, contains the parameter for the difference between the populations, where a represents the lower limit and b represents the upper limit. An interpretation of a confidence interval for the difference between two population proportions includes a reference to the parameter with the details about the populations it represents in the context of the study.

3.11.B—Justify a claim based on a confidence interval for the difference between two population proportions

Justify a claim based on a confidence interval for the difference between two population proportions.

  • A confidence interval for the difference between two population proportions provides an interval of values that may provide convincing evidence to support a particular claim about the difference between the two population proportions. For example, if the interval contains 0, then there is insufficient evidence to conclude there is a difference between the two population proportions. If the interval does not contain 0, there is sufficient evidence to conclude there is a difference between the two population proportions. Inference for Categorical Data: Proportions UNIT 3 106

Topic 3.12

3.12 Setting Up a Test for the Difference Between Two Population Proportions

Objectives in this topic

3.12.A—Identify an appropriate testing method for the difference between population proportions including the parameters

Identify an appropriate testing method for the difference between population proportions including the parameters.

  • The appropriate testing method for the difference between two population proportions is a two-sample z-test for the difference between two population proportions.
  • The parameters for a hypothesis test for the difference between two population proportions should reference the population parameters, the response variables, and the populations in context.

3.12.B—Identify the null and alternative hypotheses for the difference between population proportions

Identify the null and alternative hypotheses for the difference between population proportions.

  • For a two-sample z-test for the difference between two population proportions, the null hypothesis indicates no difference. The null hypothesis for the difference between two population proportions can be written as either H: 01 2pp or H: 001 2pp . A onesided alternative hypothesis for the difference bet ween two population proportions can be written as either H: a pp12 or equivalently H: a pp12 0, or H: a pp12 or equivalently H: a pp12 0. A two-sided alternative hypothesis for the difference bet ween two population proportions can be written as either H: a pp12 or equivalently H: a pp12 0 .

3.12.C—Justify the appropriateness of a hypothesis test for the difference between two population proportions by…

Justify the appropriateness of a hypothesis test for the difference between two population proportions by verifying conditions.

  • A two-sample z-test for a difference between two population proportions requires that three conditions be met:
    • i. The randomization condition—the data should be collected using two independent random samples or a randomized experiment.
    • ii. The 10% condition—when sampling without replacement, the size of each sample should be less than or equal to 10% of the respective population size: n s t N11 10% and n2 ulati N210% , where N1 i he size of pop on 1 and N2 is the size of population 2. The sample sizes are represented as n1 and n2. (Note: This condition is unnecessary when the data are from a randomized experiment.)
    • iii. The normality condition—n pc1 , n pc1 1  , n pc2  , and n pc2 1  , must all be at least 10, with p np np nn c   1 1 2 2 12 being the combined (or pooled) proportion assuming that H0 is true (H01: pp 2 or H:01 pp 2 0). Inference for Categorical Data: Proportions UNIT 3 108

Topic 3.13

3.13 Carrying Out a Test for the Difference Between Two Population Proportions

Objectives in this topic

3.13.A—Calculate an appropriate test statistic and p-value for testing a hypothesis for the difference between two…

Calculate an appropriate test statistic and p-value for testing a hypothesis for the difference between two population proportions.

  • The test statistic for the difference between two population proportions is z pp pp nn cc   12 12 0 1 11 , where p np np nn c   1 1 2 2 12 is the proportion of successes for the two groups combined. The z -statistic has a standard normal distribution when the null hypothesis is true.
  • The p-value for a two-sample z-test for the difference between two population proportions can be found from the standard normal distribution using a table or technology.

3.13.B—Interpret the p-value of a hypothesis test for the difference between two population proportions

Interpret the p-value of a hypothesis test for the difference between two population proportions.

  • The p-value is the probability of obtaining a test statistic as extreme or more extreme than the test statistic that was observed (i.e., in the direction of the alternative hypothesis) given that the null hypothesis is true. An interpretation of the p-value of a hypothesis test for a difference between two population proportions should include a statement that the p-value is computed assuming the null hypothesis is true (i.e., by assuming that the true population proportions are equal to each other in context).

3.13.C—Justify a claim about the populations based on the results of a hypothesis test for the difference between two…

Justify a claim about the populations based on the results of a hypothesis test for the difference between two population proportions.

  • A formal decision in a hypothesis test for the difference between two population proportions explicitly compares the p-value to the significance level, α. If the p-value , then reject the null hypothesis, H: 0 pp12 or H: 01 2pp 0. If the p-value , then fail to reject the null hypothesis.
  • The results of a hypothesis test for the difference between two population proportions can serve as the statistical reasoning to support the answer to an investigative question about the two populations that were sampled.
  • A conclusion for the hypothesis test for the difference between two population proportions is stated in context consistent with, and in terms of, the alternative hypothesis using nondefinitive language. The conclusion should contain a reference to the parameters and the populations. Inference for Categorical Data: Proportions UNIT 3 110

Topic 3.14

3.14 Setting Up a Chi-Square Test for Homogeneity or Independence

Objectives in this topic

3.14.A—Describe chi-square distributions

Describe chi-square distributions.

  • The chi-square statistic measures the distance between observed and expected counts relative to expected counts.
  • Chi-square distributions have positive values and are skewed right. Within this family of density curves, the skew becomes less pronounced with increasing degrees of freedom.

3.14.B—Identify an appropriate testing method for comparing distributions in two-way tables of categorical data…

Identify an appropriate testing method for comparing distributions in two-way tables of categorical data including the populations and variables.

  • To determine whether the distributions of a categorical variable for two or more populations are different, the appropriate test is the chi-square test for homogeneity.
  • A chi-square test for homogeneity should reference the categorical variable and the populations in context.
  • To determine whether row and column variables in a two-way table of categorical data might be associated in the single population from which the data were sampled, the appropriate test is the chi-square test for independence.
  • A chi-square test for independence should reference the categorical variables and the population in context.

3.14.C—Identify the null and alternative hypotheses for a chi-square test for homogeneity or independence

Identify the null and alternative hypotheses for a chi-square test for homogeneity or independence.

  • The appropriate null hypothesis for a chisquare test for homogeneity is H0: there is no difference in the distributions of the categorical v ariable across populations or treatments. The appropriate alternative hypothesis for a chisquare test for homogeneity is Ha: there is a difference in the distributions of the categorical v ariable across populations or treatments.
  • The appropriate null hypothesis for a chisquare test for independence is H0 : there is no association between two categorical variables in a given population or the two categorical variables in a given population are independent of each other. The appropriate alternative hypothesis for a chi-square test for independence is Ha: there is an association between two categorical variables in a given population or the two categorical variables in a given population are not independent of each other.

3.14.D—Justify the appropriateness of a chi-square test for independence or homogeneity by verifying conditions

Justify the appropriateness of a chi-square test for independence or homogeneity by verifying conditions.

  • A chi-square test for homogeneity or independence requires that three conditions must be met:
    • i. The randomization condition—the test of independence states that the data should be collected using a random sample. The test for homogeneity states that the data should be collected using independent random samples or a randomized experiment.
    • ii. The 10% condition—when sampling without replacement, check that n 10%N, where N is the size of the population and n is the sample size. (Note: This condition is unnecessary when the data are from a randomized experiment.)
    • iii. The expected counts condition—all expected counts should be greater than 5. Inference for Categorical Data: Proportions UNIT 3 112

Topic 3.15

3.15 Carrying Out a Chi-Square Test for Homogeneity or Independence

Objectives in this topic

3.15.A—Calculate expected counts for two-way tables of categorical data

Calculate expected counts for two-way tables of categorical data.

  • The expected counts (under the null hypothesis) in a particular cell of a two-way table of categorical data can be calculated using the formula row total column totalexpected count total table= .

3.15.B—Calculate the appropriate test statistic and p-value for a chisquare test for homogeneity or independence

Calculate the appropriate test statistic and p-value for a chisquare test for homogeneity or independence.

  • The appropriate test statistic for a chi-square test for homogeneity or independence is the chi-square statistic 2 2 Observed Count Expec t Expected Cou ted Count nχ − , where the sum is taken over all cells of the two-way table. The chi-square statistics have a chi-square distribution with degrees of freedom equal to number of rows −⋅1 number of columns −1 when the null hypothesis is true.
  • The p-value for a chi-square test for independence or homogeneity is found from a chi-square distribution using a table or technology.

3.15.C—Interpret the p-value for the chi-square test for homogeneity or independence

Interpret the p-value for the chi-square test for homogeneity or independence.

  • The p-value is the probability of obtaining a test statistic as extreme or more extreme than the test statistic that was observed (i.e., in the direction of the alternative hypothesis) given that the null hypothesis is true. An interpretation of the p-value for the chi-square test for homogeneity or independence should include a statement that the p-value is computed by assuming the null hypothesis is true in context.

3.15.D—Justify a claim about the population(s) based on the results of a chi-square test for homogeneity or independence

Justify a claim about the population(s) based on the results of a chi-square test for homogeneity or independence.

  • A formal decision in a hypothesis test explicitly compares the p-value to the significance level, . If the p-value , then reject the null hypothesis for the appropriate chi-square test. If the p-value , then fail to reject the null hypothesis.
  • The results of a chi-square test for homogeneity or independence can serve as the statistical reasoning to support the answer to an investigative question about the population that was sampled (independence) or the populations that were sampled (homogeneity).
  • A conclusion for a chi-square test for homogeneity or independence is stated in context consistent with, and in terms of, the alternative hypothesis using non-definitive language. The conclusion should contain a reference to the population(s).
ConceptAP Statistics