AP Statistics 1.11: Random Sampling
Identify random sampling methods and justify whether simple random, stratified, cluster, or systematic selection fits the population.
- Syllabus
- Effective Fall 2025
- Course
- AP Statistics
Identify random sampling methods and justify whether simple random, stratified, cluster, or systematic selection fits the population.
Aphids are tiny insects that feed on plants such as cabbage plants. A farmer wants to reduce the
number of aphids in a cabbage field. A river is located 100 meters south of the cabbage field.
The farmer divides the field into 25 regions of equal size, as shown in the diagram. Each region
has approximately the same number of cabbage plants.

Farmer's House and Cabbage Field
The farmer would like to estimate the proportion of cabbage plants in the field that are affected
by aphids and believes that the extent of aphid damage is greater for the regions in the cabbage
field closer to the river. To obtain the estimate, the farmer is considering three sampling
methods.
- Sampling method I: Select region 3, which is closest to the farmer's house and farthest
from the river. Examine every cabbage plant in the region for aphid damage.
- Sampling method II: Randomly select one row (A, B, C, D, or E). For every region in the
selected row, examine every cabbage plant for aphid damage.
- Sampling method III: Randomly select one region from each of rows A, B, C, D, and E. For
each selected region, examine every cabbage plant for aphid damage.
Using the information provided in the diagram of the cabbage field, describe how to
implement sampling method III, which requires a random selection of one region from each
of rows A, B, C, D, and E.
C The farmer should write the region numbers
from row A, 1 through 5, onto same-size slips of
paper, then put the numbers into a hat, mix well,
and select one of the numbers. The farmer should
repeat this process for the region numbers of
each of the other rows (i.e., row B, 6 through 10;
row C, 11 through 15; row D, 16 through 20; row
E, 21 through 25) and select one number from
each row. This process will result in the
selection of one region from each row. The
farmer will examine every cabbage plant in each
of the selected regions for aphid damage to
determine the proportion of cabbage plants in the
selected regions that are damaged by aphids.
Alternative Solution:
The farmer should use a random number
generator to generate one two-digit integer from
01 to 05, one two-digit integer from 06 to 10, one
two-digit integer from 11 to 15, one two-digit
integer from 16 to 20, and one two-digit integer
from 21 to 25. For each integer selected, the
farmer should select the corresponding
numbered region and examine every cabbage
plant in each of the selected regions for aphid
damage to determine the proportion of cabbage
plants in the selected regions that are damaged
by aphids.
Scoring
Essentially correct (E) if the response satisfies the
following four components:
1. Describes a random selection process that
indicates the groupings used for the strata (by
rows)
2. Describes how to correctly implement a random
selection process for which the selection of
regions within each stratum is equally likely
3. Describes a random selection process for which
the selections across strata are independent
4. Describes a random selection process that results
in the selection of one region from each stratum
Partially correct (P) if the response satisfies only
two or three of the four components required for E.
Incorrect (I) if the response does not meet the
criteria for E or P.
Scoring Notes:
- A response may satisfy component 1 by referring to the five regions in the rows instead of using row letters.
- For responses that use cards or slips of paper:
○ If the number of slips of paper (number of cards) does not equal 5 for each random selection, then
component 1 is not satisfied. Slips of paper (cards) do not need to be specifically identified as
equally sized.
○ If the response does not describe a thorough mixing (shuffling) of the slips of paper (cards), then
component 2 is not satisfied.
- For responses that use a random number generator or table of random digits:
○ If it is not clear that a random selection process allows the selection of regions within each stratum to
be equally likely, then component 2 is not satisfied.
○ If the response does not clearly indicate that a random number is generated from the region numbers
within each of the five rows (e.g., only describes the generation of five random numbers from the
two-digit integers from 01 to 25), then component 4 is not satisfied.
- For responses that use a fair die:
○ If a five-sided fair die is rolled for each of the five rows and the response clearly indicates the region
numbers assigned to the values on the die for each roll, then component 4 is satisfied.
○ If a six-sided die is rolled for each of the five rows and the response clearly indicates which number
is excluded (e.g., "if a 6 is rolled, roll again until a non-6 number is achieved") AND the response
clearly indicates the region numbers (or columns) assigned to the values on the die for each roll, then
component 4 is satisfied.
○ If a 25-sided fair die is rolled for each of the five selections, then component 4 is not satisfied
without sufficient further justification because the selection of one region from each stratum is not
guaranteed.
- If a response describes two separate random selection processes in detail (e.g., describes how to use a
random number generator and slips of paper in a hat), score both descriptions according to the four
components and use the lower score.
- If a response indicates that separate samples were taken from each of the strata, then component 3 is
satisfied.
- A response that selects columns instead of regions within strata cannot satisfy component 3.
Corn tortillas are made at a large facility that produces 100,000 tortillas per day on each of its two production lines. The distribution of the diameters of the tortillas produced on production line A is approximately normal with mean 5.9 inches, and the distribution of the diameters of the tortillas produced on production line B is approximately normal with mean 6.1 inches. The figure below shows the distributions of diameters for the two production lines.

The tortillas produced at the factory are advertised as having a diameter of 6 inches. For the purpose of quality control, a sample of 200 tortillas is selected and the diameters are measured. From the sample of 200 tortillas, the manager of the facility wants to estimate the mean diameter, in inches, of the 200,000 tortillas produced on a given day. Two sampling methods have been proposed.
Method 1: Take a random sample of 200 tortillas from the 200,000 tortillas produced on a given day. Measure the diameter of each selected tortilla.
Method 2: Randomly select one of the two production lines on a given day. Take a random sample of 200 tortillas from the 100,000 tortillas produced by the selected production line. Measure the diameter of each selected tortilla.
Will a sample obtained using Method 2 be representative of the population of all tortillas made that day, with respect to the diameters of the tortillas? Explain why or why not.
Part (a):
No, a sample obtained using Method 2 will not be representative of all tortillas made that day. The sample obtained using Method 2 will only represent the tortillas from one production line, not from the entire population because the distributions of diameters for the two production lines are different.
Which of the two sampling methods, Method 1 or Method 2, will result in less variability in the diameters of the 200 tortillas in the sample on a given day? Explain.
Each day, the distribution of the 200,000 tortillas made that day has mean diameter 6 inches with standard deviation 0.11 inch.
Part (c):
Method 2 would result in less variability in the sample of 200 tortillas on a given day because the sample comes from only one production line. Because the distributions of diameters are not the same for the two production lines, selecting tortillas from both lines as in Method 1 would result in more variable sample data.