Aimathic
Login | English | Deutsch

Free math worksheets

Build your own math worksheets from 30,000+ problems for grades 3 to 12, from fractions to AP Calculus. Every problem comes with step-by-step solutions.

Sampling methods and bias

Click problems to add them to your worksheet.

54152211
Two surveys use the same random sample of residents. Survey 1 asks, “Do you support the costly plan to replace perfectly usable playground equipment?” Survey 2 asks, “Do you support the plan to replace playground equipment?” Which survey is more likely to measure residents' opinions accurately? Explain.

Hints

- Sampling is only one part of collecting trustworthy data. - Compare the emotional or judgmental words in the two questions. - Choose the wording that does not suggest a preferred answer.

Solution

1. Both surveys use the same random sample, so their sampling methods are equally strong. 2. Survey 1 contains loaded words that may push respondents toward opposition. 3. Survey 2 uses more neutral wording and is more likely to measure opinions accurately.

Answer

Survey 2. Its neutral wording is less likely to influence responses.
54921911
Poll A receives \(18{,}000\) responses from a link posted on a library's social-media page. Poll B surveys a simple random sample of \(600\) households from the city address list and follows up with initially selected nonrespondents. Which poll gives the more credible estimate of citywide support for extending library hours, and why does Poll A's larger response count not solve its main problem?

Hints

- Compare how people enter each poll, not only the sample sizes. - A large sample reduces random variation but not systematic selection bias. - Check whether all city households have a realistic chance to be selected.

Solution

1. Poll B is more credible because households are selected randomly from a citywide frame and selected nonrespondents are followed up. 2. Poll A is a voluntary-response sample and may also underrepresent residents who do not follow the library online. 3. A larger number of self-selected responses can reduce neither voluntary-response bias nor undercoverage.

Answer

Poll B is more credible. Poll A's \(18{,}000\) self-selected responses can still be systematically unrepresentative, so its large count does not remove voluntary-response bias or social-media undercoverage.
54922011
A factory measures every cable produced during one shift, but technicians later discover that the device added \(0.7\,\text{cm}\) to every reading. The reported mean is \(81.4\,\text{cm}\). Find the corrected mean and explain why measuring every cable does not prevent this error.

Hints

- Separate sampling error from measurement error. - The same device error affected every observation and therefore the mean. - Undo an additive error by subtracting it.

Solution

1. Measuring every cable makes the data a census of that shift's production, so there is no sampling error for that defined population. 2. The device introduces systematic measurement bias into every value. 3. The corrected mean is \(81.4-0.7=80.7\,\text{cm}\).

Answer

The corrected mean is \(80.7\,\text{cm}\). A census removes sampling variability for the shift, but it does not remove systematic measurement bias.
54922511
A subscription service surveys only customers who have remained subscribed for at least three years and reports that \(91\%\) are satisfied. During those three years, \(38\%\) of the original customers canceled. a) Identify the selection problem. b) Explain why the reported percentage is likely biased upward as an estimate for all original customers. c) Describe a better sampling frame for estimating satisfaction with the service over the full three-year period.

Hints

- Ask which members of the original population are still eligible to answer. - Consider whether leaving the population could be related to the outcome being measured. - A better frame should retain people regardless of what happened after they entered the population.

Solution

1. The survey has survivorship bias because it includes only customers who remained and excludes those who canceled. 2. Dissatisfied customers are more likely to cancel, so the survivors are likely more satisfied than the original customer population. The \(91\%\) therefore tends to overstate satisfaction. 3. Sample randomly from the complete list of customers who began the period, including both current and former customers, and make structured follow-up efforts for both groups.

Answer

a) Survivorship bias. b) Customers who canceled, potentially because of dissatisfaction, are excluded, so satisfaction is likely overestimated. c) Use the original-customer list as the frame and sample both retained and canceled customers.
54922911
A random sample of \(250\) licensed drivers in a state reports a mean of \(7.8\) years since their last defensive-driving course. The study aims to estimate the corresponding mean for all licensed drivers in the state. Identify the population, sample, parameter, and statistic. Then state which quantity is known and which is being estimated.

Hints

- Separate the full group of interest from the people actually measured. - A numerical summary of a sample has a different role from the corresponding unknown population quantity. - Use the wording of the study's goal to identify what is being estimated.

Solution

1. The population is all licensed drivers in the state. 2. The sample is the \(250\) randomly selected licensed drivers. 3. The parameter is the unknown population mean number of years since the last course for all licensed drivers. 4. The statistic is the sample mean, \(7.8\) years. 5. The statistic is known from the data and is used to estimate the parameter.

Answer

Population: all licensed drivers in the state. Sample: the \(250\) selected drivers. Parameter: the population mean time since the last course. Statistic: the sample mean, \(7.8\) years. The statistic is known; the parameter is estimated.
54923111
A household survey asks one available adult to report the weekly work hours of every adult in the home. Later checks show that people report their own hours accurately but tend to underestimate other household members' hours by \(3.5\) hours per week. a) Identify the source of bias. b) In a household where one respondent reports for three other adults, estimate the total downward error in the reported household work hours. c) Describe a redesign that reduces this bias.

Hints

- Identify who supplies the measurement and whose behavior is being measured. - Apply the stated average error only to the people reported by someone else. - A stronger redesign obtains information closer to the original source.

Solution

1. The bias comes from proxy response: one person reports measurements for other people and systematically underestimates them. 2. The total expected downward error for three other adults is \(3\cdot 3.5=10.5\) hours per week. 3. Ask each adult to report personally, or use verified employment records with consent. If proxy reports are unavoidable, validate and adjust them using a representative subsample.

Answer

a) Proxy-response measurement bias. b) An expected underestimate of \(10.5\) hours per week. c) Collect a response from each adult or validate proxy reports against a more accurate source.
54923311
A city app sends a one-time survey notification to all \(42{,}000\) registered users. Of those users, \(3600\) open the survey and \(2100\) complete it. The city has \(180{,}000\) adult residents. Find the completion rate among registered users and explain why the completed surveys are not a simple random sample of city adults.

Hints

- Use completed surveys divided by all registered users. - Compare the app-user frame with the population of all city adults. - A notification sent to everyone in one frame is not random sampling from a larger population.

Solution

1. The completion rate is \(\frac{2100}{42000}=0.05=5\%\). 2. Adults who are not registered app users are undercovered. 3. Among users, opening and completing the survey are voluntary, so inclusion depends on self-selection rather than random selection from all city adults.

Answer

The completion rate is \(5\%\). The completed surveys are not a simple random sample because nonusers have no chance to respond and registered users decide for themselves whether to participate.
54923511
To estimate nitrate concentration in a large lake, Team A samples \(20\) easy-to-reach shoreline points next to roads. Team B divides the lake into equal-area grid cells, randomly selects \(20\) cells, and samples a random point within each selected cell. Road runoff tends to raise nitrate levels nearby. Which design better represents the lake, and what direction of bias is likely for Team A?

Hints

- Compare the geographic coverage of the two sampling frames. - Use the stated effect of road runoff to determine the direction of bias. - Equal-area cells give different parts of the lake comparable opportunities for selection.

Solution

1. Team B better represents the lake because locations throughout the equal-area frame can be selected randomly. 2. Team A overrepresents road-adjacent shoreline locations. 3. Because runoff raises nitrate levels near roads, Team A is likely to overestimate the lakewide mean concentration.

Answer

Team B is the better design. Team A is likely biased upward because its convenient road-adjacent locations tend to have unusually high nitrate levels.
54923911
A student survey asks, “Should the school extend lunch and reduce homework on weekends?” Respondents may answer only “Yes” or “No.” a) Identify the measurement problem. b) Explain why the resulting percentage cannot reveal support for either policy separately. c) Write one valid pair of separate, neutral survey questions. Answers may vary.

Hints

- Count how many distinct decisions the respondent is being asked to make. - Consider a person who agrees with one proposal but not the other. - A repair should allow each policy to receive its own response.

Solution

1. The item is double-barreled because it combines two different policies in one question. 2. A respondent may support one change but oppose the other, so either available response hides that distinction. The combined percentage cannot be assigned to either policy alone. 3. Ask separately: “Do you support or oppose extending the lunch period?” and “Do you support or oppose reducing weekend homework?”

Answer

a) It is a double-barreled question. b) One response combines two opinions that may differ. c) One valid pair is: “Do you support or oppose extending the lunch period?” and “Do you support or oppose reducing weekend homework?”
54924111
A questionnaire asks, “How many whole hours did you work last week?” with these response choices: \(0\)–\(10\), \(10\)–\(20\), \(22\)–\(30\), and more than \(30\). a) Identify two problems with the response categories. b) Explain how each problem can create measurement error. c) Give one corrected set of categories for whole-number hours.

Hints

- Check every boundary value against the choices. - A sound category system should assign each possible response to exactly one place. - Make the revised categories exhaustive as well as nonoverlapping.

Solution

1. The first two categories overlap at \(10\), so that value fits two choices. The categories also omit \(21\), so that value fits no choice. 2. Overlap lets respondents with the same value choose different categories, while a gap forces some respondents to choose an inaccurate category or skip the item. 3. One valid set is \(0\)–\(9\), \(10\)–\(19\), \(20\)–\(29\), and \(30\) or more.

Answer

a) The categories overlap at \(10\) and omit \(21\). b) The overlap permits inconsistent classification, and the gap leaves one valid response with no accurate category. c) One valid correction is \(0\)–\(9\), \(10\)–\(19\), \(20\)–\(29\), and \(30\) or more.
54924511
A public online survey does not require login or a unique invitation code, so one person may submit the form more than once. Organizers receive \(6400\) submissions and describe them as \(6400\) independent respondents. Explain why that description is invalid and give a redesign that addresses both self-selection and duplicate submissions.

Hints

- Distinguish a submitted form from a unique sampled person. - Precision depends on independent information, not merely the number of rows. - A redesign must control both who is invited and how many times each invitee can respond.

Solution

1. The survey has no defined probability-sampling frame, participation is voluntary, and submissions need not represent distinct people. 2. Duplicate or coordinated responses increase the row count without adding equivalent independent information, so the apparent precision is exaggerated. 3. A sound redesign draws a random sample from a defined list, sends one unique code to each selected person, accepts one response per code, and follows up with selected nonrespondents.

Answer

The submissions are not necessarily independent respondents because participation is self-selected and one person can respond repeatedly. Use random invitations from a defined frame with one unique, single-use code per selected person.
54924711
A school creates a simple random sample of \(40\) students from \(800\) by randomly shuffling all \(800\) student IDs and taking the first \(40\). a) What is each student's probability of inclusion? b) Explain why the order produced by the shuffle does not favor students whose IDs are numerically small. c) Why should duplicate IDs be removed before shuffling?

Hints

- For a simple random sample, compare the requested sample size with the population size. - A randomized position is not determined by the label printed on an ID. - Count how many entries in the frame can lead to each person being selected.

Solution

1. Each student has probability \(\frac{40}{800}=0.05\), or \(5\%\), of appearing in the first \(40\) positions. 2. A valid random shuffle makes every ID equally likely to occupy every position, so the numerical size of an ID is irrelevant. 3. Duplicate IDs would give some students multiple positions in the shuffled list and therefore larger inclusion probabilities.

Answer

a) \(0.05\), or \(5\%\). b) A random shuffle makes every ID equally likely to occupy each position. c) Duplicates would give some students extra chances of selection.
55015111
Match each sampling plan with its method: simple random sample, stratified random sample, or cluster sample. 1) Choose \(100\) students at random from one complete school list. 2) Divide students by grade and randomly choose \(25\) from each grade. 3) Randomly choose \(5\) homerooms and survey every student in those homerooms. Then explain why Plan 2 is not a cluster sample.

Hints

- Ask whether individuals or whole groups are selected. - In stratified sampling, every important subgroup contributes selected individuals. - In cluster sampling, all members of selected groups are included.

Solution

1. Plan 1 is a simple random sample because students are selected directly from one complete list. 2. Plan 2 is a stratified random sample because every grade is represented and students are randomly selected within each grade. 3. Plan 3 is a cluster sample because entire randomly selected homerooms are surveyed. 4. Plan 2 samples some students from every grade, while a cluster sample selects whole groups and leaves other groups unselected.

Answer

1) Simple random sample. 2) Stratified random sample. 3) Cluster sample. Plan 2 samples within every grade rather than selecting whole grades as groups.
55015211
Two simple random samples of \(400\) city residents estimate support for a proposal as \(48\%\) and \(52\%\). A voluntary website poll with \(5000\) responses reports \(78\%\) support. a) Explain why the two random-sample estimates can differ. b) Identify the main problem with the website poll. c) Which issue can usually be reduced by taking a larger valid random sample, and which issue cannot?

Hints

- Separate random differences between valid samples from systematic differences in who participates. - Compare how residents enter each study rather than only the sample sizes. - More observations reduce random noise only when the sampling method is appropriate.

Solution

1. Different random samples contain different residents, so their estimates can differ through sampling variability. 2. The website poll has voluntary-response bias because people choose whether to participate and website visitors may not represent all city residents. 3. A larger valid random sample usually reduces sampling variability. It does not remove systematic voluntary-response bias from a self-selected poll.

Answer

a) Random samples vary because they contain different people. b) The website poll is a voluntary-response sample and may be systematically unrepresentative. c) A larger valid random sample reduces sampling variability, but a larger self-selected sample does not remove selection bias.
53758311
In a population, \(40\%\) support a proposal. Of the supporters, \(80\%\) participate in an optional online poll. Of the people who do not support the proposal, only \(30\%\) participate. What percentage of poll participants support the proposal, and why is the poll result biased as an estimate of population support?
Figure for problem 537583

Hints

- Find the population proportion in each opinion group that actually participates. - Add those proportions to find the total participation rate. - Compare the composition of participants with the composition of the full population.

Solution

1. The proportion of the population who support the proposal and participate is \(0.40\cdot 0.80=0.32\). 2. The proportion who do not support the proposal and participate is \(0.60\cdot 0.30=0.18\). 3. The total participation rate is \(0.32+0.18=0.50\). 4. Therefore, the support rate among participants is \(\frac{0.32}{0.50}=0.64\), or \(64\%\). 5. Supporters are more likely to participate, so the respondent group overrepresents supporters. Because participation is optional and related to opinion, this creates voluntary-response bias.

Answer

\(64\%\) of participants support the proposal. The poll overestimates population support because supporters participate at a much higher rate than non-supporters.
54154511
A random sample of students is asked, “Which school club do you belong to?” The only answer choices are art, music, robotics, and athletics. Explain why the results may still be misleading, and give a better set of response choices.

Hints

- Separate problems with selecting people from problems with recording their answers. - Check whether every possible situation has an answer choice. - Consider students with more than one valid response.

Solution

1. The sample selection may be random, but the response choices omit students in other clubs and students in no club. 2. Some students belong to more than one club, yet the wording asks for only one. 3. A better survey allows multiple selections and includes “another club” and “no club.”

Answer

The question has incomplete response choices and does not handle multiple memberships. Allow students to select all that apply, including “another club” and “no club.”
54920711
A school has \(480\) students numbered \(001\) through \(480\). To select a simple random sample of \(8\) students, read the following three-digit groups from left to right: \(431, 072, 518, 214, 214, 009, 480, 305, 601, 177, 042, 399, 487, 263\) a) State the rejection rule. b) List the selected student numbers in order. c) Explain why accepted numbers should not be replaced with nearby valid numbers.

Hints

- Match the possible random-number labels to the actual range of student IDs. - Track accepted values in order and check both range and repetition. - Consider whether changing an invalid result would preserve equal selection chances.

Solution

1. Reject \(000\), every number greater than \(480\), and every duplicate of a number already accepted. 2. Reading left to right gives the accepted numbers \(431, 072, 214, 009, 480, 305, 177, 042\). 3. Replacing an invalid number with a nearby valid number would give some students extra ways to be selected and would destroy the equal-chance property of the simple random sample.

Answer

a) Reject \(000\), values above \(480\), and duplicates. b) \(431, 072, 214, 009, 480, 305, 177, 042\). c) Substitution would make selection probabilities unequal; invalid groups must simply be skipped.
54920911
A warehouse has \(2400\) shipping records in chronological order. An auditor wants a systematic sample of \(30\) records. A random start of \(37\) is chosen. a) Find the sampling interval and list the first five selected record numbers. b) Explain one condition under which this design could be biased. c) Give one change that would protect against the issue in part b).

Hints

- Divide the population size by the desired sample size to determine the spacing. - Look for a possible connection between the fixed spacing and the way the list is ordered. - A protective redesign should break any alignment between selection positions and a repeating pattern.

Solution

1. The sampling interval is \(2400\div 30=80\). Starting at \(37\), the first five records are \(37, 117, 197, 277, 357\). 2. The design can be biased if a relevant pattern in the ordered list has period \(80\), or a factor closely related to \(80\), so the sample repeatedly hits the same point in that pattern. 3. Randomly shuffling the records before systematic selection or using a simple random sample would reduce the risk.

Answer

a) The interval is \(80\); the first five records are \(37, 117, 197, 277, 357\). b) A repeating pattern aligned with the interval could make the sample unrepresentative. c) Randomize the order first or use a different random sampling design.
54921011
The chart shows how adults in a town can be reached by telephone. A survey company samples only numbers in a landline directory. Mobile-only adults are \(52\%\) of the population and support a transit proposal at a rate of \(70\%\); adults with landline access are \(48\%\) of the population and support it at a rate of \(40\%\). Determine the support percentage the landline-frame survey would tend to report, the actual population support percentage, and the direction and size of the expected bias.
Figure for problem 549210

Hints

- Identify which population group can appear in the landline directory. - Weight each group's support rate by its population share. - Subtract the actual population percentage from the frame-based estimate to determine bias.

Solution

1. The landline frame reaches only adults in the \(40\%\)-support group, so the survey tends to report \(40.0\%\). 2. The actual population support is \(0.52\cdot0.70+0.48\cdot0.40=0.556\), or \(55.6\%\). 3. The expected bias is \(40.0\%-55.6\%=-15.6\) percentage points, so the landline-frame survey underestimates support by \(15.6\) percentage points.

Answer

The survey would tend to report \(40.0\%\), while actual population support is \(55.6\%\). The expected bias is \(-15.6\) percentage points, an underestimate of \(15.6\) points.
54921111
A simple random sample of \(1200\) registered voters receives a mail survey. Of the \(720\) respondents, \(432\) support a proposed bond. Researchers believe the support rate among the \(480\) nonrespondents could be anywhere from \(30\%\) to \(50\%\). a) Find the respondent support rate. b) Find the resulting range of possible support rates among all \(1200\) sampled voters. c) Explain why reporting only the respondent rate can be misleading.

Hints

- Separate the observed respondent calculation from assumptions about those who did not answer. - Use the lowest and highest plausible nonrespondent rates to create two endpoint scenarios. - Compare the resulting full-sample range with the single percentage based only on replies.

Solution

1. The respondent support rate is \(\frac{432}{720}=0.60\), or \(60\%\). 2. At the lower nonresponse rate, total support is \(432+0.30\cdot 480=576\), giving \(\frac{576}{1200}=0.48\), or \(48\%\). 3. At the upper nonresponse rate, total support is \(432+0.50\cdot 480=672\), giving \(\frac{672}{1200}=0.56\), or \(56\%\). 4. Nonrespondents may differ systematically from respondents, so the observed \(60\%\) need not represent the full random sample or the population.

Answer

a) \(60\%\). b) From \(48\%\) to \(56\%\). c) The missing responses can shift the estimate substantially because response is not guaranteed to be unrelated to opinion.
54921211
A city is considering a fee to fund park maintenance. Two survey questions are proposed. Question 1: “Do you support the small fee needed to keep our parks clean and safe?” Question 2: “Do you support a \(\$4\) monthly household fee dedicated to park maintenance?” a) Which question is more likely to produce response bias? Identify the wording that creates the problem. b) Write one acceptable neutral revision that gives the amount and purpose without suggesting a preferred response. Answers may vary. c) Explain why randomly selecting respondents would not fix biased wording.

Hints

- Look for adjectives or consequences that make one answer sound more responsible. - A neutral version should state the policy accurately while balancing the available responses. - Distinguish a flaw in who is selected from a flaw in how information is measured.

Solution

1. Question 1 is more likely to create response bias. The words “small,” “needed,” and “clean and safe” frame support as desirable and minimize the cost. 2. A neutral revision is: “Do you support or oppose a \(\$4\) monthly household fee dedicated to park maintenance?” 3. Random selection can make respondents representative of the population, but every selected person would still receive the same leading prompt. The measurement process itself would remain biased.

Answer

a) Question 1; “small,” “needed,” and “clean and safe” lead respondents toward support. b) One acceptable revision is: “Do you support or oppose a \(\$4\) monthly household fee dedicated to park maintenance?” c) Random sampling cannot remove bias built into the question asked of every respondent.
54921411
A district has \(12\) high schools, each with \(20\) English classes. Researchers randomly select \(4\) schools, then randomly select \(3\) English classes within each selected school, and survey every student in those classes. a) Name the sampling design. b) What is the probability that a particular English class is selected? c) State one reason students in the final sample may be more similar to one another than students in a simple random sample of the same size.

Hints

- Follow the selection process one stage at a time. - A class must survive both random stages before it appears in the sample. - Consider what characteristics classmates or schoolmates naturally share.

Solution

1. The design is a multistage cluster sample: schools are selected first, then classes within schools, and all students in selected classes are included. 2. A class is selected only if its school is selected and then the class is selected: \(\frac{4}{12}\cdot\frac{3}{20}=\frac{1}{20}=0.05\). 3. Students in the same school or class share teachers, schedules, and local conditions, so observations are clustered rather than spread independently across the district.

Answer

a) A multistage cluster sample. b) \(0.05\), or \(5\%\). c) Students within the same selected class or school may have correlated experiences, reducing the effective diversity of the sample.
54921511
A high school district has the following enrollment. <table><tr><th>Grade</th><th>Enrollment</th></tr><tr><td>9</td><td>\(960\)</td></tr><tr><td>10</td><td>\(840\)</td></tr><tr><td>11</td><td>\(720\)</td></tr><tr><td>12</td><td>\(680\)</td></tr></table> A proportional stratified sample of \(160\) students will be selected by grade. a) Find the sample size from each grade. b) Explain why proportional allocation allows the four grade sample results to be combined without additional weighting.

Hints

- First compare the desired total sample with the total enrollment. - Apply the same sampling fraction to each grade. - Ask whether each sampled student represents the same number of students in the population.

Solution

1. The total enrollment is \(960+840+720+680=3200\), so the sampling fraction is \(\frac{160}{3200}=0.05\). 2. The grade sample sizes are \(0.05\cdot 960=48\), \(0.05\cdot 840=42\), \(0.05\cdot 720=36\), and \(0.05\cdot 680=34\). 3. Every grade is sampled at the same \(5\%\) rate, so each sampled student represents the same number of population students. The combined sample therefore already matches the grade distribution.

Answer

a) Grade 9: \(48\); Grade 10: \(42\); Grade 11: \(36\); Grade 12: \(34\). b) Each grade is sampled at the same rate, so the sample has the same grade proportions as the population.
54921611
A county is \(80\%\) urban and \(20\%\) rural. To ensure enough rural responses, a survey randomly samples \(100\) urban residents and \(100\) rural residents. In the sample, \(62\%\) of urban residents and \(48\%\) of rural residents favor a recycling program. a) Why is the unweighted average of the two percentages not a countywide estimate? b) Compute the properly weighted countywide estimate. c) Explain why oversampling the rural group is not itself a bias when weights are used correctly.

Hints

- Compare the composition of the sample with the composition of the county. - Let each group contribute according to its population share, not its sample share. - Separate unequal sample sizes from unequal selection quality within each group.

Solution

1. The sample is \(50\%\) urban and \(50\%\) rural, but the county is \(80\%\) urban and \(20\%\) rural, so an equal average gives rural responses too much influence. 2. The weighted estimate is \(0.80\cdot 0.62+0.20\cdot 0.48=0.592\), or \(59.2\%\). 3. Random sampling within each group protects representation inside the strata, and population weights restore the correct county composition. Oversampling improves information about the smaller group without biasing the weighted total.

Answer

a) The sample proportions do not match the county proportions. b) \(59.2\%\). c) Oversampling is valid because random selection occurs within strata and population weights correct the unequal sampling rates.
54921711
A survey contains two policy questions, A and B. A random half of respondents receive A before B, and the other half receive B before A. Support for B is \(58\%\) when B is asked first and \(44\%\) when B is asked second. Identify and quantify the effect being studied, then explain what random assignment of question order does—and does not—justify.

Hints

- Compare the percentage for B under the two order conditions. - Random assignment addresses comparability between treatment conditions. - Random sampling, not random assignment, determines population generalizability.

Solution

1. The design studies question-order response bias. 2. The observed order effect is \(58\%-44\%=14\) percentage points. 3. Random assignment makes the two order groups comparable on average, so the difference can reasonably be attributed to question order. 4. Random assignment of order does not make respondents representative of the population; population generalization also requires an appropriate random sample and adequate treatment of nonresponse.

Answer

The study finds a \(14\)-percentage-point question-order effect on support for B. Random assignment helps isolate the effect of order, but a representative random sample is still needed to generalize either percentage to the population.
54921811
A sampling frame contains \(1050\) records representing \(1000\) households. Exactly \(50\) households appear twice, and every other household appears once. One record is selected at random. a) Find the selection probability for a duplicated household and for a household listed once. b) By what factor is a duplicated household overrepresented? c) Describe the correct repair before drawing a household sample.

Hints

- Count how many records can lead to each household being chosen. - Compare probabilities using a ratio rather than only their numerical sizes. - A valid repair should make one sampling-frame entry correspond to one population unit.

Solution

1. A duplicated household has two records, so its selection probability is \(\frac{2}{1050}\). A household listed once has probability \(\frac{1}{1050}\). 2. The ratio is \(\frac{2/1050}{1/1050}=2\), so a duplicated household is twice as likely to be selected. 3. Match and remove duplicate records so each household appears exactly once, then sample from the cleaned frame.

Answer

a) Duplicated: \(\frac{2}{1050}\); listed once: \(\frac{1}{1050}\). b) A factor of \(2\). c) Deduplicate the frame before random selection.
54922111
A school wants to estimate weekly homework time. It can stratify students either by locker-number range or by course load. Previous records show homework time varies strongly with course load but has no meaningful relationship to locker number. a) Which stratification variable should be used? Explain. b) Why can a poorly chosen stratification variable fail to improve precision even when random samples are taken within every stratum? c) Would stratifying by course load automatically remove nonresponse bias?

Hints

- Choose the grouping that is connected to the measurement of interest. - Think about what makes observations within a stratum similar. - Do not confuse a strong sampling design with a guarantee that everyone selected will respond.

Solution

1. Course load should be used because it creates groups that are internally more similar in the variable being estimated and meaningfully different from one another. 2. Locker-number strata would not reduce within-stratum variability because locker numbers are unrelated to homework time. Randomness protects against selection bias but does not guarantee a precision gain from irrelevant grouping. 3. No. If response rates or homework times differ between respondents and nonrespondents, nonresponse bias can remain within the strata.

Answer

a) Stratify by course load. b) An unrelated grouping does not make values within strata more homogeneous, so it offers little precision benefit. c) No; nonresponse bias must still be addressed.
54922211
A researcher randomly selects one household from a list, then randomly selects one adult within that household for an interview. Compare two adults: Alex lives alone, while Brianna lives in a household with four adults. a) Conditional on their households being selected, find each adult's probability of being chosen. b) Because each household has the same first-stage selection probability, how many times as likely is Alex to be selected overall as Brianna? c) Explain why this procedure is not a self-weighting sample of adults.

Hints

- Follow the selection process after the household has already been chosen. - Count how many adults share the second-stage chance in each household. - A self-weighting design gives every target individual the same overall inclusion probability.

Solution

1. If Alex's household is selected, Alex is chosen with probability \(1\). If Brianna's household is selected, Brianna is chosen with probability \(\frac{1}{4}\). 2. The households have equal first-stage selection probabilities, so the overall inclusion-probability ratio is \(1\div\frac{1}{4}=4\). Alex is four times as likely to be selected overall. 3. Households have equal first-stage probabilities, but adults in larger households split that probability among more people. Adult inclusion probabilities therefore differ unless weights or a different design are used.

Answer

a) Alex: \(1\); Brianna: \(\frac{1}{4}\). b) Alex is \(4\) times as likely to be selected. c) Adults in households of different sizes have unequal inclusion probabilities.
54922411
A study asks students to report yesterday's physical activity in minutes. A validation subsample also provides activity-tracker records. The averages are shown. <table><tr><th>Measure</th><th>Mean minutes</th></tr><tr><td>Self-report</td><td>\(74\)</td></tr><tr><td>Tracker record</td><td>\(58\)</td></tr></table> a) Estimate the average self-reporting error in the validation subsample. b) Identify the likely type and direction of bias if only self-reports are used. c) Explain why subtracting \(16\) minutes from every survey response may still be an imperfect correction.

Hints

- Compare the two measurements made on the same validation group. - Determine whether the less objective measure is systematically high or low. - An average discrepancy does not describe how error behaves for every individual.

Solution

1. The average difference is \(74-58=16\) minutes, with self-reports higher. 2. The survey is subject to measurement or response bias that tends to overestimate activity. 3. The \(16\)-minute value is an average discrepancy. Individual errors may vary by true activity level, memory, or respondent characteristics, so a constant correction may not remove the bias for every group or preserve the distribution accurately.

Answer

a) Self-reports average \(16\) minutes higher. b) Upward measurement or response bias. c) The error need not be constant across people or activity levels, so an average correction may not generalize.
54922811
A school estimates commute time by randomly sampling from students present in homeroom on one morning. Usually \(8\%\) of enrolled students are absent. Present students have a mean commute of \(18\) minutes, while absent students have a mean commute of \(35\) minutes. a) Identify the target population and the actual sampling frame. b) Compute the true enrollment-wide mean commute under these figures. c) Find the expected direction and size of the bias from sampling only present students.

Hints

- Distinguish who the school wants to describe from who is available for selection. - Combine the two group means using their enrollment proportions. - Define bias as the sampling-frame estimate minus the target-population value.

Solution

1. The target population is all enrolled students; the sampling frame contains only students present in homeroom that morning. 2. The enrollment-wide mean is \(0.92\cdot 18+0.08\cdot 35=19.36\) minutes. 3. The attendance-frame estimate tends toward \(18\) minutes, so the expected bias is \(18-19.36=-1.36\) minutes. It underestimates the mean by \(1.36\) minutes.

Answer

a) Target population: all enrolled students; frame: students present that morning. b) \(19.36\) minutes. c) The estimate is biased downward by \(1.36\) minutes.
54923011
Researchers want to study workers in a rare occupation for which no complete membership list exists. They begin with \(12\) known workers and ask each person to refer coworkers, continuing through several waves. a) Name this sampling method. b) Explain why it can be useful here. c) Give two reasons the resulting sample should not be treated as a simple random sample of all workers in the occupation. d) State one conclusion the study could report cautiously without claiming a precise population percentage.

Hints

- Consider how each new participant enters the sample. - Ask whether all population members have a known chance of being reached. - Match the strength of the conclusion to the strength of the selection design.

Solution

1. The method is snowball or chain-referral sampling. 2. It can reach members of a hard-to-list population through existing social connections. 3. Workers with larger networks are more likely to be referred, and referral chains tend to remain within connected groups similar to the initial participants. Inclusion probabilities are unknown and depend on the starting seeds. 4. The study can describe patterns, experiences, or themes among participants while clearly limiting claims about exact prevalence in the full occupation.

Answer

a) Snowball sampling. b) It helps locate members of a population without a usable frame. c) Network size and initial referral chains create unequal, unknown inclusion probabilities. d) Report descriptive findings about participants, not a precise population percentage.
54923211
A nutrition survey asks participants to recall the number of sugary drinks consumed during the previous \(30\) days. In a random validation subsample, recall reports average \(11.2\) drinks while daily diaries average \(15.7\). Estimate the mean recall bias and explain why the validation subsample should be selected randomly.

Hints

- Subtract the diary mean from the recall mean in a consistent order. - Use the sign to describe the direction of bias. - A validation study needs a representative selection process of its own.

Solution

1. The mean recall error is \(11.2-15.7=-4.5\) drinks. 2. Recall reports therefore underestimate monthly consumption by \(4.5\) drinks on average. 3. Random selection helps the validation difference represent all survey participants; volunteers willing to keep diaries could differ in consumption or reporting accuracy.

Answer

The estimated recall bias is \(-4.5\) drinks, an average underestimate of \(4.5\). A random validation subsample is needed so the measured error is not based only on unusually cooperative volunteers.
54923611
A bakery operates continuously for \(18\) hours. Oven temperature tends to drift upward during a shift. An inspector needs a sample of \(30\) loaves to estimate the day's mean loaf mass. Compare these plans: Plan A selects the first \(30\) loaves produced. Plan B divides production into three \(6\)-hour periods and randomly selects \(10\) loaves from each period. a) Name the main weakness of Plan A. b) Explain why Plan B is preferable. c) If production volume is much larger in the final period than in the first two, what adjustment is needed for an overall daily mean?

Hints

- Relate the sampling times to the known process drift. - A stronger plan should cover the full production period rather than one convenient segment. - Equal sample counts do not imply equal population sizes across time periods.

Solution

1. Plan A is a convenience sample restricted to the beginning of production and undercovers later conditions affected by temperature drift. 2. Plan B is a stratified random sample across time, so every part of the production day is represented and random selection occurs within each period. 3. Equal sample sizes from periods with unequal production volumes require weighting each period's sample mean by that period's share of total daily production.

Answer

a) Time-of-day undercoverage from a beginning-of-shift convenience sample. b) Plan B represents all three periods through stratified random sampling. c) Weight period estimates by their shares of total production.
54923711
In a door-to-door survey, interviewers receive a random list of selected addresses. The instructions say, “If no one answers, interview the household next door instead.” a) Explain why the substitution rule destroys the original random-sample design. b) Identify a group that may become overrepresented. c) Give a valid nonresponse protocol that preserves the original selections.

Hints

- Track whether the final respondent was actually chosen by the stated random procedure. - Consider which households are easiest to reach under the substitution rule. - A valid protocol keeps the sampled unit fixed even when contact is difficult.

Solution

1. The neighbor was not selected by the random mechanism, so substitution gives unselected households a chance to enter and removes selected households that are initially unavailable. 2. Households with someone home at the interview time, or households living next to often-absent selected residents, can be overrepresented. 3. Keep the original address in the sample, make callbacks at varied days and times, offer alternate response modes, and record it as nonresponse if contact ultimately fails. Do not replace it with a convenient household.

Answer

a) The replacement household has no selection probability defined by the original random draw. b) Households with someone home during interviewing times may be overrepresented. c) Make repeated, varied contact attempts to the selected address and retain unresolved cases as nonresponse.
54924011
A random sample of employees is randomly divided into two survey-mode groups. In an anonymous online form, \(27\%\) report having ignored a safety rule during the past month. In a face-to-face interview, \(12\%\) report doing so. a) Estimate the observed mode effect. b) Explain why random assignment supports attributing the difference to survey mode. c) Which mode is more likely to reduce social-desirability response bias? Explain. d) Does this study reveal the exact true violation rate?

Hints

- Take the difference between the two reported percentages in a consistent direction. - Identify what the random assignment changes and what it balances. - Separate comparing two measurement methods from knowing the unobserved truth.

Solution

1. The observed mode effect is \(27\%-12\%=15\) percentage points. 2. Random assignment makes the groups comparable on average before the survey, so the mode is the main systematic difference. 3. The anonymous online form is more likely to reduce social-desirability bias because respondents have less pressure to give an approved answer. 4. No. Both modes may still contain reporting error, so the experiment compares modes but does not establish the exact true rate.

Answer

a) \(15\) percentage points. b) Random assignment balances other employee characteristics between modes. c) The anonymous online form is more likely to reduce social-desirability bias. d) No; a less biased mode is not necessarily perfectly accurate.
54924211
A roster of \(1200\) students is arranged in repeating groups of four by class period: Period 1, Period 2, Period 3, Period 4, then the pattern repeats. A researcher selects every fourth name, always beginning with the first name. a) Which class period will appear in the sample? b) Explain why this is not a representative systematic sample. c) Give two valid repairs.

Hints

- Track the selected positions modulo the length of the repeating pattern. - A systematic design is vulnerable when its spacing matches an ordered cycle. - A repair should break the alignment rather than merely increase the sample size.

Solution

1. Every selected position is congruent to \(1\) in the four-name cycle, so only Period 1 students appear. 2. The interval matches the list's period and the start is fixed, producing complete alignment with one subgroup. 3. Two valid repairs are to randomly shuffle the full roster and then use a random-start systematic sample, or to select a simple random sample of \(300\) students from the roster. A stratified random sample from all four periods would also be valid.

Answer

a) Only Period 1. b) The fixed interval and start align with the list's repeating structure. c) Two valid repairs are to shuffle the roster before a random-start systematic sample, or to take a simple random sample of \(300\) students.
54924311
After randomly selecting a household, a survey needs one adult respondent. Compare two rules. Rule A interviews the adult who answers the door. Rule B lists all eligible adults, assigns them numbers, and uses a random-number generator to select one. a) Which rule gives adults within the household equal selection chances? b) Explain the likely bias in Rule A. c) State one practical condition needed for Rule B to work as intended.

Hints

- Focus on the second-stage selection after a household has already been chosen. - Ask whether availability at the contact moment is equally likely for all adults. - A within-household rule fails if interviewers replace the selected person.

Solution

1. Rule B gives eligible adults within the household equal chances because the random-number selection is made from a complete list. 2. Rule A favors adults who are home, available, and willing to answer the door at the contact time. Those traits may be related to employment, age, or the survey outcome. 3. The interviewer must list all eligible adults accurately and contact the randomly selected adult rather than substituting whoever is easiest to reach.

Answer

a) Rule B. b) Rule A overrepresents adults who are home and available when contact occurs. c) All eligible adults must be listed, and the randomly selected adult must actually be contacted without substitution.
54924411
An auditor selects a random dollar from all sales revenue and audits the transaction containing that dollar. Compare a \(\$10\) transaction with a \(\$50\) transaction. a) How many times as likely is the \(\$50\) transaction to be selected? b) Is this a simple random sample of transactions? c) For what audit goal might this unequal-probability design be useful?

Hints

- Count how many selectable revenue units belong to each transaction. - A simple random sample requires equal chances for the target units named in the question. - Consider whether the study cares equally about transactions or about dollars.

Solution

1. The \(\$50\) transaction contains five times as many revenue dollars as the \(\$10\) transaction, so it is \(5\) times as likely to contain the selected dollar. 2. No. Transactions do not have equal selection probabilities; probability is proportional to transaction amount. 3. The design can be useful for estimating total monetary misstatement because larger transactions have more potential financial impact and receive appropriately greater selection chance.

Answer

a) \(5\) times as likely. b) No; selection probability is proportional to transaction value. c) It can be useful for auditing monetary totals or detecting financially important errors.
54924811
A city draws a random sample from a property-tax mailing list to survey current residents. The list still includes some vacant properties and people who moved away. a) Which frame error is definitely present? What additional fact would establish undercoverage as well? b) Explain how the outdated records affect the contact rate. c) Under what condition could the outdated records also bias the survey estimate rather than merely reduce efficiency? d) Give one frame-improvement step.

Hints

- Compare every type of record in the list with the definition of the target population. - A frame error can affect both who is reachable and how many attempts are wasted. - Bias requires a systematic relationship, not merely a smaller number of completed interviews.

Solution

1. Overcoverage is definitely present because the frame contains records that are not current residents. Undercoverage would also be present if some current residents are absent from the tax list and therefore cannot be selected. 2. Vacant and outdated addresses create failed contacts, wasting sample slots and lowering the usable response rate. 3. Bias can occur if the missing or unreachable current residents differ systematically on the survey outcome from those successfully contacted, especially if replacements are made conveniently. 4. Update the frame with current occupancy records, remove confirmed vacant properties, and avoid substituting unselected addresses.

Answer

a) Overcoverage is definite. Undercoverage would also exist if some current residents are missing from the list. b) Invalid records increase failed contacts and reduce efficiency. c) Bias occurs when frame errors are related to the outcome or when convenient replacements are used. d) Clean and update the address frame before selection.
54924911
A random sample receives a survey in three contact waves. Support for a proposal is \(64\%\) among first-wave respondents, \(55\%\) among second-wave respondents, and \(48\%\) among people who respond only after the third reminder. a) What pattern do the response waves show? b) Why does the pattern raise concern about nonresponse bias? c) Does it prove that the remaining nonrespondents support the proposal at less than \(48\%\)? Explain. d) Name one stronger follow-up strategy.

Hints

- Compare the outcome across groups defined by how difficult they were to contact. - A systematic trend can signal bias without determining the unseen values exactly. - A stronger strategy gathers information specifically about those still missing.

Solution

1. Support decreases among later and harder-to-reach respondents: \(64\%\), then \(55\%\), then \(48\%\). 2. Response timing is associated with the outcome, suggesting that people who never respond may differ from early respondents. Using only early responses would likely overestimate support. 3. No. Late respondents provide evidence of a trend but do not reveal the unobserved opinions of remaining nonrespondents. 4. Use targeted callbacks, alternate contact modes, a short nonresponse survey, or administrative comparisons between respondents and nonrespondents.

Answer

a) Support declines with each later response wave. b) Response difficulty appears related to opinion, so early respondents may not represent the full sample. c) No; the remaining opinions are still unobserved. d) Use targeted follow-up or a nonresponse study.
54925011
A health study begins with a random sample of \(1000\) adults. Five years later, only \(620\) remain in the study. Participants who leave had a baseline mean stress score of \(7.1\), while those who remain had a baseline mean of \(5.4\). a) Name the sampling problem that develops over time. b) Explain the likely direction of bias if the final report estimates five-year stress using only those who remain. c) Why does the original random sample not automatically protect the final analysis? d) Give one design or analysis response.

Hints

- Track how the observed group changes between the beginning and the end of the study. - Use the baseline difference to infer which direction the remaining sample shifts. - Random selection at one time does not prevent later selective loss.

Solution

1. The study develops attrition bias, a form of nonresponse bias in a longitudinal sample. 2. Because people who leave began with higher stress, the remaining group is lower-stress than the original sample. A final complete-case estimate is likely biased downward if the relationship persists. 3. Random selection protected the baseline sample, but differential loss after selection changes the composition of the observed group. 4. Maintain intensive follow-up, compare leavers with remainers using baseline data, apply justified attrition weights, or report sensitivity analyses.

Answer

a) Attrition bias. b) The final complete-case estimate is likely biased downward. c) Differential dropout changes the composition after random sampling. d) Use follow-up, justified weighting, or sensitivity analysis.
54925111
A randomly sampled wage survey has a \(92\%\) overall response rate, but \(25\%\) of respondents skip the income question. Records show that income-item nonresponse is especially common among top-level managers. a) Distinguish unit nonresponse from item nonresponse in this study. b) Explain the likely direction of bias in the mean income computed only from completed income items. c) Why is the \(92\%\) overall response rate not enough to dismiss the problem? d) State one useful remedy.

Hints

- Identify whether missingness occurs for the entire survey or only one question. - Use the stated characteristics of item nonrespondents to infer which values are missing more often. - Evaluate response quality for the actual variable being analyzed, not only the form as a whole.

Solution

1. Unit nonresponse occurs when a selected person does not complete the survey at all; item nonresponse occurs when a respondent completes the survey but leaves the income item blank. 2. Because high-income managers disproportionately skip the item, the complete-item mean is likely biased downward. 3. The high overall response rate concerns survey participation, not completeness of the specific variable used for the income estimate. 4. Use confidential collection methods, targeted follow-up for the missing item, or a justified imputation or weighting method based on job level and other observed information.

Answer

a) Unit nonresponse is failure to participate; item nonresponse is a missing answer within an otherwise returned survey. b) The complete-case mean is likely biased downward. c) Overall response does not measure completeness of the income variable. d) Improve confidentiality and follow up, or use a justified adjustment based on observed characteristics.
54925311
A city asks people entering a weekend fitness expo whether they support building more public exercise stations. In the city, \(35\%\) of adults exercise regularly and \(65\%\) do not. At the expo, \(80\%\) exercise regularly. Support is \(70\%\) among regular exercisers and \(40\%\) among other adults. a) Compute the citywide support rate under these figures. b) Compute the support rate expected from the expo crowd. c) Find the direction and size of the venue-selection bias. d) Describe a better sampling method.

Hints

- Weight support rates first by the city composition and then by the venue composition. - Compare the two weighted results in a consistent direction. - A better frame should not be tied to interest in the survey topic.

Solution

1. The citywide rate is \(0.35\cdot 0.70+0.65\cdot 0.40=0.505\), or \(50.5\%\). 2. The expo-crowd rate is \(0.80\cdot 0.70+0.20\cdot 0.40=0.640\), or \(64.0\%\). 3. The bias is \(64.0\%-50.5\%=13.5\) percentage points, so the venue sample is biased upward. 4. Select adults randomly from a citywide address or resident frame and follow up with those selected.

Answer

a) \(50.5\%\). b) \(64.0\%\). c) Upward by \(13.5\) percentage points. d) Use a citywide random sample rather than a fitness-event convenience sample.
54920811
A district wants to estimate the mean one-way bus ride for all high school students. Travel times differ greatly among the district's four high schools. Plan A randomly selects \(60\) students from each high school. Plan B randomly selects one high school and surveys every student at that school. a) Name each sampling method. b) Which plan is more likely to represent the districtwide pattern of travel times? Explain. c) What weighting issue must be handled if the four schools have different enrollments?

Hints

- Identify whether the population is divided and sampled within every group or whether an entire group is selected. - Use the stated source of variation to judge which design controls it better. - Compare equal sample sizes with unequal population sizes before combining group results.

Solution

1. Plan A is a stratified random sample with high school as the stratum. Plan B is a one-stage cluster sample with a high school as the cluster. 2. Plan A is more likely to represent the districtwide pattern because every school is included and the important between-school differences are controlled. Plan B can be far from the district mean if the selected school is unusual. 3. Because equal numbers are taken from schools of unequal size, the four sample means cannot simply be averaged equally. Each school’s result must be weighted by its share of district enrollment.

Answer

a) Plan A is stratified sampling; Plan B is cluster sampling. b) Plan A is preferable because it represents every school despite large school-to-school differences. c) Weight each school's estimate according to its proportion of total district enrollment.
54921311
To estimate streaming-service use, interviewers must collect exactly \(50\) responses from each of four age groups. Within each age group, they may approach any people they find at a downtown shopping center until the quota is filled. a) Explain why this is not a stratified random sample even though every age group has a quota. b) Name two sources of bias that could remain. c) Redesign the study as a stratified random sample.

Hints

- Ask how each individual within an age group actually gets chosen. - Separate representation of group labels from representation of people inside those groups. - A valid redesign needs both defined strata and a random selection mechanism within each one.

Solution

1. A stratified random sample requires random selection within every stratum. Here, interviewers choose convenient people at one location, so individuals within an age group do not have known or equal chances of selection. 2. The shopping-center location creates undercoverage of people who do not visit it, and interviewer choice can create selection bias within each quota. 3. Divide a complete population list into the four age strata, then use a random method to select the required number from each stratum. If estimating an overall rate, combine strata using population-proportion weights rather than automatically weighting each quota equally.

Answer

a) Quotas alone do not make sampling random; selection within each age group is by convenience. b) Possible biases include location undercoverage and interviewer selection bias. c) Randomly select people from a complete list within each age stratum, then weight strata appropriately for an overall estimate.
54922311
To reduce response bias on a sensitive yes-or-no question, each respondent privately spins the fair spinner shown. On A, the respondent answers the sensitive question truthfully. On B, the respondent answers “Yes” regardless of the truth. The researcher sees only the final answer, not the spinner result. In a large sample, \(68\%\) of responses are “Yes.” Estimate the actual percentage for whom the truthful answer is “Yes,” and explain how the procedure protects privacy.
Figure for problem 549223

Hints

- Separate the responses produced by each spinner sector. - Represent the unknown truthful proportion with a variable and combine the two equally likely routes to “Yes.” - Distinguish what can be learned about the group from what can be learned about one person.

Solution

1. Let \(p\) be the actual truthful “Yes” proportion. 2. Half the responses come from A and are “Yes” with probability \(p\); half come from B and are always “Yes.” Thus \(0.68=0.5p+0.5\). 3. Solving gives \(0.5p=0.18\) and \(p=0.36\), or \(36\%\). 4. An individual “Yes” could be a truthful answer or a forced answer, so the researcher cannot infer that person's status even though the group rate can be estimated.

Answer

The estimated truthful “Yes” rate is \(36\%\). Privacy is protected because any individual “Yes” may have been forced by the spinner rather than revealing the person's true answer.
54922611
A transit agency wants to survey riders at a downtown station. Ridership patterns differ by weekday versus weekend and by morning, afternoon, and evening. Design one valid time-location sample using \(12\) survey sessions. Your plan must represent all six daypart combinations and use randomization. State how many sessions each combination receives and how riders are selected during a session. More than one valid design is possible.

Hints

- Cross the two stated sources of variation to define the groups that need representation. - Distribute the available sessions so no combination is omitted. - Specify a reproducible rule for choosing people after a session begins.

Solution

1. Form six strata: weekday morning, weekday afternoon, weekday evening, weekend morning, weekend afternoon, and weekend evening. 2. Assign two sessions to each stratum, for \(6\cdot2=12\) sessions. Randomly select the dates and start times within each stratum. 3. During each session, choose a random start among the first few entering riders and then approach every fixed-numbered rider, such as every fifth rider, with a rule for handling refusals rather than substituting convenient people. 4. Record the selection rate and relevant ridership volume for each stratum. For an overall rider estimate, weight the stratum results by their ridership shares or by the inverse of the resulting selection probabilities. 5. This design represents all important time periods and uses randomization within both session selection and rider selection.

Answer

One valid plan is to use six daypart strata and assign \(2\) randomly scheduled sessions to each. Within each session, use a random start and a fixed systematic interval to select riders, recording nonresponse instead of replacing refusals conveniently. For an overall rider estimate, weight the strata using ridership shares or known selection probabilities.
54922711
A survey sample contains \(300\) adults ages 18–39 and \(200\) adults ages 40 and older. In the younger group, \(150\) respond and \(90\) of those respondents support a proposal. In the older group, \(160\) respond and \(80\) support it. a) Find the unweighted support rate among respondents. b) Weight each respondent by the inverse of the response rate in that age group. Find the weighted support estimate for the original sample. c) State the key assumption behind this adjustment.

Hints

- Compute the observed result before making any adjustment. - Find how many sampled people each respondent represents within the person's group. - A weighting correction works only if the grouping captures the important response differences.

Solution

1. There are \(90+80=170\) supporters among \(150+160=310\) respondents, so the unweighted rate is \(\frac{170}{310}\approx 54.8\%\). 2. The younger response rate is \(\frac{150}{300}=0.5\), so its weight is \(2\). The older response rate is \(\frac{160}{200}=0.8\), so its weight is \(1.25\). 3. Weighted support is \(90\cdot 2+80\cdot 1.25=280\), and the weighted total is \(150\cdot 2+160\cdot 1.25=500\). The estimate is \(\frac{280}{500}=56.0\%\). 4. The adjustment assumes that, within each age group, respondents are representative of nonrespondents regarding support.

Answer

a) Approximately \(54.8\%\). b) \(56.0\%\). c) Within each age group, response status must be unrelated to support after accounting for the group.
54923411
An election office plans an exit poll in a county with urban, suburban, and rural precincts. Voting patterns differ across these areas, and turnout varies by time of day. Describe one valid sampling plan that uses both precinct type and voting time. Explain how your plan reduces two different sources of bias. More than one valid plan is possible.

Hints

- Identify each stated source of variation and make sure the design represents it. - Randomization is needed both when choosing locations and when choosing people at a location. - A complete plan should state how refusals are handled without convenient substitution.

Solution

1. Divide precincts into urban, suburban, and rural strata, then randomly select precincts within each stratum with probabilities proportional to expected turnout. If unequal selection probabilities are used, retain them for weighting. 2. At each selected precinct, randomly schedule interviewing blocks across morning, afternoon, and evening so each voting period has a known chance of selection. 3. Within each block, choose a random start and approach every fixed-numbered exiting voter, recording refusals instead of replacing them conveniently. 4. Stratifying precincts prevents geographic underrepresentation, and sampling across the day prevents undercoverage of voters whose arrival times differ. Systematic selection within a block limits interviewer choice. Weighting by the known selection probabilities preserves a valid countywide estimate when probabilities differ.

Answer

One valid plan is to randomly select precincts within urban, suburban, and rural strata using known probabilities, randomly cover morning, afternoon, and evening blocks, and use a random-start systematic rule for exiting voters. This addresses geographic and time-of-day undercoverage while reducing interviewer selection bias; use the known selection probabilities as weights when needed.
54923811
Panel a) shows the city's adult population, and panel b) shows the survey respondents. The charts compare the age distribution of a city's adult population with the age distribution of respondents to a survey. Support for a proposal is \(70\%\) among ages 18–34, \(55\%\) among ages 35–54, and \(40\%\) among ages 55 and older. a) Which age group is most overrepresented among respondents? b) Compute the unweighted support estimate using the respondent distribution. c) Compute the age-adjusted estimate using the population distribution. d) State the direction and size of the compositional bias.
Figure for problem 549238

Hints

- Compare the same age category across the two panels. - Use respondent shares for the unweighted calculation and population shares for the adjusted calculation. - The group with the lowest support rate has extra influence in the respondent mix.

Solution

1. Adults ages 55 and older are most overrepresented: \(50\%\) of respondents versus \(30\%\) of the population. 2. The unweighted estimate is \(0.15\cdot 0.70+0.35\cdot 0.55+0.50\cdot 0.40=0.4975\), or \(49.75\%\). 3. The age-adjusted estimate is \(0.30\cdot 0.70+0.40\cdot 0.55+0.30\cdot 0.40=0.5500\), or \(55.00\%\). 4. The compositional bias is \(49.75\%-55.00\%=-5.25\) percentage points, so the unweighted result is biased downward.

Answer

a) Ages 55 and older. b) \(49.75\%\). c) \(55.00\%\). d) Downward by \(5.25\) percentage points.
54924611
A library randomly selects \(8\) shelves and records the publication year of every book on those shelves. Books on a shelf tend to belong to the same subject and were often acquired at similar times. a) Name the sampling design. b) Explain why the estimate of the library's mean publication year can have greater sampling variability than a simple random sample with the same number of books. c) Is the clustering necessarily a source of bias if shelves are selected randomly from a complete shelf list?

Hints

- Identify the naturally occurring groups that are selected as wholes. - Compare information from many similar observations with information spread across the population. - Keep bias and sampling variability as separate properties of a design.

Solution

1. This is a cluster sample with shelves as clusters. 2. Books within a shelf are similar in subject and acquisition period, so many observations from one shelf provide less independent information than the same number of books spread across the library. Different selected shelf sets can therefore produce more variable means. 3. No. Random selection from a complete shelf frame can avoid systematic selection bias, although unequal shelf sizes may require design attention. Clustering mainly affects precision here.

Answer

a) Cluster sampling. b) Books within a shelf are correlated, so the sample contains less diverse information than an equally large SRS. c) No; random cluster selection can be unbiased while still being less precise.
54925211
A county wants to estimate the percentage of households that used a public recreation facility during the past year. The county includes one large city, three small towns, and a rural area. It can contact at most \(500\) households. Design one valid probability sample and contact plan. Your answer must address geographic representation, unequal population sizes, nonresponse, and how to combine the results. More than one valid plan is possible.

Hints

- Turn each important geographic area into a group that cannot be omitted. - Random selection must occur within every represented group. - The final county estimate should reflect population shares rather than raw sample shares.

Solution

1. Divide the county into five geographic strata: the city, each of the three towns, and the rural area. 2. Randomly select addresses within every stratum. Allocate most of the \(500\) selections proportionally to household counts, while allowing a modest oversample of small strata if reliable separate estimates are needed. 3. Use several contact modes and repeated attempts at varied times for the originally selected addresses; do not replace nonrespondents with convenient neighbors. 4. Compute each stratum estimate and combine them using household-population weights. If oversampling or response rates differ, adjust weights accordingly and report remaining nonresponse limitations.

Answer

One valid plan is to use a stratified random address sample covering the city, each town, and the rural area; follow up repeatedly with the selected households; and combine stratum results using household-population weights, adjusted for any deliberate oversampling and justified nonresponse corrections.

All problems may be used, copied and printed free of charge for school and tutoring, including paid tutoring. Commercial adaptations as well as publication or redistribution on the internet are not permitted.