Everything below prints as one AP AP Statistics practice paper set: papers A & B with their answer keys, plus the full-length study package exam. Use the Download PDF / Print button (or Cmd/Ctrl+P) to save it.
Paper A
AP Statistics — Practice Paper A
Original unofficial practice questions · paper A · answer key on the last page
Total time: see section headers · No guessing penalty
| Section | Questions | Format |
|---|---|---|
| Section I: Multiple Choice | ||
| Section II: Free Response |
Section I — Multiple Choice
The mean of 2, 4, 6, 8 is
A. 5B. 4C. 6D. 20A z-score measures distance from the mean in units of
A. standard deviationB. varianceC. medianD. rangeIf P(A) = 0.3 and P(B) = 0.4 and A,B are independent, P(A ∩ B) =
A. 0.12B. 0.7C. 0.3D. 0.1The Central Limit Theorem says the sampling distribution of the sample mean is approximately normal when
A. n is largeB. n is smallC. the population is skewedD. the sample is biasedA 95% confidence interval means that over many samples, the interval
A. captures the parameter 95% of the timeB. contains 95% of dataC. is 95% accurateD. rejects the null 95% of the timeCorrelation r measures
A. linear associationB. causationC. slopeD. spreadThe p-value is the probability of
A. obtaining results as extreme as observed given H₀ trueB. H₀ being trueC. H₁ being trueD. Type II errorIf data are right-skewed, the median is typically
A. less than the meanB. greater than the meanC. equal to the meanD. undefinedThe standard deviation of a data set is always
A. nonnegativeB. negativeC. a percentageD. equal to the rangeIn a regression, the residual is
A. actual - predictedB. predicted - meanC. actual - meanD. slope-interceptTwo events with no outcomes in common are
A. disjointB. independentC. conditionalD. complimentaryA Type I error is
A. rejecting a true H₀B. accepting a false H₀C. increasing nD. decreasing αThe interquartile range is
A. Q3 - Q1B. max - minC. σ²D. mean - modeIncreasing sample size generally
A. narrows the confidence intervalB. widens the confidence intervalC. increases biasD. changes the parameterThe exponential distribution models
A. waiting timesB. countsC. scoresD. categoriesSection II — Free Response
Describe the sampling distribution of x̄ for a population with μ=50, σ=10, and n=40. Justify assumptions and give the standard error.
5 points · rubric: Shape 2 pts, mean 1 pt, standard error 2 pts.
A study found r = -0.85 between screen time and reading scores. Interpret r, and explain why this does not imply causation.
4 points · rubric: Interpretation 2 pts, causation caveat 2 pts.
You collect {3, 5, 8, 10, 9}. Find the mean, median, and standard deviation, and state which is more resistant to outliers.
6 points · rubric: Mean 1 pt, median 1 pt, SD 2 pts, resistance 2 pts.
Test the claim that the mean exam score is 70 using a sample mean of 74, n=25, s=8, α=0.05. State hypotheses, compute the test statistic, and conclude.
7 points · rubric: Hypotheses 2 pts, test statistic 2 pts, p-value/critical 2 pts, conclusion 1 pt.
Answer Key
1. 5 — (2+4+6+8)/4 = 5.
2. standard deviation — z = (x-μ)/σ.
3. 0.12 — Product of independent probabilities.
4. n is large — n≲30 generally.
5. captures the parameter 95% of the time — Confidence level interpretation.
6. linear association — Direction/strength of linear association.
7. obtaining results as extreme as observed given H₀ true — Definition of p-value.
8. less than the mean — Right tail pulls mean up.
9. nonnegative — SD ≥ 0.
10. actual - predicted — Residual = y - ŵ.
11. disjoint — Mutually exclusive.
12. rejecting a true H₀ — False positive.
13. Q3 - Q1 — Spread of middle 50%.
14. narrows the confidence interval — Lower standard error.
15. waiting times — Time between events.
Free response — rubric notes
1. Shape 2 pts, mean 1 pt, standard error 2 pts. · model: Approx normal (CLT, n=40); mean = 50; standard error = 10/√40 ≈ 1.58.
2. Interpretation 2 pts, causation caveat 2 pts. · model: Strong negative linear association; confounding variables (socioeconomic status) and directionness abound, so no causation.
3. Mean 1 pt, median 1 pt, SD 2 pts, resistance 2 pts. · model: Mean 7, median 8, SD ≈ 2.83 (population); median resists outliers.
4. Hypotheses 2 pts, test statistic 2 pts, p-value/critical 2 pts, conclusion 1 pt. · model: t = (74-70)/(8/5) = 2.5, df=24, p ≈ 0.02 < 0.05; reject H₀, evidence mean differs from 70.
Paper B
AP Statistics — Practice Paper B
Original unofficial practice questions · paper B · answer key on the last page
Total time: see section headers · No guessing penalty
| Section | Questions | Format |
|---|---|---|
| Section I: Multiple Choice | ||
| Section II: Free Response |
Section I — Multiple Choice
The mean of 2, 4, 6, 8 is
A. 20B. 4C. 5D. 6A z-score measures distance from the mean in units of
A. varianceB. medianC. rangeD. standard deviationIf P(A) = 0.3 and P(B) = 0.4 and A,B are independent, P(A ∩ B) =
A. 0.7B. 0.12C. 0.1D. 0.3The Central Limit Theorem says the sampling distribution of the sample mean is approximately normal when
A. the sample is biasedB. the population is skewedC. n is largeD. n is smallA 95% confidence interval means that over many samples, the interval
A. is 95% accurateB. captures the parameter 95% of the timeC. contains 95% of dataD. rejects the null 95% of the timeCorrelation r measures
A. slopeB. spreadC. linear associationD. causationThe p-value is the probability of
A. H₀ being trueB. Type II errorC. H₁ being trueD. obtaining results as extreme as observed given H₀ trueIf data are right-skewed, the median is typically
A. undefinedB. less than the meanC. equal to the meanD. greater than the meanThe standard deviation of a data set is always
A. negativeB. a percentageC. nonnegativeD. equal to the rangeIn a regression, the residual is
A. predicted - meanB. slope-interceptC. actual - meanD. actual - predictedTwo events with no outcomes in common are
A. complimentaryB. disjointC. independentD. conditionalA Type I error is
A. decreasing αB. accepting a false H₀C. increasing nD. rejecting a true H₀The interquartile range is
A. mean - modeB. σ²C. Q3 - Q1D. max - minIncreasing sample size generally
A. widens the confidence intervalB. changes the parameterC. narrows the confidence intervalD. increases biasThe exponential distribution models
A. countsB. categoriesC. scoresD. waiting timesSection II — Free Response
Describe the sampling distribution of x̄ for a population with μ=50, σ=10, and n=40. Justify assumptions and give the standard error.
5 points · rubric: Shape 2 pts, mean 1 pt, standard error 2 pts.
A study found r = -0.85 between screen time and reading scores. Interpret r, and explain why this does not imply causation.
4 points · rubric: Interpretation 2 pts, causation caveat 2 pts.
You collect {3, 5, 8, 10, 9}. Find the mean, median, and standard deviation, and state which is more resistant to outliers.
6 points · rubric: Mean 1 pt, median 1 pt, SD 2 pts, resistance 2 pts.
Test the claim that the mean exam score is 70 using a sample mean of 74, n=25, s=8, α=0.05. State hypotheses, compute the test statistic, and conclude.
7 points · rubric: Hypotheses 2 pts, test statistic 2 pts, p-value/critical 2 pts, conclusion 1 pt.
Answer Key
1. 5 — (2+4+6+8)/4 = 5.
2. standard deviation — z = (x-μ)/σ.
3. 0.12 — Product of independent probabilities.
4. n is large — n≲30 generally.
5. captures the parameter 95% of the time — Confidence level interpretation.
6. linear association — Direction/strength of linear association.
7. obtaining results as extreme as observed given H₀ true — Definition of p-value.
8. less than the mean — Right tail pulls mean up.
9. nonnegative — SD ≥ 0.
10. actual - predicted — Residual = y - ŵ.
11. disjoint — Mutually exclusive.
12. rejecting a true H₀ — False positive.
13. Q3 - Q1 — Spread of middle 50%.
14. narrows the confidence interval — Lower standard error.
15. waiting times — Time between events.
Free response — rubric notes
1. Shape 2 pts, mean 1 pt, standard error 2 pts. · model: Approx normal (CLT, n=40); mean = 50; standard error = 10/√40 ≈ 1.58.
2. Interpretation 2 pts, causation caveat 2 pts. · model: Strong negative linear association; confounding variables (socioeconomic status) and directionness abound, so no causation.
3. Mean 1 pt, median 1 pt, SD 2 pts, resistance 2 pts. · model: Mean 7, median 8, SD ≈ 2.83 (population); median resists outliers.
4. Hypotheses 2 pts, test statistic 2 pts, p-value/critical 2 pts, conclusion 1 pt. · model: t = (74-70)/(8/5) = 2.5, df=24, p ≈ 0.02 < 0.05; reject H₀, evidence mean differs from 70.
Full-length study package exam
AP Statistics — Full Practice Exam
Section I: Multiple-Choice Questions (40 questions, 90 minutes)
1. A dataset has mean 50 and standard deviation 8. After adding 5 to every value, the new mean and standard deviation are:
(A) Mean = 55, SD = 13
(B) Mean = 55, SD = 8
(C) Mean = 50, SD = 8
(D) Mean = 55, SD = 40
(E) Mean = 50, SD = 13
2. Which of the following sampling methods is most likely to produce a representative sample of a city's population?
(A) Surveying people at a shopping mall
(B) Using a random digit dialing system to call phone numbers
(C) Posting a survey on social media
(D) Asking customers at a local restaurant
(E) Surveying every 10th person on a list of registered voters
3. A scatterplot shows a moderate positive linear association. Which value of the correlation coefficient r is most consistent with this description?
(A) -0.82
(B) -0.30
(C) 0.05
(D) 0.55
(E) 0.98
4. Events A and B are mutually exclusive. P(A) = 0.35 and P(B) = 0.45. What is P(A or B)?
(A) 0.80
(B) 0.1575
(C) 0.10
(D) 0.90
(E) 0.35
5. A 95% confidence interval for a population mean is (23.4, 28.6). Which of the following is the best interpretation?
(A) 95% of the data are between 23.4 and 28.6.
(B) There is a 95% probability that mu is between 23.4 and 28.6.
(C) We are 95% confident that mu is between 23.4 and 28.6.
(D) 95% of all possible sample means fall between 23.4 and 28.6.
(E) If we repeated this procedure, 95% of the data would fall in this interval.
6. A researcher computes a 90% confidence interval for a proportion and a 95% confidence interval for the same proportion from the same data. How do the margins of error compare?
(A) The 90% interval has a larger margin of error.
(B) The 95% interval has a larger margin of error.
(C) They have the same margin of error.
(D) It depends on the sample proportion.
(E) It depends on the sample size.
7. In an experiment, the purpose of a control group is to:
(A) Increase the sample size.
(B) Provide a baseline for comparison.
(C) Ensure double-blinding.
(D) Reduce nonresponse bias.
(E) Eliminate the need for random assignment.
8. A histogram of a quantitative variable shows a roughly symmetric, bell-shaped distribution. Which of the following is the best measure of center?
(A) Mode
(B) Median
(C) Mean
(D) Both mean and median
(E) Range
9. A z-score of 2.3 means that a value is:
(A) 2.3 points above the mean
(B) 2.3 standard deviations above the mean
(C) In the 23rd percentile
(D) 2.3 times larger than the mean
(E) 2.3% above the mean
10. Which of the following is a condition for constructing a one-sample z-interval for a proportion?
(A) The population standard deviation is known.
(B) The population distribution is normal.
(C) np-hat >= 10 and n(1 - p-hat) >= 10.
(D) The sample size is at least 50.
(E) The data are quantitative.
11. A two-way table has 3 rows and 4 columns. For a chi-square test for independence, the degrees of freedom are:
(A) 12
(B) 7
(C) 6
(D) 5
(E) 11
12. The sampling distribution of the sample mean has a mean equal to the population mean. This means that x-bar is:
(A) Biased
(B) Unbiased
(C) Efficient
(D) Consistent
(E) Normal
13. In a regression analysis, the residual standard error s is reported as 4.2. This means:
(A) The slope is 4.2.
(B) The correlation is 4.2.
(C) The typical distance of data points from the regression line is 4.2.
(D) The standard deviation of x is 4.2.
(E) The margin of error is 4.2.
14. A study randomly assigns patients to Drug A or Drug B, and measures blood pressure after 4 weeks. What type of study is this?
(A) Observational study
(B) Survey
(C) Randomized experiment
(D) Census
(E) Case study
15. P(A) = 0.6 and P(B | A) = 0.4. What is P(A and B)?
(A) 1.0
(B) 0.24
(C) 0.20
(D) 0.40
(E) 0.67
16. A paired t-test with 12 pairs of observations has how many degrees of freedom?
(A) 10
(B) 11
(C) 12
(D) 22
(E) 24
17. The power of a hypothesis test is the probability of:
(A) Making a Type I error
(B) Making a Type II error
(C) Correctly rejecting a false null hypothesis
(D) Correctly failing to reject a true null hypothesis
(E) Rejecting the null hypothesis regardless of whether it is true
18. A chi-square goodness of fit test is used to determine whether:
(A) Two categorical variables are independent
(B) The distribution of one categorical variable matches a specified distribution
(C) The mean of one group differs from the mean of another
(D) A regression slope is significant
(E) A proportion equals a specified value
19. Which of the following will decrease the margin of error of a confidence interval for a proportion?
(A) Increasing the confidence level
(B) Decreasing the sample size
(C) Increasing the sample size
(D) Using p-hat = 0.50
(E) Decreasing the population size
20. A researcher tests H0: mu = 100 vs. Ha: mu > 100 and gets a p-value of 0.07. At alpha = 0.05, the correct decision is:
(A) Reject H0
(B) Fail to reject H0
(C) Accept H0
(D) Accept Ha
(E) The test is inconclusive
21. Two variables have a correlation of r = -0.85. Which statement is correct?
(A) There is a weak negative association.
(B) There is a strong positive association.
(C) There is a strong negative linear association.
(D) The slope of the regression line is positive.
(E) 85% of the data points fall on a line.
22. A random variable X has E(X) = 8 and SD(X) = 3. What are E(2X + 5) and SD(2X + 5)?
(A) E = 21, SD = 11
(B) E = 16, SD = 6
(C) E = 21, SD = 6
(D) E = 16, SD = 11
(E) E = 21, SD = 9
23. In which situation would stratified sampling be most appropriate?
(A) When the population naturally forms distinct groups that may differ
(B) When a complete list of the population is unavailable
(C) When cost must be minimized
(D) When the population is homogeneous
(E) When only one sample is needed
24. For a t-distribution with 10 degrees of freedom, approximately what percentage of the area is within t = +/- 2.23 (the 95% critical value)?
(A) 90%
(B) 95%
(C) 99%
(D) 68%
(E) 5%
25. A 99% confidence interval for the difference in two proportions (p1 - p2) is (0.02, 0.18). Which conclusion is correct?
(A) p1 = p2
(B) p1 < p2
(C) p1 > p2
(D) There is not convincing evidence of a difference.
(E) The test would fail to reject H0 at alpha = 0.01.
26. The Law of Large Numbers states that as the sample size increases:
(A) The sampling distribution becomes more normal.
(B) The sample statistic approaches the population parameter.
(C) The margin of error increases.
(D) The p-value decreases.
(E) The confidence level increases.
27. A residual plot shows a clear U-shaped pattern. This suggests that:
(A) The linear model is appropriate.
(B) There is an outlier.
(C) A nonlinear model would be more appropriate.
(D) The residuals are normally distributed.
(E) The correlation is strong.
28. An experiment uses a randomized block design. The purpose of blocking is to:
(A) Increase the sample size.
(B) Reduce variability from a known source.
(C) Eliminate the need for a control group.
(D) Ensure double-blinding.
(E) Make the experiment observational.
29. A hypothesis test for a regression slope yields t = 2.89 with df = 22. What is the approximate p-value for a two-sided test?
(A) 0.004
(B) 0.008
(C) 0.996
(D) 0.992
(E) 0.050
30. Which of the following is the most important difference between a confidence interval and a hypothesis test?
(A) A CI gives a range; a test gives a yes/no decision.
(B) A CI uses z; a test uses t.
(C) A CI requires a larger sample.
(D) A test is more accurate.
(E) They always give the same conclusion.
31. A sample of 200 voters finds that 55% support a candidate. The standard error of p-hat is:
(A) 0.55
(B) 0.035
(C) 0.050
(D) 0.003
(E) 0.245
32. In a boxplot, the box represents:
(A) The range of the data.
(B) The middle 50% of the data (from Q1 to Q3).
(C) One standard deviation from the mean.
(D) The 95% confidence interval.
(E) The modal category.
33. Which of the following is NOT a source of bias in sampling?
(A) Undercoverage
(B) Nonresponse
(C) Random selection
(D) Response bias
(E) Voluntary response
34. The Central Limit Theorem guarantees that the sampling distribution of x-bar is approximately normal when:
(A) The population is normal.
(B) n >= 30.
(C) The sample is random.
(D) Both (A) and (B).
(E) Both (B) and (C) together with (A) not required.
35. A matched pairs design is a special case of:
(A) Completely randomized design.
(B) Randomized block design.
(C) Stratified sampling.
(D) Cluster sampling.
(E) Observational study.
36. If a 95% CI for mu is (10.2, 14.8), what is the point estimate?
(A) 10.2
(B) 14.8
(C) 4.6
(D) 12.5
(E) 2.3
37. The expected count in a chi-square test for a cell in a 2x2 table with row total 80, column total 60, and grand total 200 is:
(A) 24
(B) 40
(C) 30
(D) 20
(E) 48
38. A two-sample t-test assumes:
(A) Equal variances
(B) Known population standard deviations
(C) Independent random samples
(D) Paired observations
(E) A normal population distribution for both
39. The complement of P(A | B) is:
(A) P(A^c | B)
(B) P(A | B^c)
(C) 1 - P(A and B)
(D) P(B | A)
(E) P(A^c and B)
40. In regression inference, the t-distribution used for testing the slope has degrees of freedom equal to:
(A) n - 1
(B) n
(C) n - 2
(D) n/2
(E) 2n
Section II: Free-Response Questions (5 questions, 90 minutes)
Question 1 (Investigative Task)
A large university wants to study the relationship between the number of hours students work at part-time jobs per week and their cumulative GPA. A random sample of 200 students is selected. The university collects the following data for each student: hours worked per week, cumulative GPA, year in school (freshman, sophomore, junior, senior), and whether the student receives financial aid (yes/no).
Computer output from the regression of GPA on hours worked:
- n = 200
- y-hat = 3.42 - 0.025x
- SE_b = 0.008
- s = 0.38
- r^2 = 0.18
(a) Interpret the slope of the regression line in context. (1 point)
(b) Is there a statistically significant negative linear relationship between hours worked and GPA? Perform an appropriate test at alpha = 0.05. (4 points)
(c) Construct and interpret a 95% confidence interval for the true slope. (3 points)
(d) The r^2 value is 0.18. Interpret this value in context. (1 point)
(e) A critic says, "This study proves that working more hours causes lower GPA." Evaluate this claim. Address whether the study design supports a causal conclusion and discuss what the r^2 value tells us about the strength of the relationship. (3 points)
(f) The researcher wants to check whether the conditions for regression inference are met. Describe how to check the LINE conditions (Linear, Normal, Independent, Equal variance). Be specific about what graphs or calculations you would examine. (4 points)
(g) Suppose the researcher discovers that year in school is a lurking variable: seniors tend to work fewer hours and have higher GPAs. Explain how this lurking variable could account for the observed negative association between hours worked and GPA. (2 points)
(h) The researcher decides to control for year in school by including it in the analysis. She creates four separate regressions (one for each year) and finds that the slope of hours worked is -0.008 (SE = 0.015) for freshmen, -0.010 (SE = 0.012) for sophomores, -0.012 (SE = 0.011) for juniors, and -0.005 (SE = 0.014) for seniors. Based on these results, what conclusion would you draw about the relationship between hours worked and GPA? Justify using statistical reasoning. (2 points)
Question 2
A company claims that its new energy-efficient light bulbs last an average of 10,000 hours. A consumer group tests a random sample of 40 bulbs and finds a mean life of 9,620 hours with a standard deviation of 1,200 hours.
(a) Do these data provide convincing evidence that the mean life of the bulbs is less than 10,000 hours? Perform a complete hypothesis test at alpha = 0.05. (5 points)
(b) Construct and interpret a 95% confidence interval for the true mean life of the bulbs. (3 points)
(c) Explain how your results in parts (a) and (b) are consistent with each other. (2 points)
Question 3
A school district is considering changing its lunch menu. They survey a random sample of 300 high school students and 250 middle school students about their satisfaction with the current menu. The results are:
- High school: 180 satisfied, 120 not satisfied
- Middle school: 175 satisfied, 75 not satisfied
(a) Construct a 95% confidence interval for the difference in the proportion of students who are satisfied (high school - middle school). (4 points)
(b) Based on your interval, is there convincing evidence that the satisfaction rates differ between high school and middle school students? Explain. (2 points)
(c) A school board member says, "The difference is not statistically significant, so the two groups have the same satisfaction rate." Is this statement correct? Explain. (2 points)
Question 4
A die is suspected of being unfair. It is rolled 120 times with the following results:
| Face | 1 | 2 | 3 | 4 | 5 | 6 |
|---|---|---|---|---|---|---|
| Observed | 12 | 18 | 22 | 25 | 20 | 23 |
(a) State the null and alternative hypotheses. (1 point)
(b) Calculate the expected count for each face. (1 point)
(c) Calculate the chi-square test statistic. (3 points)
(d) The p-value for this test is 0.082. State an appropriate conclusion at alpha = 0.05. (2 points)
(e) At alpha = 0.10, what would the conclusion be? (1 point)
Question 5
A researcher wants to determine whether listening to classical music while studying improves exam performance. She recruits 60 student volunteers and randomly assigns them to two groups: one studies with classical music and the other studies in silence. After one week, all students take the same exam.
- Music group (n = 30): mean = 78.5, SD = 8.2
- Silence group (n = 30): mean = 74.8, SD = 9.1
(a) Is this an experiment or an observational study? Explain how you can tell. (2 points)
(b) Perform a two-sample t-test to determine whether there is a significant difference in mean exam scores between the two groups. Use alpha = 0.05. (5 points)
(c) Construct a 90% confidence interval for the difference in means (music - silence) and interpret it. (3 points)
(d) The researcher originally considered using a matched pairs design where each student would take two exams (one with music, one without). Explain one advantage and one disadvantage of the paired design compared to the two-independent-groups design used. (2 points)
END OF EXAM
Answer Key & Rubric
AP Statistics — Full Practice Exam: Answer Key and Scoring Rubrics
Section I: Multiple-Choice Answer Key
- B — Adding a constant changes the mean but not the standard deviation.
- E — Systematic sampling from a voter list is the most structured method listed.
- D — r = 0.55 indicates a moderate positive linear association.
- A — P(A or B) = P(A) + P(B) = 0.35 + 0.45 = 0.80 for mutually exclusive events.
- C — The standard interpretation of a confidence interval.
- B — Higher confidence level requires a larger critical value, hence larger margin of error.
- B — A control group provides a baseline to measure the treatment effect against.
- D — For symmetric data, mean and median are approximately equal; report both.
- B — A z-score measures standard deviations from the mean.
- C — Large Counts condition: np-hat >= 10 and n(1-p-hat) >= 10.
- C — df = (3-1)(4-1) = 2(3) = 6.
- B — When the mean of a statistic's sampling distribution equals the parameter, it is unbiased.
- C — The residual standard error measures typical distance of points from the line.
- C — Random assignment makes this a randomized experiment.
- B — P(A and B) = P(A) P(B|A) = 0.6 0.4 = 0.24.
- B — df = number of pairs - 1 = 12 - 1 = 11.
- C — Power = P(correctly rejecting a false H0) = 1 - beta.
- B — GOF tests whether a single variable matches a specified distribution.
- C — Larger n decreases the standard error, decreasing the margin of error.
- B — p-value 0.07 > 0.05, so fail to reject H0.
- C — r = -0.85 indicates a strong negative linear association.
- C — E(2X+5) = 2(8)+5 = 21. SD(2X+5) = |2|(3) = 6.
- A — Stratified sampling ensures representation of distinct subgroups.
- B — By definition, the 95% critical value contains 95% of the area.
- C — The interval is entirely positive (does not contain 0), so p1 > p2.
- B — The LLN states the statistic approaches the parameter as n grows.
- C — A U-shaped pattern indicates nonlinearity.
- B — Blocking reduces variability from a known confounding variable.
- B — p-value = 2 * tcdf(2.89, 1E99, 22) = 2(0.004) = 0.008.
- A — A CI gives a range of plausible values; a test gives a binary decision.
- B — SE = sqrt(0.55*0.45/200) = sqrt(0.001238) = 0.035.
- B — The box spans from Q1 to Q3, the middle 50%.
- C — Random selection reduces bias; it is not a source of bias.
- E — The CLT requires both random sampling AND n >= 30 (when population is not normal).
- B — Matched pairs is a special case of randomized block design with blocks of size 2.
- D — Point estimate = (10.2 + 14.8)/2 = 12.5.
- A — E = (80 * 60) / 200 = 4800/200 = 24.
- C — Two-sample t-tests assume independent random samples.
- A — The complement of "A given B" is "not A given B": P(A^c | B).
- C — df = n - 2 for regression inference on the slope.
Section II: Free-Response Scoring Rubrics
Question 1 — Investigative Task (20 points total)
(a) [1 point] For each additional hour a student works per week, the predicted GPA decreases by 0.025 points.
(b) [4 points]
- Hypotheses [1 point]: H0: beta = 0 (no linear relationship). Ha: beta < 0 (negative linear relationship).
- Test statistic [1 point]: t = (b - 0) / SE_b = -0.025 / 0.008 = -3.125. df = 200 - 2 = 198.
- p-value [1 point]: p = P(T < -3.125 with df = 198) ≈ 0.001.
- Conclusion [1 point]: Since 0.001 < 0.05, we reject H0. There is convincing evidence of a negative linear relationship between hours worked per week and GPA.
(c) [3 points]
- Critical value [1 point]: t* ≈ 1.972 (for 95% CI with large df, approximately 1.96).
- Calculation [1 point]: ME = 1.972 * 0.008 = 0.0158. CI: (-0.025 - 0.016, -0.025 + 0.016) = (-0.041, -0.009).
- Interpretation [1 point]: We are 95% confident that for each additional hour worked per week, the true mean decrease in GPA is between 0.009 and 0.041 points.
(d) [1 point] Approximately 18% of the variation in GPA among students is explained by the linear relationship with hours worked per week. The remaining 82% is due to other factors.
(e) [3 points]
- Causation [1 point]: This is an observational study (no random assignment of work hours), so we cannot conclude that working more hours causes lower GPA.
- r^2 interpretation [1 point]: r^2 = 0.18 indicates a weak relationship. Even if causation were established, working hours explain only a small portion of GPA variation.
- Overall evaluation [1 point]: The claim is not supported. The observational design cannot establish causation, and the weak r^2 means the relationship, while statistically significant, has limited practical importance.
(f) [4 points — 1 point per condition]
- Linear: Create a scatterplot of GPA vs. hours worked. Check that the pattern is roughly linear (no clear curve).
- Normal: Create a normal probability plot of the residuals. Check that the points are approximately on a straight line. Alternatively, check a histogram of residuals for approximate normality.
- Independent: The sample of 200 should be less than 10% of all university students. Also, students' GPAs should not influence each other's.
- Equal variance: Create a residual plot (residuals vs. hours worked). Check that the spread of residuals is roughly the same across all x-values (no fan shape).
(g) [2 points] Seniors (high GPA, few hours) pull the regression line in a way that creates a spurious negative association. If seniors are overrepresented among low-work-hours/high-GPA students, the aggregate data show a negative trend even if within each year there is little or no relationship. This is an example of Simpson's paradox. [1 point for identifying the confounding, 1 point for explaining the mechanism.]
(h) [2 points] Within each year, the slopes are small (-0.005 to -0.012) and none are statistically significant (each has a p-value well above 0.05 since |t| < 1 for all four). This suggests that after controlling for year in school, the apparent negative relationship between hours worked and GPA largely disappears. The original observed association was likely due to the confounding effect of year in school. [1 point for noting none are significant, 1 point for connecting this to the confounding.]
Question 2 (10 points total)
(a) [5 points]
- Hypotheses [1 point]: H0: mu = 10,000. Ha: mu < 10,000.
- Conditions [1 point]: Random (stated). Normal: n = 40 >= 30, so CLT applies. Independent: 40 < 10% of all bulbs. Checked.
- Test statistic [1 point]: t = (9620 - 10000) / (1200/sqrt(40)) = -380 / 189.74 = -2.002. df = 39.
- p-value [1 point]: p = P(T < -2.002 with df = 39) ≈ 0.026.
- Conclusion [1 point]: Since 0.026 < 0.05, reject H0. There is convincing evidence that the mean life of the bulbs is less than 10,000 hours.
(b) [3 points]
- Calculation [2 points]: t for 95% CI with df = 39 is approximately 2.023. SE = 1200/sqrt(40) = 189.74. ME = 2.023 189.74 = 383.9. CI: 9620 +/- 384 = (9236, 10004).
- Interpretation [1 point]: We are 95% confident that the true mean life of the bulbs is between 9,236 and 10,004 hours.
(c) [2 points] The confidence interval barely includes 10,000 (it is essentially at the boundary). This is consistent with the hypothesis test: at alpha = 0.05 (which corresponds to a 95% CI), the result is borderline. Since 10,000 is within (or very nearly at the edge of) the interval, the test gives a p-value near 0.05, and we just barely reject H0. The CI and the test give consistent conclusions.
Question 3 (8 points total)
(a) [4 points]
- p-hat_HS = 180/300 = 0.60. p-hat_MS = 175/250 = 0.70. Difference = 0.60 - 0.70 = -0.10.
- SE = sqrt(0.600.40/300 + 0.700.30/250) = sqrt(0.000800 + 0.000840) = sqrt(0.001640) = 0.0405.
- z = 1.96. ME = 1.96 0.0405 = 0.0794.
- CI: -0.10 +/- 0.079 = (-0.179, -0.021). [Points: 1 for proportions, 1 for SE, 1 for CI calculation, 1 for conditions checked.]
(b) [2 points] Since the interval (-0.179, -0.021) does not contain 0, there is convincing evidence that the satisfaction rates differ between high school and middle school students. The interval is entirely negative, suggesting middle school students have a higher satisfaction rate. [1 point for answer, 1 point for justification.]
(c) [2 points] The statement is incorrect. Failing to reject H0 (or finding no significant difference) does not prove the proportions are equal. It only means the sample does not provide sufficient evidence to conclude they differ. The true difference could be small but nonzero. Additionally, the answer in part (b) actually shows the interval does NOT contain 0, so there IS a statistically significant difference.
Question 4 (8 points total)
(a) [1 point] H0: The die is fair (each face has probability 1/6). Ha: The die is not fair (at least one face has a different probability).
(b) [1 point] Expected count for each face = 120/6 = 20.
(c) [3 points]
- (12-20)^2/20 = 64/20 = 3.20
- (18-20)^2/20 = 4/20 = 0.20
- (22-20)^2/20 = 4/20 = 0.20
- (25-20)^2/20 = 25/20 = 1.25
- (20-20)^2/20 = 0/20 = 0
- (23-20)^2/20 = 9/20 = 0.45
X^2 = 3.20 + 0.20 + 0.20 + 1.25 + 0 + 0.45 = 5.30. [1 point for each correct term (partial credit), 1 point for sum.]
(d) [2 points] p-value = 0.082. Since 0.082 > 0.05, we fail to reject H0. There is not convincing evidence that the die is unfair at the 5% significance level. [1 point for comparison, 1 point for contextual conclusion.]
(e) [1 point] At alpha = 0.10, since 0.082 < 0.10, we reject H0. There is convincing evidence that the die is unfair at the 10% significance level.
Question 5 (12 points total)
(a) [2 points] This is an experiment. Students were randomly assigned to the two groups (music vs. silence), which is the defining feature of an experiment. The researcher imposed a treatment and measured the response. [1 point for identification, 1 point for justification.]
(b) [5 points]
- Hypotheses [1 point]: H0: mu_music = mu_silence (or mu_music - mu_silence = 0). Ha: mu_music != mu_silence (or mu_music - mu_silence != 0).
- Conditions [1 point]: Random assignment (stated). Normal: Both n >= 30, so CLT applies. Independent: The two groups are independent (different students). Each group is assumed to be less than 10% of all students.
- Test statistic [1 point]: SE = sqrt(8.2^2/30 + 9.1^2/30) = sqrt(67.24/30 + 82.81/30) = sqrt(2.241 + 2.760) = sqrt(5.001) = 2.236. t = (78.5 - 74.8) / 2.236 = 3.7 / 2.236 = 1.654. df = min(29, 29) = 29.
- p-value [1 point]: p = 2 * P(T > 1.654 with df = 29) = 2(0.0544) = 0.109.
- Conclusion [1 point]: Since 0.109 > 0.05, we fail to reject H0. There is not convincing evidence of a difference in mean exam scores between the music and silence groups.
(c) [3 points]
- Calculation [2 points]: t for 90% CI with df = 29 is 1.699. ME = 1.699 2.236 = 3.799. CI: 3.7 +/- 3.80 = (-0.10, 7.50).
- Interpretation [1 point]: We are 90% confident that the true difference in mean exam scores (music - silence) is between -0.10 and 7.50 points.
(d) [2 points]
- Advantage [1 point]: The paired design controls for individual student ability, which could reduce variability and increase the power to detect an effect.
- Disadvantage [1 point]: The paired design could introduce carryover effects (the experience from the first exam might affect performance on the second), or order effects (taking exams in a different order might matter). Also, it requires each student to take two exams.
Scoring Summary
| Question | Points |
|---|---|
| MCQ (40 questions) | 40 points (scaled) |
| FRQ 1 (Investigative Task) | 20 points (scaled to ~12.5) |
| FRQ 2 | 10 points (scaled to ~6.25) |
| FRQ 3 | 8 points (scaled to ~6.25) |
| FRQ 4 | 8 points (scaled to ~6.25) |
| FRQ 5 | 12 points (scaled to ~6.25) |
Note: On the actual AP exam, each section is worth 50% of the total score. The raw FRQ points are scaled to ensure equal weighting. This practice exam uses a simplified point system.