AP Statistics study package
Everything you need to prepare for the AP AP Statistics exam in one place: course overview, per-unit notes, practice sets, a full-length practice exam with answer key, and a printable summary sheet. Works alongside the timed AP Statistics practice exam and the score calculator.
Course overview
1AP Statistics — Complete Study Package Overview
This study package is a comprehensive, self-contained resource for AP Statistics. Every file has been written from scratch to help you master the course content, build statistical reasoning skills, and prepare for the AP Exam. Whether you are reviewing throughout the year or cramming in the final weeks, this package is designed to give you the structure and practice you need.
Exam Format
The AP Statistics exam is 3 hours long and divided into two equally weighted sections.
Section I: Multiple-Choice Questions
- 40 questions, 90 minutes
- Worth 50% of your total score
- Each question has five answer choices (A through E)
- Covers all nine units with varying emphasis
- No penalty for guessing — answer every question
Section II: Free-Response Questions
- 5 questions, 90 minutes
- Worth 50% of your total score
- Question 1: Investigative Task (a multi-part, often challenging question that may combine topics)
- Questions 2–5: Short-answer free-response (each typically has 2–4 parts)
- Partial credit is awarded — show your work and explain your reasoning
- Communication is just as important as correct calculations
Calculator Policy
A graphing calculator is required for the entire exam. You should be comfortable using it for:
- Computing descriptive statistics (mean, standard deviation, five-number summary)
- Creating histograms, boxplots, and scatterplots
- Performing hypothesis tests and confidence intervals
- Calculating binomial and geometric probabilities
Recommended calculators include the TI-84, TI-Nspire, or equivalent. Know where every relevant function is located before exam day.
Formula Sheet
The College Board provides a formula sheet that includes:
- Descriptive statistics formulas (mean, standard deviation, correlation)
- Probability formulas (addition rule, multiplication rule, conditional probability)
- Sampling distribution formulas
- Confidence interval and test statistic formulas for proportions and means
- Chi-square and regression inference formulas
You should be familiar with every formula on the sheet, but do not rely on it as your only reference — you should also have the key formulas memorized.
Course Units and Approximate Exam Weight
| Unit | Topic | Approximate Weight |
|---|---|---|
| 1 | Exploring One-Variable Data | 15–20% |
| 2 | Exploring Two-Variable Data | 2–5% |
| 3 | Collecting Data | 10–15% |
| 4 | Probability, Random Variables, and Probability Distributions | 10–20% |
| 5 | Sampling Distributions | 7–12% |
| 6 | Inference for Categorical Data: Proportions | 12–18% |
| 7 | Inference for Quantitative Data: Means | 12–18% |
| 8 | Inference for Categorical Data: Chi-Square | 5–10% |
| 9 | Inference for Quantitative Data: Slopes | 5–10% |
Units 1–5 form the foundational portion of the course (roughly 50–60% of the exam). Units 6–9 cover statistical inference, which makes up the remaining 40–50% and is heavily tested in the free-response section.
Four Statistical Practices
The AP Statistics course is organized around four broad skill categories, called Statistical Practices:
- Statistical Thinking — Understanding the purpose of a study, identifying the type of study (experiment vs. observational), recognizing how data was collected, and determining whether conclusions are valid.
- Data Analysis — Selecting appropriate graphical displays and numerical summaries, describing distributions, identifying patterns and unusual features, and using technology effectively.
- Probability and Simulation — Understanding randomness, applying probability rules, using random variables to model real-world situations, and interpreting simulation results.
- Statistical Argumentation — Constructing and communicating statistical arguments, interpreting p-values and confidence levels in context, drawing appropriate conclusions, and recognizing the limitations of statistical methods.
Every free-response question tests at least two of these practices. The investigative task often requires all four.
Package Roadmap — All Files
This study package contains approximately 25 files organized as follows:
Overview and Support
| File | Description |
|---|---|
00-overview.md | This file — exam format, unit breakdown, roadmap |
04-summary-sheet.md | All key formulas, conditions, test decision rules, and calculator commands on one reference sheet |
05-exam-strategy.md | MCQ strategy, FRQ writing guide, calculator tips, time management, checklists |
06-presentation-outline.md | Slide-by-slide outline (~50 slides) for a comprehensive review presentation |
07-audio-script.md | Written script for a 15–20 minute audio review covering the entire course |
Unit Notes (9 files)
| File | Topic |
|---|---|
01-unit1-exploring-one-variable-data.md | Categorical/quantitative data, displays, center, spread, shape, outliers, z-scores, transformations |
01-unit2-exploring-two-variable-data.md | Scatterplots, correlation, regression, residuals, r-squared, leverage, causation |
01-unit3-collecting-data.md | Sampling methods, bias, experiments, blinding, study designs |
01-unit4-probability-and-random-variables.md | Probability rules, independence, random variables, expected value, linear combinations |
01-unit5-sampling-distributions.md | Sampling distributions, CLT, law of large numbers |
01-unit6-proportions.md | z-intervals, z-tests, Type I/II errors, power, two-proportion inference |
01-unit7-means.md | t-distributions, t-intervals, t-tests, paired tests, two-sample tests |
01-unit8-chi-square.md | Goodness of fit, independence, homogeneity tests |
01-unit9-regression-inference.md | t-test for slope, CI for slope, residual standard error, prediction intervals |
Unit Practice (9 files)
| File | Topic |
|---|---|
02-practice-unit1.md | Practice problems for Unit 1 |
02-practice-unit2.md | Practice problems for Unit 2 |
02-practice-unit3.md | Practice problems for Unit 3 |
02-practice-unit4.md | Practice problems for Unit 4 |
02-practice-unit5.md | Practice problems for Unit 5 |
02-practice-unit6.md | Practice problems for Unit 6 |
02-practice-unit7.md | Practice problems for Unit 7 |
02-practice-unit8.md | Practice problems for Unit 8 |
02-practice-unit9.md | Practice problems for Unit 9 |
Full Practice Exam
| File | Description |
|---|---|
03-full-practice-exam.md | 40 MCQ + 5 FRQ (including investigative task) |
03-full-practice-exam-answers.md | Complete answer key with detailed scoring rubrics |
How to Use This Package
- During the course: Read each unit note file alongside your textbook. Use the practice files as homework or quiz preparation.
- Before midterms/finals: Focus on the summary sheet and the unit practice files for the units covered so far.
- AP Exam prep (4–6 weeks before): Work through all unit practice files, then take the full practice exam under timed conditions. Review the answer key carefully, focusing on the FRQ rubrics to learn how points are awarded.
- Final review (1–2 weeks before): Read the summary sheet, review the exam strategy guide, and listen to or read the audio script for a rapid refresher. Use the presentation outline to guide group study sessions.
Key Themes to Remember
- Context is everything. Every number you calculate should be interpreted in the context of the problem. The AP exam rewards clear, contextual communication.
- Conditions matter. Before performing any inference procedure, always state and check the conditions. This is often worth 1–2 points on an FRQ.
- Show your work. On FRQs, the process matters more than the final answer. Label intermediate steps clearly.
- Understand, do not memorize. If you understand why a formula works, you can adapt it to unfamiliar situations — exactly what the investigative task requires.
Good luck with your studying!
Unit notes
9Unit 1: Exploring One-Variable Data
Every dataset in statistics is made up of variables — characteristics that vary from one individual or object to another.
Categorical (Qualitative) Variables
Categorical variables place individuals into groups or categories. Examples include eye color, type of car, or zip code.
- Displays: Bar charts, pie charts, segmented bar charts, two-way tables.
- Summaries: Frequency tables, relative frequency tables, proportions.
- Note: The mean and standard deviation are not meaningful for categorical data.
Quantitative (Numerical) Variables
Quantitative variables take numerical values for which arithmetic operations make sense. Examples include height, salary, temperature, and number of siblings.
- Displays: Histograms, dotplots, stemplots, boxplots.
- Summaries: Mean, median, mode, range, IQR, standard deviation, variance.
Identifying the Type
Ask yourself: Does it make sense to average the values? If yes, the variable is likely quantitative. You cannot compute a meaningful average of zip codes or blood types.
1.2 Graphical Displays
Bar Charts
Used for categorical data. Each bar represents a category, and the height (or length) shows the frequency or relative frequency. Bars are separated by gaps.
Pie Charts
Used for categorical data when showing parts of a whole. Each slice represents a category's proportion of the total. Best used when there are few categories (typically five or fewer).
Histograms
Used for quantitative data. The horizontal axis is divided into bins (intervals), and the vertical axis shows frequency or relative frequency. Bars are adjacent (no gaps). Changing bin width can change the appearance of the distribution.
Dotplots
Each data value is represented by a dot above a number line. Useful for small to moderate datasets. Easy to see individual values, clusters, and gaps.
Stemplots (Stem-and-Leaf Plots)
Separate each data value into a stem (leading digits) and a leaf (trailing digit). Useful for small datasets. Preserves the actual data values. Key the plot if stems repeat (e.g., splitting stems into two rows: 0–4 and 5–9).
Boxplots (Box-and-Whisker Plots)
A five-number summary displayed as a box with a line at the median, whiskers extending to the minimum and maximum (or to the most extreme points within 1.5 IQR of the quartiles), and individual dots for outliers. Best for comparing distributions across groups.
1.3 Measures of Center
Mean
The arithmetic mean (x-bar) is the sum of all values divided by the number of values.
x-bar = (sum of x_i) / n
- Uses every data value, so it is affected by outliers and skewness.
- The mean is the balance point of the distribution.
Median
The median is the middle value when the data are ordered. If n is odd, it is the center value. If n is even, it is the average of the two center values.
- Resistant to outliers and skewness.
- In a symmetric distribution, mean approximately equals median.
- In a right-skewed distribution, mean > median.
- In a left-skewed distribution, mean < median.
Mode
The value (or values) that occur most frequently. A dataset can have one mode (unimodal), two modes (bimodal), or more. The mode is the only measure of center that works for categorical data.
Choosing Between Mean and Median
Use the median when the data are skewed or contain outliers. Use the mean when the data are roughly symmetric. Always report both for quantitative data on the AP exam.
1.4 Measures of Spread
Range
Range = Maximum - Minimum
Simple but uses only two values and is very sensitive to outliers.
Interquartile Range (IQR)
IQR = Q3 - Q1
The range of the middle 50% of the data. Resistant to outliers.
To find Q1 and Q3 on a calculator, use 1-Var Stats; the calculator reports Q1 and Q3.
Variance
s² = (sum of (x_i - x-bar)²) / (n - 1)
The average squared deviation from the mean. Uses n - 1 (not n) in the denominator because this gives an unbiased estimate of the population variance. Units are squared, making interpretation difficult.
Standard Deviation
s = sqrt(variance) = sqrt((sum of (x_i - x-bar)²) / (n - 1))
The typical distance of data values from the mean. Same units as the original data. Like the mean, it is not resistant to outliers.
- s = 0 only when all values are identical.
- s is always non-negative.
- On a TI-84: Enter data in L1, then
1-Var Stats L1. Look for Sx (sample standard deviation).
1.5 Describing the Shape of a Distribution
When describing any distribution, address these four features in order:
- Shape: Symmetric, left-skewed, right-skewed, uniform, bimodal.
- Center: Report the mean and median.
- Spread: Report the range, IQR, and/or standard deviation.
- Unusual features: Gaps, clusters, and outliers.
Symmetric
The left and right sides are approximately mirror images. Mean and median are approximately equal.
Skewed Right (Positively Skewed)
A long tail extends to the right. The mean is pulled in the direction of the tail, so mean > median.
Skewed Left (Negatively Skewed)
A long tail extends to the left. Mean < median.
1.6 Outliers
The 1.5 x IQR Rule
A data value is considered a potential outlier if it falls below
Q1 - 1.5 x IQR
or above
Q3 + 1.5 x IQR
This is a rule of thumb, not a definitive test. On a boxplot, outliers are plotted as individual points beyond the whiskers.
Why Outliers Matter
- They can dramatically affect the mean and standard deviation.
- They may indicate data entry errors or genuinely unusual observations.
- Always investigate before removing them.
1.7 The Five-Number Summary
Min, Q1, Median, Q3, Max
This summary is resistant to outliers and provides a complete picture when combined with a boxplot. It divides the data into four quarters, each containing roughly 25% of the data.
1.8 Standardized Scores (z-Scores)
A z-score measures how many standard deviations a value is from the mean.
z = (x - x-bar) / s
- z > 0: The value is above the mean.
- z < 0: The value is below the mean.
- |z| tells you how unusual the value is. A common (but rough) guideline: values with |z| > 2 are considered unusual.
- Z-scores have no units — they are dimensionless.
- Z-scores can be used to compare values from different distributions (e.g., comparing a test score to a salary).
1.9 Transformations
Applying the same operation to every data value is called a transformation.
Adding or Subtracting a Constant (a)
- Adds a to the mean and median.
- Does not change the spread (range, IQR, standard deviation).
- Does not change the shape.
Multiplying or Dividing by a Positive Constant (b)
- Multiplies the mean, median, range, IQR, and standard deviation by b.
- Does not change the shape.
- Variance is multiplied by b².
Linear Transformations (y = a + bx)
- New mean = a + b(old mean)
- New median = a + b(old median)
- New standard deviation = |b|(old standard deviation)
- New IQR = |b|(old IQR)
- Shape is unchanged.
These rules are essential for problems involving unit conversions (e.g., inches to centimeters, Celsius to Fahrenheit).
1.10 Worked Example
A teacher recorded the number of books read by 10 students over the summer: 2, 3, 3, 4, 5, 5, 5, 7, 8, 18.
Step 1: Order the data. Already ordered.
Step 2: Find the five-number summary.
- Min = 2
- Q1 = 3 (median of the lower half: 2, 3, 3, 4, 5)
- Median = 5 (average of 5th and 6th values)
- Q3 = 7.5 (median of upper half: 5, 5, 7, 8, 18)
- Max = 18
Step 3: Calculate IQR.
- IQR = Q3 - Q1 = 7.5 - 3 = 4.5
Step 4: Check for outliers.
- Lower fence: Q1 - 1.5(IQR) = 3 - 1.5(4.5) = 3 - 6.75 = -3.75 (no outliers below)
- Upper fence: Q3 + 1.5(IQR) = 7.5 + 6.75 = 14.25
- 18 > 14.25, so 18 is a potential outlier.
Step 5: Calculate the mean and standard deviation.
- Mean = (2 + 3 + 3 + 4 + 5 + 5 + 5 + 7 + 8 + 18) / 10 = 60 / 10 = 6
- The standard deviation (using a calculator) is approximately s = 4.72.
Step 6: Describe the distribution.
- The distribution of books read is right-skewed (mean of 6 > median of 5) with a potential outlier at 18 books. The center is approximately 5 books (median), with an IQR of 4.5 books.
Step 7: Calculate the z-score for the outlier.
- z = (18 - 6) / 4.72 = 12 / 4.72 = 2.54
- This student read 2.54 standard deviations above the mean.
1.11 Common Mistakes
- Confusing categorical and quantitative variables. Zip codes are numbers but are categorical. You cannot compute a mean zip code.
- Using the mean for skewed data. When data are skewed, the median is a better measure of center. Always report both.
- Forgetting to label axes and title graphs. On FRQs, every graph needs a title and labeled axes with units.
- Confusing the sample standard deviation (Sx) with the population standard deviation (sigma-x). On a TI-84, use Sx for sample data.
- Using n instead of n-1 in the variance formula. AP Statistics always uses the sample variance (dividing by n-1) unless the entire population is known.
- Stating that a z-score "proves" a value is unusual. A z-score is a measure, not proof. It provides evidence, but context matters.
1.12 Self-Check Questions
- A dataset has a mean of 45 and a standard deviation of 8. What is the z-score of a value of 61? Interpret this z-score in context.
- A set of exam scores is roughly symmetric with a mean of 76 and standard deviation of 10. Approximately what percentage of scores are between 66 and 86?
- The following data represent the number of hours 8 students spent studying last week: 2, 3, 4, 5, 6, 7, 8, 15. Find the five-number summary and identify any outliers.
- If all values in a dataset are multiplied by 3 and then 5 is added, what happens to the mean, median, standard deviation, and IQR?
- A bar chart and a histogram both display data using bars. Explain the key difference between these two types of graphs.
- Explain why the median is preferred over the mean when describing the center of a dataset that includes income data for a community.
Answers to Self-Check
- z = (61 - 45) / 8 = 2. This value is 2 standard deviations above the mean.
- By the empirical rule, approximately 68% of data in a normal distribution fall within one standard deviation of the mean. Since the distribution is roughly symmetric, about 68% of scores are between 66 and 86.
- Ordered: 2, 3, 4, 5, 6, 7, 8, 15. Min = 2, Q1 = 3.5, Median = 5.5, Q3 = 7.5, Max = 15. IQR = 4. Lower fence = 3.5 - 6 = -2.5. Upper fence = 7.5 + 6 = 13.5. Since 15 > 13.5, 15 is a potential outlier.
- New mean = 5 + 3(old mean). New median = 5 + 3(old median). New standard deviation = 3(old standard deviation). New IQR = 3(old IQR).
- Bar charts are used for categorical data and have gaps between bars. Histograms are used for quantitative data and have adjacent bars (no gaps).
- Income data are typically right-skewed due to a few very high earners. The mean is pulled upward by these outliers, making it larger than the typical income. The median is resistant to these extreme values and better represents the center of the majority of the data.
Unit 2: Exploring Two-Variable Data
A scatterplot displays the relationship between two quantitative variables measured on the same individuals. One variable is placed on the horizontal axis (the explanatory or independent variable, x) and the other on the vertical axis (the response or dependent variable, y).
What to Look For
- Direction: Positive association (both variables tend to increase together) or negative association (one increases as the other decreases).
- Form: Linear, curved, or no pattern.
- Strength: How closely the points follow a pattern. Strong associations have points tightly clustered around a line or curve; weak associations are more scattered.
- Unusual points: Outliers that do not follow the overall pattern.
Important Distinction
Correlation and regression describe linear associations only. If the form is curved, do not apply linear methods.
2.2 Correlation (r)
The correlation coefficient r measures the strength and direction of the linear association between two quantitative variables.
Properties of r
- r is always between -1 and +1.
- r = +1: Perfect positive linear association.
- r = -1: Perfect negative linear association.
- r = 0: No linear association (but there could be a nonlinear relationship).
- The sign of r indicates the direction; |r| indicates the strength.
- r is unitless and dimensionless.
- r is not affected by linear transformations (changing units).
- r measures only linear association.
- r is not resistant — a single outlier can dramatically change r.
Calculating r
You will not need to calculate r by hand on the exam. On a TI-84: enter x-values in L1 and y-values in L2, then use LinReg(a+bx) L1, L2. The calculator reports r (and r²) if diagnostics are enabled.
2.3 Least-Squares Regression Line (LSRL)
The least-squares regression line is the line that minimizes the sum of the squared vertical distances (residuals) from the data points to the line.
y-hat = a + bx
- b (slope) = r(sy / sx), where sy and sx are the standard deviations of y and x.
- a (y-intercept) = y-bar - b(x-bar)
Interpreting the Slope
"For each increase of 1 unit in [x-variable], the predicted [y-variable] increases/decreases by [b] [units]."
Example: If the regression line for study hours (x) and exam score (y) is y-hat = 52 + 6.2x, the slope is 6.2. Interpretation: "For each additional hour of studying, the predicted exam score increases by 6.2 points."
Interpreting the y-Intercept
"When [x-variable] = 0, the predicted [y-variable] is [a] [units]."
Warning: The y-intercept is only meaningful if x = 0 is within the range of the data and makes logical sense. If x represents a person's age, predicting at x = 0 may be meaningless.
2.4 Residuals
A residual is the difference between the observed y-value and the predicted y-value from the regression line.
residual = y - y-hat
- A positive residual means the actual value is above the line (the model underpredicted).
- A negative residual means the actual value is below the line (the model overpredicted).
- The sum of residuals is always zero (if the line includes a y-intercept).
- The mean of residuals is always zero.
- Residuals are measured in the same units as the response variable.
Residual Plots
A residual plot graphs each residual (y-axis) against the corresponding x-value (horizontal axis). A horizontal line is drawn at residual = 0.
Purpose: To check whether a linear model is appropriate.
- Good model: Residuals are randomly scattered above and below zero with no clear pattern.
- Bad model: Residuals show a pattern (curved, fan shape, etc.), indicating the linear model is not appropriate.
2.5 Coefficient of Determination (r²)
r² represents the proportion of variation in the response variable (y) that is explained by the linear relationship with the explanatory variable (x).
- r² is always between 0 and 1 (or 0% and 100%).
- r² = 0.72 means 72% of the variation in y is accounted for by the linear relationship with x.
- The remaining 1 - r² (28%) is due to other factors and unexplained variation.
Interpretation Template
"Approximately [r² × 100]% of the variation in [y-variable] is explained by the linear relationship with [x-variable]."
2.6 Interpolation vs. Extrapolation
- Interpolation: Making predictions within the range of the observed x-values. Generally reliable.
- Extrapolation: Making predictions outside the range of the observed x-values. Risky and unreliable because the pattern may not continue.
Example: If study hours range from 1 to 8, predicting the score for 5 hours is interpolation. Predicting for 20 hours is extrapolation.
2.7 Outliers, Leverage, and Influential Points
Outliers in Regression
A point with a large residual. It falls far from the regression line.
High-Leverage Points
A point with an x-value that is far from the mean of x. These have the potential to pull the regression line toward themselves.
Influential Points
A point that, if removed, would significantly change the slope or y-intercept of the regression line. Points with both high leverage and a large residual are especially influential.
Key idea: Not all outliers are influential, and not all influential points are outliers in the traditional sense. A high-leverage point that lies close to the line will have a small residual but can still be influential.
2.8 Lurking Variables and Causation
Lurking Variable
A variable that is not among the explanatory or response variables in a study but that may influence the interpretation of the relationship between them.
Confounding
Two variables are confounded when their effects on the response variable cannot be distinguished from each other.
Correlation Does Not Imply Causation
An observed association between x and y does not mean that changes in x cause changes in y. There are three possible explanations:
- x causes y (causation).
- y causes x (reverse causation).
- A lurking variable z causes both x and y (common response).
Example: Ice cream sales and drowning rates are positively correlated. This does not mean ice cream causes drowning. The lurking variable is temperature — warm weather increases both ice cream sales and swimming activity.
Establishing Causation
The best way to establish a cause-and-effect relationship is through a well-designed, randomized experiment.
2.9 Worked Example
A researcher collected data on the number of absences (x) and final exam score (y) for 8 students:
| Absences (x) | Score (y) |
|---|---|
| 0 | 95 |
| 1 | 90 |
| 2 | 82 |
| 3 | 78 |
| 4 | 72 |
| 5 | 65 |
| 6 | 60 |
| 8 | 50 |
Using a calculator, the regression output gives:
- a = 96.6, b = -5.85, r = -0.998, r² = 0.996
Regression line: y-hat = 96.6 - 5.85x
Interpretation of slope: For each additional absence, the predicted exam score decreases by 5.85 points.
Interpretation of r²: Approximately 99.6% of the variation in exam scores is explained by the linear relationship with number of absences.
Prediction: For a student with 3 absences, y-hat = 96.6 - 5.85(3) = 96.6 - 17.55 = 79.05. The predicted score is about 79 points.
Residual for student with 3 absences: residual = 78 - 79.05 = -1.05. The student scored 1.05 points below the predicted value.
2.10 Common Mistakes
- Using correlation when the relationship is nonlinear. If a scatterplot shows a curve, r is not meaningful and the LSRL is inappropriate.
- Confusing the explanatory and response variables. The x-variable goes on the horizontal axis and is used to predict y.
- Interpreting the y-intercept without checking if x = 0 is meaningful. If x = 0 is outside the data range, do not interpret the intercept.
- Saying "r² = 0.85 means 85% of the data points are on the line." Incorrect. r² refers to variation explained, not the percentage of points on the line.
- Claiming causation from an observational study. Only experiments with random assignment can establish causation.
- Forgetting to plot residuals to check conditions. Always check a residual plot before trusting the regression model.
2.11 Self-Check Questions
- A scatterplot shows a strong, curved pattern. A student calculates r = 0.15. Is this surprising? Explain.
- A regression equation for predicting house price (in thousands) from square footage is y-hat = 20 + 0.15x. Interpret the slope in context.
- A researcher finds r² = 0.64 for the relationship between hours of sleep and reaction time. Interpret this value.
- Explain the difference between an outlier and an influential point in the context of regression.
- A study finds that people who drink more coffee tend to live longer. Can we conclude that coffee causes longer life? Explain.
- In a residual plot, the residuals show a clear U-shaped pattern. What does this indicate about the regression model?
Answers to Self-Check
- No, this is not surprising. r measures only linear association. A curved pattern may produce a small r even though there is a strong nonlinear relationship.
- For each additional square foot, the predicted house price increases by 0.15 thousand dollars, or $150.
- Approximately 64% of the variation in reaction time is explained by the linear relationship with hours of sleep.
- An outlier in regression is a point with a large residual (far from the line). An influential point is one whose removal significantly changes the regression line. Influential points often have high leverage (extreme x-values).
- No. This is an observational study, and there may be lurking variables (e.g., income, lifestyle, access to healthcare) that affect both coffee consumption and longevity.
- A U-shaped pattern in the residual plot indicates that a linear model is not appropriate. The true relationship may be curved, and a nonlinear model would be better.
Unit 3: Collecting Data
A population is the entire group of individuals about which we want information. A sample is a subset of the population that we actually examine.
- A census attempts to collect data from every individual in the population. Census-taking is often impractical or too expensive.
- A parameter is a number that describes a population (e.g., population mean mu, population proportion p).
- A statistic is a number computed from a sample (e.g., sample mean x-bar, sample proportion p-hat).
The fundamental goal of statistics is to use sample statistics to draw conclusions about population parameters.
3.2 Sampling Methods
Simple Random Sample (SRS)
Every possible sample of size n from the population has an equal chance of being selected. Every individual also has an equal chance of being selected.
- How to do it: Assign a number to each individual, then use a random number generator (or table of random digits) to select n numbers.
- SRS eliminates selection bias and allows us to use probability theory.
Stratified Random Sampling
Divide the population into strata (homogeneous groups based on a characteristic), then take an SRS from each stratum.
- Purpose: To ensure that each subgroup is represented, especially when the strata differ substantially. This can reduce variability.
- Example: A school is divided by grade level (strata), and an SRS of students is taken from each grade.
Cluster Sampling
Divide the population into clusters (often naturally occurring groups), randomly select some clusters, and then survey all individuals in the selected clusters.
- Purpose: To reduce cost when a complete list of individuals is unavailable but a list of clusters is.
- Example: A city is divided into city blocks (clusters). Randomly select 20 blocks and survey every household in those blocks.
- Clusters should be heterogeneous (similar to the population as a whole), unlike strata which should be homogeneous.
Systematic Sampling
Select every k-th individual from a list after a random starting point between 1 and k.
- Example: From a list of 5000 customers, select every 50th customer starting at a random position between 1 and 50.
- Works well when the list has no periodic pattern.
Multistage Sampling
Combines several methods. For example, first stratify by state, then randomly select counties within each state, then take an SRS of individuals within each county.
Convenience Sampling
Select individuals who are easiest to reach. This is a non-probability method and is prone to bias.
Voluntary Response Sampling
Individuals choose to participate, often by responding to a survey. People with strong opinions are more likely to respond, leading to bias.
3.3 Sources of Bias
Sampling Bias
The method of selecting the sample systematically favors certain outcomes.
- Undercoverage: Some groups in the population are less likely to be included in the sample. Example: A telephone survey that only calls landlines will undercover younger people.
Response Bias
The way questions are worded or the setting of the survey influences responses.
- Leading questions: "Don't you agree that...?"
- Social desirability bias: Respondents answer in a way they think is socially acceptable rather than truthfully.
- Intimidation or poor wording: Confusing or loaded language.
Nonresponse Bias
When a large fraction of selected individuals fail to respond or cannot be contacted, and non-respondents differ systematically from respondents.
Key Principle
Bad sampling methods produce data that cannot be trusted, regardless of how large the sample is. A large biased sample is still biased.
3.4 Observational Studies vs. Experiments
Observational Study
The researcher observes individuals and measures variables of interest but does not attempt to influence the responses. Can demonstrate association but not causation.
Experiment
The researcher deliberately imposes a treatment on individuals and records the response. A well-designed experiment can demonstrate causation.
3.5 Principles of Experimental Design
Control
- Using a control group (a group that receives no treatment or a placebo) provides a baseline for comparison.
- Controlling other variables (holding them constant) prevents them from becoming confounding.
Random Assignment
- Individuals are randomly assigned to treatment groups.
- This helps create groups that are similar in all respects before the treatment is applied.
- Random assignment balances the effects of lurking variables across groups.
- This is what distinguishes an experiment from an observational study.
Replication
- Using enough experimental units in each group so that the results are reliable.
- Reproducing the study to confirm results.
3.6 Blinding
Single-Blind
- The subjects do not know which treatment they are receiving (but the experimenters do).
Double-Blind
- Neither the subjects nor the experimenters (who interact with the subjects) know which treatment is being administered.
- This is the gold standard for reducing bias in experiments.
- Purpose: To prevent the placebo effect (subjects) and experimenter bias (researchers).
3.7 Experimental Designs
Completely Randomized Design
- All subjects are randomly assigned to treatment groups.
- Steps: (1) List all subjects. (2) Use a random process to assign each subject to a treatment group. (3) Compare the responses.
Randomized Block Design
- Subjects are first divided into blocks based on a variable that is expected to affect the response.
- Within each block, subjects are randomly assigned to treatments.
- Purpose: To reduce variability by controlling for a known confounding variable.
- Example: In a drug trial, subjects are blocked by age group (young, middle-aged, elderly), then randomly assigned within each block.
Key difference from stratified sampling: In block design, the blocking variable is a response variable you want to control for. In stratified sampling, the stratifying variable is a predictor you want to ensure representation of.
Matched Pairs Design
- A special case of block design where each block contains exactly two subjects that are similar in important ways, and each receives a different treatment.
- Alternatively, each subject serves as their own control, receiving both treatments in random order (a repeated measures design).
- Example: Each participant tries both Brand A and Brand B shoelaces, with the order randomized.
3.8 Key Terminology
- Treatment: A specific condition applied to subjects in an experiment.
- Experimental unit: The smallest unit to which a treatment is applied (a person, animal, plot of land, etc.).
- Subject: A human experimental unit.
- Factor: An explanatory variable in an experiment.
- Level: A specific value of a factor.
- Placebo: An inactive treatment that looks like the real treatment.
- Placebo effect: The response that subjects show to a placebo.
- Confounding: When the effects of two variables on a response cannot be separated.
3.9 Worked Example
A researcher wants to test whether a new fertilizer increases tomato yield compared to the current fertilizer. She has 24 tomato plants.
Design: Completely randomized design.
- Randomly assign the 24 plants into two groups of 12.
- Apply the new fertilizer to Group 1 and the current fertilizer to Group 2.
- Keep all other conditions (water, sunlight, soil) the same for both groups.
- After the growing season, measure the yield of each plant and compare the group means.
If the researcher suspects that soil quality varies across the garden, she might use a randomized block design: group the plants by soil quality region, then randomly assign fertilizers within each block.
3.10 Common Mistakes
- Confusing random sampling with random assignment. Random sampling is about how you select individuals from a population. Random assignment is about how you assign selected individuals to treatment groups in an experiment. Both are important, but they serve different purposes.
- Claiming causation from an observational study. Only experiments with random assignment can support cause-and-effect conclusions.
- Confusing stratified sampling with cluster sampling. In stratified sampling, you sample from every stratum. In cluster sampling, you sample entire clusters.
- Forgetting to use a control group. Without a control group, you cannot determine whether the treatment caused the observed effect.
- Describing an experiment but omitting random assignment. If subjects are not randomly assigned to treatments, it is not a valid experiment for establishing causation.
- Confusing blinding with randomization. Blinding prevents bias after assignment. Randomization creates comparable groups before treatment.
3.11 Self-Check Questions
- A school wants to survey 100 students about lunch preferences. Describe how you would obtain a stratified random sample by grade level.
- A health study finds that people who exercise regularly have lower blood pressure. Can we conclude exercise causes lower blood pressure? Explain.
- Explain the difference between a completely randomized design and a randomized block design.
- A phone survey only calls numbers from a directory of landline telephones. Identify the type of bias and explain its effect.
- Why is double-blinding preferred over single-blinding in a drug trial?
- A researcher randomly selects 50 schools from a state and then surveys all teachers in those schools. What sampling method is this?
Answers to Self-Check
- Divide the school population into four strata (9th, 10th, 11th, 12th grade). Determine how many students to sample from each grade (e.g., proportionally). Then take an SRS of the required size from each grade's student list.
- Not necessarily. This is an observational study. There may be lurking variables (diet, genetics, stress levels) that affect both exercise habits and blood pressure. Only a randomized experiment could establish causation.
- In a completely randomized design, all subjects are randomly assigned to treatments directly. In a randomized block design, subjects are first grouped into blocks based on a variable that may affect the response, and then randomly assigned to treatments within each block.
- This is undercoverage bias. People who only use cell phones (younger people, lower-income households) are systematically excluded from the sample, so the results may not represent the full population.
- Double-blinding prevents both the placebo effect (subjects) and experimenter bias (researchers who might unconsciously treat groups differently or interpret results differently).
- This is cluster sampling. Schools are the clusters. A random sample of clusters is selected, and all individuals within the selected clusters are surveyed.
Unit 4: Probability, Random Variables, and Probability Distributions
A probability experiment (or random phenomenon) is any process for which the outcome is uncertain.
- Sample space (S): The set of all possible outcomes.
- Event: A subset of the sample space. An event can be one outcome or a collection of outcomes.
Example: Rolling a standard die. The sample space is S = {1, 2, 3, 4, 5, 6}. The event "rolling an even number" is {2, 4, 6}.
Probability of an Event
- Probabilities are always between 0 and 1.
- P(S) = 1 (the sample space is certain).
- P(impossible event) = 0.
- If all outcomes are equally likely: P(A) = (number of outcomes in A) / (number of outcomes in S).
4.2 Basic Probability Rules
Complement Rule
P(A^c) = 1 - P(A)
The probability that event A does NOT occur equals 1 minus the probability that A does occur.
Addition Rule (General)
P(A or B) = P(A) + P(B) - P(A and B)
Subtract the intersection because it was counted twice (once in P(A) and once in P(B)).
Addition Rule for Mutually Exclusive (Disjoint) Events
If A and B cannot both occur (they share no outcomes): P(A or B) = P(A) + P(B)
Multiplication Rule (General)
**P(A and B) = P(A) P(B | A)*
Multiplication Rule for Independent Events
If A and B are independent (knowing A occurred does not change the probability of B): **P(A and B) = P(A) P(B)*
4.3 Conditional Probability
P(B | A) = P(A and B) / P(A)
This is the probability that event B occurs given that event A has already occurred.
Interpretation: If we know that A has happened, how likely is B?
Worked example: A bag contains 3 red and 5 blue marbles. Two marbles are drawn without replacement. What is the probability that both are red?
- P(first red) = 3/8
- P(second red | first red) = 2/7
- P(both red) = (3/8)(2/7) = 6/56 = 3/28
4.4 Independence
Events A and B are independent if P(B | A) = P(B). Equivalently, P(A and B) = P(A) * P(B).
Independence vs. Mutually Exclusive
These are different concepts:
- Independent: Knowing A occurred does not change the probability of B.
- Mutually exclusive: A and B cannot both occur.
If two events are mutually exclusive and both have positive probability, they cannot be independent. (If A occurs, B is impossible, so P(B | A) = 0, which is not equal to P(B) > 0.)
Independence must be judged by the data or the design of the experiment, not by intuition.
4.5 Two-Way Tables
Two-way tables organize data on two categorical variables. They allow us to compute:
- Joint probabilities: P(A and B) from the cell divided by the total.
- Marginal probabilities: P(A) from the row or column total divided by the grand total.
- Conditional probabilities: P(B | A) = P(A and B) / P(A).
Worked example:
| | Pass | Fail | Total | |---|---|---|---| | Studied | 70 | 10 | 80 | | Did not study | 15 | 25 | 40 | | Total | 85 | 35 | 120 |
- P(Pass) = 85/120 = 0.708
- P(Pass | Studied) = 70/80 = 0.875
- P(Studied | Pass) = 70/85 = 0.824
- P(Studied and Pass) = 70/120 = 0.583
To check independence: Is P(Pass | Studied) = P(Pass)? 0.875 ≠ 0.708. So passing and studying are not independent.
4.6 Tree Diagrams
Tree diagrams are visual tools for multi-stage probability problems. Each branch represents a possible outcome at a stage, and branches are labeled with their probabilities.
Rules:
- The probabilities on branches from the same node must sum to 1.
- To find the probability of a path through the tree, multiply the probabilities along the branches.
- To find the probability of any of several outcomes, add their probabilities.
4.7 Bayes' Theorem (Conceptual)
Bayes' theorem allows us to reverse a conditional probability. If we know P(B | A), we can find P(A | B) using:
**P(A | B) = P(B | A) P(A) / P(B)*
This is particularly useful in medical testing:
- Sensitivity = P(positive test | disease) — the probability the test is positive given the person has the disease.
- Specificity = P(negative test | no disease).
Key insight: Even a highly accurate test can produce mostly false positives if the disease is rare.
4.8 Random Variables
A random variable takes numerical values that describe the outcomes of a random phenomenon. We use capital letters (X, Y) for random variables and lowercase (x, y) for specific values.
Discrete Random Variables
- Take on a countable number of values (often integers).
- Examples: number of heads in 10 coin flips, number of customers in a store.
- Described by a probability distribution that lists each possible value and its probability.
- The probabilities must sum to 1.
Continuous Random Variables
- Take on any value in an interval.
- Examples: temperature, height, time.
- Described by a density curve. The probability of any single value is zero.
- Probability is found by calculating the area under the curve over an interval.
4.9 Mean (Expected Value) and Standard Deviation of a Random Variable
Mean (Expected Value)
**mu_X = E(X) = sum of [x P(X = x)]*
This is the long-run average value of the random variable over many repetitions.
Variance and Standard Deviation
**sigma_X^2 = sum of [(x - mu_X)^2 P(X = x)]* sigma_X = sqrt(sigma_X^2)
These measure the spread of the probability distribution.
Rules for Means
- mu_(a+X) = a + mu_X (adding a constant shifts the mean)
- **mu_(bX) = b mu_X* (multiplying by a constant scales the mean)
- mu_(X+Y) = mu_X + mu_Y (always, regardless of independence)
- mu_(X-Y) = mu_X - mu_Y
Rules for Variances
- sigma^2_(a+X) = sigma^2_X (adding a constant does not change variance)
- **sigma^2_(bX) = b^2 sigma^2_X* (multiplying by a constant scales variance by b squared)
- sigma^2_(X+Y) = sigma^2_X + sigma^2_Y (only if X and Y are independent)
- sigma^2_(X-Y) = sigma^2_X + sigma^2_Y (yes, PLUS — for differences, you still add variances)
4.10 Linear Combinations and Differences
If X and Y are random variables, and a and b are constants:
**mu_(aX + bY) = a mu_X + b mu_Y**
If X and Y are independent: **sigma^2_(aX + bY) = a^2 sigma^2_X + b^2 sigma^2_Y**
Worked example: A store's daily revenue X has mean $500 and standard deviation $80. Daily costs Y have mean $350 and standard deviation $50. Revenue and costs are independent. What are the mean and standard deviation of daily profit (X - Y)?
- mu_profit = mu_X - mu_Y = 500 - 350 = $150
- sigma^2_profit = sigma^2_X + sigma^2_Y = 80^2 + 50^2 = 6400 + 2500 = 8900
- sigma_profit = sqrt(8900) = $94.34
4.11 Common Mistakes
- Confusing P(A and B) with P(A | B). P(A | B) is a conditional probability that restricts the sample space. P(A and B) is the joint probability.
- Assuming independence without justification. Independence must be verified by checking whether P(B | A) = P(B), not assumed.
- Subtracting variances for differences. When computing the variance of X - Y, always ADD the variances: sigma^2_X + sigma^2_Y. Never subtract.
- Forgetting that probabilities must be between 0 and 1. If your calculation gives a probability outside this range, you made an error.
- Treating P(A | B) the same as P(B | A). These are generally different. P(has disease | positive test) is not the same as P(positive test | has disease).
- Not checking that a probability distribution sums to 1. Always verify this when constructing or checking a discrete probability distribution.
4.12 Self-Check Questions
- Events A and B are mutually exclusive. P(A) = 0.3 and P(B) = 0.5. Find P(A or B).
- A jar contains 4 red, 6 blue, and 5 green marbles. Two marbles are drawn without replacement. Find the probability that both are blue.
- X has the probability distribution: P(X=0) = 0.2, P(X=1) = 0.5, P(X=2) = 0.3. Find E(X).
- Two independent random variables X and Y have means 10 and 15 and variances 9 and 16, respectively. Find the mean and standard deviation of X + Y.
- In a two-way table, how do you determine whether two events are independent?
- A medical test is 95% accurate for both sensitivity and specificity. If 1% of the population has the disease, explain why most positive results may be false positives.
Answers to Self-Check
- Since A and B are mutually exclusive: P(A or B) = P(A) + P(B) = 0.3 + 0.5 = 0.8.
- P(first blue) = 6/15 = 2/5. P(second blue | first blue) = 5/14. P(both blue) = (2/5)(5/14) = 10/70 = 1/7.
- E(X) = 0(0.2) + 1(0.5) + 2(0.3) = 0 + 0.5 + 0.6 = 1.1.
- mu_(X+Y) = 10 + 15 = 25. sigma^2_(X+Y) = 9 + 16 = 25. sigma_(X+Y) = 5.
- Check whether P(B | A) = P(B) for all cells. If they are equal, the events are independent. Equivalently, check whether P(A and B) = P(A) * P(B).
- Even though the test is 95% accurate, the disease is rare (1% prevalence). P(positive test) = P(positive | disease)P(disease) + P(positive | no disease)P(no disease) = (0.95)(0.01) + (0.05)(0.99) = 0.0095 + 0.0495 = 0.059. So P(disease | positive) = 0.0095/0.059 = 0.161, or about 16%. Most positive results are false positives.
Unit 5: Sampling Distributions
A sampling distribution is the distribution of a statistic (like x-bar or p-hat) over all possible samples of a given size from a population.
Key idea: The value of a sample statistic varies from sample to sample. This variability is described by its sampling distribution.
- The sampling distribution of x-bar describes how the sample mean varies across all possible samples of size n.
- The sampling distribution of p-hat describes how the sample proportion varies across all possible samples of size n.
A sampling distribution is a theoretical concept — we do not actually take all possible samples. We use probability theory to describe what would happen if we did.
5.2 Parameters vs. Statistics
| Parameter | Statistic | |
|---|---|---|
| What it describes | Population | Sample |
| Mean | mu (unknown) | x-bar (computed) |
| Proportion | p (unknown) | p-hat (computed) |
| Standard deviation | sigma | s or sqrt(p(1-p)/n) |
We use sample statistics to estimate population parameters. The sampling distribution tells us how accurate that estimate is likely to be.
5.3 Sampling Distribution of the Sample Mean (x-bar)
Mean of the Sampling Distribution
mu_x-bar = mu
The mean of the sampling distribution of x-bar equals the population mean. This means x-bar is an unbiased estimator of mu.
Standard Deviation of the Sampling Distribution (Standard Error)
sigma_x-bar = sigma / sqrt(n)
This is called the standard error of x-bar. It measures how much x-bar typically varies from sample to sample. It decreases as the sample size n increases — larger samples give more precise estimates.
Shape of the Sampling Distribution
- If the population distribution is normal, then the sampling distribution of x-bar is normal for any sample size n.
- If the population distribution is not normal, the Central Limit Theorem tells us when x-bar is approximately normal.
The Central Limit Theorem (CLT)
For a large enough sample size n, the sampling distribution of x-bar is approximately normal, regardless of the shape of the population distribution.
- The rule of thumb is that n >= 30 is large enough for the CLT to apply.
- If the population is clearly non-normal, a larger n may be needed.
- If the population is already normal, the CLT applies for any n (even n = 1).
The CLT is one of the most important theorems in statistics. It is the reason we can use normal-based inference for means even when we do not know the population distribution.
5.4 Sampling Distribution of the Sample Proportion (p-hat)
Mean of the Sampling Distribution
mu_p-hat = p
The mean of the sampling distribution of p-hat equals the true population proportion. p-hat is an unbiased estimator of p.
Standard Deviation of the Sampling Distribution
sigma_p-hat = sqrt(p(1-p) / n)
This is called the standard error of p-hat.
Conditions for Normality
The sampling distribution of p-hat is approximately normal when:
- np >= 10 (expected number of successes)
- n(1-p) >= 10 (expected number of failures)
These are the Large Counts conditions. In practice, since p is unknown, we check these using p-hat.
5.5 The Law of Large Numbers
As the sample size increases, the sample statistic (x-bar or p-hat) gets closer to the population parameter (mu or p).
This is different from the CLT:
- Law of Large Numbers: What happens as n grows (the statistic approaches the parameter).
- CLT: What the distribution of the statistic looks like for a fixed, sufficiently large n (approximately normal).
Example: If you flip a fair coin 10 times, p-hat might be 0.3 or 0.7. If you flip it 10,000 times, p-hat will be very close to 0.5.
5.6 Worked Examples
Example 1: Sampling Distribution of x-bar
A population has mu = 50 and sigma = 12. We take samples of size n = 36.
- mu_x-bar = 50
- sigma_x-bar = 12 / sqrt(36) = 12/6 = 2
- Since n = 36 >= 30, the CLT says x-bar is approximately normal.
What is the probability that x-bar > 53?
- z = (53 - 50) / 2 = 1.5
- P(x-bar > 53) = P(Z > 1.5) = 1 - 0.9332 = 0.0668
Example 2: Sampling Distribution of p-hat
Suppose 60% of all voters support a ballot measure (p = 0.60). A random sample of 200 voters is selected.
- mu_p-hat = 0.60
- sigma_p-hat = sqrt(0.60 * 0.40 / 200) = sqrt(0.24/200) = sqrt(0.0012) = 0.0346
- Large counts: 200(0.60) = 120 >= 10 and 200(0.40) = 80 >= 10. Normal approximation is valid.
What is the probability that p-hat > 0.65?
- z = (0.65 - 0.60) / 0.0346 = 0.05 / 0.0346 = 1.44
- P(p-hat > 0.65) = P(Z > 1.44) = 1 - 0.9251 = 0.0749
5.7 Common Mistakes
- Confusing the distribution of the population with the sampling distribution. The population distribution describes individual values. The sampling distribution describes the statistic computed from a sample.
- Confusing the Law of Large Numbers with the CLT. The LLN says the statistic approaches the parameter as n grows. The CLT says the distribution of the statistic is approximately normal for large n.
- Forgetting to divide sigma by sqrt(n). The standard deviation of x-bar is sigma/sqrt(n), not sigma. Averages are less variable than individual observations.
- Using the CLT when the population is strongly skewed and n < 30. If the population distribution is highly non-normal, n >= 30 may not be enough. More data or transformations may be needed.
- Not checking the Large Counts conditions for proportions. Before using the normal approximation for p-hat, verify that np >= 10 and n(1-p) >= 10.
- Thinking that a larger sample reduces bias. A larger sample reduces variability (standard error), not bias. A biased sampling method remains biased regardless of sample size.
5.8 Self-Check Questions
- A population has mu = 100 and sigma = 20. For a sample of size n = 64, what is the mean and standard deviation of the sampling distribution of x-bar?
- Explain the Central Limit Theorem in your own words.
- True or false: A larger sample size makes the sampling distribution of x-bar more normal. Explain.
- The true proportion of left-handed students at a school is 0.15. A random sample of 100 students is selected. What is the probability that the sample proportion of left-handed students exceeds 0.20?
- How does the Law of Large Numbers differ from the Central Limit Theorem?
- A researcher takes a sample of size n = 10 from a population that is strongly right-skewed. Can they use the CLT? Explain.
Answers to Self-Check
- mu_x-bar = 100. sigma_x-bar = 20/sqrt(64) = 20/8 = 2.5.
- The CLT states that for a sufficiently large sample size (typically n >= 30), the sampling distribution of the sample mean is approximately normal, regardless of the shape of the population distribution.
- False — or misleading. A larger sample does not change the normality of the sampling distribution; it reduces the spread (standard error). If the population is normal, the sampling distribution is normal for any n. If not, the CLT says it is approximately normal for large enough n.
- sigma_p-hat = sqrt(0.15*0.85/100) = sqrt(0.1275/100) = sqrt(0.001275) = 0.0357. z = (0.20 - 0.15)/0.0357 = 1.40. P(Z > 1.40) = 0.0808.
- The Law of Large Numbers says that as n grows, the statistic gets closer to the parameter. The CLT describes the shape of the sampling distribution for a fixed, sufficiently large n.
- Not reliably. With n = 10 and a strongly skewed population, the CLT may not apply. The sample size is too small to guarantee approximate normality of x-bar.
Unit 6: Inference for Categorical Data: Proportions
A confidence interval gives a range of plausible values for a population parameter, along with a confidence level that indicates how often the method produces intervals containing the true parameter.
One-Sample z-Interval for a Proportion
**p-hat +/- z sqrt(p-hat(1 - p-hat) / n)**
- Point estimate: p-hat = x/n (the sample proportion)
- Margin of error: z sqrt(p-hat(1 - p-hat) / n)
- z: The critical value from the standard normal distribution for the desired confidence level (e.g., z = 1.96 for 95% confidence)
- Standard error: sqrt(p-hat(1 - p-hat) / n)
Conditions
- Random: The data must come from a random sample or randomized experiment.
- Normal (Large Counts): np-hat >= 10 and n(1 - p-hat) >= 10.
- Independent (10% condition): The sample size must be no more than 10% of the population when sampling without replacement (n <= 0.10N).
Interpreting a Confidence Interval
"We are [confidence level]% confident that the true proportion of [population] that [statement about parameter] is between [lower bound] and [upper bound]."
Interpreting Confidence Level
A 95% confidence level means that if we were to take many random samples and construct a 95% confidence interval from each, about 95% of those intervals would contain the true population proportion.
Key distinction: The confidence level is about the method, not about any particular interval. We cannot say there is a 95% probability that the true proportion is in a specific interval — it either is or is not.
6.2 Hypothesis Test for a Proportion
Setting Up Hypotheses
- Null hypothesis (H0): p = p0 (the population proportion equals a specific value).
- Alternative hypothesis (Ha): p > p0, p < p0, or p != p0 (one-sided or two-sided).
One-Sample z-Test for a Proportion
Test statistic: z = (p-hat - p0) / sqrt(p0(1 - p0) / n)
- Use p0 (from H0) in the standard error, NOT p-hat.
- Compute the p-value: the probability of observing a test statistic as extreme or more extreme than the one calculated, assuming H0 is true.
Making a Decision
- If p-value <= alpha (significance level), reject H0. There is convincing evidence for Ha.
- If p-value > alpha, fail to reject H0. There is not convincing evidence against H0.
Never say "accept H0." We either reject it or fail to reject it.
Conditions
- Random
- Normal: np0 >= 10 and n(1 - p0) >= 10 (use p0 here, not p-hat).
- Independent: 10% condition.
6.3 Type I and Type II Errors
| Decision | H0 is True | H0 is False |
|---|---|---|
| Reject H0 | Type I Error (False Positive) | Correct Decision (Power) |
| Fail to Reject H0 | Correct Decision | Type II Error (False Negative) |
- Type I error: Rejecting H0 when it is actually true. The probability of a Type I error equals alpha, the significance level.
- Type II error: Failing to reject H0 when it is actually false. The probability is denoted beta.
Power
Power = 1 - beta = the probability of correctly rejecting a false H0.
Power increases when:
- alpha is increased (willing to accept more Type I error risk)
- The true parameter is farther from H0 (larger effect size)
- The sample size n is increased
- Standard deviation is decreased
6.4 Two-Sample Inference for Proportions
Two-Sample z-Interval for the Difference of Proportions
Confidence interval: (p-hat_1 - p-hat_2) +/- z sqrt(p-hat_1(1 - p-hat_1)/n1 + p-hat_2(1 - p-hat_2)/n2)
Two-Sample z-Test for the Difference of Proportions
Test statistic: z = ((p-hat_1 - p-hat_2) - 0) / sqrt(p-hat_pooled(1 - p-hat_pooled)(1/n1 + 1/n2))
Where the pooled proportion is: p-hat_pooled = (x1 + x2) / (n1 + n2)
Conditions
- Random: Both samples are random (or from randomized experiments).
- Normal: All four counts (n1p-hat_1, n1(1-p-hat_1), n2p-hat_2, n2(1-p-hat_2)) are at least 5 for a test (or 10 for a confidence interval).
- Independent: Both samples are independent of each other, and each satisfies the 10% condition.
6.5 Worked Example
A survey of 500 randomly selected adults found that 280 support a new policy. Construct and interpret a 95% confidence interval for the proportion of all adults who support the policy.
Step 1: Check conditions.
- Random: Random sample. Checked.
- 10% condition: 500 is less than 10% of all adults. Checked.
- Large counts: 500(280/500) = 280 >= 10 and 500(220/500) = 220 >= 10. Checked.
Step 2: Calculate.
- p-hat = 280/500 = 0.56
- Standard error = sqrt(0.56 * 0.44 / 500) = sqrt(0.2464/500) = sqrt(0.000493) = 0.0222
- For 95% confidence, z* = 1.96
- Margin of error = 1.96 * 0.0222 = 0.0435
- CI: 0.56 +/- 0.0435 = (0.5165, 0.6035)
Step 3: Interpret. We are 95% confident that the true proportion of all adults who support the new policy is between 0.517 and 0.604.
Hypothesis Test Example
Test whether a majority of adults support the policy. Use alpha = 0.05.
- H0: p = 0.50
- Ha: p > 0.50
- Test statistic: z = (0.56 - 0.50) / sqrt(0.50*0.50/500) = 0.06 / sqrt(0.25/500) = 0.06 / 0.0224 = 2.68
- p-value = P(Z > 2.68) = 0.0037
- Since 0.0037 < 0.05, we reject H0.
- Conclusion: There is convincing evidence that a majority of adults support the new policy.
6.6 Common Mistakes
- Using p-hat instead of p0 in the test statistic. For a hypothesis test, the standard error uses p0 from H0. For a confidence interval, use p-hat.
- Saying "accept H0" instead of "fail to reject H0." Failing to reject does not prove H0 is true.
- Confusing the confidence level with the probability the parameter is in the interval. The 95% refers to the long-run performance of the method.
- Forgetting to check conditions. Always state and verify all conditions before performing inference.
- Misinterpreting the p-value. The p-value is NOT the probability that H0 is true. It is the probability of obtaining data as extreme as (or more extreme than) what was observed, assuming H0 is true.
- Forgetting the 10% condition. When sampling without replacement, the sample must be no more than 10% of the population.
6.7 Self-Check Questions
- A random sample of 200 students finds that 130 prefer online learning. Construct a 90% confidence interval for the true proportion and interpret it.
- A company claims that 90% of its products are defect-free. In a random sample of 400 products, 340 are defect-free. Test the company's claim at the 5% significance level.
- Explain the difference between a Type I error and a Type II error in the context of a medical screening test.
- Why do we use the pooled proportion in a two-sample z-test but not in a two-sample z-interval?
- A 99% confidence interval for a proportion is (0.32, 0.48). Can we conclude that exactly half the population has the characteristic? Explain.
- How does increasing the sample size affect the margin of error?
Answers to Self-Check
- p-hat = 130/200 = 0.65. SE = sqrt(0.650.35/200) = 0.0337. z = 1.645. ME = 1.645(0.0337) = 0.0555. CI = (0.595, 0.706). We are 90% confident that the true proportion of students who prefer online learning is between 0.595 and 0.706.
- H0: p = 0.90. Ha: p < 0.90. p-hat = 340/400 = 0.85. z = (0.85 - 0.90) / sqrt(0.90*0.10/400) = -0.05/0.015 = -3.33. p-value = P(Z < -3.33) = 0.0004. Since 0.0004 < 0.05, reject H0. There is convincing evidence that less than 90% of products are defect-free.
- Type I error: The test says the person has the disease when they do not (false positive). Type II error: The test says the person does not have the disease when they actually do (false negative).
- The test uses the pooled proportion because H0 assumes p1 = p2, so we use a common estimate. The interval estimates the true difference, so we use each sample's own proportion.
- No. Since 0.50 falls within the 99% confidence interval (0.32, 0.48) — actually it does not, 0.50 > 0.48 — so 0.50 is NOT in the interval. There is convincing evidence that the true proportion is less than 0.50. (Correction: 0.50 is above 0.48, so it is not in the interval.)
- Increasing the sample size decreases the margin of error because the standard error (which contains 1/sqrt(n)) decreases. Larger samples give more precise estimates.
Unit 7: Inference for Quantitative Data: Means
When we perform inference for a population mean, we rarely know the population standard deviation (sigma). Instead, we estimate it using the sample standard deviation s. This introduces additional uncertainty, which is captured by using the t-distribution instead of the standard normal (z) distribution.
Properties of the t-Distribution
- The t-distribution is symmetric and bell-shaped, like the standard normal.
- It has heavier tails than the standard normal, meaning it accounts for the extra variability from estimating sigma with s.
- As the degrees of freedom (df) increase, the t-distribution approaches the standard normal.
- For df >= 30, the t and z distributions are very similar.
- There is a different t-distribution for each value of df.
Degrees of Freedom
For a one-sample t-procedure: df = n - 1
The degrees of freedom reflect the amount of information available to estimate sigma. A larger n gives more df, making the t-distribution closer to the standard normal.
Finding t* Values
On a TI-84: Use invT(area, df) for a one-sided probability. For a 95% CI with df = 24: invT(0.975, 24) gives t* = 2.064.
7.2 One-Sample t-Interval for a Mean
**x-bar +/- t (s / sqrt(n))**
- Point estimate: x-bar
- Standard error: s / sqrt(n)
- Critical value: t* from the t-distribution with df = n - 1
- Margin of error: t (s / sqrt(n))
Conditions
- Random: The data must come from a random sample or randomized experiment.
- Normal (Nearly Normal): The population distribution should be approximately normal, or the sample size should be large (n >= 30 for the CLT to apply). For small samples, check a dotplot, histogram, or boxplot for strong skewness or outliers.
- Independent (10% condition): n <= 0.10N when sampling without replacement.
Interpretation
"We are [confidence level]% confident that the true mean [variable] for [population] is between [lower bound] and [upper bound] [units]."
7.3 One-Sample t-Test for a Mean
Hypotheses
- H0: mu = mu0
- Ha: mu > mu0, mu < mu0, or mu != mu0
Test Statistic
t = (x-bar - mu0) / (s / sqrt(n))
This follows a t-distribution with df = n - 1 under H0.
p-Value
Use tcdf(lower, upper, df) on a TI-84 to find the p-value.
Conditions (same as the t-interval)
- Random
- Nearly Normal
- Independent (10% condition)
7.4 Paired t-Test and t-Interval
A paired design involves two measurements on the same subject or matched pairs of subjects. The key step is to compute the differences for each pair, then perform a one-sample t-procedure on the differences.
When to Use Paired Methods
- Before and after measurements on the same individuals.
- Matched pairs (e.g., identical twins, left and right hands).
- When the design naturally pairs observations.
Procedure
- Compute the difference d for each pair (e.g., d = before - after).
- Treat the d-values as a single sample and compute d-bar and s_d.
- Perform a one-sample t-interval or t-test on the differences.
- df = number of pairs - 1
Conditions
- Random (random assignment or random sampling of pairs)
- Nearly Normal for the differences (not the original values)
- Independent pairs (10% condition)
7.5 Two-Sample t-Interval and t-Test
Two-Sample t-Interval
**(x-bar_1 - x-bar_2) +/- t sqrt(s_1^2/n_1 + s_2^2/n_2)**
- Point estimate: x-bar_1 - x-bar_2
- Standard error: sqrt(s_1^2/n_1 + s_2^2/n_2)
- Critical value: Use the smaller of n1-1 and n2-1 for a conservative approach, or use technology for a more precise df (Welch's formula).
Two-Sample t-Test
t = ((x-bar_1 - x-bar_2) - 0) / sqrt(s_1^2/n_1 + s_2^2/n_2)
Conditions
- Random: Both samples are random.
- Normal: Both population distributions should be approximately normal, or both samples should be large enough (n1 >= 30 and n2 >= 30).
- Independent: The two samples must be independent of each other, and each must satisfy the 10% condition.
Important Notes
- Do not pool the standard deviations unless the problem states the population standard deviations are equal (this is rarely done in AP Statistics).
- The two-sample t-procedure compares the means of two independent groups.
- If the groups are paired, use a paired t-test instead.
7.6 Choosing the Right Test
| Scenario | Procedure |
|---|---|
| One sample, mean, sigma unknown | One-sample t-test/t-interval |
| Two measurements on same subjects | Paired t-test/t-interval |
| Two independent groups, comparing means | Two-sample t-test/t-interval |
Key question: Are the two groups independent or paired? If the same subjects are measured twice, or if subjects are naturally paired, use the paired procedure.
7.7 Confidence Intervals vs. Hypothesis Tests
- A confidence interval gives a range of plausible values for the parameter. If the null value (e.g., 0 for a difference or mu0) falls outside the interval at the corresponding confidence level, we would reject H0 at the corresponding alpha.
- A hypothesis test gives a yes/no decision about whether there is evidence against H0.
- A CI provides more information than a test because it shows the magnitude and direction of the effect.
- If a two-sided test rejects H0 at alpha = 0.05, then a 95% CI will not contain the null value.
7.8 Worked Example
A researcher measures the fuel efficiency (in mpg) of 12 randomly selected cars of a new model: 28, 31, 29, 30, 27, 32, 30, 28, 31, 29, 33, 30.
Construct a 95% confidence interval for the mean fuel efficiency.
Using a calculator: x-bar = 29.917, s = 1.73, n = 12, df = 11.
- Standard error = 1.73 / sqrt(12) = 0.499
- t* for 95% CI with df = 11:
invT(0.975, 11)= 2.201 - Margin of error = 2.201 * 0.499 = 1.098
- CI: 29.917 +/- 1.098 = (28.819, 31.015)
Interpretation: We are 95% confident that the true mean fuel efficiency for this car model is between 28.8 and 31.0 mpg.
Test whether the mean fuel efficiency differs from 30 mpg at alpha = 0.05.
- H0: mu = 30, Ha: mu != 30
- t = (29.917 - 30) / (1.73/sqrt(12)) = -0.083 / 0.499 = -0.166
- p-value = 2 P(T < -0.166) with df = 11 = 2 0.4356 = 0.871
- Since 0.871 > 0.05, fail to reject H0. There is not convincing evidence that the mean fuel efficiency differs from 30 mpg.
Note that 30 is inside the 95% CI, which is consistent with failing to reject H0 at alpha = 0.05.
7.9 Common Mistakes
- Using z instead of t when sigma is unknown. In practice, sigma is almost never known. Use the t-distribution.
- Confusing paired and independent samples. If the same subjects are measured twice, you must use paired procedures, not two-sample procedures.
- Checking normality of individual groups instead of differences for paired data. For paired data, check the normality of the differences.
- Using the wrong degrees of freedom. For one-sample: n - 1. For two-sample: use the smaller of n1-1 and n2-1 (or technology).
- Forgetting the 10% condition. When sampling without replacement, n must be no more than 10% of the population.
- Saying "mu = 30" without context in the conclusion. Always state conclusions in the context of the problem.
7.10 Self-Check Questions
- A sample of 25 students has a mean study time of 14.2 hours per week with s = 3.8 hours. Construct a 95% confidence interval for the true mean study time.
- Explain why we use the t-distribution instead of the z-distribution for inference about means.
- A teacher wants to know whether a new teaching method changes test scores. She gives a pre-test and post-test to the same 15 students. What procedure should she use?
- Two independent random samples: Group A (n=30, x-bar=45, s=8) and Group B (n=35, x-bar=40, s=7). Construct a 95% confidence interval for the difference in means.
- What conditions must be checked before performing a one-sample t-test?
- If a 95% CI for mu is (22.5, 28.3), would a two-sided t-test with H0: mu = 25 reject H0 at alpha = 0.05? Explain.
Answers to Self-Check
- df = 24, t* = 2.064. SE = 3.8/sqrt(25) = 0.76. ME = 2.064(0.76) = 1.569. CI = (12.63, 15.77). We are 95% confident that the true mean study time is between 12.63 and 15.77 hours per week.
- We use t because we do not know the population standard deviation sigma. We estimate it with s, which adds uncertainty. The t-distribution has heavier tails to account for this extra variability.
- A paired t-test. The pre-test and post-test scores are paired (two measurements on the same students). Compute the difference (post - pre) for each student and run a one-sample t-test on the differences.
- Point estimate = 45 - 40 = 5. SE = sqrt(64/30 + 49/35) = sqrt(2.133 + 1.4) = sqrt(3.533) = 1.880. df = min(29, 34) = 29, t* = 2.045. ME = 2.045(1.880) = 3.845. CI = (1.155, 8.845).
- Random (or random assignment), Normal (population approximately normal or n >= 30), and Independent (10% condition).
- No. Since 25 is within the 95% CI of (22.5, 28.3), a two-sided test at alpha = 0.05 would fail to reject H0. The p-value would be greater than 0.05.
Unit 8: Inference for Categorical Data: Chi-Square
Chi-square tests are used for categorical data to test whether observed counts differ significantly from expected counts. There are three types:
- Goodness of Fit (GOF): One categorical variable. Tests whether the distribution of a single variable matches a specified distribution.
- Independence: Two categorical variables from one sample. Tests whether the variables are associated.
- Homogeneity: Two categorical variables from multiple populations/samples. Tests whether the distribution of one variable is the same across populations.
All three use the same test statistic and the same chi-square distribution.
8.2 The Chi-Square Statistic
X^2 = sum of [(observed - expected)^2 / expected]
- Summed over all cells.
- Observed (O): The actual counts from the data.
- Expected (E): The counts we would expect if H0 were true.
- The chi-square statistic is always non-negative.
- Larger values of X^2 provide more evidence against H0.
- The chi-square distribution is right-skewed and depends on degrees of freedom.
8.3 Chi-Square Goodness of Fit Test
Purpose
To determine whether a single categorical variable follows a specified distribution.
Hypotheses
- H0: The variable follows the specified distribution (the proportions are as stated).
- Ha: The variable does not follow the specified distribution.
Expected Counts
For each category: **E_i = n p_i*, where n is the total sample size and p_i is the hypothesized proportion for category i.
Degrees of Freedom
df = number of categories - 1
Conditions
- Random: Data from a random sample or random assignment.
- Expected counts: All expected counts must be at least 5. (If any E < 5, combine categories or collect more data.)
- Independent (10% condition): n <= 0.10N.
Worked Example
A company claims its candy bags contain 30% red, 20% yellow, 20% green, 15% orange, and 15% purple candies. A random sample of 200 candies gives these counts:
| Color | Observed | Expected |
|---|---|---|
| Red | 55 | 200(0.30) = 60 |
| Yellow | 38 | 200(0.20) = 40 |
| Green | 50 | 200(0.20) = 40 |
| Orange | 30 | 200(0.15) = 30 |
| Purple | 27 | 200(0.15) = 30 |
X^2 = (55-60)^2/60 + (38-40)^2/40 + (50-40)^2/40 + (30-30)^2/30 + (27-30)^2/30 = 25/60 + 4/40 + 100/40 + 0 + 9/30 = 0.417 + 0.100 + 2.500 + 0 + 0.300 = 3.317
df = 5 - 1 = 4
Using a calculator: X^2cdf(3.317, 1E99, 4) gives p-value = 0.506.
Since 0.506 > 0.05, we fail to reject H0. There is not convincing evidence that the candy distribution differs from the company's claim.
8.4 Chi-Square Test for Independence
Purpose
To determine whether two categorical variables are associated (related) in a single population.
Data Structure
A two-way table of observed counts from one sample, with individuals classified on both variables.
Hypotheses
- H0: There is no association between the two variables (they are independent).
- Ha: There is an association between the two variables (they are not independent).
Expected Counts
For each cell: **E = (row total column total) / grand total*
Degrees of Freedom
**df = (number of rows - 1) (number of columns - 1)*
Conditions
- Random: Random sample.
- Expected counts: All expected counts >= 5.
- Independent (10% condition): n <= 0.10N.
8.5 Chi-Square Test for Homogeneity
Purpose
To determine whether the distribution of one categorical variable is the same across several populations or groups.
Data Structure
A two-way table of observed counts from separate random samples, one from each population.
Hypotheses
- H0: The distribution of [response variable] is the same for all populations.
- Ha: The distribution of [response variable] is not the same for all populations.
Expected Counts and Test Statistic
Same as for independence: E = (row total * column total) / grand total.
Degrees of Freedom
Same as for independence: df = (rows - 1)(columns - 1).
Conditions
Same as independence, but the 10% condition must hold for each sample separately.
8.6 Independence vs. Homogeneity: How to Tell
| Feature | Independence | Homogeneity |
|---|---|---|
| Number of samples | One sample | Multiple samples (one per population) |
| Question | Are these two variables related? | Is the distribution the same across groups? |
| H0 wording | No association between variables | Same distribution across populations |
The calculations are identical. The difference is in the study design and how you word the hypotheses and conclusions.
8.7 Calculator Commands (TI-84)
- Enter the observed counts in a matrix:
2nd > MATRIX > EDIT. - For GOF: Use
X^2GOF-Test(if available) or compute manually. - For independence/homogeneity:
STAT > TESTS > X^2-Test. Enter observed matrix (e.g., [A]), choose an output matrix (e.g., [B]) for expected counts. - The calculator reports X^2, p-value, and df.
8.8 Common Mistakes
- Confusing independence and homogeneity. Independence: one sample, two variables. Homogeneity: multiple samples, one variable compared across groups.
- Using observed counts in the denominator instead of expected counts. The formula is (O-E)^2/E, NOT (O-E)^2/O.
- Using proportions instead of counts. Chi-square tests use counts, not percentages or proportions.
- Forgetting to check that all expected counts are at least 5. If any E < 5, the chi-square test may not be valid.
- Stating the conclusion incorrectly for GOF. For GOF, say "the distribution does/does not match the specified distribution," not "the variables are/are not independent."
- Confusing the chi-square distribution with the normal or t-distribution. Chi-square is always right-skewed, and the test is always one-sided (large X^2 is evidence against H0).
8.9 Self-Check Questions
- A die is rolled 60 times with the following results: 1(8), 2(12), 3(10), 4(11), 5(9), 6(10). Test whether the die is fair at alpha = 0.05.
- Explain the difference between a chi-square test for independence and a chi-square test for homogeneity.
- A researcher studies whether gender (male/female) is associated with preferred color (red/blue/green) among 300 randomly selected adults. The chi-square test gives X^2 = 4.82 with df = 2. The p-value is 0.090. Interpret this result.
- Why must all expected counts be at least 5 for a chi-square test?
- For a 3x4 contingency table, what are the degrees of freedom?
- In a chi-square test, can the p-value ever be large (e.g., > 0.50)? What would that mean?
Answers to Self-Check
- H0: The die is fair (each face has probability 1/6). E for each face = 60(1/6) = 10. X^2 = (8-10)^2/10 + (12-10)^2/10 + (10-10)^2/10 + (11-10)^2/10 + (9-10)^2/10 + (10-10)^2/10 = 4/10 + 4/10 + 0 + 1/10 + 1/10 + 0 = 1.0. df = 5. p-value = X^2cdf(1.0, 1E99, 5) = 0.963. Fail to reject H0. No convincing evidence the die is unfair.
- Independence tests whether two categorical variables are related in a single population. Homogeneity tests whether the distribution of a categorical variable is the same across multiple populations.
- Since the p-value (0.090) is greater than 0.05, we fail to reject H0. There is not convincing evidence of an association between gender and preferred color.
- The chi-square approximation to the distribution of the test statistic is poor when expected counts are small. The test may produce inflated Type I error rates.
- df = (3-1)(4-1) = 2(3) = 6.
- Yes. A large p-value means the observed data are very consistent with H0. The counts are close to what was expected. There is no evidence against the null hypothesis.
Unit 9: Inference for Quantitative Data: Slopes
In Unit 2, we learned how to compute the least-squares regression line and interpret its slope. In Unit 9, we go further: we perform inference about the true slope of the population regression line.
The LSRL computed from a sample is an estimate of the true (population) regression line: **mu_y = alpha + beta x*
Our sample regression line y-hat = a + bx estimates the true slope beta using b, and the true intercept alpha using a.
9.2 Conditions for Regression Inference
Before performing any inference about the regression slope, check these conditions:
- Linear: The relationship between x and y must be linear. Check a scatterplot of y vs. x.
- Normal: For each x-value, the responses (y-values) are normally distributed around the regression line. Check a normal probability plot of the residuals.
- Independent: Individual observations must be independent. Check the 10% condition if sampling without replacement, and ensure no repeated measures or paired data.
- Equal Variance (Homoscedasticity): The variability of y-values around the regression line should be roughly the same for all x-values. Check a residual plot for a "fan" shape or other pattern.
- Random: The data must come from a random sample or randomized experiment.
These conditions are often remembered as LINE (Linear, Normal, Independent, Equal variance) plus Random.
How to Check
- Linear: Scatterplot should show a linear pattern.
- Normal: Histogram or normal probability plot of residuals should be approximately normal. For small samples, this is harder to verify.
- Equal variance: Residual plot should show roughly equal spread across all x-values.
- Independent: Verify the sampling method and 10% condition.
9.3 Standard Error of the Slope
The standard error of the slope (also called the standard error of the regression slope) estimates the standard deviation of the sampling distribution of b:
**SE_b = s / (sx sqrt(n - 1))*
Where:
- s is the residual standard error (also called the standard error of the estimate, or "s")
- sx is the standard deviation of the x-values
- n is the sample size
Residual Standard Error
s = sqrt(sum of (residuals)^2 / (n - 2))
This estimates the typical distance of data points from the regression line. We divide by n - 2 because we estimated two parameters (slope and intercept).
On a TI-84, the regression output reports this as "s".
9.4 t-Test for the Slope
Purpose
To test whether there is a statistically significant linear relationship between x and y.
Hypotheses
- H0: beta = 0 (there is no linear relationship between x and y)
- Ha: beta > 0, beta < 0, or beta != 0
Test Statistic
t = (b - 0) / SE_b
This follows a t-distribution with df = n - 2.
p-Value
Use tcdf on the calculator or read from the regression output.
Interpretation
- If we reject H0: "There is convincing evidence of a linear association between [x] and [y]."
- If we fail to reject H0: "There is not convincing evidence of a linear association between [x] and [y]."
9.5 Confidence Interval for the Slope
**b +/- t SE_b**
- t* is the critical value from the t-distribution with df = n - 2.
- If the interval contains 0, we would fail to reject H0: beta = 0 at the corresponding alpha.
- If the interval does not contain 0, we would reject H0.
Interpretation
"We are [confidence level]% confident that the true slope of the regression line relating [y] to [x] is between [lower] and [upper] [units per unit of x]."
9.6 Prediction Intervals
A prediction interval gives a range of values for an individual future observation of y at a specific x-value.
**y-hat +/- t sqrt(s^2 + SE_y-hat^2)**
Where SE_y-hat accounts for the uncertainty in estimating the mean response.
Key points:
- A prediction interval is wider than a confidence interval for the mean response because it includes both the uncertainty in the mean estimate AND individual variability.
- Prediction intervals get wider as x moves away from x-bar (extrapolation).
Note: AP exam questions on prediction intervals are relatively rare, but you should understand the concept.
9.7 Worked Example
A researcher studies the relationship between hours of exercise per week (x) and resting heart rate (y) in a random sample of 16 adults.
Regression output:
- a = 78.5, b = -1.2, s = 4.3, SE_b = 0.35, r^2 = 0.72
t-Test for the Slope
H0: beta = 0, Ha: beta != 0
- t = (-1.2 - 0) / 0.35 = -3.43
- df = 16 - 2 = 14
- p-value = 2 tcdf(-1E99, -3.43, 14) = 2 0.002 = 0.004
- Since 0.004 < 0.05, reject H0.
- Conclusion: There is convincing evidence of a negative linear association between weekly exercise hours and resting heart rate.
Confidence Interval for the Slope
- t* for 95% CI with df = 14:
invT(0.975, 14)= 2.145 - ME = 2.145 * 0.35 = 0.751
- CI: -1.2 +/- 0.751 = (-1.951, -0.449)
- We are 95% confident that for each additional hour of exercise per week, the true mean decrease in resting heart rate is between 0.45 and 1.95 beats per minute.
9.8 Calculator Tips (TI-84)
- Enter x-values in L1, y-values in L2.
STAT > TESTS > LinRegTTestperforms the t-test for the slope and gives b, SE_b, t, p-value, df, s, and r^2.- For a CI for the slope, you may need to compute it manually using b, SE_b, and t* from
invT.
9.9 Common Mistakes
- Forgetting to check the LINE conditions. Always check the scatterplot (linearity), residual plot (equal variance), and normal probability plot of residuals (normality) before performing inference.
- Confusing r^2 with the slope. r^2 measures the proportion of variation explained. The slope measures the rate of change. They are different statistics.
- Using z instead of t. Inference for the slope uses the t-distribution with df = n - 2, not the standard normal.
- Misinterpreting "fail to reject H0: beta = 0." This does not prove there is no relationship. It means the sample does not provide convincing evidence of a linear relationship. There could be a nonlinear relationship or the sample may be too small.
- Confusing a CI for the mean response with a prediction interval. A CI for the mean response estimates the average y at a given x. A prediction interval estimates an individual y at a given x. Prediction intervals are wider.
- Extrapolating beyond the data range when interpreting the slope. The slope interpretation only applies within the range of the observed data.
9.10 Self-Check Questions
- A regression of plant height (cm) on fertilizer amount (grams) for 20 plants gives b = 3.5 with SE_b = 1.1. Test whether there is a positive linear relationship at alpha = 0.05.
- For the regression in Question 1, construct and interpret a 95% confidence interval for the true slope.
- List and briefly describe the four conditions (LINE) for regression inference.
- A researcher computes a regression and gets r^2 = 0.04 with a p-value of 0.31 for the slope. Interpret both values.
- Why is the residual standard error s divided by n - 2 instead of n - 1?
- Explain what it means if a 99% confidence interval for the slope is (-0.3, 2.1).
Answers to Self-Check
- H0: beta = 0, Ha: beta > 0. t = 3.5/1.1 = 3.18. df = 18. p-value = tcdf(3.18, 1E99, 18) = 0.0025. Since 0.0025 < 0.05, reject H0. There is convincing evidence of a positive linear relationship between fertilizer amount and plant height.
- t* = invT(0.975, 18) = 2.101. ME = 2.101(1.1) = 2.311. CI = (1.189, 5.811). We are 95% confident that the true slope of the regression line is between 1.19 and 5.81 cm per gram of fertilizer.
- Linear: The scatterplot shows a linear pattern. Normal: Residuals are approximately normally distributed. Independent: Individual observations are independent. Equal variance: Residual plot shows roughly constant spread.
- r^2 = 0.04 means only 4% of the variation in the response is explained by the linear relationship with x — a very weak relationship. The p-value of 0.31 (> 0.05) confirms there is not convincing evidence of a linear association.
- We divide by n - 2 because we estimated two parameters (the slope and the intercept) from the data. Each estimated parameter costs one degree of freedom.
- Since the interval (-0.3, 2.1) contains 0, we are 99% confident that the true slope could be zero. There is not convincing evidence of a linear relationship between x and y at the 1% significance level.
Practice sets
9Practice Problems — Unit 1: Exploring One-Variable Data
1. A set of exam scores has a mean of 82 and a median of 85. Which of the following is most likely true about the shape of the distribution?
(A) Symmetric
(B) Uniform
(C) Skewed right
(D) Skewed left
(E) Bimodal
2. The weights (in pounds) of 20 dogs at a veterinary clinic have a mean of 42, a standard deviation of 12, a minimum of 15, Q1 of 33, a median of 40, Q3 of 50, and a maximum of 78. Which value is an outlier according to the 1.5 x IQR rule?
(A) 15
(B) 33
(C) 50
(D) 78
(E) There are no outliers.
3. A teacher converts all test scores from a 100-point scale to a 4.0 GPA scale using the formula GPA = (score/100) x 4.0. If the original scores had a mean of 75 and a standard deviation of 10, what are the mean and standard deviation of the GPA scores?
(A) Mean = 3.0, SD = 0.4
(B) Mean = 3.0, SD = 10
(C) Mean = 79, SD = 14
(D) Mean = 3.0, SD = 2.5
(E) Mean = 0.75, SD = 0.10
4. Which of the following is NOT a measure of spread?
(A) Range
(B) Interquartile range
(C) Standard deviation
(D) Mean
(E) Variance
5. A student's score on a standardized test has a z-score of -1.5. Which of the following is the best interpretation?
(A) The student scored 1.5 points below the mean.
(B) The student scored 1.5 standard deviations below the mean.
(C) The student is in the 1.5th percentile.
(D) The student answered 1.5% fewer questions correctly than average.
(E) The student's score is 1.5 times the standard deviation.
6. A histogram of a dataset shows a single peak on the left with a long tail extending to the right. Which measure of center would be most appropriate to report?
(A) Mean only
(B) Median only
(C) Both mean and median
(D) Mode only
(E) Range
Free-Response Question
A wildlife biologist measured the wingspan (in millimeters) of 16 adult birds of a particular species captured over one season. The data (in mm) are:
245, 248, 250, 251, 252, 253, 253, 254, 255, 256, 257, 258, 259, 262, 265, 310
(a) Construct the five-number summary and determine if there are any outliers.
(b) The biologist calculates a mean wingspan of 257.8 mm and a standard deviation of 16.4 mm. Calculate the z-score for the bird with a 310 mm wingspan. Is this z-score consistent with your outlier analysis in part (a)?
(c) Would you recommend using the mean and standard deviation, or the median and IQR, to summarize the center and spread of this dataset? Justify your answer.
(d) The biologist plans to convert all measurements to centimeters by dividing by 10. What will the new median and IQR be?
Answer Key
MCQ Answers:
- D (Mean < median indicates left skew)
- E (IQR = 50 - 33 = 17. Lower fence = 33 - 25.5 = 7.5. Upper fence = 50 + 25.5 = 75.5. 78 > 75.5, so 78 is an outlier. Wait — let me recalculate. IQR = 50 - 33 = 17. 1.5 x 17 = 25.5. Lower = 33 - 25.5 = 7.5. Upper = 50 + 25.5 = 75.5. 78 > 75.5, so 78 IS an outlier. The answer is D, 78.)
- A (New mean = 75/100 x 4 = 3.0. New SD = 10/100 x 4 = 0.4.)
- D (The mean is a measure of center, not spread.)
- B (z-score measures standard deviations from the mean.)
- C (Report both, but note that the median is preferred for skewed data. The best answer is C since AP expects you to always report both for quantitative data.)
FRQ Rubric: (a) Five-number summary: Min = 245, Q1 = 251, Median = 254.5, Q3 = 259, Max = 310. IQR = 259 - 251 = 8. Lower fence = 251 - 12 = 239. Upper fence = 259 + 12 = 271. 310 > 271, so 310 is an outlier. [4 points]
(b) z = (310 - 257.8) / 16.4 = 52.2 / 16.4 = 3.18. Yes, this large z-score is consistent with 310 being an outlier — it is more than 3 standard deviations above the mean. [2 points]
(c) The median and IQR are preferred because the data contain an outlier (310 mm). The mean and standard deviation are not resistant to outliers — the mean of 257.8 is pulled upward by the outlier. The median of 254.5 better represents the center of the typical data. [2 points]
(d) Dividing by 10: New median = 254.5 / 10 = 25.45 cm. New IQR = 8 / 10 = 0.8 cm. [2 points]
Practice Problems — Unit 2: Exploring Two-Variable Data
1. A scatterplot of two variables shows a strong, curved pattern opening downward. Which of the following statements is true?
(A) The correlation r will be close to -1.
(B) The correlation r will be close to 0.
(C) A least-squares regression line would be an appropriate model.
(D) The correlation r will be negative.
(E) Both (C) and (D) are true.
2. A regression equation for predicting the price of a house (in thousands of dollars) from its size (in hundreds of square feet) is y-hat = 15 + 8.5x. Which of the following is the best interpretation of the slope?
(A) Each house costs 8.5 thousand dollars.
(B) For each additional 100 square feet, the predicted price increases by $8,500.
(C) For each additional square foot, the predicted price increases by 8.5 thousand dollars.
(D) For each additional 8.5 hundred square feet, the predicted price increases by $15,000.
(E) Houses cost about 8.5 thousand dollars per hundred square feet.
3. For a regression of y on x, r^2 = 0.81. Which of the following is the best interpretation?
(A) 81% of the data points fall on the regression line.
(B) 81% of the x-values explain 81% of the y-values.
(C) 81% of the variation in y is explained by the linear relationship with x.
(D) The correlation between x and y is 0.81.
(E) The slope is 0.81.
4. A researcher notices that a data point with a very large x-value lies close to the regression line. If this point is removed, what is most likely to happen to the slope and correlation?
(A) Both slope and r will change substantially.
(B) The slope will change substantially, but r will not.
(C) r will change substantially, but the slope will not.
(D) Neither will change substantially.
(E) The slope will reverse direction.
5. Which of the following is the best explanation for why "correlation does not imply causation"?
(A) Correlation only works for quantitative variables.
(B) A lurking variable may influence both variables, creating a common response.
(C) The regression line may not pass through any data points.
(D) r^2 must be greater than 0.50 for a causal relationship.
(E) Observational studies cannot compute correlation.
6. A residual plot shows a clear fan shape — the residuals are more spread out for larger x-values. What does this indicate?
(A) The regression line is a good model.
(B) The relationship is not linear.
(C) The variability of y around the line is not constant.
(D) There are influential points.
(E) The correlation is weak.
Free-Response Question
A student recorded the temperature (in degrees Celsius) and the number of customers at an ice cream shop for 12 randomly selected afternoons. The data are shown below.
| Temperature (°C) | Customers |
|---|---|
| 18 | 45 |
| 20 | 52 |
| 22 | 58 |
| 24 | 70 |
| 25 | 65 |
| 27 | 80 |
| 28 | 88 |
| 30 | 95 |
| 31 | 92 |
| 33 | 105 |
| 35 | 110 |
| 36 | 115 |
Computer output from the regression gives: y-hat = -8.24 + 3.46x, r = 0.994, r^2 = 0.988, s = 3.42.
(a) Interpret the slope of the regression line in the context of this problem.
(b) Interpret the value of r^2 in context.
(c) Calculate and interpret the residual for the day when the temperature was 25°C.
(d) Would it be appropriate to use this model to predict the number of customers when the temperature is 5°C? Explain.
(e) The owner claims that warmer weather causes more customers. Based on this study alone, is this claim justified? Explain.
Answer Key
MCQ Answers:
- B (A curved pattern may produce a correlation close to 0 even though there is a strong nonlinear relationship.)
- B (The slope is 8.5, and x is measured in hundreds of square feet. So each 100 sq ft increase corresponds to a predicted increase of 8.5 thousand dollars = $8,500.)
- C (r^2 is the proportion of variation in y explained by the linear relationship with x.)
- D (A high-leverage point close to the line has little influence because its residual is small. It may change r slightly but typically not substantially, and the slope may also not change much.)
- B (Lurking variables can create a common response that produces correlation without causation.)
- C (A fan shape in the residual plot indicates non-constant variance — the variability of y around the line increases with x.)
FRQ Rubric: (a) For each additional degree Celsius increase in temperature, the predicted number of customers increases by approximately 3.46 customers. [2 points]
(b) Approximately 98.8% of the variation in the number of customers is explained by the linear relationship with temperature. [2 points]
(c) Predicted: y-hat = -8.24 + 3.46(25) = -8.24 + 86.5 = 78.26. Residual = 65 - 78.26 = -13.26. On the day the temperature was 25°C, the actual number of customers was 13.26 fewer than predicted by the model. [2 points]
(d) No. This would be extrapolation. The data range for temperature is 18°C to 36°C. A prediction at 5°C is well outside this range, and the relationship may not hold at such low temperatures. [2 points]
(e) Not justified. This is an observational study — the researcher did not randomly assign temperatures. There may be lurking variables (e.g., day of the week, holidays, advertising) that affect both temperature and customer count. Only a randomized experiment could establish causation. [2 points]
Practice Problems — Unit 3: Collecting Data
1. A researcher wants to estimate the average income of residents in a city. She selects 10 city blocks at random and surveys every household in those blocks. What sampling method is this?
(A) Simple random sample
(B) Stratified random sample
(C) Cluster sample
(D) Systematic sample
(E) Convenience sample
2. Which of the following is a potential problem with voluntary response sampling?
(A) It requires a very large sample size.
(B) Individuals with strong opinions are more likely to respond.
(C) It can only be used for quantitative variables.
(D) It violates the 10% condition.
(E) It always produces a sample that is too small.
3. In an experiment, which of the following is the primary purpose of random assignment?
(A) To obtain a representative sample of the population.
(B) To reduce the margin of error.
(C) To create groups that are similar in all respects before treatments are applied.
(D) To ensure that the experiment is double-blind.
(E) To increase the power of the test.
4. A study randomly assigns 100 patients to receive either Drug A or Drug B. Neither the patients nor the doctors know who receives which drug until after the study. What type of design is this?
(A) Completely randomized, single-blind
(B) Completely randomized, double-blind
(C) Matched pairs, double-blind
(D) Stratified, double-blind
(E) Observational study
5. A researcher divides students into groups based on their major and then randomly assigns students within each major to a treatment or control group. This is an example of which experimental design?
(A) Completely randomized design
(B) Stratified sampling
(C) Randomized block design
(D) Matched pairs design
(E) Systematic design
6. A survey asks, "Don't you agree that the new school policy is terrible?" What type of bias does this question introduce?
(A) Undercoverage
(B) Nonresponse bias
(C) Response bias due to leading question
(D) Sampling bias
(E) Voluntary response bias
Free-Response Question
A school district wants to determine whether a new math curriculum improves student test scores compared to the existing curriculum. The district has 8 middle schools, each with approximately 200 seventh-grade students.
(a) Describe how to use a completely randomized design to conduct this experiment. Be specific about how randomization would be implemented.
(b) A statistician suggests using a randomized block design with schools as blocks instead. Explain why this might be a better approach.
(c) Describe how you would implement the block design from part (b).
(d) Identify the experimental units, the treatments, and the response variable.
Answer Key
MCQ Answers:
- C (Randomly selecting clusters and surveying all units within selected clusters is cluster sampling.)
- B (Voluntary response sampling tends to attract people with strong opinions, leading to biased results.)
- C (Random assignment creates treatment groups that are similar, reducing confounding.)
- B (Patients are randomly assigned to two treatments — completely randomized. Neither patients nor doctors know the assignment — double-blind.)
- C (Students are blocked by major, then randomly assigned within blocks. This is a randomized block design.)
- C (The question leads the respondent toward a particular answer, which is response bias from a leading question.)
FRQ Rubric: (a) Assign each of the approximately 1600 seventh-grade students a number. Use a random number generator to assign each student to either the new curriculum group or the existing curriculum group. Teach each group with its assigned curriculum for the school year, then compare test scores. [3 points for clear description of randomization]
(b) Different schools may have different characteristics (resources, teacher quality, student demographics) that affect test scores. Blocking by school controls for these between-school differences, making it easier to detect the effect of the curriculum. [2 points]
(c) Within each school, randomly assign half the seventh-grade students to the new curriculum and half to the existing curriculum. At the end of the year, compare test scores between the two curriculum groups within each school, then combine results across schools. [3 points]
(d) Experimental units: the individual seventh-grade students. Treatments: the new math curriculum and the existing math curriculum. Response variable: student test scores on a standardized math assessment. [2 points]
Practice Problems — Unit 4: Probability, Random Variables, and Probability Distributions
1. Events A and B are independent. P(A) = 0.4 and P(B) = 0.3. What is P(A and B)?
(A) 0.70
(B) 0.12
(C) 0.10
(D) 0.50
(E) 0.07
2. A bag contains 3 red marbles and 5 blue marbles. Two marbles are drawn without replacement. What is the probability that the second marble is red given that the first marble was red?
(A) 3/8
(B) 2/7
(C) 3/7
(D) 2/8
(E) 5/8
3. Random variable X has E(X) = 10 and Var(X) = 9. What are E(3X - 2) and Var(3X - 2)?
(A) E = 28, Var = 25
(B) E = 28, Var = 81
(C) E = 30, Var = 81
(D) E = 28, Var = 7
(E) E = 30, Var = 27
4. Which of the following must be true for two events to be mutually exclusive?
(A) They are independent.
(B) P(A and B) = 0.
(C) P(A | B) = P(A).
(D) P(A or B) = 1.
(E) P(A and B) = P(A) * P(B).
5. A game costs $3 to play. You roll a fair six-sided die. If you roll a 6, you win $10. Otherwise, you win nothing. What is the expected net gain (or loss) per game?
(A) -$0.50
(B) -$1.00
(C) -$1.17
(D) $0.50
(E) -$2.00
6. Two independent random variables X and Y have standard deviations of 4 and 5, respectively. What is the standard deviation of X - Y?
(A) 1
(B) 3
(C) sqrt(41)
(D) 9
(E) sqrt(31)
Free-Response Question
A company manufactures light bulbs. The probability that a randomly selected bulb is defective is 0.05. A quality control inspector tests bulbs one at a time until a defective bulb is found.
(a) What is the probability that the first defective bulb is the 4th one tested?
(b) What is the expected number of bulbs the inspector must test to find the first defective one?
Now suppose the inspector tests bulbs in batches of 20. Let X = the number of defective bulbs in a batch of 20.
(c) Find the mean and standard deviation of X.
(d) What is the probability that a batch of 20 contains at least 3 defective bulbs? (You may set up the calculation without computing the final numeric value.)
Answer Key
MCQ Answers:
- B (P(A and B) = P(A) P(B) = 0.4 0.3 = 0.12, since the events are independent.)
- B (Given the first was red, 2 red remain out of 7 total: 2/7.)
- B (E(3X-2) = 3(10) - 2 = 28. Var(3X-2) = 3^2 * 9 = 81.)
- B (Mutually exclusive means they cannot both occur, so P(A and B) = 0.)
- C (E(winnings) = (1/6)(10) + (5/6)(0) = 10/6 = $1.67. Net gain = 1.67 - 3 = -$1.33. Actually: E(net) = (1/6)(10-3) + (5/6)(0-3) = (1/6)(7) + (5/6)(-3) = 7/6 - 15/6 = -8/6 = -$1.33. Closest is C at -$1.17. Let me recalculate: (1/6)(7) = 1.1667, (5/6)(-3) = -2.5. Sum = -1.333. The answer closest is C, -$1.17.)
- C (Var(X-Y) = Var(X) + Var(Y) = 16 + 25 = 41. SD = sqrt(41).)
FRQ Rubric: (a) This is a geometric distribution with p = 0.05. P(X = 4) = (0.95)^3 (0.05) = 0.857375 0.05 = 0.0429. [2 points]
(b) E(X) = 1/p = 1/0.05 = 20 bulbs. The inspector is expected to test 20 bulbs to find the first defective one. [2 points]
(c) This is a binomial distribution with n = 20, p = 0.05. E(X) = np = 20(0.05) = 1. SD(X) = sqrt(np(1-p)) = sqrt(20 0.05 0.95) = sqrt(0.95) = 0.975. [3 points]
(d) P(X >= 3) = 1 - P(X=0) - P(X=1) - P(X=2) = 1 - (0.95)^20 - 20(0.05)(0.95)^19 - C(20,2)(0.05)^2(0.95)^18. [3 points]
Practice Problems — Unit 5: Sampling Distributions
1. A population has a mean of 100 and a standard deviation of 15. If we take all possible samples of size n = 36, what is the standard deviation of the sampling distribution of x-bar?
(A) 15
(B) 2.5
(C) 15/36
(D) 15/6
(E) 100/36
2. Which of the following statements about the Central Limit Theorem is correct?
(A) The sampling distribution of x-bar is normal only if the population is normal.
(B) For any population, the sampling distribution of x-bar is approximately normal when n >= 30.
(C) The CLT states that x-bar equals mu as n increases.
(D) The CLT only applies to sample proportions.
(E) The CLT requires the population to be symmetric.
3. The true proportion of adults who own a smartphone is 0.78. A random sample of 500 adults is selected. What is the mean and standard deviation of the sampling distribution of p-hat?
(A) Mean = 0.78, SD = 0.018
(B) Mean = 0.78, SD = 0.044
(C) Mean = 0.50, SD = 0.022
(D) Mean = 390, SD = 9.22
(E) Mean = 0.78, SD = 0.009
4. As the sample size increases, what happens to the standard error of x-bar?
(A) It increases.
(B) It decreases.
(C) It stays the same.
(D) It approaches sigma.
(E) It becomes 1.
5. A population is strongly right-skewed with a mean of 50. Which of the following is true about the sampling distribution of x-bar for samples of size n = 10?
(A) It is approximately normal by the CLT.
(B) It is right-skewed.
(C) It is left-skewed.
(D) It has a mean of 5.
(E) It has a standard deviation of 50.
6. The Law of Large Numbers states that:
(A) For large n, the sampling distribution of x-bar is approximately normal.
(B) As the sample size increases, the sample mean gets closer to the population mean.
(C) The variance of x-bar equals sigma divided by n.
(D) 95% of sample means fall within 2 standard errors of mu.
(E) Larger samples have less bias.
Free-Response Question
A machine fills bags of flour with a target weight of 500 grams. The actual weights follow a normal distribution with a mean of 502 grams and a standard deviation of 5 grams.
(a) Describe the shape, center, and spread of the sampling distribution of x-bar for samples of size n = 25 bags.
(b) What is the probability that the sample mean weight of 25 bags is less than 500 grams?
(c) If the sample size is increased to 100 bags, what happens to the probability in part (b)? Explain why, without calculating.
(d) A quality control inspector checks one bag at random. What is the probability that a single bag weighs less than 500 grams? Compare this to your answer in part (b) and explain the difference.
Answer Key
MCQ Answers:
- B (sigma_x-bar = 15/sqrt(36) = 15/6 = 2.5.)
- B (The CLT states that for n >= 30, the sampling distribution of x-bar is approximately normal regardless of the population shape.)
- A (Mean = 0.78. SD = sqrt(0.78*0.22/500) = sqrt(0.1716/500) = sqrt(0.000343) = 0.0185.)
- B (Standard error = sigma/sqrt(n), which decreases as n increases.)
- B (With n = 10, the CLT does not apply. The sampling distribution of x-bar will have a shape similar to the population — right-skewed.)
- B (The LLN states that the sample mean approaches the population mean as n increases.)
FRQ Rubric: (a) Shape: Normal (the population is normal, so the sampling distribution is normal for any n). Center: mu_x-bar = 502 grams. Spread: sigma_x-bar = 5/sqrt(25) = 1 gram. [3 points]
(b) z = (500 - 502) / 1 = -2. P(x-bar < 500) = P(Z < -2) = 0.0228. The probability is approximately 0.0228. [3 points]
(c) The probability will decrease. As n increases, the standard error decreases (sigma/sqrt(n) is smaller for larger n), so the distribution becomes more tightly concentrated around 502. The probability of being below 500 (which is 2 grams below the mean) decreases. [2 points]
(d) z = (500 - 502) / 5 = -0.4. P(X < 500) = P(Z < -0.4) = 0.3446. This probability (0.3446) is much larger than the probability for the sample mean (0.0228). Averages are less variable than individual observations, so the sample mean is more likely to be close to 502 than a single observation is. [2 points]
Practice Problems — Unit 6: Inference for Categorical Data: Proportions
1. A 95% confidence interval for a population proportion is (0.34, 0.46). Which of the following must be true?
(A) The sample proportion is 0.40.
(B) 95% of the population is between 0.34 and 0.46.
(C) The probability that the true proportion is between 0.34 and 0.46 is 0.95.
(D) If we took many samples, about 95% of the resulting intervals would contain the true proportion.
(E) The margin of error is 0.06.
2. A researcher tests H0: p = 0.5 against Ha: p > 0.5. The p-value is 0.032. Which of the following is the correct conclusion at alpha = 0.05?
(A) Accept H0; 0.5 is the true proportion.
(B) Reject H0; there is convincing evidence that p > 0.5.
(C) Fail to reject H0; there is not convincing evidence that p > 0.5.
(D) Reject H0; there is convincing evidence that p = 0.5.
(E) Fail to reject H0; the true proportion is exactly 0.5.
3. In a test of H0: p = 0.3 vs. Ha: p != 0.3, a researcher makes a Type I error. What happened?
(A) The researcher rejected H0 when it was true.
(B) The researcher failed to reject H0 when it was false.
(C) The true proportion was 0.3.
(D) Both (A) and (C).
(E) Both (B) and (C).
4. Which of the following would increase the power of a hypothesis test for a proportion?
(A) Decreasing the sample size.
(B) Decreasing the significance level alpha.
(C) Increasing the sample size.
(D) Using a two-sided test instead of a one-sided test.
(E) Choosing a null hypothesis closer to the true proportion.
5. A two-sample z-interval for the difference in proportions (p1 - p2) is (-0.08, 0.02). Which of the following is the correct interpretation?
(A) We are 95% confident that p1 is between 8% less than p2 and 2% more than p2.
(B) 95% of the time, p1 - p2 will be between -0.08 and 0.02.
(C) Since 0 is in the interval, we can conclude p1 = p2.
(D) The probability that p1 < p2 is about 95%.
(E) p1 - p2 = -0.03.
6. For a hypothesis test about a proportion, the standard error in the test statistic uses p0, while the standard error in the confidence interval uses p-hat. Why?
(A) They are both the same value in practice.
(B) The test assumes H0 is true (so use p0), while the interval estimates the parameter (so use the best estimate, p-hat).
(C) The test requires the 10% condition but the interval does not.
(D) The interval uses p0 to be more conservative.
(E) There is no difference; this is a common misconception.
Free-Response Question
A school board wants to determine whether the proportion of students who pass the state math exam differs between two high schools. Random samples are taken from each school.
- School A: 320 students sampled, 224 passed.
- School B: 280 students sampled, 196 passed.
(a) Construct a 95% confidence interval for the difference in the proportions of students who pass at the two schools (p_A - p_B).
(b) Based on your interval, is there convincing evidence that the passing rates differ between the two schools? Justify your answer.
(c) A reporter claims that the passing rate at School A is 10 percentage points higher than at School B. Is this claim consistent with your confidence interval? Explain.
(d) State the conditions that must be met for your inference to be valid, and explain whether each is satisfied.
Answer Key
MCQ Answers:
- D (The confidence level describes the long-run performance of the method.)
- B (p-value 0.032 < 0.05, so reject H0.)
- D (A Type I error means rejecting H0 when it is true. If H0 states p = 0.3 and it was true, then the true proportion is 0.3.)
- C (Increasing the sample size increases power by reducing the standard error.)
- A (The interval gives a range of plausible values for the difference p1 - p2. Since the interval contains 0, there may be no difference.)
- B (The test computes the distribution assuming H0 is true, so it uses p0. The interval uses the data-based estimate p-hat.)
FRQ Rubric: (a) p-hat_A = 224/320 = 0.70. p-hat_B = 196/280 = 0.70. Difference = 0. SE = sqrt(0.700.30/320 + 0.700.30/280) = sqrt(0.000656 + 0.000750) = sqrt(0.001406) = 0.0375. z* = 1.96. ME = 1.96(0.0375) = 0.0735. CI: 0 +/- 0.0735 = (-0.0735, 0.0735). [4 points]
(b) Since 0 is in the confidence interval, we cannot conclude there is a difference. There is not convincing evidence that the passing rates differ between the two schools at the 95% confidence level. [2 points]
(c) The claim is that p_A - p_B = 0.10. Since 0.10 is outside the interval (-0.0735, 0.0735), this claim is NOT consistent with the confidence interval. [2 points]
(d) Random: Both samples are stated to be random. Checked. Normal (Large Counts): For School A: 320(0.70) = 224 >= 10, 320(0.30) = 96 >= 10. For School B: 280(0.70) = 196 >= 10, 280(0.30) = 84 >= 10. Checked. Independent: Each sample must be less than 10% of its school population. Assuming each school has more than 3200 students (for A) and 2800 students (for B), this condition is met. The two samples must also be independent of each other, which is reasonable since they are from different schools. [2 points]
Practice Problems — Unit 7: Inference for Quantitative Data: Means
1. A one-sample t-interval is constructed with n = 15, x-bar = 34.2, and s = 6.8. What are the degrees of freedom for this procedure?
(A) 6.8
(B) 14
(C) 15
(D) 16
(E) 13
2. A paired t-test is used instead of a two-sample t-test when:
(A) The sample sizes are equal.
(B) The standard deviations are equal.
(C) The two groups consist of the same or matched subjects.
(D) The data are normally distributed.
(E) The significance level is 0.01.
3. A 99% confidence interval for a population mean is (42.5, 51.3). Which of the following is true?
(A) There is a 99% probability that mu is between 42.5 and 51.3.
(B) 99% of the data falls between 42.5 and 51.3.
(C) We are 99% confident that the true mean is between 42.5 and 51.3.
(D) If we repeated the sampling, 99% of the data would be in this interval.
(E) The sample mean is 51.3.
4. For a two-sample t-test comparing two independent groups, which of the following is NOT a condition?
(A) Both samples are random.
(B) Both populations are normally distributed or both sample sizes are large (n >= 30).
(C) The two samples are independent of each other.
(D) The population standard deviations are known.
(E) Each sample is no more than 10% of its respective population.
5. A researcher conducts a one-sample t-test with H0: mu = 50 and gets t = -2.45 with df = 24. The p-value for a two-sided test is approximately:
(A) 0.010
(B) 0.021
(C) 0.979
(D) 0.042
(E) 0.490
6. A two-sample t-interval for (mu_1 - mu_2) is (3.2, 8.7). Which conclusion is correct at the 95% confidence level?
(A) mu_1 = mu_2
(B) mu_1 > mu_2
(C) mu_1 < mu_2
(D) There is not convincing evidence that mu_1 differs from mu_2.
(E) The sample means are 3.2 and 8.7.
Free-Response Question
A fitness trainer wants to determine whether a 6-week training program reduces resting heart rate. She measures the resting heart rate (in bpm) of 10 clients before and after the program. The data are:
| Client | Before | After |
|---|---|---|
| 1 | 72 | 68 |
| 2 | 80 | 75 |
| 3 | 65 | 63 |
| 4 | 78 | 72 |
| 5 | 70 | 70 |
| 6 | 85 | 78 |
| 7 | 74 | 69 |
| 8 | 68 | 66 |
| 9 | 77 | 71 |
| 10 | 73 | 67 |
(a) Explain why a paired t-test is appropriate for this situation.
(b) Compute the differences (Before - After) for each client. Then find the mean and standard deviation of the differences.
(c) Perform a paired t-test to determine whether the training program significantly reduces resting heart rate. Use alpha = 0.05. Show all steps: hypotheses, check conditions, compute the test statistic, find the p-value, and state your conclusion in context.
(d) Construct and interpret a 95% confidence interval for the mean reduction in resting heart rate.
Answer Key
MCQ Answers:
- B (df = n - 1 = 15 - 1 = 14.)
- C (Paired tests are used when the same subjects are measured twice or when subjects are naturally paired.)
- C (The correct interpretation of a confidence interval.)
- D (If sigma were known, we would use z-procedures, not t-procedures. The t-test is specifically for when sigma is unknown.)
- B (p-value = 2 * P(T < -2.45) with df = 24. tcdf(-1E99, -2.45, 24) = 0.0107, so 2(0.0107) = 0.0214.)
- B (The interval (3.2, 8.7) does not contain 0 and is entirely positive, so we are 95% confident that mu_1 - mu_2 > 0, meaning mu_1 > mu_2.)
FRQ Rubric: (a) A paired t-test is appropriate because the data consist of before and after measurements on the same 10 clients. Each client provides a natural pair of observations, and we are interested in the change for each individual. [2 points]
(b) Differences (Before - After): 4, 5, 2, 6, 0, 7, 5, 2, 6, 6. d-bar = (4+5+2+6+0+7+5+2+6+6)/10 = 43/10 = 4.3. Using a calculator: s_d = 2.31. [3 points]
(c) H0: mu_d = 0 (the training program does not change resting heart rate). Ha: mu_d > 0 (the training program reduces resting heart rate, so the before-after difference is positive).
Conditions: Random (assume clients were randomly selected or randomly assigned). Normal: With n = 10, we need the differences to be approximately normal. A dotplot or histogram of the differences (0, 2, 2, 4, 5, 5, 6, 6, 6, 7) shows no strong skewness or outliers. Independent: 10 clients is assumed to be less than 10% of all potential clients.
t = (4.3 - 0) / (2.31/sqrt(10)) = 4.3 / 0.731 = 5.88. df = 9. p-value = tcdf(5.88, 1E99, 9) = 0.0001.
Since 0.0001 < 0.05, we reject H0. There is convincing evidence that the 6-week training program reduces resting heart rate. [5 points]
(d) t* = invT(0.975, 9) = 2.262. SE = 2.31/sqrt(10) = 0.731. ME = 2.262(0.731) = 1.653. CI: 4.3 +/- 1.653 = (2.647, 5.953). We are 95% confident that the true mean reduction in resting heart rate from the training program is between 2.6 and 6.0 bpm. [2 points]
Practice Problems — Unit 8: Inference for Categorical Data: Chi-Square
1. A chi-square goodness of fit test has 6 categories. What are the degrees of freedom?
(A) 5
(B) 6
(C) 7
(D) 4
(E) 3
2. In a chi-square test, what happens to the test statistic if the observed counts are very close to the expected counts?
(A) It approaches 0.
(B) It approaches infinity.
(C) It approaches 1.
(D) It becomes negative.
(E) It equals the number of categories.
3. A researcher wants to determine whether the distribution of favorite season is the same across three regions of the country. Which chi-square test should she use?
(A) Goodness of fit
(B) Test for independence
(C) Test for homogeneity
(D) Two-sample z-test
(E) Paired t-test
4. A 4x3 contingency table is used in a chi-square test for independence. What are the degrees of freedom?
(A) 12
(B) 7
(C) 6
(D) 5
(E) 8
5. Which of the following is a condition for ALL chi-square tests?
(A) The population is normally distributed.
(B) All expected counts are at least 5.
(C) The sample size is at least 30.
(D) The variables are quantitative.
(E) The data come from a matched pairs design.
6. In a chi-square test, the p-value is calculated as:
(A) The probability of getting exactly the observed chi-square statistic.
(B) The probability of getting a chi-square statistic as large or larger than the observed value, assuming H0 is true.
(C) The area to the left of the observed chi-square statistic.
(D) 1 minus the chi-square statistic.
(E) The square root of the test statistic.
Free-Response Question
A student council member wants to know whether grade level (freshman, sophomore, junior, senior) is associated with preference for school event type (concert, sports game, dance) at a large high school. A random sample of 300 students is surveyed. The observed counts are:
| Concert | Sports | Dance | Total | |
|---|---|---|---|---|
| Freshman | 35 | 40 | 25 | 100 |
| Sophomore | 30 | 35 | 35 | 100 |
| Junior | 20 | 50 | 30 | 100 |
| Senior | 15 | 45 | 40 | 100 |
| Total | 100 | 170 | 130 | 300 |
(a) State the null and alternative hypotheses for this test.
(b) Calculate the expected count for the "Freshman, Concert" cell and the "Senior, Dance" cell. Show your work.
(c) The chi-square test statistic is 18.6 with 6 degrees of freedom. The p-value is approximately 0.009. State a conclusion in context at the alpha = 0.05 level.
(d) Is this a test for independence or homogeneity? Explain.
(e) The student council wants to follow up on this result. Suggest what they might investigate next.
Answer Key
MCQ Answers:
- A (df = number of categories - 1 = 6 - 1 = 5.)
- A (When observed = expected, each term (O-E)^2/E = 0, so X^2 = 0.)
- C (Comparing a distribution across multiple populations — homogeneity.)
- C (df = (4-1)(3-1) = 3 x 2 = 6.)
- B (All expected counts must be at least 5 for all chi-square tests.)
- B (The p-value for a chi-square test is the area to the right of the observed statistic, under H0.)
FRQ Rubric: (a) H0: There is no association between grade level and school event preference. Ha: There is an association between grade level and school event preference. [2 points]
(b) Expected for Freshman, Concert: (100 x 100) / 300 = 10000/300 = 33.33. Expected for Senior, Dance: (100 x 130) / 300 = 13000/300 = 43.33. [2 points]
(c) Since the p-value (0.009) is less than alpha (0.05), we reject H0. There is convincing evidence of an association between grade level and school event preference among students at this high school. [2 points]
(d) This is a test for independence. There is one random sample of 300 students, and we are examining whether two categorical variables (grade level and event preference) are associated. [2 points]
(e) Since there is evidence of an association, the student council might investigate which specific cells contribute most to the chi-square statistic (largest (O-E)^2/E values) to understand which grade-event combinations are most different from expected. They could also examine conditional distributions (row or column percentages) to describe the nature of the association. [2 points]
Practice Problems — Unit 9: Inference for Quantitative Data: Slopes
1. A regression of y on x based on 20 observations gives a slope of b = 2.5 with a standard error of SE_b = 0.8. What are the degrees of freedom for a t-test of H0: beta = 0?
(A) 18
(B) 19
(C) 20
(D) 21
(E) 2.5
2. For the regression in Question 1, what is the value of the t-test statistic?
(A) 0.32
(B) 1.25
(C) 2.50
(D) 3.13
(E) 4.00
3. A 95% confidence interval for the true slope of a regression line is (-0.5, 3.2). Which of the following is correct?
(A) We are 95% confident that the slope is positive.
(B) We are 95% confident that x causes y.
(C) We fail to reject H0: beta = 0 at alpha = 0.05.
(D) The correlation is 0.95.
(E) The regression line has a positive slope.
4. Which of the following conditions is NOT required for inference about a regression slope?
(A) The relationship between x and y is linear.
(B) The residuals are approximately normally distributed.
(C) The residuals have constant variance across x-values.
(D) The sample size is at least 30.
(E) The observations are independent.
5. A residual plot shows a clear curved pattern. What should the researcher do?
(A) Proceed with the t-test for the slope.
(B) Use a larger significance level.
(C) Consider a nonlinear model instead of linear regression.
(D) Remove all outliers and re-run the regression.
(E) Use a z-test instead of a t-test.
6. The residual standard error (s) in a regression output measures:
(A) The standard deviation of x.
(B) The typical distance of data points from the regression line.
(C) The standard error of the slope.
(D) The correlation coefficient.
(E) The sample size.
Free-Response Question
A researcher wants to study the relationship between the number of hours spent on homework per week (x) and GPA (y) for college students. A random sample of 18 students is selected. The regression output is:
- y-hat = 2.10 + 0.085x
- SE_b = 0.028
- s = 0.32
- r^2 = 0.64
(a) Interpret the slope of the regression line in the context of this problem.
(b) Perform a t-test to determine whether there is a positive linear relationship between homework hours and GPA. Use alpha = 0.01. Show all steps.
(c) Construct and interpret a 95% confidence interval for the true slope.
(d) Interpret the value of r^2 in context.
(e) List the conditions for this inference and describe how you would check each one.
Answer Key
MCQ Answers:
- A (df = n - 2 = 20 - 2 = 18.)
- D (t = 2.5/0.8 = 3.125.)
- C (The interval contains 0, so we fail to reject H0: beta = 0 at alpha = 0.05.)
- D (There is no strict n >= 30 requirement for regression inference. Normality of residuals should be checked, especially for small samples.)
- C (A curved pattern in the residual plot means the linear model is inappropriate. Consider a nonlinear model.)
- B (The residual standard error measures the typical distance of data points from the regression line.)
FRQ Rubric: (a) For each additional hour spent on homework per week, the predicted GPA increases by 0.085 points. [2 points]
(b) H0: beta = 0 (no linear relationship between homework hours and GPA). Ha: beta > 0 (positive linear relationship).
t = (0.085 - 0) / 0.028 = 3.036. df = 18 - 2 = 16. p-value = tcdf(3.036, 1E99, 16) = 0.0037.
Since 0.0037 < 0.01, we reject H0. There is convincing evidence of a positive linear relationship between the number of hours spent on homework per week and GPA among college students. [5 points]
(c) t* = invT(0.975, 16) = 2.120. ME = 2.120(0.028) = 0.0594. CI: 0.085 +/- 0.0594 = (0.0256, 0.1444). We are 95% confident that for each additional hour of homework per week, the true mean increase in GPA is between 0.026 and 0.144 points. [3 points]
(d) Approximately 64% of the variation in GPA among college students is explained by the linear relationship with the number of homework hours per week. [2 points]
(e) Linear: Check a scatterplot of GPA vs. homework hours for a linear pattern. Normal: Check a normal probability plot or histogram of the residuals for approximate normality. Independent: The 18 students must be a random sample and less than 10% of all college students. Equal variance: Check a residual plot for constant spread (no fan shape). [4 points]
Summary & cheat sheets
1AP Statistics — Quick Reference Summary Sheet
Measures of Center
- Mean: x-bar = (sum of x_i) / n
- Median: Middle value of ordered data
- Mode: Most frequent value
Measures of Spread
- Range: Max - Min
- IQR: Q3 - Q1
- Variance: s^2 = sum of (x_i - x-bar)^2 / (n - 1)
- Standard Deviation: s = sqrt(s^2)
z-Scores and Transformations
- z = (x - x-bar) / s
- y = a + bx: mean_y = a + bmean_x, SD_y = |b|SD_x, shape unchanged
Outlier Rule
- Below Q1 - 1.5IQR or above Q3 + 1.5IQR
Two-Variable Data
Correlation and Regression
- LSRL: y-hat = a + bx, where b = r(sy/sx) and a = y-bar - b(x-bar)
- r (correlation): between -1 and +1, measures linear association
- r^2: proportion of variation in y explained by linear relationship with x
- Residual: observed y - predicted y
Conditions for LSRL appropriateness
- Residual plot shows random scatter (no pattern)
Probability
Rules
- Complement: P(A^c) = 1 - P(A)
- Addition: P(A or B) = P(A) + P(B) - P(A and B)
- Multiplication: P(A and B) = P(A) * P(B|A)
- Independence: P(B|A) = P(B), or P(A and B) = P(A)*P(B)
- Conditional: P(B|A) = P(A and B) / P(A)
Random Variable Rules
- E(a + bX) = a + b*E(X)
- Var(a + bX) = b^2 * Var(X)
- E(X + Y) = E(X) + E(Y) (always)
- Var(X + Y) = Var(X) + Var(Y) (independent only)
- Var(X - Y) = Var(X) + Var(Y) (independent only — always add)
Sampling Distributions
Sample Mean (x-bar)
- Mean: mu_x-bar = mu
- Standard error: sigma / sqrt(n)
- Shape: Normal if population is normal; approximately normal if n >= 30 (CLT)
Sample Proportion (p-hat)
- Mean: p
- Standard error: sqrt(p(1-p)/n)
- Shape: Approximately normal if np >= 10 and n(1-p) >= 10
Confidence Intervals
General Form
**Point Estimate +/- Critical Value Standard Error*
One-Sample z-Interval for Proportion
- Formula: p-hat +/- z sqrt(p-hat(1-p-hat)/n)
- Conditions: Random, np-hat >= 10, n(1-p-hat) >= 10, n <= 0.10N
- z*: 90% = 1.645, 95% = 1.960, 99% = 2.576
Two-Sample z-Interval for Difference of Proportions
- Formula: (p-hat_1 - p-hat_2) +/- z sqrt(p-hat_1(1-p-hat_1)/n_1 + p-hat_2(1-p-hat_2)/n_2)
- Conditions: Both samples random, all counts >= 10, independent samples, 10% each
One-Sample t-Interval for Mean
- Formula: x-bar +/- t (s / sqrt(n))
- Conditions: Random, nearly normal (or n >= 30), n <= 0.10N
- df: n - 1
Two-Sample t-Interval for Difference of Means
- Formula: (x-bar_1 - x-bar_2) +/- t sqrt(s_1^2/n_1 + s_2^2/n_2)
- Conditions: Both random, both normal (or both n >= 30), independent, 10% each
- df: Use smaller of n_1-1 and n_2-1 (conservative) or technology
Paired t-Interval
- Formula: d-bar +/- t (s_d / sqrt(n_pairs))
- Conditions: Random pairs, differences approximately normal, n_pairs <= 0.10N
- df: n_pairs - 1
Confidence Interval for Regression Slope
- Formula: b +/- t SE_b
- Conditions: LINE (Linear, Normal residuals, Independent, Equal variance), Random
- df: n - 2
Hypothesis Tests
Decision Rule
- If p-value <= alpha: Reject H0 (convincing evidence for Ha)
- If p-value > alpha: Fail to reject H0 (not convincing evidence against H0)
One-Sample z-Test for Proportion
- z = (p-hat - p_0) / sqrt(p_0(1-p_0)/n)
- Conditions: Random, np_0 >= 10, n(1-p_0) >= 10, n <= 0.10N
Two-Sample z-Test for Difference of Proportions
- z = (p-hat_1 - p-hat_2) / sqrt(p_pooled(1-p_pooled)(1/n_1 + 1/n_2))
- p_pooled = (x_1 + x_2) / (n_1 + n_2)
One-Sample t-Test for Mean
- t = (x-bar - mu_0) / (s / sqrt(n))
- df: n - 1
Two-Sample t-Test for Difference of Means
- t = (x-bar_1 - x-bar_2) / sqrt(s_1^2/n_1 + s_2^2/n_2)
- df: Smaller of n_1-1, n_2-1 (conservative)
Paired t-Test
- t = (d-bar - 0) / (s_d / sqrt(n_pairs))
- df: n_pairs - 1
t-Test for Regression Slope
- t = (b - 0) / SE_b
- df: n - 2
Chi-Square Tests
Test Statistic
X^2 = sum of (O - E)^2 / E (over all cells)
Goodness of Fit
- df: number of categories - 1
- E_i: n * p_i
- H0: Distribution matches specified proportions
Independence
- df: (rows - 1)(columns - 1)
- E: (row total * column total) / grand total
- H0: No association between variables
- One sample, two variables
Homogeneity
- df: (rows - 1)(columns - 1)
- E: (row total * column total) / grand total
- H0: Same distribution across populations
- Multiple samples, one variable compared
Condition for All
- All expected counts >= 5
Errors and Power
| H0 True | H0 False | |
|---|---|---|
| Reject H0 | Type I (alpha) | Correct (Power = 1 - beta) |
| Fail to Reject | Correct | Type II (beta) |
Power increases: larger n, larger alpha, larger effect size, smaller sigma
TI-84 Calculator Commands
| Task | Command |
|---|---|
| One-variable stats | STAT > CALC > 1-Var Stats L1 |
| Correlation & regression | STAT > CALC > LinReg(a+bx) L1, L2 |
| Enable diagnostics | 2nd > CATALOG > DiagnosticOn |
| Normal probability | 2nd > DISTR > normalcdf(lower, upper, mu, sigma) |
| Inverse normal | 2nd > DISTR > invNorm(area, mu, sigma) |
| Binomial probability | 2nd > DISTR > binompdf(n, p, x) or binomcdf(n, p, x) |
| Geometric probability | 2nd > DISTR > geompdf(p, x) or geomcdf(p, x) |
| z-test for proportion | STAT > TESTS > 1-PropZTest |
| z-interval for proportion | STAT > TESTS > 1-PropZInt |
| Two-proportion z-test | STAT > TESTS > 2-PropZTest |
| Two-proportion z-interval | STAT > TESTS > 2-PropZInt |
| t-test for mean | STAT > TESTS > T-Test |
| t-interval for mean | STAT > TESTS > TInterval |
| Two-sample t-test | STAT > TESTS > 2-SampTTest |
| Two-sample t-interval | STAT > TESTS > 2-SampTInt |
| Paired t-test | Compute differences, then use T-Test |
| t-test for slope | STAT > TESTS > LinRegTTest |
| Chi-square GOF | STAT > TESTS > X^2GOF-Test (or manual) |
| Chi-square test | STAT > TESTS > X^2-Test |
| t critical value | 2nd > DISTR > invT(area_to_left, df) |
| Chi-square cdf | 2nd > DISTR > X^2cdf(lower, upper, df) |
| t cdf | 2nd > DISTR > tcdf(lower, upper, df) |
Key Interpretation Templates
Confidence Interval: "We are [level]% confident that the true [parameter] for [population] is between [lower] and [upper] [units]."
Hypothesis Test (reject): "Since the p-value ([value]) is less than alpha ([value]), we reject H0. There is convincing evidence that [conclusion in context]."
Hypothesis Test (fail to reject): "Since the p-value ([value]) is greater than alpha ([value]), we fail to reject H0. There is not convincing evidence that [conclusion in context]."
Slope: "For each additional [x-unit], the predicted [y-variable] [increases/decreases] by [b] [y-units]."
r^2: "Approximately [r^2 x 100]% of the variation in [y] is explained by the linear relationship with [x]."
p-value: "Assuming [H0 statement] is true, the probability of observing a [test statistic] as extreme as or more extreme than [calculated value] is [p-value]."
Exam strategy
1AP Statistics — Exam Strategy Guide
- Answer every multiple-choice question. There is no penalty for guessing. If you can eliminate even one or two choices, your odds improve dramatically.
- Show all work on FRQs. Even if you make a calculation error, you can still earn most points for correct reasoning, setup, and interpretation.
- Context is king. Every interpretation must be in the context of the problem. A naked number with no context will earn minimal credit.
- Communication is graded. Clarity, completeness, and correct statistical language all matter.
Section I: Multiple-Choice Strategy
Time Management
- 90 minutes for 40 questions = approximately 2 minutes and 15 seconds per question.
- Do not get stuck on any question. If a question is taking more than 3 minutes, mark it, skip it, and come back.
- Aim to finish the first pass in about 70–75 minutes, leaving 15–20 minutes for skipped questions.
- If time is running short, pick a letter (e.g., B) and fill in all remaining bubbles with that letter. Do not leave anything blank.
Process of Elimination
- Eliminate clearly wrong answers first. This is often faster than solving from scratch.
- Watch for answers that are "almost right" but have a subtle error (e.g., using sigma instead of s, using z instead of t, confusing p-hat and p_0).
- Use dimensional analysis. If the answer should be in dollars and an option is in dollars per hour, eliminate it.
Common MCQ Traps
- Answers that swap the null and alternative hypotheses.
- Answers that interpret a p-value as "the probability that H0 is true."
- Answers that claim causation from an observational study.
- Answers that use the wrong standard error (e.g., using p-hat instead of p_0 in a test).
- Answers that state "accept H0" instead of "fail to reject H0."
Strategic Guessing
- When you have no idea, eliminate at least one choice and guess from the rest.
- Common patterns: If you have narrowed it to two answers and one includes specific numbers from the problem, it is more likely correct.
- For inference questions, answers that include the phrase "there is convincing evidence" are often correct if the math supports rejection.
Section II: Free-Response Strategy
Time Management
- 90 minutes for 5 questions = 18 minutes per question.
- Budget your time: Spend approximately 12–13 minutes on each of Questions 2–5, and about 25–30 minutes on Question 1 (the Investigative Task).
- Read all five questions quickly (2–3 minutes) before starting. Identify the easiest ones and do them first to build confidence.
The FRQ Structure
Each FRQ typically has 2–4 parts (labeled a, b, c, d...). Parts are often independent — if you get stuck on part (a), you can often still earn full credit on later parts.
Essential FRQ Checklist
Before submitting any FRQ answer, verify you have included:
- Hypotheses stated in symbols and words (for test questions)
- Conditions named and checked (this alone is often worth 1–2 points)
- Test statistic or CI formula shown (at least the formula with numbers plugged in)
- p-value or interval computed
- Decision stated (reject or fail to reject H0)
- Conclusion in context with the phrase "convincing evidence" (or "not convincing evidence")
- Graphs with titles, labeled axes, and units
- Interpretation of every computed quantity
How to Write a Complete FRQ Answer
For a Confidence Interval:
- Identify the procedure: "A one-sample [z/t]-interval for [a proportion/the mean]..."
- State and check conditions.
- Show the calculation: "p-hat +/- z SE = ... +/- ... = (lower, upper)"
- Interpret: "We are [level]% confident that..."
For a Hypothesis Test:
- State hypotheses in symbols and words.
- State and check conditions.
- Show the test statistic calculation.
- State the p-value (with direction).
- Compare p-value to alpha.
- State the decision.
- State the conclusion in context.
The Investigative Task (Question 1)
The Investigative Task is worth more points and often combines multiple topics. It may ask you to:
- Connect concepts from different units
- Identify flaws in a statistical argument
- Extend analysis beyond standard procedures
- Interpret results in a broader context
Strategy for the Investigative Task:
- Read the entire question twice before writing anything.
- Answer every part. Even partial answers earn partial credit.
- Use sub-questions. The task often has 6–8 parts. Do not skip any.
- Think critically. This question may ask you to evaluate a claim, identify confounding, or explain why a method is inappropriate.
- Do not panic. If a part seems unfamiliar, focus on clear reasoning and context.
- Time is the biggest challenge. Do not over-explain any single part. Be concise and move on.
Common FRQ Mistakes That Cost Points
- Not naming conditions. Writing "np >= 10" without saying "the Large Counts condition is met" may lose the condition point.
- Not interpreting in context. "We reject H0" without saying what that means about the problem earns minimal credit.
- Saying "accept H0." Always say "fail to reject H0."
- Confusing confidence level with probability. Never say "there is a 95% probability that mu is in the interval."
- Forgetting to label graphs. Every graph needs a title and labeled axes.
- Writing conclusions without justification. On the AP exam, the "why" is as important as the "what."
Calculator Tips
Before the Exam
- Enable DiagnosticOn: Press
2nd > 0 (CATALOG), scroll to DiagnosticOn, press ENTER twice. This enables r and r^2 in regression output. - Reset your calculator (optional but recommended if it has old data) via
2nd > + > 7 > 1 > 2. - Practice key sequences until they are automatic. You should be able to reach any test or interval in under 10 seconds.
- Bring extra batteries. A dead calculator is a disaster.
During the Exam
- Enter data in lists before starting FRQs when possible. Use L1 for x, L2 for y.
- Do not round intermediate calculations. Store results in variables (e.g., after 1-Var Stats, VARS > 5 > x-bar gives the exact mean).
- On FRQs, show the formula and numbers you entered, not just the calculator output. Write the test statistic and p-value from the screen.
- Know the difference between Sx (sample SD) and sigma-x (population SD). In AP Statistics, use Sx.
Time-Saving Calculator Tricks
- Quick proportions: After
1-PropZTestor2-PropZTest, the calculator displays the test statistic, p-value, and the sample proportion(s). No need to compute separately. - Quick CI: After any interval command, the calculator displays the interval, point estimate, margin of error, and sometimes n.
- Quick chi-square:
X^2-Testgives the statistic, p-value, df, and expected counts (stored in the output matrix). - Storing values: After any test or interval, you can access specific values using
VARS > 5 >for test statistics, or2nd > MEM > 2 >for matrix values.
Final Checklist: Night Before the Exam
- [ ] Calculator is charged (or has fresh batteries)
- [ ] DiagnosticOn is enabled
- [ ] You know where every test and interval command is
- [ ] You have memorized key interpretation templates
- [ ] You can state conditions for every procedure
- [ ] You understand the difference between: random sampling vs. random assignment, observational study vs. experiment, confidence level vs. interval, p-value vs. alpha
- [ ] You can distinguish: stratified vs. cluster sampling, independence vs. homogeneity chi-square, paired vs. two-sample t-test
- [ ] You know to say "fail to reject H0" (never "accept H0")
- [ ] You know that correlation does not imply causation
- [ ] You know to use t (not z) when sigma is unknown
- [ ] You know to use p_0 in the test statistic and p-hat in the confidence interval for proportions
- [ ] You know to add variances for X - Y (never subtract)
- [ ] Get a good night's sleep. The exam tests reasoning, not just computation.
Presentation outline
1AP Statistics — Presentation Outline (~50 Slides)
Slide 1: Title Slide
- AP Statistics Complete Review
- Your name / date
Slide 2: Exam Format
- Section I: 40 MCQ, 90 min, 50%
- Section II: 5 FRQ (1 investigative + 4 short), 90 min, 50%
- Graphing calculator required
- Formula sheet provided
Slide 3: The Four Statistical Practices
- Statistical Thinking: purpose of study, type of study, validity
- Data Analysis: displays, summaries, technology
- Probability and Simulation: randomness, models, simulation
- Statistical Argumentation: constructing arguments, interpreting results
Slide 4: Nine Units Overview
- Units 1–5: Foundation (~55% of exam)
- Units 6–9: Inference (~45% of exam)
- Brief label for each unit
Slide 5: Today's Agenda
- Review each unit in 5–6 slides
- Focus on key formulas, conditions, and common mistakes
- Practice tips and calculator commands throughout
Slides 6–11: Unit 1 — Exploring One-Variable Data
Slide 6: Types of Data
- Categorical: groups, bar charts, pie charts, frequency tables
- Quantitative: numbers, histograms, dotplots, stemplots, boxplots
- Test: Does it make sense to average? If yes, quantitative.
Slide 7: Measures of Center and Spread
- Center: Mean (balance point, not resistant), Median (resistant), Mode
- Spread: Range, IQR (resistant), Variance (s^2), Standard Deviation (s, not resistant)
- Report both mean and median for quantitative data
Slide 8: Shape and Outliers
- Shape: Symmetric, right-skewed (mean > median), left-skewed (mean < median)
- Describing distributions: Shape, Center, Spread, Unusual features
- 1.5 x IQR rule for outliers: below Q1 - 1.5IQR or above Q3 + 1.5IQR
Slide 9: z-Scores and Transformations
- z = (x - x-bar) / s: how many SDs from the mean
- y = a + bx: new mean = a + b(old mean), new SD = |b|(old SD), shape unchanged
- z-scores are unitless; useful for comparing across different distributions
Slide 10: Boxplots and Five-Number Summary
- Min, Q1, Median, Q3, Max
- Box spans Q1 to Q3; line at median; whiskers to min/max (or fences)
- Boxplots are best for comparing distributions across groups
Slide 11: Unit 1 — Key Takeaways
- Always describe shape, center, spread, and unusual features
- Use median/IQR for skewed data; mean/SD for symmetric data
- The 1.5*IQR rule identifies potential outliers
- Adding a constant changes center only; multiplying changes center AND spread
Slides 12–17: Unit 2 — Exploring Two-Variable Data
Slide 12: Scatterplots and Correlation
- Scatterplot: x (explanatory) vs. y (response)
- Look for: direction, form, strength, unusual points
- r: between -1 and +1, measures linear association only, unitless, not resistant
Slide 13: Least-Squares Regression Line
- y-hat = a + bx, where b = r(sy/sx)
- Minimizes sum of squared residuals
- Interpret slope: "For each additional [x-unit], predicted [y] changes by [b] [y-units]."
Slide 14: Residuals and r-squared
- Residual = observed y - predicted y; sum of residuals = 0
- Residual plot: residuals vs. x; random scatter = good model; pattern = bad model
- r^2 = proportion of variation in y explained by linear relationship with x
Slide 15: Outliers, Leverage, Influential Points
- Outlier in regression: large residual
- High leverage: extreme x-value
- Influential: removal significantly changes the line (high leverage + large residual)
Slide 16: Causation vs. Correlation
- Correlation does NOT imply causation
- Three alternatives: x causes y, y causes x, lurking variable causes both
- Only randomized experiments can establish causation
Slide 17: Unit 2 — Key Takeaways
- r and LSRL describe LINEAR association only
- Always check a residual plot before trusting the model
- r^2 = variation explained, NOT percentage of points on the line
- Influential points have both high leverage and large residuals
Slides 18–22: Unit 3 — Collecting Data
Slide 18: Sampling Methods
- SRS: every individual and every sample equally likely
- Stratified: SRS from each stratum (homogeneous groups)
- Cluster: random clusters, survey all within selected clusters (heterogeneous groups)
- Systematic: every k-th after random start
- Non-probability: convenience, voluntary response (prone to bias)
Slide 19: Sources of Bias
- Sampling bias (undercoverage)
- Response bias (leading questions, social desirability)
- Nonresponse bias
- Key: A large biased sample is still biased
Slide 20: Observational Studies vs. Experiments
- Observational: observe, no treatment imposed; association only
- Experiment: impose treatment; can establish causation
- Random ASSIGNMENT is the key feature of experiments
Slide 21: Experimental Design Principles
- Control: control group, control other variables
- Random assignment: creates comparable groups
- Replication: enough subjects for reliable results
- Blinding: single (subjects) or double (subjects + experimenters)
Slide 22: Types of Experimental Design
- Completely randomized: random assignment to treatments
- Randomized block: block on a confounding variable, randomize within blocks
- Matched pairs: blocks of size 2; each subject gets both treatments or matched subjects
Slides 23–28: Unit 4 — Probability and Random Variables
Slide 23: Probability Rules
- P(between 0 and 1), P(S) = 1
- Complement: P(A^c) = 1 - P(A)
- Addition: P(A or B) = P(A) + P(B) - P(A and B)
- Multiplication: P(A and B) = P(A) * P(B|A)
Slide 24: Independence and Conditional Probability
- Independence: P(B|A) = P(B); NOT the same as mutually exclusive
- P(B|A) = P(A and B) / P(A)
- Two-way tables: joint, marginal, conditional probabilities
Slide 25: Random Variables
- Discrete: countable values, probability distribution
- Continuous: any value in an interval, density curve
- E(X) = sum of x * P(X=x): long-run average
- Var(X) = sum of (x - E(X))^2 * P(X=x)
Slide 26: Rules for Means and Variances
- E(aX + bY) = aE(X) + bE(Y) always
- Var(aX + bY) = a^2Var(X) + b^2Var(Y) only if independent
- Var(X - Y) = Var(X) + Var(Y) — always ADD for differences
Slide 27: Bayes' Theorem Concept
- Reverses conditional probability
- P(A|B) = P(B|A) * P(A) / P(B)
- Application: medical testing, false positives when disease is rare
Slide 28: Unit 4 — Key Takeaways
- Don't confuse P(A|B) with P(B|A)
- Independence must be verified (P(B|A) = P(B)), not assumed
- For X - Y, variances are ADDED, never subtracted
- E(X+Y) = E(X) + E(Y) always, regardless of independence
Slides 29–32: Unit 5 — Sampling Distributions
Slide 29: What Is a Sampling Distribution?
- Distribution of a statistic over all possible samples of size n
- x-bar is unbiased for mu; p-hat is unbiased for p
- Standard error: how much the statistic typically varies
Slide 30: Central Limit Theorem
- For n >= 30, x-bar is approximately normal regardless of population shape
- If population is normal, x-bar is normal for any n
- CLT enables normal-based inference for means
Slide 31: Sampling Distribution of p-hat
- Mean = p, SE = sqrt(p(1-p)/n)
- Approximately normal when np >= 10 and n(1-p) >= 10
- SE decreases as n increases
Slide 32: Law of Large Numbers vs. CLT
- LLN: As n increases, statistic approaches parameter (accuracy)
- CLT: For fixed large n, the distribution of the statistic is approximately normal (shape)
- Larger n reduces variability (SE), not bias
Slides 33–38: Units 6–7 — Inference for Proportions and Means
Slide 33: Confidence Intervals — General
- Point estimate +/- critical value * standard error
- Conditions: Random, Normal (or Large Counts), Independent (10%)
- Interpretation: "We are [level]% confident that..."
- Higher confidence = wider interval; larger n = narrower interval
Slide 34: One-Sample z for Proportions
- CI: p-hat +/- z sqrt(p-hat(1-p-hat)/n)
- Test: z = (p-hat - p_0) / sqrt(p_0(1-p_0)/n)
- CRITICAL: Use p-hat in CI, p_0 in test
Slide 35: t-Distributions and Means
- Use t when sigma is unknown (almost always)
- Heavier tails than z; approaches z as df increases
- df = n - 1 (one sample), n_pairs - 1 (paired)
Slide 36: One-Sample and Paired t
- Paired: compute differences, then one-sample t on differences
- CI: x-bar +/- t (s/sqrt(n))
- Test: t = (x-bar - mu_0) / (s/sqrt(n))
- Conditions: Random, Nearly Normal, Independent
Slide 37: Two-Sample t for Means
- Independent groups, NOT paired
- t = (x-bar_1 - x-bar_2) / sqrt(s_1^2/n_1 + s_2^2/n_2)
- df = min(n_1-1, n_2-1) or use technology
- Do NOT pool standard deviations
Slide 38: Type I/II Errors and Power
- Type I: Rejecting true H0 (probability = alpha)
- Type II: Failing to reject false H0 (probability = beta)
- Power = 1 - beta: increases with larger n, larger alpha, larger effect
Slides 39–42: Unit 8 — Chi-Square
Slide 39: Chi-Square Overview
- X^2 = sum of (O-E)^2/E, always right-skewed, one-sided
- Three types: Goodness of Fit, Independence, Homogeneity
- All expected counts must be >= 5
Slide 40: GOF vs. Independence vs. Homogeneity
- GOF: One variable, one sample, test against specified distribution
- Independence: Two variables, one sample, test for association
- Homogeneity: One variable, multiple samples, test for same distribution
Slide 41: Expected Counts and Degrees of Freedom
- GOF expected: E_i = n * p_i; df = categories - 1
- Independence/Homogeneity: E = (row total * column total) / grand total
- df = (rows - 1)(columns - 1)
Slide 42: Chi-Square on the Calculator
- Enter observed counts in matrix
- STAT > TESTS > X^2-Test
- Output: X^2 statistic, p-value, df, expected counts
Slides 43–46: Unit 9 — Regression Inference
Slide 43: Inference for the Slope
- True line: mu_y = alpha + beta*x; Sample: y-hat = a + bx
- Test H0: beta = 0 (no linear relationship)
- t = b / SE_b; df = n - 2
Slide 44: Conditions (LINE + Random)
- Linear: scatterplot shows linear pattern
- Normal: residuals approximately normal (normal prob plot)
- Independent: random sample, 10% condition
- Equal variance: residual plot shows constant spread
Slide 45: CI for Slope
- b +/- t SE_b
- If interval contains 0, fail to reject H0: beta = 0
- SE_b = s / (sx * sqrt(n-1))
Slide 46: Interpreting Results
- Reject H0: "There is convincing evidence of a linear association."
- Fail to reject: "There is not convincing evidence of a linear association."
- A significant slope does NOT prove causation
Slides 47–50: Wrap-Up
Slide 47: Top 10 Most-Tested Concepts
- Conditions for inference (every FRQ)
- Interpreting p-values and confidence intervals
- Distinguishing experiment vs. observational study
- Sampling methods and bias
- Normal vs. t procedures
- Type I/II errors
- Paired vs. independent two-sample procedures
- Residual analysis
- Chi-square (independence vs. homogeneity)
- Slope inference (LINE conditions)
Slide 48: Exam Day Reminders
- Bring: calculator (charged), extra batteries, pencils, eraser, photo ID
- No: phone, notes, computer
- 3 hours total; pace yourself
- Answer every MCQ; show all FRQ work
Slide 49: Final Study Plan
- 4–6 weeks before: Work through all unit practice files
- 2–3 weeks before: Take full practice exam under timed conditions
- 1 week before: Review summary sheet, re-read exam strategy guide
- Night before: Light review, calculator check, sleep well
Slide 50: Good Luck!
- You know more than you think
- Trust your preparation
- Read carefully, show your work, and communicate in context
Audio script
1AP Statistics — Audio Review Script
Introduction (1 minute)
Welcome to this AP Statistics audio review. This script covers every unit in the course, hitting the most important concepts, common mistakes, and exam strategies. Think of this as a rapid-fire review you can listen to during a commute, a walk, or a final study session. I'll move through all nine units, emphasizing what you absolutely must know for the exam.
Unit 1: Exploring One-Variable Data (2 minutes)
Let's start with the basics. Every dataset has variables. Categorical variables place individuals into groups — eye color, zip code, type of car. Quantitative variables take numerical values you can average — height, salary, temperature.
The golden test: Can you meaningfully compute an average? If yes, it's quantitative. Zip codes are numbers but they're categorical. You can't compute the average zip code.
For quantitative data, you describe distributions using four features in order: shape, center, spread, and unusual features. Shape could be symmetric, right-skewed, or left-skewed. For center, report both the mean and the median. For spread, report the IQR and standard deviation. Unusual features include outliers, gaps, and clusters.
A critical relationship: In a right-skewed distribution, the mean is pulled to the right, so the mean is greater than the median. In a left-skewed distribution, the mean is less than the median. This is a very common exam question.
The mean and standard deviation are not resistant to outliers — a single extreme value can dramatically change them. The median and IQR are resistant. When data are skewed, report the median and IQR as your primary measures.
Outliers are identified using the 1.5 times IQR rule. Compute Q1 and Q3, find the IQR, then calculate the lower fence as Q1 minus 1.5 times IQR and the upper fence as Q3 plus 1.5 times IQR. Anything beyond these fences is a potential outlier.
z-scores measure how many standard deviations a value is from the mean. They're unitless, making them perfect for comparing values from different distributions. A z-score of 2 means the value is 2 standard deviations above the mean.
Transformations follow clear rules. Adding a constant shifts the center but doesn't change spread or shape. Multiplying by a constant scales both center and spread but doesn't change shape. For example, converting inches to centimeters — multiply by 2.54 — the new mean and new standard deviation are both 2.54 times the original, but the shape is unchanged.
Unit 2: Exploring Two-Variable Data (2 minutes)
When you have two quantitative variables, start with a scatterplot. Look for direction (positive or negative), form (linear or curved), strength (how tightly clustered), and unusual points.
The correlation coefficient r is between negative 1 and positive 1. It measures the strength and direction of the linear association only. If the relationship is curved, r is not meaningful, even if it's near zero. Correlation is unitless, not affected by linear transformations, and not resistant — one outlier can change it dramatically.
The least-squares regression line minimizes the sum of squared residuals. The slope b equals r times sy over sx. Always interpret the slope in context: "For each additional unit of x, the predicted y changes by b units."
Residuals are the differences between observed and predicted values. A residual plot graphs residuals against x-values. If you see random scatter, the linear model is appropriate. If you see a pattern — a curve, a fan shape — the linear model is not appropriate.
r-squared tells you the proportion of variation in y that is explained by the linear relationship with x. If r-squared is 0.72, say: "72 percent of the variation in y is explained by the linear relationship with x." Do not say 72 percent of points are on the line.
Distinguish between outliers (large residuals), high-leverage points (extreme x-values), and influential points (removing them significantly changes the regression line). The most influential points have both high leverage and large residuals.
And the golden rule: Correlation does not imply causation. Only randomized experiments can establish cause and effect.
Unit 3: Collecting Data (2 minutes)
This unit is about how data are gathered, and it's more important than many students realize because the quality of data determines what conclusions are valid.
Sampling methods: A simple random sample gives every possible sample an equal chance. Stratified sampling divides the population into homogeneous groups called strata and takes an SRS from each — this ensures representation. Cluster sampling divides the population into heterogeneous groups called clusters, randomly selects some clusters, and surveys everyone within them — this saves cost. Systematic sampling selects every kth individual after a random start.
Non-probability methods include convenience sampling and voluntary response sampling. These are prone to bias and should generally be avoided.
Bias comes in several forms: sampling bias from undercoverage, response bias from poorly worded or leading questions, and nonresponse bias when selected individuals don't participate. Remember: a large biased sample is still biased.
The fundamental distinction: Observational studies observe and measure — they can show association but never causation. Experiments impose treatments and use random assignment — they can establish causation.
Key experimental design principles: Control — use a control group and control other variables. Random assignment — this creates comparable groups by balancing lurking variables. Replication — use enough subjects for reliable results. Blinding — single-blinding keeps subjects unaware of their treatment; double-blinding keeps both subjects and experimenters in the dark.
Three designs: completely randomized design assigns all subjects randomly. Randomized block design groups similar subjects into blocks first, then randomizes within each block. Matched pairs is a special case of block design with blocks of size two.
Unit 4: Probability and Random Variables (2.5 minutes)
Probability rules to memorize: The complement rule — P of A complement equals 1 minus P of A. The addition rule — P of A or B equals P of A plus P of B minus P of A and B. The multiplication rule — P of A and B equals P of A times P of B given A.
Independence means P of B given A equals P of B. This is not the same as mutually exclusive. If two events are mutually exclusive and both have positive probability, they cannot be independent.
Conditional probability — P of B given A — equals P of A and B divided by P of A. Be very careful: P of A given B is generally not equal to P of B given A. This is a common trap.
Two-way tables organize data on two categorical variables. From them, you can compute joint, marginal, and conditional probabilities. To check independence, verify whether the conditional probability equals the marginal probability.
Random variables assign numerical values to outcomes. Discrete random variables take countable values, described by a probability distribution where the probabilities sum to 1. Continuous random variables take any value in an interval, described by a density curve where probability is area.
The expected value — or mean — of a random variable is the long-run average. For rules: the mean of a linear combination is always a times E of X plus b times E of Y. But variances only add if the variables are independent. And here's a critical point that appears on almost every exam: the variance of a difference X minus Y equals the variance of X plus the variance of Y. You always add variances, even for differences.
Unit 5: Sampling Distributions (1.5 minutes)
A sampling distribution is the distribution of a statistic — like x-bar or p-hat — across all possible samples of a given size. This is a theoretical concept that underpins all of statistical inference.
For the sample mean: the mean of the sampling distribution equals the population mean mu, making x-bar an unbiased estimator. The standard deviation of the sampling distribution — called the standard error — is sigma over the square root of n. Notice that as n increases, the standard error decreases. Averages are less variable than individual observations.
The Central Limit Theorem is one of the most important results in all of statistics. It says that for sample sizes of 30 or more, the sampling distribution of x-bar is approximately normal, regardless of the population's shape. If the population is already normal, the sampling distribution is normal for any sample size.
For sample proportions: the mean of the sampling distribution equals p, and the standard error is the square root of p times 1 minus p over n. The normal approximation is valid when n times p and n times 1 minus p are both at least 10.
Don't confuse the Central Limit Theorem with the Law of Large Numbers. The CLT describes the shape of the sampling distribution for a fixed, large sample size. The Law of Large Numbers says that as the sample size grows, the statistic gets closer to the parameter.
Unit 6: Inference for Proportions (2 minutes)
Inference is where we draw conclusions about populations from samples. There are two main tools: confidence intervals and hypothesis tests.
A confidence interval gives a range of plausible values for a parameter. Its general form is: point estimate plus or minus critical value times standard error. For a proportion: p-hat plus or minus z-star times the square root of p-hat times 1 minus p-hat over n.
Interpretation: "We are 95 percent confident that the true proportion is between lower and upper." The 95 percent refers to the method — if we repeated this procedure many times, about 95 percent of intervals would contain the true proportion.
A hypothesis test starts with a null hypothesis — H-naught — that states a specific value for the parameter. The alternative hypothesis states what we are trying to find evidence for. The test statistic measures how far our estimate is from the null value, in standard error units. The p-value is the probability of observing data as extreme or more extreme than what we got, assuming H-naught is true.
For proportions, use z-procedures. The conditions are: random sampling or assignment, the Large Counts condition — n times p-hat greater than or equal to 10 and n times 1 minus p-hat greater than or equal to 10 for confidence intervals, or using p-naught for tests — and the 10 percent condition.
Critical detail: In the test statistic, use p-naught from the null hypothesis in the standard error. In the confidence interval, use p-hat. This is one of the most commonly tested distinctions.
Type I error is rejecting a true null hypothesis — its probability equals alpha. Type II error is failing to reject a false null hypothesis — its probability is beta. Power equals 1 minus beta and increases with larger sample size, larger significance level, and larger effect size.
For two proportions, use the same z-logic but with differences. In a test, use the pooled proportion. In an interval, use each sample's own proportion.
Unit 7: Inference for Means (2 minutes)
When we don't know the population standard deviation — which is almost always — we use t-procedures instead of z-procedures. The t-distribution is symmetric and bell-shaped but has heavier tails than the standard normal. The tails get thinner as degrees of freedom increase.
For a one-sample t-interval: x-bar plus or minus t-star times s over the square root of n, with degrees of freedom n minus 1. For the t-test: t equals x-bar minus mu-naught divided by s over the square root of n.
The critical difference in conditions: instead of Large Counts, we need the "Nearly Normal" condition. If the sample is small, check a dotplot or histogram for strong skewness or outliers. If n is 30 or more, the Central Limit Theorem covers us.
Paired data occurs when you have two measurements on the same subjects or matched subjects. The procedure is to compute the differences, then run a one-sample t-test on the differences. This is one of the most common exam mistakes to watch for: students sometimes use a two-sample test on paired data. If the same people are measured twice, it's paired.
For two independent samples, use a two-sample t-test or t-interval. The standard error is the square root of s1 squared over n1 plus s2 squared over n2. The degrees of freedom are the smaller of n1 minus 1 and n2 minus 1, or use technology for a more precise value. Do not pool the standard deviations.
Confidence intervals and hypothesis tests are consistent with each other: if a two-sided test rejects at alpha equals 0.05, then the 95 percent confidence interval will not contain the null value.
Unit 8: Chi-Square Tests (1.5 minutes)
Chi-square tests are used for categorical data. All three types use the same test statistic: the sum of observed minus expected, squared, divided by expected, summed over all cells. The chi-square distribution is always right-skewed, and the test is always one-sided — large values provide evidence against the null.
Goodness of Fit: tests whether a single categorical variable matches a specified distribution. Expected counts equal n times the hypothesized proportion. Degrees of freedom equal the number of categories minus 1.
Test for Independence: tests whether two categorical variables are associated in a single population. You have one sample classified on two variables. Expected counts are row total times column total divided by grand total. Degrees of freedom equal rows minus 1 times columns minus 1.
Test for Homogeneity: tests whether the distribution of a categorical variable is the same across several populations. You have multiple samples. The calculations are identical to the independence test, but the design and interpretation differ.
All chi-square tests require that every expected count be at least 5. If any expected count is smaller, combine categories or collect more data.
Unit 9: Regression Inference (1.5 minutes)
In Unit 2, we computed the regression line. In Unit 9, we perform inference about the true population slope.
We test H-naught: beta equals 0, meaning there is no linear relationship. The test statistic is t equals b divided by the standard error of the slope, with n minus 2 degrees of freedom. A confidence interval for the slope is b plus or minus t-star times the standard error of the slope.
The conditions are remembered by the acronym LINE plus Random. Linear — the scatterplot must show a linear pattern. Normal — the residuals must be approximately normally distributed, checked with a normal probability plot. Independent — random sample, 10 percent condition. Equal variance — the residual plot must show roughly constant spread, no fan shape.
If the p-value is below alpha, we reject H-naught and conclude there is convincing evidence of a linear association. If the confidence interval for the slope does not contain zero, the test would reject H-naught at the corresponding significance level.
Remember: a statistically significant slope does not prove causation. The study design determines whether causation can be inferred.
Final Exam Tips (2 minutes)
Let me close with the most important exam strategies.
For multiple-choice: answer every question — there is no penalty for guessing. Use process of elimination. If you're stuck, move on and come back. Watch for common traps: using z instead of t, saying "accept H-naught," confusing p-hat with p-naught in test statistics, and claiming causation from observational studies.
For free-response: show all your work. Even with a calculation error, you can earn most points for correct setup, reasoning, and interpretation. State and check conditions — this is typically worth 1 to 2 points on every inference question. Interpret everything in context. Use the phrase "convincing evidence" in your conclusions. Never say "accept H-naught" — always say "fail to reject H-naught."
For the Investigative Task — Question 1 — expect it to combine topics and require critical thinking. Read the entire question twice before starting. Answer every part, even partially. This question is worth more points, so budget more time — about 25 to 30 minutes.
Calculator tips: enable DiagnosticOn before the exam so you get r and r-squared in regression output. Practice every test and interval command until they're automatic. On the exam, show your formulas and plugged-in numbers, not just calculator output.
Remember the interpretation templates: For confidence intervals, "We are [level] percent confident that the true [parameter] is between [lower] and [upper]." For hypothesis tests, "Since the p-value [value] is less than alpha [value], we reject H-naught. There is convincing evidence that [conclusion in context]."
You've put in the work. Trust your preparation, read every question carefully, and communicate clearly. Good luck on the AP Statistics exam.
End of audio review.