ScienceHub
← AP® Statistics · All units
Notes & flashcards are free to explore. Create an account for practice and saved progress.Sign up free
On this page0% through guide
UNIT 4About 11 min + practice

Inference for Quantitative Data: Means

Use t procedures to reason about population averages and paired changes.

What you’ll learn

  • Choose one-sample, paired, or independent two-sample mean inference.
  • Check design and distribution conditions.
  • Interpret intervals, P-values, and study limitations.
01

Before you begin

A mean summarizes quantitative observations. When the population standard deviation is unknown, the sample standard deviation estimates it. A paired response is one numerical difference per matched pair.

Explain these starting ideas in your own words. Revisit them whenever a later step feels unclear.

02

A mean parameter needs context and units

Define μ as the population mean of the measured quantity. The sample mean estimates μ, and the sample standard deviation estimates individual-level variation. The standard error describes the estimated variability of sample means, not the variability of individual observations.

When the population standard deviation is unknown, t procedures account for estimating it. A t distribution is centered at zero with heavier tails than the standard normal, especially at small degrees of freedom. For a one-sample procedure, degrees of freedom are n − 1.

PAUSE & TRY IT

What are one-sample t degrees of freedom for n = 25?

Reveal answer

24.

03

Conditions concern design and shape

Random sampling or appropriate random assignment supports the inferential model. When sampling without replacement, the 10% condition is commonly used for approximate independence. Examine quantitative data for strong skew and outliers, especially for small samples.

A nearly normal population supports small-sample t methods. With larger samples, the sampling distribution of the mean can become more nearly normal, but extreme outliers, severe heavy tails, or dependence still require care. A large n does not repair selection bias or turn a time-dependent series into independent observations.

04

One-sample intervals and tests use related but different questions

A t interval estimates plausible values for μ using x̄ ± t*. A test compares the observed mean with a specified null mean using t = . The alternative determines which tail or tails define the P-value.

Interpret an interval for the population mean, not individual outcomes. A narrow interval can estimate the mean precisely even while individuals vary widely. If the scientific question concerns individual prediction rather than a mean, a confidence interval for μ is not the correct answer.

t =
interval: x̄ ± t*
The standard error follows a square-root ruleHolding sample standard deviation at 12 units, SE=12/√n. Increasing sample size improves precision with diminishing returns.
The standard error follows a square-root rule01234062.5125187.5250 Sample size nStandard error (response units)SE=12/√n
Read figure values as text

SE=12/√n: 10: 3.794733192202055; 16: 3; 22: 2.5584085962673253; 28: 2.2677868380553634; 34: 2.057983021710106; 40: 1.8973665961010275; 46: 1.7693034738587656; 52: 1.6641005886756874; 58: 1.5756771943166705; 64: 1.5; 70: 1.4342743312012722; 76: 1.3764944032233704; 82: 1.3251783128981585; 88: 1.2792042981336627; 94: 1.237705495510552; 100: 1.2; 106: 1.165543034828717; 112: 1.1338934190276817; 118: 1.104689541477988; 124: 1.0776318121606494; 130: 1.0524696231684352; 136: 1.028991510855053; 142: 1.007017629956027; 148: 0.9863939238321437; 154: 0.9669875568304563; 160: 0.9486832980505138; 166: 0.9313806308475994; 172: 0.914991421995628; 178: 0.8994380267950337; 184: 0.8846517369293828; 190: 0.870571500132014; 196: 0.8571428571428571; 202: 0.8443170536763502; 208: 0.8320502943378437; 214: 0.8203031124295959; 220: 0.8090398349558905; 226: 0.7982281262852872; 232: 0.7878385971583353; 238: 0.7778444682625973; 244: 0.7682212795973759; 250: 0.7589466384404111

05

Paired data become one sample of differences

Before-and-after measurements on the same individuals or deliberately matched pairs are dependent within pairs. Define a difference direction, calculate one difference per pair, and perform a one-sample t procedure on those differences. The relevant standard deviation is the standard deviation of differences.

Check distribution conditions for the differences, not separately for each original measurement set. Pairing can reduce variation by removing stable individual differences, but it does not guarantee a causal conclusion if time trends or other changes are uncontrolled. Randomized order or a suitable control may be needed.

Pairing makes within-person change visibleIllustrative model, not collected experimental data. Illustrative scores differ at baseline but all increase by four. A paired analysis examines each difference; treating the measurements as unrelated loses that structure.
Pairing makes within-person change visible025507510001234 StudentScoreBeforeAfter
Read figure values as text

Before: 1: 50; 2: 65; 3: 80; 4: 90 • After: 1: 54; 2: 69; 3: 84; 4: 94

PAUSE & TRY IT

Why is a randomized controlled trial stronger than an uncontrolled before/after study for causation?

Reveal answer

Random assignment and a concurrent control help separate treatment effects from other differences and time-related changes.

06

Independent groups need a two-sample model

For independent groups, estimate μ1 − μ2 with x̄1 − x̄2. The standard error combines and under a square root. The usual unpooled two-sample t procedure does not assume equal population variances. Technology may use approximate Welch degrees of freedom.

Keep the group order consistent in hypotheses, calculations, and conclusions. An interval containing zero does not prove the means are identical; it indicates that zero remains compatible with the interval procedure at that confidence level. A meaningful effect can be missed by a low-power small study.

SE(x̄1 − x̄2) = √( + )

PAUSE & TRY IT

Does an interval for μ1 − μ2 containing zero prove equality?

Reveal answer

No. It indicates insufficient precision to rule out zero at that confidence level, not proof of exact equality.

07

Report effect size and uncertainty together

A small P-value can occur for a very small effect with a large sample, while a large practical effect may be uncertain in a small sample. Report the estimated difference and units as well as evidence against the null. Distinguish no evidence of an effect from evidence that any effect is practically negligible.

A complete response names the procedure, checks conditions using the study, shows the relevant calculation or technology output, and concludes about the population in context. Avoid a generic ending such as “reject H0” with no explanation of what that means for the original question.

08

Why a t distribution appears

Estimating population spread introduces extra uncertainty. A t procedure accounts for this with heavier tails than a standard Normal distribution, especially at small degrees of freedom. For a one-sample mean, degrees of freedom equal n−1. The standard error is s divided by the square root of n, not s divided by n.

Check randomization and independence, then examine sample size, shape, and outliers. With small samples, strong skewness or outliers can make a t procedure unreliable. A large sample improves robustness to non-Normal shape, but it does not fix dependence or biased data collection. Include the data’s context and units in the conclusion.

PAUSE & TRY IT

Why are 20 people measured twice not 40 independent observations?

Reveal answer

Each person’s measurements are related. There are 20 pairs and therefore 20 differences for a paired analysis.

09

Recognize paired data before calculating

Measurements before and after on the same person are paired, as are deliberately matched subjects. Define one difference direction, such as after minus before, and compute a difference for each pair. Analyze those differences with a one-sample t procedure. The sample size is the number of pairs, not the total number of measurements.

Conditions concern the distribution of differences. Even if the two raw measurement distributions are skewed, their differences may behave differently. Reversing the subtraction changes the sign of the estimate, test statistic, and directional hypothesis; it does not change a two-sided p-value.

PAUSE & TRY IT

What information besides statistical significance should be reported?

Reveal answer

The estimated effect size, its uncertainty and units, practical importance, and the study design’s limitations.

10

Compare independent groups and report meaning

Two independent groups require a two-sample method. The standard error of their mean difference combines the two estimated variances divided by their sample sizes. Do not treat two independent groups as pairs merely because the sample sizes match. Conversely, discarding genuine pairing can lose useful information.

State whether evidence concerns an association or a causal effect based on the design. Report an estimated difference and uncertainty, not just “significant.” An interval from 0.1 to 0.3 minutes may exclude zero but represent a small practical gain. A wide interval can include both valuable and negligible effects, revealing a need for greater precision.

11

Identify whether the data are paired

A quantitative response calls for mean inference when the target is a population average. If each subject contributes before and after measurements, analyze the within-subject differences with a one-sample t procedure. State the subtraction order so that the sign has a clear meaning.

Two independent groups require a two-sample procedure. Equal sample sizes do not create pairing, and unequal sample sizes do not by themselves invalidate independent-group inference. Pairing comes from the design and relationship between observations. Treating paired measurements as independent discards their dependence and changes the standard error.

The t model accounts for estimating population standard deviation from sample data. Degrees of freedom depend on the procedure; for one-sample or paired differences, df=n−1 where n is the number of independent observations or pairs. Use the appropriate software output for a two-sample method rather than inventing a single common n.

PAUSE & TRY IT

Should you use 32 as n when 16 students each have before and after scores?

Reveal answer

No. There are 16 independent paired differences. Counting both measurements as independent doubles the apparent replication incorrectly.

12

Check shape and independence before trusting a t result

Inspect quantitative data for strong skewness and outliers, especially with small samples. A large sample can make mean inference more robust, but cannot repair dependent observations or biased sampling. For paired data, examine the distribution of differences, not merely the two separate measurement distributions.

Standard error measures uncertainty in a sample mean, while s measures variation among individual observations. A confidence interval for a mean does not predict where most individuals fall. Keeping these roles separate prevents a common error in interpreting a narrow interval as evidence that all observations are similar.

A one-sample test compares the sample mean with a null mean using the estimated standard error. The alternative determines the tail direction. Choose that direction from the question before inspecting the result; changing it after seeing the data changes the evidential meaning of the procedure.

13

Connect statistical and practical conclusions

A small P-value indicates incompatibility with the null model under the assumptions, not necessarily a large or useful effect. With a very large sample, a tiny difference can be statistically detectable. Report the estimated difference and a confidence interval to discuss practical importance.

A Type I error rejects a true null; a Type II error fails to reject a false null. Power is the probability of rejecting for a specified true alternative. Increasing sample size generally improves power for the same effect and significance level, but the practical meaning depends on which effect matters.

For a study conclusion, separate the numerical inference from the scope of inference. Random assignment can support causality; random sampling can support generalization. A statistically convincing difference from an observational sample still may be confounded.

FROM IDEA TO APPLICATION

Worked examples

EXAMPLE 1

A one-sample interval

A random sample of 16 measurements has mean 52 mg and standard deviation 8 mg. Assume the population is approximately normal. Use t* = 2.131 for a 95% interval.

Reveal worked solution
  1. SE = = 2 mg.
  2. Margin = 2.131 × 2 = 4.262 mg.
  3. Endpoints = 52 ± 4.262.
Result & interpretation

Approximately (47.74, 56.26) mg for the population mean, with 15 degrees of freedom.

EXAMPLE 2

Choose the paired analysis

Twenty people are measured before and after a training program. Why is treating the 40 measurements as two independent samples inappropriate?

Reveal worked solution
  1. Each person’s before and after measurements are linked.
  2. Within-person correlation violates the independent-groups structure.
  3. Analyze 20 consistently defined changes instead.
Result & interpretation

Use a paired t procedure on the 20 differences if its conditions are met; a before/after design alone does not eliminate other time-related causes.

EXAMPLE 3

A paired interval

For 16 randomly selected pairs, after−before has mean 3.0 and standard deviation 4.0. Using t*=2.131, calculate a 95% interval for the population mean difference.

Reveal worked solution
  1. Use the 16 differences as the observations.
  2. SE==1.
  3. Margin of error=2.131×1=2.131.
  4. Compute 3.0±2.131 and retain the after−before direction.
Result & interpretation

(0.869, 5.131), assuming the randomization, independence, and shape conditions are satisfied. The units are those of the original measurement.

EXAMPLE 4

Analyze paired improvement

For 16 independent students, after-minus-before scores have mean 4 and standard deviation 6. Find the standard error and test statistic for zero mean change.

Reveal worked solution
  1. The analysis variable is each student’s difference.
  2. SE = = 1.5 points.
  3. t = ≈ 2.67 with 15 degrees of freedom.
Result & interpretation

Use a paired t analysis, subject to design and difference-distribution conditions; the statistic alone is not the full conclusion.

MAKE THE DISTINCTION

Common mistakes, clearer reasoning

The traps and s/√n measure the same variability.

The better explanations describes individual observations; s/√n estimates variability of sample means.

The trapFor paired inference, check each original group’s normality instead of differences.

The better explanationThe analyzed variable is the within-pair difference, so its distribution is the relevant one.

RETRIEVE BEFORE YOU REVEAL

Practice checkpoints

Revisit the quick checks from this guide without looking back. Explain why, then reveal the answer.

1. What are one-sample t degrees of freedom for n = 25?

Reveal answer

24.

2. Does an interval for μ1 − μ2 containing zero prove equality?

Reveal answer

No. It indicates insufficient precision to rule out zero at that confidence level, not proof of exact equality.

3. Why is a randomized controlled trial stronger than an uncontrolled before/after study for causation?

Reveal answer

Random assignment and a concurrent control help separate treatment effects from other differences and time-related changes.

4. Why are 20 people measured twice not 40 independent observations?

Reveal answer

Each person’s measurements are related. There are 20 pairs and therefore 20 differences for a paired analysis.

5. What information besides statistical significance should be reported?

Reveal answer

The estimated effect size, its uncertainty and units, practical importance, and the study design’s limitations.

6. Should you use 32 as n when 16 students each have before and after scores?

Reveal answer

No. There are 16 independent paired differences. Counting both measurements as independent doubles the apparent replication incorrectly.

Key language

t distribution
A symmetric reference distribution accounting for estimated variability.
Paired difference
A within-pair subtraction with a defined direction.
Degrees of freedom
A parameter describing the reference distribution used by a procedure.
Practical significance
The substantive importance of an effect in its context.
Connect it to the course

Mean inference extends sampling-distribution reasoning while preserving the same design-based limits on conclusions.

Reading marks could not be saved in this browser. They do not affect your practice score.

Written for ScienceHub · Original instructional material. Course framework reference ↗. These notes are independently authored and are not College Board materials. External photographs retain their credited licenses.

YOUR EXPERIENCE MATTERS

How’s your study space?

Sign in to share a review of ScienceHub.

Sign in