ScienceHub
← AP® Statistics · All units
Notes & flashcards are free to explore. Create an account for practice and saved progress.Sign up free
On this page0% through guide
UNIT 3About 13 min + practice

Inference for Categorical Data: Proportions

Estimate population proportions and assess claims with explicit assumptions.

What you’ll learn

  • Construct and interpret proportion intervals.
  • Carry out one- and two-proportion tests.
  • Analyze categorical association with chi-square methods in current course scope.
01

Before you begin

A parameter describes a population; a statistic describes a sample. A proportion measures the fraction in a specified category. Define the success category and population before choosing a procedure.

Explain these starting ideas in your own words. Revisit them whenever a later step feels unclear.

02

Inference begins with a population parameter

Define p as the true proportion of a stated population with a stated characteristic. A sample proportion estimates it. Random sampling or a relevant randomized design supports the probability model; large counts support a normal approximation. For sampling without replacement, a sample no more than about 10% of the population is a common independence check.

For a one-proportion interval, check at least 10 observed successes and 10 failures. For a one-proportion null test, check expected successes and failures using the null proportion. State the actual values so the reader can see whether the condition holds.

  1. Define the targetName the population proportion and success category.
  2. Check the designRandomness, independence, and appropriate count conditions.
  3. Interpret evidenceEstimate or test in context; separate evidence from causality.
03

A confidence interval describes an estimation procedure

A one-proportion z interval uses the sample proportion plus or minus a critical value times its estimated standard error. Higher confidence uses a wider critical-value multiplier; larger sample size generally narrows the interval. Neither change repairs a biased sampling method.

A 95% confidence interpretation concerns the long-run success rate of the method. For one realized interval, say you are 95% confident it captures the population proportion. Do not say 95% of individuals lie in the interval, and do not describe the fixed parameter as randomly moving between interval endpoints.

p̂ ± z*√[p̂]
Larger samples improve precisionFor a sample proportion fixed at 0.5, the approximate 95% margin of error is 1.96√(0.25/n). Quadrupling n halves the margin. This does not correct sampling bias.
Larger samples improve precision00.0250.050.0750.1040080012001600 Sample size nApproximate margin of error95% margin at p-hat=.5
Read figure values as text

95% margin at p-hat=.5: 100: 0.098; 137.5: 0.0835746808113993; 175: 0.07408103670980853; 212.5: 0.06722744537586346; 250: 0.06198064213930023; 287.5: 0.05779724681271967; 325: 0.05436061923007245; 362.5: 0.05147212168101124; 400: 0.049; 437.5: 0.04685296148590823; 475: 0.0449654838386301; 512.5: 0.04328915822133985; 550: 0.04178734040569965; 587.5: 0.0404317128533447; 625: 0.0392; 662.5: 0.03807440580440475; 700: 0.03704051835490427; 737.5: 0.036086525021614274; 775: 0.035202639197247886; 812.5: 0.03438067435683555; 850: 0.03361372268793173; 887.5: 0.03289590924523021; 925: 0.032222201511850034; 962.5: 0.03158826018979491; 1000: 0.030990321069650113; 1037.5: 0.030425100607688247; 1075: 0.029889719785190515; 1112.5: 0.02938164220863777; 1150: 0.028898623406359836; 1187.5: 0.028438669004312456; 1225: 0.027999999999999997; 1262.5: 0.02758102375342744; 1300: 0.027180309615036227; 1337.5: 0.0267965683393668; 1375: 0.026428634608559088; 1412.5: 0.026075452125319382; 1450: 0.02573606084050562; 1487.5: 0.025409585963244843; 1525: 0.02509522846684761; 1562.5: 0.024792256855720094; 1600: 0.0245

04

A significance test asks what the null model predicts

State a null hypothesis with equality and an alternative matching the question before using the data to choose a direction. A one-proportion z test standardizes the observed difference from p0 using the null standard error. The P-value is the probability, assuming the null model, of a statistic at least as extreme in the alternative’s direction.

A small P-value provides evidence against the null, not the probability that the null is true. Compare it with a prespecified significance level. “Fail to reject” means insufficient evidence for the alternative, not proof of no effect. Statistical significance does not establish practical importance.

z = [p0]
05

Compare two proportions using the right standard error

Define p1 − p2 in a clear group order. For an interval, estimate variability separately in the two groups. For a test of equality, pool successes and sample sizes because the null assumes a shared population proportion. The pooled estimate is total successes divided by total observations, not the simple average of proportions when group sizes differ.

The groups must be independent for the ordinary two-proportion procedures. Paired before/after categorical outcomes are not automatically two independent samples. Check each group’s relevant success/failure counts and the study design. Interpret an interval in percentage points, not as a relative percent change unless that is explicitly calculated.

PAUSE & TRY IT

When is the pooled proportion used?

Reveal answer

In the usual two-proportion z test of equal population proportions, not in the ordinary two-proportion interval.

PAUSE & TRY IT

What does an interval for p1 − p2 entirely above zero suggest?

Reveal answer

It supports p1 being greater than p2 at the corresponding confidence level under the procedure’s assumptions.

06

Chi-square compares observed and expected table counts

For a test of independence, one sample is classified by two categorical variables. For a test of homogeneity, independent samples or randomized groups are compared across a categorical response. Both use expected counts based on row and column totals under the null and a chi-square statistic measuring departures.

Expected counts are row total × column total / grand total. Check that all expected counts meet the usual threshold of at least 5. Degrees of freedom are (rows − 1)(columns − 1). A significant result indicates association or differing distributions, not which cell alone caused it or that the relation is causal. Follow up with the pattern of proportions and the design.

χ2 = Σ(O − E)

PAUSE & TRY IT

Why does a significant chi-square result not establish causation?

Reveal answer

Causal interpretation depends on study design, especially random assignment, not the statistic alone.

07

Errors and power depend on a decision rule

A Type I error rejects a true null; a Type II error fails to reject a false null for a particular alternative. Power is the probability of rejecting the null at a specified true alternative. Larger samples generally increase power to detect a fixed effect.

Raising the significance level can increase power but also increases the allowed Type I error rate. Context determines the consequences of each error. Report the actual effect estimate and uncertainty rather than using a significance label as the whole conclusion.

08

Build a proportion interval from the design

A confidence interval estimates an unknown population proportion with a sample proportion plus or minus a margin of error. Check that the data come from a suitable random process, observations are independent, and the sample has enough observed successes and failures for the Normal approximation. When sampling without replacement, check the 10% condition against the population size.

The confidence level describes the long-run success rate of the interval-producing method. After an interval is computed, the fixed population parameter is either inside it or outside it. An AP-level interpretation identifies the population and parameter and states confidence in the interval; it does not say that the specified fraction of observations falls between its endpoints.

Confidence intervals vary from sample to sampleIllustrative model, not collected experimental data. This schematic displays five interval midpoints with error bars. Compare each interval with the same true parameter p=0.50; one displayed interval misses it. Five examples are not a calibration of a 95% method.
Confidence intervals vary from sample to sample00.20.40.60.801.252.53.755 Sample numberProportionEstimate ± margin
Read figure values as text

Estimate ± margin: 1: 0.48; 2: 0.53; 3: 0.46; 4: 0.62; 5: 0.5

PAUSE & TRY IT

A p-value is 0.08 and α=0.05. Is there proof that the null is true?

Reveal answer

No. Fail to reject the null; the evidence is insufficient for the alternative at this significance level.

09

Use the null model for a significance test

State the null and alternative in population parameters before analyzing results. The null gives a reference value; the alternative specifies greater than, less than, or different from. For a one-proportion z test, the standard error and large-count check use the null proportion, because the p-value describes results assuming that null model is true.

A small p-value indicates that the observed statistic, or one more extreme in the direction of the alternative, would be unusual under the null model. It is not the probability that the null is true. Compare the p-value with a previously chosen significance level, then state a contextual conclusion. Failing to reject is not proof that the null is exactly correct.

PAUSE & TRY IT

What does an interval entirely above zero imply about p1−p2?

Reveal answer

Its plausible values are positive, supporting that the first population proportion is larger under the procedure’s conditions.

10

Keep two-proportion procedures distinct

For two independent samples, define p1 and p2 and keep the order of subtraction consistent through hypotheses, interval, and conclusion. An interval for p1−p2 estimates the difference using separate sample proportions in its standard error. A test of p1=p2 uses a pooled estimate under the equality null. Do not substitute one standard error automatically into the other procedure.

A statistically significant difference can be too small to matter practically. Examine the estimate, interval, and context alongside the p-value. A large sample can detect a tiny effect. Also distinguish a Type I error, rejecting a true null, from a Type II error, failing to reject a false null; describe the real-world mistaken decision in each case.

11

Choose the parameter and procedure before the formula

A population proportion describes a categorical outcome. State the population and success category. For two populations, define the order p1−p2 and keep it consistent in hypotheses, calculations, and conclusions. A negative estimated difference is meaningful only relative to that order.

A confidence interval estimates a parameter; a significance test evaluates evidence against a specific null model. For a one-proportion interval, estimated standard error uses the observed proportion. A one-proportion test uses the null proportion in its null standard error and success–failure check. Do not move formulas between purposes without checking their assumptions.

For a two-proportion test of equality, the pooled estimate reflects the null assumption of a common proportion. A two-proportion confidence interval generally uses separate sample estimates in its standard error. Independent groups are required for that procedure; paired categorical outcomes are not automatically two independent samples.

12

Write conditions and interpretations in context

Random sampling or a defensible randomized design supports the intended inference. When sampling without replacement, check that the sample is at most about 10% of the population for the usual independence approximation. Check success and failure counts using values appropriate to the procedure, and name what those counts represent.

A 95% confidence level describes the long-run capture rate of the interval method under its assumptions. For one computed interval, say that you are 95% confident the population parameter lies between its endpoints. Do not say 95% of individual people have proportions in that interval, or that the parameter changes from sample to sample.

A P-value is the probability, assuming the null model and conditions, of a statistic at least as extreme in the direction of the alternative. It is not the probability that the null hypothesis is true. Failure to reject does not prove equality; sample size and variability may limit the evidence.

PAUSE & TRY IT

Does P=0.02 mean the null hypothesis has a 2% chance of being true?

Reveal answer

No. It describes the probability of results at least as extreme under the null model, not a probability assigned to the hypothesis.

13

Use chi-square analysis for the appropriate categorical question

For a two-way table, a chi-square procedure can assess association between categorical variables or compare distributions across populations, depending on the design. Expected counts come from row total × column total / grand total under the null. Degrees of freedom are (rows−1)(columns−1).

The statistic adds (observed−expected) over cells, so large departures increase it regardless of sign. Examine counts and the sampling design before using the reference distribution. The test identifies evidence of a pattern, but does not by itself quantify which category is responsible or establish a causal mechanism.

An association test can be followed by contextual comparison of conditional proportions to describe the pattern. This unit follows the revised course scope; do not substitute a one-variable goodness-of-fit procedure for a two-way-table question. A correct numerical result still needs a conclusion about the stated populations and variables.

FROM IDEA TO APPLICATION

Worked examples

EXAMPLE 1

A proportion interval

In a random sample of 400 eligible voters, 240 support a proposal. Estimate a 95% interval assuming a sufficiently large population.

Reveal worked solution
  1. p̂ = = 0.60; success and failure counts are 240 and 160.
  2. SE = √(0.60 × ) ≈ 0.0245.
  3. Margin ≈ 1.96(0.0245) ≈ 0.048.
Result & interpretation

Approximately (0.552, 0.648). We are 95% confident that the population support proportion is between 55.2% and 64.8%, subject to the design assumptions.

EXAMPLE 2

Expected table count

A contingency table has row total 80, column total 150, and grand total 300 for a particular cell. Find its expected count under independence.

Reveal worked solution
  1. E = .
Result & interpretation

40 expected observations in that cell.

EXAMPLE 3

Interpret an interval correctly

A random sample produces a 95% confidence interval of (0.41, 0.53) for the proportion of district students who walk to school. Interpret it.

Reveal worked solution
  1. Identify the target as the population proportion of district students who walk.
  2. Use the interval endpoints as plausible values for that proportion.
  3. Attach 95% confidence to the estimation method rather than to individual students.
Result & interpretation

We are 95% confident that between 41% and 53% of district students walk to school.

EXAMPLE 4

Interpret a difference interval

A 95% interval for p1−p2 is (0.03, 0.11).

Reveal worked solution
  1. The order is population 1 minus population 2.
  2. Every plausible difference in this interval is positive.
  3. Express endpoints as percentage points, not a relative percentage increase.
Result & interpretation

We are 95% confident population 1’s proportion is 3–11 percentage points higher than population 2’s, under the procedure’s assumptions.

MAKE THE DISTINCTION

Common mistakes, clearer reasoning

The trapThe P-value is the probability the null hypothesis is true.

The better explanationIt is a probability of data at least as extreme under the assumed null model.

The trapA confidence interval for a proportion describes the middle 95% of people.

The better explanationIt estimates one population proportion, not a distribution of individual people.

RETRIEVE BEFORE YOU REVEAL

Practice checkpoints

Revisit the quick checks from this guide without looking back. Explain why, then reveal the answer.

1. When is the pooled proportion used?

Reveal answer

In the usual two-proportion z test of equal population proportions, not in the ordinary two-proportion interval.

2. What does an interval for p1 − p2 entirely above zero suggest?

Reveal answer

It supports p1 being greater than p2 at the corresponding confidence level under the procedure’s assumptions.

3. Why does a significant chi-square result not establish causation?

Reveal answer

Causal interpretation depends on study design, especially random assignment, not the statistic alone.

4. A p-value is 0.08 and α=0.05. Is there proof that the null is true?

Reveal answer

No. Fail to reject the null; the evidence is insufficient for the alternative at this significance level.

5. What does an interval entirely above zero imply about p1−p2?

Reveal answer

Its plausible values are positive, supporting that the first population proportion is larger under the procedure’s conditions.

6. Does P=0.02 mean the null hypothesis has a 2% chance of being true?

Reveal answer

No. It describes the probability of results at least as extreme under the null model, not a probability assigned to the hypothesis.

Key language

Confidence level
The long-run capture rate of an interval procedure under its assumptions.
P-value
The probability of a result at least as extreme under the null model.
Power
The probability of rejecting a null at a specified alternative truth.
Pooled proportion
Total successes divided by total observations across groups.
Connect it to the course

The logic of parameter, conditions, estimate/test, and contextual conclusion also governs mean inference.

Reading marks could not be saved in this browser. They do not affect your practice score.

Written for ScienceHub · Original instructional material. Course framework reference ↗. These notes are independently authored and are not College Board materials. External photographs retain their credited licenses.

YOUR EXPERIENCE MATTERS

How’s your study space?

Sign in to share a review of ScienceHub.

Sign in