Glossary

Confidence interval

Published 2 min read

What is a confidence interval?

A confidence interval is a range calculated from data to express uncertainty about an estimated quantity, such as a mean or a difference in attendance rates. A 95% confidence procedure would produce intervals containing the true value in about 95% of repeated samples under its assumptions. This is a property of the procedure, as NIST explains, not a statement that 95% of people fall inside the interval.

Once a frequentist interval has been calculated, its endpoints are fixed. Strictly, the procedure does not assign a 95% probability to the fixed population value lying inside that particular interval. The practical reading is that the interval shows a range of effect sizes reasonably compatible with the data and model at the chosen level. It is not a wall separating possible from impossible values.

Calculate an interval for an attendance difference

Use a hypothetical experiment with independent participants, random assignment, one attendance outcome per person and complete follow-up. At the same event, 1,000 of 5,000 people in the usual-invitation group attend, compared with 1,100 of 5,000 in the reminder group. The estimated effect is 0.22 − 0.20 = 0.02, or two percentage points.

For this large-sample illustration, estimate the standard error of the difference as:

SE = √[(0.22 × 0.78 ÷ 5,000) + (0.20 × 0.80 ÷ 5,000)] ≈ 0.00814.

The standard error describes the sampling variability of the estimated difference. It is 0.814 percentage points here, not 0.814% of the effect. A two-sided normal-approximation 95% interval is the estimate plus or minus 1.96 standard errors:

0.02 ± 1.96 × 0.00814 ≈ 0.00404 to 0.03596.

On the attendance scale, that is approximately 0.4 to 3.6 percentage points. The estimate is positive and the interval excludes zero under this procedure. It still leaves a substantial difference between the smaller and larger plausible benefits. Neither endpoint tells us how much a specific individual changed.

What makes an interval wider?

With only 500 people in each group and the same observed rates, the same approximation gives roughly −3.0 to 7.0 percentage points. There is now too little precision to distinguish a modest decline from a useful increase. The interval including zero does not prove that the reminder has no effect.

More independent observations generally improve precision. Greater variability, a higher confidence level, or dependence between observations can require wider intervals. If whole workplaces were assigned together, thousands of employees would not behave like thousands of independent assignment units. The analysis must account for that clustering. Paired observations, very rare outcomes and small samples also need methods suited to their design; this simple formula is not universal.

An interval is only as useful as its comparison

A narrow interval around a biased estimate is still misleading. If reminder recipients volunteered while the comparison group did not, differences in prior interest might explain attendance. A larger sample would not remove that problem. Missing outcomes and inconsistent check-in records can create other errors that this formula does not capture.

Finally, compare the interval with a worthwhile effect, not only with zero. If an expensive program needs a five-percentage-point increase to justify itself, even the upper end of this example falls short. If half a percentage point would cover its cost, the interval leaves that decision less settled. Confidence intervals make the uncertainty visible; they do not choose your priorities.