Skip to content

Sample Size for Paired Means

Plan how many matched pairs, such as subjects measured before and after an intervention, you need to reliably detect a target change, expressed relative to how variable that change is across subjects.

What this answers

This calculator answers "how many matched pairs do I need to reliably detect a change of this standardized size?" using dz, the mean difference divided by the standard deviation of the differences, not the original measurements' own standard deviation, which is the correct scale for a paired design.

How it is calculated

This calculator uses the normal-approximation formula for a paired design, explicitly documented rather than hidden: the required number of pairs is the squared sum of the alpha-based and power-based z critical values, divided by dz squared, rounded up to the next whole pair. If you expect some subjects to drop out before completing both measurements, an attrition-adjusted total is also shown, inflating the raw requirement to compensate.

Worked example

For a standardized difference of .5, a common medium effect size, at the conventional 95% confidence and 80% power, this calculator requires 32 pairs, a well known benchmark result that matches published statistical software's normal-approximation output for this same design.

Assumption audit

Calculated from your data: the raw required pair count from the normal-approximation formula, and the attrition-adjusted count if you specified an expected dropout rate.
Evidence to review: where your assumed dz came from; a dz borrowed from a different population or measurement instrument may not transfer accurately to your own study.
You must verify: that your planned analysis genuinely is a paired comparison (the same subjects or matched units measured twice), not two independent groups, which would need a different calculator entirely.

What this result does not mean

This sample size guarantees the stated power only if your assumed dz turns out to be accurate; an overly optimistic assumed effect size, a common planning mistake, produces an underpowered study even though the arithmetic here is correct.

Common mistakes

Using the original, unpaired standard deviation of your measurements instead of the standard deviation of the paired differences is a common and consequential error; those two numbers can differ substantially whenever pairs are correlated, and dz specifically needs the latter. A second mistake is planning for zero attrition on a design that genuinely expects some subjects to drop out between the two measurements, which quietly understates the number of pairs you need to actually recruit.

Limitations

This calculator uses a normal approximation rather than the exact noncentral t distribution a paired t-test actually follows; the normal approximation is standard practice for planning purposes and differs from the exact result by at most a pair or two in typical cases.