Statistical Power and Required Sample Size Calculator

Power observatory for one mean or paired mean difference

Statistical Power Calculator

Estimate the probability that a z-approximation test rejects its null under a specified true mean shift. Enter the shift in original units, a planning standard deviation, alpha, sample size, direction, and target power; the calculator returns achieved power and the smallest integer n reaching the target.

Define the planned test

Power orbit

Achieved power at current n83.80%
Minimum n for target46
Type II error β16.20%
Standardized effect |δ|/σ0.4167
Noncentral shift at n2.9463

Two-sided rejection boundaries under H₀: Z < −1.9600 or Z > 1.9600. Under δ = 5, the test statistic is modeled as Normal(mean 2.9463, SD 1).

If the planning assumptions are correct, this design rejects the zero-shift null in about 83.8% of repeated studies when the true shift is 5 units.

Scope of this statistical power result

This calculator covers a one-sample mean test or a paired-mean-difference test using a normal approximation and a known or planning standard deviation. For a paired design, δ is the true mean of the within-pair differences and σ is the standard deviation of those differences—not either original measurement’s SD. The observations or pairs are assumed independent.

The calculation does not automatically apply to two independent means, proportions, correlations, regression coefficients, ANOVA, survival outcomes, clustered trials, noninferiority, equivalence, or multiple endpoints. Those tests have different sampling distributions and design inputs. A broad “power” label is useful only when its test family is stated, so the result panel repeats the rejection rule and modeled shift.

Power is conditional. It is the probability of rejecting the null when the true effect equals the entered δ and all planning assumptions hold. It is not the probability that a significant result is true, the chance that an observed estimate equals δ, or a guarantee that one study succeeds.

Six planning quantities

Effect δ

Choose the smallest true shift that matters for the decision, expressed in outcome units. A larger assumed shift raises power but may ignore important smaller effects.

Noise σ

Use a comparable historical, pilot, or conservative standard deviation. Understating σ overstates signal-to-noise and can underpower the study.

Sample n

Power grows because the standard error falls as σ/√n. For paired data, n is complete pairs after exclusions and missingness.

Alpha and direction

Lower alpha raises the evidential boundary and reduces power at fixed n. One-sided testing concentrates alpha in one prespecified direction.

The target power determines the required-n search. Common values such as 80% and 90% are conventions, not universal ethical or scientific rules. Higher target power reduces Type II error under the chosen δ but requires more observations. Feasibility, decision losses, participant burden, and uncertainty in assumptions all belong in the choice.

How achieved power is calculated

Under the null, the z statistic has a standard normal distribution. Under a true shift δ, its mean moves to μZ = δ√n/σ while its modeled standard deviation remains one. For a two-sided test, the null is rejected when Z < −z1−α/2 or Z > z1−α/2. Power is the alternative-distribution probability falling in either rejection tail.

For an upper one-sided test, the rejection boundary is z1−α and power is P(Z above that boundary under the shifted distribution). For a lower test, the boundary is −z1−α. The sign of δ matters in directional tests: a positive shift does not create high power for a lower alternative. The calculator rejects a required-n request when the entered shift points opposite the selected direction.

Required n is found by evaluating exact normal-approximation power at successive integers from 2 upward until the target is met. This avoids relying only on a closed-form two-sided approximation. All counts are complete analyzable observations; inflate separately for attrition, clustering, or unusable data.

Worked example: detect a five-unit shift

The default specifies δ = 5, σ = 12, n = 50, α = 0.05, and a two-sided alternative. The standardized planning effect is 5/12 = 0.4167. At n = 50, the alternative mean of the z statistic is 5√50/12 ≈ 2.9463.

The two-sided null boundaries are ±1.9600. Under the shifted Normal(2.9463,1) distribution, most rejection probability lies above +1.96, with a tiny contribution below −1.96. Total power is about 83.8%, so β is about 16.2%. In repeated valid studies under exactly this true shift, roughly 84 of 100 would be expected to reject the zero-shift null.

Searching integer n for an 80% target finds 46. The familiar planning approximation [z1−α/2+zpower]²(σ/δ)² gives a similar starting point. The calculator reports the exact normal-tail result for its selected integer rather than treating the approximation as the final power.

Choosing an effect worth detecting

A power calculation is only as meaningful as δ. Avoid selecting the effect observed in a small pilot without considering upward sampling error. Define a minimum important difference from clinical, operational, educational, engineering, or policy consequences. If several values are plausible, report a power curve or scenarios rather than one optimistic point.

Standardized effects aid comparison but do not replace original units. A standardized 0.4 may be consequential on one outcome and trivial on another. Report both δ and δ/σ, and explain the source of σ. For paired studies, account for the correlation through the SD of differences; pairing can improve or worsen precision depending on that variability.

If the meaningful effect is an interval of practically equivalent values around zero, an equivalence design is needed. Failing to reject a two-sided zero-effect null at low power is not evidence of equivalence.

Assumptions and sensitivity checks

Independent, valid measurements

Clusters, repeated outcomes, family units, or serial dependence change the effective information. Design-based or multilevel power methods may be required.

Distribution and variance

The z model treats σ as known and uses a normal approximation. Small samples with estimated σ need noncentral-t calculations for more accurate power.

Complete-case count

Missing outcomes and exclusions reduce n and can cause bias. Divide the required complete count by a defensible retention rate for recruitment planning.

Multiplicity and analysis choices

Adjusted alpha for several primary tests changes the boundary. Flexible stopping, subgroup searches, or outcome switching invalidate the advertised power and Type I error.

Run sensitivity cases for smaller δ, larger σ, lower retention, and stricter alpha. A robust plan should not collapse under modest departures. Simulation is often better when the estimator, missingness, allocation, or stopping rule is complex.

Prospective power versus observed power

Use power prospectively, before results, to choose a design and expose assumptions. “Observed” or post hoc power calculated from the observed effect is generally a one-to-one transformation of the p-value and adds little insight. After the study, emphasize the effect estimate, confidence interval, data quality, and whether the interval excludes effects that matter.

Document the test, alternative, alpha, δ, σ, target power, required complete n, attrition inflation, software method, and date.

Frequently asked questions

Is 80% power always enough?

No. It is a common convention. The appropriate target depends on the cost of missed effects, burden, feasibility, uncertainty, and field standards. Report the chosen rationale.

Can I use a negative δ?

Yes for a two-sided or lower-direction calculation. For an upper alternative, a negative true shift points away from the rejection region and cannot reach a high target merely by increasing n.

Is sigma the standard error?

No. Enter the standard deviation of individual observations or paired differences. The calculator forms the standard error as σ/√n.

Does required n include dropouts?

No. It is the analyzable complete count. Inflate recruitment for anticipated attrition and address whether missingness could bias the estimate.

Can I calculate two independent groups here?

Not accurately. Allocation and both group variances affect the standard error, and a t or z distribution specific to that design is needed.

Why not compute power from my observed result?

Post hoc power based on the observed effect largely restates the p-value. Confidence intervals and effect estimates provide more useful information after data are known.

Protocol-ready power statement

A concise planning statement should name the estimand and every assumption: “A two-sided one-sample test at α=.05 requires 46 complete observations to provide at least 80% normal-approximation power for a five-unit mean shift when the population standard deviation is 12.” Add the recruitment count after attrition inflation, the source of the variability estimate, the software or page version, and the intended primary analysis.

Do not write only “the study was powered at 80%.” That sentence hides which effect, variance, alternative, and sample count produced the number. If the design changes after enrollment begins, retain both calculations and explain the change. If a blinded variance review is planned, state its method and decision rule in advance. Transparent assumptions let another analyst reproduce the calculation and judge whether the selected effect remains meaningful.

For several scenarios, create a small sensitivity table with rows for effect size and columns for variability or retention. The largest scientifically defensible complete-n requirement can guide capacity planning, while the full table shows stakeholders the tradeoffs that one preferred number conceals.

If the planned study compares three or more independent group means, review the analysis structure with the one-way ANOVA statistical analysis.

References

The planning relationship follows the U.S. National Institute of Standards and Technology guidance on sample sizes required for tests of a mean. NIST gives two-sided N = (z1−α/2+z1−β)²(σ/δ)² and the corresponding one-sided expression when σ is assumed known. This calculator additionally evaluates the normal rejection-tail probability directly.

Scroll to Top