← Back to Blog

Your Power Analysis Is Probably Wrong — Here's Why

Power analysis is usually the first piece of statistics a funder, IRB, or dissertation committee scrutinizes — and in my experience, it's also the piece most often done on autopilot: plug numbers into G*Power, get a sample size, move on. That's not necessarily wrong, but it skips three assumptions that quietly determine whether the number you land on actually means anything.

1. Where did the effect size come from?

The single most common issue I see is an effect size pulled from a published study without asking whether that study's context resembles yours. Published effects are frequently inflated — by publication bias, by small original samples, by researcher degrees of freedom in how the effect was estimated. If your power analysis borrows an effect size from a paper with an N of 40, you are very likely powering your study to detect an effect that doesn't exist at that magnitude in the population you're actually studying. Where possible, use a conservative estimate: the lower bound of a confidence interval, an average across several comparable studies, or — if you have pilot data — your own estimate, treated cautiously given its own uncertainty.

2. Does the analysis match the design?

A power analysis has to be run for the model you're actually going to fit, not a simplified stand-in. If your real analysis is a three-level model — students nested in classrooms nested in schools — but your power analysis assumes a simple two-group t-test, the sample size that comes out the other end is not the sample size your study needs. Clustering reduces effective sample size through the intraclass correlation; ignoring it produces power estimates that are systematically too optimistic. The same applies to planned covariates, repeated measures, and any interaction effects that are central to your hypotheses rather than the main effect alone — interactions typically require substantially larger samples than main effects to detect at the same power.

3. What's your realistic attrition, and did you power for it?

A power analysis that produces a target N is describing your analytic sample, not your recruitment target. If you expect 20% attrition over a two-year longitudinal study — which is not unusual — your recruitment number needs to be inflated accordingly, and ideally your power analysis should account for the fact that attrition is rarely random (a sensitivity analysis under a "worse than expected" attrition scenario is worth including).

A quick gut check before you submit

  • Can you name exactly which published study (or pilot data) your effect size came from, and would you defend it as conservative?
  • Does your power analysis use the same model structure — clustering, repeated measures, covariates — as your actual planned analysis?
  • Have you inflated your recruitment target for expected attrition, separately from your analytic sample size?
  • If your key hypothesis is an interaction, did you power for the interaction specifically, not just the main effects?

If any of those give you pause, it's worth a second look before the number goes into a proposal — a power analysis that doesn't hold up is one of the fastest ways to lose a reviewer's confidence in everything else in the design. Happy to take a look.