9 min read

How to Calculate Sample Size with Power Simulation

An underpowered study can't be rescued after the fact. Here's how to calculate sample size by simulation — before you collect a single data point.
phd-compass graphic "Power your study before you start" — power-simulation steps: Effect size, Simulate, Repeat, Find N.
An underpowered study is a study you cannot publish.

Somewhere between the research question and the first data point, one number gets decided — sometimes deliberately, sometimes by accident. That number is your sample size, and it decides whether your study can answer the question you are asking.

Get it right and everything downstream is interpretable. Get it wrong and no analysis, however sophisticated, can rescue the study.

I am a clinical psychologist and statistician. I finished my PhD on schedule, I have published around twenty papers, and I would describe myself as a fairly average researcher. That is exactly why I take this one number seriously: I have watched careful, well-run studies come apart at the seams because of a step that should have taken an afternoon at the start.

This post makes the case for taking sample size seriously, then shows you how — using simulation, the most general and most honest way to calculate power. The stories here are about people who skipped it.

An underpowered study is a study you cannot publish

There is a version of this mistake that plays out slowly and painfully.

You spend a year collecting data. You run your analyses. Many of your results are non-significant. You start wondering whether your effects are real, or whether you simply did not have enough participants to detect them. Then someone checks the power. It is forty percent.

At that point the study has a problem no amount of rewriting will fix. Non-significant results from an underpowered study cannot be interpreted. You cannot conclude that the effect does not exist — only that your study was not in a position to find it if it did. Journals know this. Reviewers know this. The paper is very likely to be rejected, or at best reframed into something far more modest than what you set out to do.

This is why power calculations belong at the beginning of the research process, not the end.

A power calculation is a commitment device. It forces you to specify, before you collect a single data point, three things:

  • What effect size you expect to find.
  • How confident you want to be in your conclusions.
  • How many participants you need to get there.

That specificity is not bureaucracy — it is what makes the study answerable. If the numbers tell you that you need two thousand participants and you have access to three hundred, that is important information to have before you spend two years finding it out the hard way.

Power also helps you choose your analysis. A study with a modest sample size and a continuous outcome might have adequate power for a simple regression but not for a moderation analysis. Knowing this early shapes the design in useful ways — sometimes toward a different analysis, sometimes toward a different primary outcome, sometimes toward a narrower, honest research question.

📄 Free: I have built a ready-to-run Power Simulation R Script that implements the 5-step workflow below for the most common clinical and psychology designs. No extra packages, adapt it to your own analysis. Download it here.

Two mistakes I see constantly

Mistake one: letting a general-purpose AI tool run the calculation. These tools produce numbers that look plausible but rest on assumptions the user never specified and cannot verify. Power calculations require careful choices — the expected effect size, the design, the specific statistical test — and those choices must be grounded in the literature or in prior data. Delegating them to a tool that cannot make those judgements reliably is a risk not worth taking.

Mistake two: treating power as something you use after the fact. When results come back non-significant, some researchers calculate the power their study had to detect the effect and report it as an explanation. This misunderstands what power is. Post-hoc power is circular — it is determined by the observed effect size, so it tells you nothing the p-value did not already tell you. And it rescues nothing. A non-significant result from an underpowered study stays uninterpretable. The calculation needed to happen before the data existed, not after.

I watched this go wrong on a project I collaborated on. It was a well-designed RCT — good protocol, careful implementation, a meaningful clinical question. When the analyses came in, many of the key findings were non-significant. I looked at the power: around forty percent for most of the primary analyses. The team had used an AI tool at the design stage and had misread the output. When we went back through it, the numbers did not support what they had assumed.

The consequence was real. A study designed and presented as a definitive RCT had to be reframed as an initial pilot — not because the work was poor, but because it had never been in a position to answer the question it set out to answer. Years of work, substantially reframed, because of a step that should have taken an afternoon.

Do the power calculation before you design the study. Do it carefully, with numbers grounded in the literature. And do not expect it to do anything useful once the data are already collected.

How to calculate sample size with simulation

Most people approach power calculations as if they need a special formula or a dedicated tool. Some reach for G*Power. Some ask an AI. Both can work — and both can quietly fail you if you do not understand what is happening underneath.

The most reliable way to understand and calculate power is also the most intuitive: simulation. Instead of looking up a formula, you ask a simpler question:

If the effect I am trying to detect were real, how often would my analysis catch it?

That is all statistical power is. Simulate the data, run the test, repeat a thousand times, count the hits. The proportion of times you get a significant result at your chosen sample size is your power. This power analysis simulation approach works for any design and any test — not just the ones that happen to have a bespoke formula in a software menu.

Here are the five steps.

Step 1 — Specify the effect you expect to find. This is the hardest step, and where most calculations go wrong. You need a number: an expected effect size, a difference between means, an odds ratio, a correlation. Get it from the literature — comparable studies, meta-analyses, or clinically meaningful thresholds. If the literature is sparse, be conservative. An effect smaller than you expect is exactly the scenario that leaves you underpowered.

Step 2 — Define what your data looks like. Is your outcome continuous or binary? How many groups? What do your predictors look like? This is your data-generating process — the set of assumptions that, if true, would produce the effect you specified in Step 1.

Step 3 — Simulate one dataset and run your test. Generate n rows of fake data in which the effect is genuinely real. Run the exact statistical test you plan to use on your real data. Note whether you get p < 0.05.

Step 4 — Repeat 1,000 times. Run Steps 1–3 a thousand times at the same sample size. Count how many times you got p < 0.05. That proportion is your estimated power at that N.

Step 5 — Find your N. Repeat Steps 1–4 across a range of sample sizes — say 50, 100, 150, 200 — and find the point where power crosses 80%. That is your target sample size.

That is the whole method. The Power Simulation R Script implements exactly this workflow for the most common clinical and psychology designs, so you can start from a working example rather than a blank file.

Why simulation beats the shortcuts

Two errors are worth naming again here.

The first is the AI shortcut from the previous section — and it earns a second mention, because simulation is the antidote. AI tools fail when the inputs are not precisely specified: the true effect size, the design, the exact test. Working through the simulation yourself means you know what you are specifying, and you can sanity-check whether the result makes sense.

The second is believing you need G*Power or a similar tool for every design. G*Power is excellent when your design fits its menu. But many clinical research designs do not. Simulation scales to anything you can analyse — mixed models, cluster randomisation, custom composites — because it runs your actual analysis, not an approximation of it.

A colleague of mine gathered data for sixteen years before discovering the study was underpowered for the question it set out to answer. Sixteen years of work, substantially constrained by a calculation that would have taken an afternoon at the start.

Spend the hour.

When the numbers do not work: power as a creative constraint

Most researchers treat a power calculation as a verdict. They run it, get a number, compare it to their available sample, and either feel relieved or feel stuck.

What they miss is that the calculation is not fixed. The inputs can change — not by inflating the numbers to get a result you like, but by changing the question.

The same research interest can almost always be framed in more than one way. You might want to know whether a treatment works. That can be a t-test comparing two group means, a regression adjusting for several covariates, or a moderation analysis testing whether the treatment works differently for different subgroups. Each framing needs a different amount of power. The simplest is often the most powerful.

So when your first calculation says you need 400 participants and you have 120, the question is not how can I justify a larger effect size? The question is: which version of this question can 120 participants actually answer?

The wrong answer: motivated optimism

The literature says the effect is 0.4. You decide that, really, your effect is probably closer to 0.8 — or at least it could be. Dropout will be lower than the literature suggests because your setting is different. You adjust the inputs until the simulation returns the number you can achieve, and you write that down.

This is not a power calculation. It is a wish written in statistical notation. Reviewers who check your assumptions will notice — and even if they do not, the underpowered study will.

The honest answer: redesign around the constraint

Assume things will go badly, and choose methods that still give you power when they do. Assume your effect size is small. Assume dropout is what the literature says, or worse. Then redesign. That usually goes in one of two directions.

1. Method simplification. A t-test comparing two group means is generally more powerful than a regression with five covariates at the same N — a covariate only earns its place when it strongly predicts the outcome. A single primary outcome beats a composite. If your core question can be answered by the simpler method, even approximately, the simpler method is the right choice. You may be able to ask a genuinely interesting question with the participants you have — just not the most ornate version of it.

2. Scope reduction — reframe as a pilot. A pilot has a different mandate than a confirmatory study. Its job is not to detect a final effect; it is to sharpen the hypothesis, estimate realistic effect sizes and dropout rates, and generate the data you need to design the next study well. This is not a consolation prize. A well-designed pilot that produces a sharp, credible hypothesis for a funded follow-up is a real scientific contribution.

I have lost count of the times I have sat with a researcher fixated on a four-group design with multiple moderators, who had 80 participants and was trying to make it work. In almost every case, a simpler version of the same question — two groups, one outcome, no moderation — would have answered the core scientific interest well enough, with a fraction of the required sample. The complex design was not more rigorous; it was just more expensive. The attachment to it was usually about what the study felt like it should be, not what the data could support.

When the numbers do not work, the answer is redesign. Be creative under constraints. That is the skill.

The takeaway

Power is not a statistical formality. It is the difference between a study that can speak and a study that cannot.

The calculation forces you to commit — to an effect size grounded in the literature, to a design, to an analysis — before the data exists to tempt you into flexibility. And when the numbers do not cooperate, the honest response is not more optimistic assumptions but a better-framed question: a simpler method, a sharper hypothesis, or a pilot with a clear mandate.

Simulation is how you do all of this without pretending. Five steps: specify the effect, define the data, simulate and test once, repeat a thousand times, find your N.

📄 Free: Don't build the script from scratch. My ready-to-run Power Simulation R Script already implements the full 5-step workflow for common clinical and psychology designs — no extra packages, easy to adapt to your own analysis. Download it here.

What design are you trying to power right now? Hit reply and tell me the shape of it — how many groups, what outcome, roughly how many participants you can realistically recruit. I read every reply, and the shape of the design usually tells me where the power is hiding.