* Total Points: 0
Back to Lessons The Underpowered Study Problem 0 pts Module 6 · Lesson 4
Introduction

The Underpowered Study Problem

Your database has 74 patients. It will never have more. Now what?

The Retrospective Reality

The power lessons so far assume you're designing a study — choosing n before data exist. But most resident research doesn't work that way. You inherit a database: every carotid endarterectomy since 2015, all 74 of them. The sample size was decided by history, not by a power calculation.

Two things follow. First, an honest question: what can 74 patients actually tell us? — sometimes the answer is "not what we hoped," and knowing that before six months of chart review is a gift. Second, a temptation: after the analysis comes back non-significant, to compute "post-hoc power" and let it do your interpreting. That maneuver is a fallacy dressed as diligence, and reviewers still ask for it. This lesson arms you for both.

When n is fixed, power analysis changes jobs: it stops being a design tool and becomes an honesty tool — run before the analysis to set expectations, never after to explain away a p-value. Interpretation of a finished study belongs to the confidence interval.

The Post-Hoc Power Fallacy

"Observed power" sounds rigorous. It is the p-value wearing a fake mustache.

What people do

The study finds 18% vs 12% complications, p = 0.35. A reviewer asks: "What was your power?" The authors plug the observed rates (18% and 12%) and their n back into a power calculator, get "power = 15%," and write: "the analysis may have been underpowered to detect the observed difference."

Why it's circular

Observed power is computed from the observed effect and n — the exact ingredients of the p-value. It is a deterministic transformation of the p-value: a non-significant p always yields low observed power, and a significant p always yields high observed power. It cannot tell you anything the p-value didn't already say. "We found p = 0.35 and observed power was low" is one fact stated twice.

The deeper error: it treats the observed effect as if it were the true effect. But the observed effect in a small study is mostly noise — which is the whole reason the study is hard to interpret in the first place.

What is legitimate

Power computed against a prespecified, clinically meaningful effect size — the MCID from Lesson 2 — is fine at any time, because it doesn't depend on your results. "With 74 patients, we had 30% power to detect a 10-point difference in complications" is honest context. The difference isn't when you compute power; it's which delta you compute it against: a clinical threshold chosen independently of the data, versus the data themselves.

One-line rule: power calculations may use the effect you set out to find, never the effect you happened to observe. If a methods section computes power from its own results, you're watching the p-value interview itself.

Let the Confidence Interval Do the Talking

A finished study's interpretation lives in its CI — specifically, in where the CI sits relative to the effects that matter.

From Module 3 you know a 95% CI is the range of effects compatible with the data. For a "negative" study, the question is never "was p > 0.05?" but: does the CI exclude the minimal clinically important difference? Three studies, all "negative," all completely different:

ABSOLUTE RISK DIFFERENCE (COMPLICATIONS) — three "p > 0.05" studies, MCID = 5 points
Study A (n = 1,200/arm): difference +0.4 points, 95% CI −1.8 to +2.6
Informative negative: the entire CI sits below the MCID. Any effect that exists is too small to change practice.
Study B (n = 60/arm): difference +3.0 points, 95% CI −7.2 to +13.2
Uninformative: the CI spans "harmful," "nothing," and "twice the MCID." This study cannot distinguish the possibilities anyone cares about.
Study C (n = 150/arm): difference +4.1 points, 95% CI −0.3 to +8.5
"Negative" but suggestive: the estimate is near the MCID and most of the CI is above zero. Calling this "no difference" buries the lead — it's a signal awaiting a bigger study.
null (0) MCID (5 points)

So what do you do with 74 patients?

Before analyzing: run the feasibility check (Lesson 3's calculator) against the MCID. If power is 25%, decide — with your mentor — whether the study is worth doing as-is, worth reframing, or worth expanding (more years, more centers, a registry).

If you proceed: frame the study as estimation, not testing. Report effect sizes with CIs, describe them honestly ("compatible with anything from modest benefit to moderate harm"), call the work hypothesis-generating, and resist the equivalence conclusion at all costs.

Legitimate small-study purposes: describing local outcomes against benchmarks, feasibility and safety signals, estimating event rates to power the next study. Small studies aren't worthless — they're just not hypothesis tests.

Small-Study Pitfalls

Four moves that turn an underpowered study from "limited" into "misleading."

"Although underpowered, we found no difference in mortality, supporting the safety of the endovascular-first approach."

The sentence refutes itself. "Underpowered" means "unable to detect realistic differences" — so finding none is the expected result under both hypotheses and supports neither. Safety claims need either an adequately powered comparison or a non-inferiority design with a prespecified margin. Watch for "no difference" quietly morphing into "safe" or "equivalent" between the results and the conclusion.

"Post-hoc power analysis using the observed effect showed 22% power, explaining the non-significant result."

Circular by construction. Observed power is a re-expression of the p-value; low observed power "explains" a non-significant result the way a thermometer explains a fever. What would help: the CI against a prespecified MCID, or power against the effect the study was designed around. If no effect was prespecified — that's the actual limitation to report.

"Given our limited sample, we tested 14 secondary outcomes to maximize the information gained, and found significantly reduced ICU stay (p = 0.03)."

Low power plus multiplicity is the false-positive factory. Fourteen tests at α = 0.05 give a ~51% chance of at least one false alarm, and in a small study the winner's curse (Lesson 1) guarantees any "hit" is inflated. Small samples argue for fewer, prespecified outcomes, not more. A p = 0.03 from an outcome buffet is a coincidence with a byline.

"To increase power, we combined death, MI, stroke, reintervention, and readmission into a single composite endpoint."

Composites buy power by blurring the question. More events do mean more power — that part is real (it's why trials use MACE). But the composite is driven by its most frequent component, usually the least serious one: a "significant" composite can mean more readmissions and nothing else, while the conclusion implies lives saved. Always check the component table, and be suspicious of composites assembled after the primary outcome came up short — you learned to spot that switch in Module 1; now you know the power motive behind it.

An underpowered study handled honestly — estimation framing, CIs against the MCID, restrained conclusions — is a real contribution. The same data oversold as a hypothesis test subtracts from the literature. The difference is entirely in the writing, which means it's entirely in your control.

Exercise: Interpreting Without Power

Post-hoc power, CI-based interpretation, and honest small-study framing.

Question 1 of 8

Lesson Complete!

0
Total Points Earned
Exercise (0/8 correct) +0 pts
Lesson Completed +100 pts

The Underpowered Study Problem

Module 6 - Lesson 4 complete

Key Takeaways

  • Fixed n changes power's job: from design tool to honesty tool — run it before analysis to set expectations, against the MCID.
  • Post-hoc "observed power" is circular: it's computed from the same numbers as the p-value and always agrees with it. Power against a prespecified delta is fine; power against your own results never is.
  • Interpret finished studies by the CI: a negative study is informative only if its CI excludes the MCID; a CI spanning harm-to-benefit is a shrug, not a finding.
  • Watch the near-miss: an estimate near the MCID with a mostly-positive CI is a signal, not "no difference."
  • Small-study sins compound: outcome buffets, post-hoc composites, and "underpowered but safe" conclusions each convert low power into active misinformation.
  • Small studies have honest jobs: estimation, feasibility, event rates for powering the next study — just not hypothesis testing.

Next lesson: the capstone — a full power-and-sample-size audit of a realistic vascular surgery manuscript, start to finish.