* Total Points: 0
Back to Lessons Where Does Delta Come From? 0 pts Module 6 · Lesson 2
Introduction

Where Does Delta Come From?

Every sample size calculation contains exactly one judgment call — and it's the one that decides everything

The Judgment Hiding in the Math

Alpha is fixed by convention. Power is fixed by convention. The event rate in your control group is (roughly) a fact of the world. That leaves one number a human must choose: the effect size the study is designed to detect — the delta (δ).

Choose δ = a 15-point absolute reduction in complications (30% → 15%) and your trial needs ~120 patients per arm. Choose 5 points (30% → 25%) and it needs ~1,250. Same question, same outcome, same statistics — a tenfold difference in cost, driven entirely by one editable number. Which is exactly why δ is where sample size calculations get quietly rigged.

The right delta is the smallest effect that would change clinical practice — not the effect you expect, not the effect your pilot study showed, and definitely not the effect that makes your feasible n look adequate.

MCID: The Effect Worth Finding

Power the study for the smallest effect that matters, and every result becomes interpretable.

The Minimal Clinically Important Difference (MCID)

The MCID is the smallest change in an outcome that patients or clinicians would consider meaningful — the threshold below which a real difference wouldn't change what anyone does. For patient-reported outcomes, MCIDs are formally estimated (anchor-based methods tie score changes to patients saying "I feel better"; distribution-based methods use fractions of the outcome's variability). For hard outcomes like death, amputation, or stroke, the "MCID" is a clinical judgment: what absolute risk reduction would justify the new operation's cost, learning curve, and harms?

Why it's the right delta: if you power for the MCID and the result is negative with a tight confidence interval, you can genuinely say "any effect that exists is too small to care about." Power for anything larger and a negative result is ambiguous forever.

Ask: what's worth detecting?

"A 5% absolute reduction in amputation would change our practice, so we power to detect 5%." The study is sized to answer the clinical question, whatever the answer turns out to be.

Ask: what do we expect?

"We think the new graft reduces amputation by 15%, so we power for 15%." If the truth is a still-important 7%, the study is underpowered for it — and a negative result answers nothing.

The distinction sounds subtle but drives everything: expected effect is a bet about reality; worth-detecting effect is a statement about decisions. Trials powered on optimistic expected effects are how the literature fills with "negative" studies of interventions that actually work modestly well.

Absolute vs. Relative: Say Which One

"A 33% reduction in surgical site infection" means very different things depending on the baseline: from 30% to 20% (10 absolute points — a big, detectable effect) or from 3% to 2% (1 absolute point — needing thousands of patients). Sample size calculations run on absolute rates. Any power statement quoting only a relative reduction without the control event rate is incomplete — and the control rate assumption deserves scrutiny too, because overestimating it silently deflates the required n.

Bad Sources of Delta

Three places deltas commonly come from — and why each one biases the study before it starts.

Source 1: The Pilot Study Point Estimate

Your 30-patient pilot showed a 14% absolute reduction, so you power the trial for 14%. The problem: with 30 patients, that estimate has a confidence interval running from roughly −2% to +30%. You've anchored the entire trial to a number that is mostly noise — and if the pilot caught attention because its result was impressive, the estimate is inflated by selection on top of noise (the winner's curse from Lesson 1). Pilots are for testing feasibility — recruitment, protocol, data collection — not for estimating effect sizes.

If you must use pilot data: use the plausible lower end of its confidence interval, or better, use the pilot's control-arm event rate (which it estimates less badly) and pair it with an MCID-based delta.

Source 2: Reverse-Engineering from Feasible n

You can recruit 80 patients in two years. Run the calculation backwards and 80 patients gives 80% power to detect... a 20% absolute reduction. So the protocol declares, with a straight face, that the study is "powered to detect a 20% difference" — an effect so large that if it were real, you wouldn't need a trial to notice it. The calculation is arithmetically valid and scientifically hollow.

How to spot it as a reader: the tell is a delta far larger than any effect reported in the field, in a study whose n is suspiciously round. Ask: "would the authors have called a difference half this size clinically irrelevant?" If not, the delta was chosen for the n, not the other way around.

Source 3: Cohen's Benchmarks as a Default

For continuous outcomes you'll meet the standardized effect size, Cohen's d = difference in means ÷ standard deviation, with folklore labels: 0.2 small, 0.5 medium, 0.8 large. Useful when you truly know nothing about the scale — but "we powered for a medium effect (d = 0.5)" is a confession that no one asked what change in this outcome matters to these patients. A d of 0.5 in pain scores and a d of 0.5 in length of stay are unrelated clinical claims wearing the same number.

Prefer natural units: "detect a 2-day difference in length of stay, assuming SD 5 days" is checkable clinical reasoning. "Detect d = 0.4" is a shrug in mathematical notation.

To make the stakes concrete — required n per arm for a two-arm trial with a 30% control event rate (α = 0.05, 80% power):

Target absolute reductionTreatment raten per arm
15 points (30% → 15%)15%~120
10 points (30% → 20%)20%~290
5 points (30% → 25%)25%~1,250
3 points (30% → 27%)27%~3,600

Halving the delta roughly quadruples the required sample size. This is why the choice of δ is not a technical detail — it IS the study design decision, and every other number in the power paragraph orbits around it.

Delta Pitfalls

Four power statements from real-world methods sections. Find the problem.

"Based on our pilot study showing a reduction from 28% to 12%, we calculated that 95 patients per arm provide 80% power."

The delta is a noisy, probably inflated pilot estimate. A 16-point reduction from a small pilot carries a huge confidence interval, and pilots that lead to trials are selected for impressive results. If the true effect is a still-meaningful 8 points (28% → 20%), this trial has roughly 25% power — three-to-one odds of a false negative. The honest approach: power for the smallest reduction that would change practice.

"We powered the study to detect a 50% relative risk reduction in graft occlusion."

Relative reduction without a control event rate is uninterpretable. Halving a 40% occlusion rate is a moderate-sized trial; halving a 4% rate needs thousands. And a 50% relative reduction is an enormous ask of any intervention — deltas this optimistic usually signal a study sized to its budget. Look for the assumed control rate, then ask where it came from and what happens to power if it's lower.

"A sample of 40 patients per group was calculated to provide 80% power to detect a difference of 1.5 days in length of stay (SD 2.0 days)."

Check the SD assumption as hard as the delta. The arithmetic here is fine (d = 0.75 needs ~29/arm; 40 gives margin) — but length-of-stay data are right-skewed, and real SDs often run 4–6 days, not 2. Recompute with SD 4 and this study's power collapses to ~40% for the same 1.5-day difference. Optimistic variability assumptions are the quieter sibling of optimistic deltas: same effect on power, less likely to be questioned.

"With the 90 patients available in our database, we had 80% power to detect a 19% absolute difference in reintervention."

This is the feasible-n calculation dressed up as design. The sample existed first; the delta was computed from it. Nothing dishonest has been stated — but note what it implies: for any realistic effect (say 5–10 points), the study is badly underpowered, and a negative result is uninformative. Retrospective studies can't fix their n, but they can be honest that power constrains interpretation — which is Lesson 4's whole topic.

Reading a power statement is a three-question audit: Where did the delta come from — a decision threshold or a hope? Is the control event rate (or SD) realistic? And would a difference just below the delta still matter clinically? If yes, the study can't distinguish "no effect" from "an effect worth having."

Exercise: Choosing and Critiquing Delta

MCID vs expected effects, pilot traps, and reading power statements.

Question 1 of 8

Lesson Complete!

0
Total Points Earned
Exercise (0/8 correct) +0 pts
Lesson Completed +100 pts

Where Does Delta Come From?

Module 6 - Lesson 2 complete

Key Takeaways

  • Delta is the judgment call: alpha and power are convention, the control rate is (approximately) a fact — the effect size is chosen, and it drives everything.
  • Power for the MCID: the smallest effect that would change practice — not the effect you expect or hope for.
  • Pilot studies estimate feasibility, not effect sizes: their point estimates are noisy and selected for impressiveness.
  • Reverse-engineered deltas: when n was fixed first, the "detectable difference" is a description of the study's limits, not its design.
  • Absolute beats relative: a relative risk reduction means nothing without the control event rate — and check the SD assumption for continuous outcomes.
  • The quadrupling rule: halving delta roughly quadruples required n. Small effects are expensive.

Next lesson: Running the Numbers — an interactive calculator where you drive all four levers yourself and build intuition for what sample sizes actually buy.