Why statistical significance isn't the same as clinical importance
That framework is incomplete at best, misleading at worst. By the end of this lesson, you'll understand why a p=0.001 result can be meaningless, and why a p=0.09 result might be the one that should change your practice.
Both tested Drug X for claudication. Look at the results:
Which drug works better? Which result should change your practice?
Hold that thought. We'll come back to it.
The p-value is the probability of seeing this result (or something more extreme) if the null hypothesis were true.
In other words: "If there's actually no difference, how often would we see data like this by chance?"
The p-value does not tell you:
• The probability that your finding is true
• The size of the effect
• Whether the effect matters clinically
"Is this likely due to chance?"
That's it. Nothing about importance. Nothing about magnitude.
The p-value answers the wrong question. You want to know if the effect matters. The p-value only tells you if it's probably real.
With enough patients, any difference becomes statistically significant.
Same effect (2 mg/dL LDL reduction), different sample sizes:
The effect didn't change. Your certainty about a trivial effect increased.
The p-value is a function of sample size. Effect size is not.
Effect size answers "how much?"—the question you actually care about.
• Absolute difference (e.g., 4 meters vs 85 meters)
• Relative risk or hazard ratio
• Number needed to treat (NNT)
• Odds ratio
Below this threshold, who cares if it's significant?
For 6-minute walk distance in PAD: the MCID is roughly 30-50 meters.
A 4-meter improvement? Statistically significant noise.
"Is this difference big enough to change what I do for my patient?"
If the answer is no, the p-value is irrelevant.
The 95% CI is more informative than the p-value alone.
The middle of the CI is your best guess at the true effect.
Narrow CI = large sample, precise estimate.
Wide CI = small sample, uncertain estimate.
If the 95% CI for a difference includes zero (or 1.0 for ratios), p > 0.05.
If the entire CI falls within a clinically meaningless range, the study is definitively negative—even if p < 0.05.
A study shows: Effect: 2 mg/dL, 95% CI: 1.5 to 2.5, p < 0.001
The entire confidence interval is below any meaningful LDL reduction. This is a confident null—we're certain the effect is too small to matter.
Classify each scenario into one of four categories.
Trial A is statistically significant but not clinically important. A 4-meter improvement is far below the MCID of 30-50 meters. With 10,000 patients, you've achieved high confidence in a trivial effect.
Trial B is not statistically significant but potentially important. An 85-meter improvement is clinically meaningful—nearly double the MCID. The p=0.09 reflects an underpowered study, not an absent effect.
Trial B should influence your thinking more. It suggests a meaningful effect that deserves a larger study.
How to look past p-values and evaluate what a study actually found. You won't be fooled by significant-but-trivial results, and you'll recognize potentially important findings that failed to reach significance due to sample size.