Back to Module 3
Module 3 - Lesson 4

Smelling Fragile Results

How to spot findings that probably won't replicate

Why Most Published Results Are Wrong

A sobering statistic

When researchers tried to replicate 100 psychology studies published in top journals...

Only 36% replicated

— Open Science Collaboration, 2015

Medicine isn't immune. Many "breakthrough" findings quietly disappear when larger studies are done.

The problem isn't fraud (usually). It's that certain statistical patterns are inherently fragile—they arise from noise, not signal.

The skill: Learning to "smell" fragile results before wasting time, money, or patient trust on findings that won't hold up.

Red Flag #1: Barely Significant

A p-value of 0.048 is not the same as p = 0.001.

P-value distribution in "positive" studies

p = 0.001 Very likely to replicate
p = 0.03 Moderate confidence
p = 0.048 Often fails replication
⚠️

The p = 0.04 zone

When you see p-values clustering just below 0.05 (like 0.048, 0.042, 0.038), be suspicious. This often indicates p-hacking—running multiple analyses until one "works."

Reality check: If the effect were real and robust, the p-value would typically be much smaller. Barely significant results are often noise that got lucky.

Red Flag #2: Wide Confidence Intervals

The CI width tells you how much uncertainty remains.

Two studies of the same drug

Both report "significant reduction in pain scores"

Study A
1.2 - 2.1
Precise
Study B
0.3 - 4.8
Imprecise

Both exclude zero (statistically significant), but Study B's true effect could be anywhere from tiny to huge

Study A (n=500)

Effect: 1.6 points

95% CI: 1.2 - 2.1

Narrow CI = confident estimate

Study B (n=40)

Effect: 2.5 points

95% CI: 0.3 - 4.8

Wide CI = could be almost nothing

The test: If the lower bound of the CI is close to zero (or below the minimally important difference), the "significant" finding might not be clinically meaningful.

Red Flag #3: Underpowered Studies

Small studies that find "significant" effects often massively overestimate the true effect size.

The Winner's Curse

Imagine the true effect of a treatment is 0.3 standard deviations (small but real).

A small study (n=20 per group) needs to get lucky with sampling to reach p < 0.05. The studies that "win" (find significance) will overestimate the effect.

Reported effect sizes by study size

Study Size Reported Effect True Effect Inflation
n = 20 0.8 0.3 2.7x
n = 50 0.5 0.3 1.7x
n = 200 0.32 0.3 1.1x
⚠️

The small study paradox

Counter-intuitively, small studies often show larger effects than big studies. This isn't because small studies are better—it's because only the ones that got lucky with their sample get published.

Rule of thumb

Be skeptical of any RCT with fewer than 50 participants per arm reporting a "significant" effect. The effect size is probably inflated.

Red Flag #4: Too Many Comparisons

When you test 20 outcomes, expect 1 false positive—by pure chance.

The "data dredging" study

A supplement study measures 15 different outcomes: mood, energy, sleep quality, focus, memory, anxiety, stress, creativity, motivation, happiness, calmness, alertness, confidence, optimism, and satisfaction.

Result: "Significant improvement in optimism (p=0.04)!"

Probability of at least one false positive

Tests Run P(≥1 False Positive)
1 5%
5 23%
10 40%
20 64%
The giveaway: When a study reports many non-significant outcomes and one or two "significant" findings, especially if the significant one wasn't the primary outcome—it's probably noise.
What to look for: Was the outcome pre-specified? Did they adjust for multiple comparisons (Bonferroni, FDR)? If not mentioned, assume no correction was made.

The Fragility Checklist

Before trusting a result, run through these questions.

  • 1
    Is p barely significant?

    p = 0.04 is much more fragile than p = 0.001

  • 2
    Is the confidence interval wide?

    Wide CI = uncertain estimate, even if "significant"

  • 3
    Is the sample size small?

    Under 50 per group? Effect is probably inflated

  • 4
    Were many outcomes tested?

    Cherry-picked findings from many comparisons don't replicate

  • 5
    Was this a subgroup analysis?

    "Works in women over 50 with diabetes" = fishing expedition

  • 6
    Has it been replicated independently?

    First findings are more often wrong than replicated ones

The Fragility Index

For binary outcomes, count how many patients switching from "event" to "no event" (or vice versa) would flip the p-value past 0.05. If the fragility index is 1-3, the result is extremely fragile.

Smell Test: Fragile or Robust?

For each study result, identify whether it smells fragile or robust.

Key Takeaways

1. Barely significant = barely convincing

P-values hovering just below 0.05 rarely replicate. Don't bet on p = 0.04.

2. Width matters more than significance

A wide CI that excludes zero is less informative than a narrow CI. Look at the bounds, not just whether it "crossed the line."

3. Small + significant = probably inflated

The Winner's Curse means small studies that reach significance likely overestimate the true effect.

4. One in twenty is just noise

With multiple comparisons and no correction, false positives are expected, not surprising.

5. First ≠ True

Initial findings are more likely to be wrong than replicated findings. Wait for confirmation before changing practice.