Back to Lessons Red Flags in Methods 0 pts Module 1 · Lesson 3
Introduction

Red Flags in Methods

How to spot manipulation before you're fooled

The Uncomfortable Truth

Most published research findings are false. Not because scientists are dishonest, but because the incentives of academic publishing reward positive results, and there are many ways to turn negative data into a "significant" finding.

What You'll Learn

By the end of this lesson, you'll recognize the most common red flags in methods sections: p-hacking, HARKing, and cherry-picking. These are the tricks that make bad science look good.

P-Hacking: Torturing Data Until It Confesses

P-hacking is running multiple analyses until you find p < 0.05, then reporting only that one.

1

The Problem

If you test 20 hypotheses at p = 0.05, you'll get one "significant" result by chance alone. P-hackers exploit this by testing many things but only reporting the hits.

2

Common Techniques

• Testing multiple outcomes, reporting only the significant one
• Adding or removing covariates until p drops below 0.05
• Excluding "outliers" until the result becomes significant
• Stopping data collection when p < 0.05

Red Flag Phrases

"After adjusting for multiple covariates..." (Which ones? Why those?)
"In a subgroup analysis..." (Was this planned?)
"After excluding outliers..." (How were outliers defined?)

Suspicious Example
"The primary endpoint showed no significant difference (p=0.23). However, in patients over 65 with diabetes and prior MI, a significant benefit was observed (p=0.04)."
This screams p-hacking. They sliced the data until they found something. With enough subgroups, you'll always find one.

If the primary endpoint failed but a convenient subgroup "worked," be very skeptical.

HARKing: Hypothesizing After Results Known

HARKing is presenting a post-hoc finding as if it were the original hypothesis.

1

How It Works

Researcher designs study to test X. X fails. But they notice Y looks interesting in the data. Paper is written as if they always intended to study Y.

2

Why It's Dangerous

Exploratory findings need confirmation. When presented as confirmatory, they appear more reliable than they are. The false positive rate skyrockets.

Honest Reporting

"Our primary endpoint (wound healing at 12 weeks) showed no difference. In exploratory analysis, we noted a trend toward reduced infection rates that warrants further study."

HARKing

"We hypothesized that the intervention would reduce infection rates. Our results confirm this hypothesis (p=0.03)."

How to Spot It

• Check trial registrations (clinicaltrials.gov) for original endpoints
• Look for mismatch between stated aims and reported outcomes
• "Secondary endpoint" that gets more attention than the primary
• Very specific hypothesis that perfectly matches the result

If the hypothesis seems suspiciously perfect for the data, it was probably written after seeing the data.

Cherry-Picking: Selective Reporting

Cherry-picking means reporting only the results that support your conclusion while hiding the rest.

1

Selective Outcome Reporting

Study measures 10 outcomes. Three show benefit, seven show no difference or harm. Paper reports only the three.

2

Selective Time Point Reporting

Data collected at 1, 3, 6, and 12 months. Only the 3-month data (where results looked best) makes it into the paper.

3

Selective Citation

Discussion cites 10 studies supporting the conclusion, ignores 15 that contradict it.

Warning Signs

• Methods list outcomes not reported in results
• Unusual or oddly specific time points
• Missing data on adverse events or harms
• Supplementary data that contradicts main findings

Classic Cherry-Pick
"At 90 days, limb salvage was significantly improved in the treatment group (p=0.02)."
Why 90 days? Most limb salvage studies use 1 year. If 30-day, 180-day, and 1-year data weren't reported, they probably didn't look as good.

Other Red Flags to Watch For

Tiny sample size with big claims

N=12 and claiming to prove a treatment works? Underpowered studies that find effects are likely false positives.

Exactly p=0.05 or p=0.049

Suspiciously convenient. Real data rarely lands exactly on the threshold.

No pre-registration

Without it, you can't verify the analysis plan wasn't changed after seeing data.

Industry funding, no independent analysis

Not automatic disqualification, but raises the bar for scrutiny.

Authors with obvious conflicts

Inventor of device authors the pivotal trial. Consultant for company writes the favorable review.

Implausibly perfect results

No complications, no dropouts, no missing data. Real clinical research is messy.

"Data not shown"

Why not? What are you hiding?

Nonsensical composite endpoints

"Death, MI, or need for bandage change" is real. Mixing severe and trivial events inflates rates.

The more a study's results align perfectly with the authors' interests, the more skeptical you should be.

Exercise: Spot the Red Flag

Identify the methodological problem in each scenario.

Question 1 of 8

The Bottom Line

0
Total Points Earned
Exercise (0/8 correct) +0 pts
Lesson Completed +100 pts
  • P-hacking is testing multiple things until something reaches p < 0.05. Watch for unexplained subgroup analyses and post-hoc covariate adjustments.
  • HARKing is presenting exploratory findings as if they were the original hypothesis. Check trial registrations for original endpoints.
  • Cherry-picking is reporting only favorable results. Look for missing outcomes, unusual time points, and "data not shown."
  • Trust but verify. Check trial registrations, look at supplementary data, and be suspicious of results that seem too good to be true.

Your New Superpower

You can now read a methods section and spot the tricks that make weak evidence look strong. Use this power wisely—most researchers aren't malicious, but incentives shape behavior. Your job is to evaluate the evidence, not the intentions.