Back to StatUp
Journal Club · Eur Heart J 2026

SMART-REACH2: How a Prediction Model Is Built — and How to Read One

A lifetime cardiovascular risk model for patients with established atherosclerotic disease — and a working tour of discrimination, calibration, competing risks, and individualized treatment benefit.

Holtrop J, Gynnild MN, Hageman SHJ, Dorresteijn JAN, et al. Predicting lifetime cardiovascular risk and benefits of preventive treatment in patients with established atherosclerotic cardiovascular disease: the SMART-REACH2 model. Eur Heart J 2026. PMID 42246978.
Lineage: SMART-REACH model, JAHA 2018 (Kaasenbrood et al.). Endorsed for use by the 2021 ESC CVD prevention guidelines.
Article details retrieved via PubMed. DOI: 10.1093/eurheartj/ehag400

Section 1: From the Average Patient to the One in Front of You

A randomized trial tells you what a therapy does on average. The patient in your clinic is not the average. A prediction model is the bridge — it turns a trial's average relative effect into this person's absolute benefit. SMART-REACH2, published in the European Heart Journal in June 2026, is the current guideline-referenced version of that bridge for established atherosclerotic disease.

SMART & REACH risk scores

Two observational cohorts produced short-horizon recurrence scores: the 10-year SMART score and the 20-month REACH score. Useful for ranking risk, but no competing-risk handling and no lifetime view.

2018: SMART-REACH lifetime model

Kaasenbrood et al. (JAHA) combined the cohorts into a competing-risk lifetime model estimating CVD-free life expectancy — and, crucially, life-years gained from a given therapy.

2021: ESC guideline endorsement

The ESC CVD-prevention guidelines named SMART-REACH as a tool to guide treatment decisions in patients with established ASCVD. Guideline endorsement raises the bar for transportability and calibration.

2026: SMART-REACH2 (this paper)

Re-derived and systematically recalibrated to four European and other global risk regions, then externally validated in over two million patients across 54 countries. The update is mostly about making the model travel — and proving it does.

The Central Question

For an individual with established coronary, cerebrovascular, peripheral, or aortic atherosclerotic disease: what is their lifetime risk of a recurrent cardiovascular event, and how many cardiovascular-disease-free life-years would a given preventive therapy add for them?

Note the shift in object. A trial asks "does the drug work?" A model asks "given that it works, how much does it do for this patient?"

A model is not a trial. SMART-REACH2 does not establish that statins or antihypertensives work — trials did that. It personalizes the benefit a proven therapy confers, by estimating each patient's baseline risk and remaining life expectancy. Read it as an arithmetic engine sitting on top of trial evidence, not as new efficacy evidence.

Section 2: The Core Idea — Relative Risk In, Absolute Benefit Out

The whole reason individualized models exist is one stubborn fact: relative risk reduction is roughly constant across patients, but absolute benefit is not. Absolute benefit scales with baseline risk.

The Arithmetic in One Line

Absolute Risk Reduction ≈ Baseline Risk × Relative Risk Reduction. A therapy with a fixed 25% RRR delivers a large absolute benefit to a high-risk patient and a trivial one to a low-risk patient — same drug, same RRR.

Same 25% RRR, Three Patients

Baseline 10-yr risk RRR Absolute benefit (ARR) NNT
10% 25% 2.5% 40
30% 25% 7.5% ~13
50% 25% 12.5% 8

The drug never changed. Only the patient did. The model's entire job is to estimate the first column accurately for an individual.

Treat risk, not risk factors. A single alarming LDL or blood-pressure number says little about whether intensifying therapy helps this patient. Their total baseline risk — the integration of all their risk factors — is what determines absolute benefit. That integration is exactly what a multivariable model provides.
The catch: this only works if the model's estimate of baseline risk is accurate in absolute terms. A model that ranks patients perfectly but systematically overstates everyone's risk will overstate everyone's benefit. Hold that thought — it is the difference between discrimination and calibration, and it is Section 4.

Section 3: How the Model Was Built — the Derivation Cohort

Every prediction model has a development ("derivation") dataset, an outcome definition, a set of predictors, and a statistical engine. Get these four straight and you understand most of any modeling paper.

SMART-REACH2 Derivation

  • Cohort: UCC-SMART, a single-center prospective cohort of patients with manifest arterial disease (Utrecht, Netherlands)
  • Sample: 8,708 individuals, aged 40–90, with coronary, cerebrovascular, peripheral artery disease and/or abdominal aortic aneurysm
  • Events: 2,057 recurrent CV events over a median follow-up of 8.5 years (IQR 4.3–13.0)
  • Outcome: a composite of myocardial infarction, stroke, or cardiovascular death
  • Engine: sex-stratified, cause-specific Cox models with age as the timescale

Routinely Available Predictors — on Purpose

The predictors are things already in the chart: type and number of vascular beds involved, systolic blood pressure, lipids (non-HDL cholesterol), renal function (eGFR), smoking, diabetes, duration of disease, and an inflammation marker. No exotic assays, no genomics. A model that needs a test you cannot order at the bedside is a model nobody uses.

Why "Age as the Timescale" Matters

Most Cox models use time-since-enrollment as the clock. SMART-REACH-type lifetime models use attained age as the clock instead, with left truncation (patients enter the risk set at the age they joined). This is what lets a cohort with ~8 years of follow-up generate predictions across the whole 40–90 age span — you borrow information across people at different ages rather than waiting decades to observe one person age.

What to check in any derivation: Is the cohort representative of where the model will be used? Is the outcome a hard, adjudicated composite or a soft surrogate? Are there enough events? A rough rule of thumb for these models is roughly 10–20 events per candidate predictor; with ~2,000 events, SMART-REACH2 has comfortable room for its handful of predictors.

Section 4: The Two Jobs of Any Model — Discrimination and Calibration

A prediction model has to do two different things, and they are not the same thing. Conflating them is the single most common error in reading these papers.

Discrimination

Can it rank? Does it give higher-risk patients higher scores than lower-risk patients?

Measured by the C-statistic (AUC). 0.5 = coin flip, 1.0 = perfect ranking.

Calibration

Are the numbers right? When the model says 30%, do about 30% of those patients actually have events?

Assessed by predicted-vs-observed plots and expected/observed ratios.

SMART-REACH2's Discrimination: C = 0.68

The pooled C-statistic was 0.68 (95% CI 0.66–0.69), ranging from 0.66 in the European low-risk region to 0.72 in Latin America. By the usual rough labels that is "modest." A colleague will say "0.68, that's barely better than a coin flip's cousin — useless."

They are wrong, and here is why.

Why Modest Discrimination Is Expected Here — and Acceptable

Discrimination depends on how spread out risk is in the population. In primary prevention you are separating healthy 40-year-olds from sick 70-year-olds — easy, C-statistics run high. In secondary prevention, everyone already has established disease. The population is homogeneously high-risk, so there is simply less spread to discriminate. C-statistics around 0.65–0.70 are typical and near the practical ceiling; adding more predictors has repeatedly failed to move them much.

For context: the categorical "very high risk" criteria in some guidelines discriminate at roughly 0.53–0.54. A calibrated 0.68 model is a real improvement over the yes/no labels clinicians use now.

For treatment decisions, calibration matters more than discrimination. You do not act on a patient's rank — you act on their absolute predicted risk and the absolute benefit that follows. A model could discriminate beautifully (C = 0.80) yet be miscalibrated and overstate everyone's risk by 50%, leading you to overtreat across the board. SMART-REACH2's headline achievement is not its C-statistic; it is adequate calibration across regions after recalibration.
Which one you need depends on the use: ranking/triage (who to call back first, who to enroll in a trial) leans on discrimination; deciding whether an individual's absolute benefit clears a threshold leans on calibration. SMART-REACH2 is built for the second job.

How to Actually Evaluate Calibration

"Calibration is adequate" is a claim, not a measurement. Here is what to look for so you can judge it yourself rather than take the authors' word.

perfect (45°) model over-predicts here Predicted risk → Observed rate →

A calibration plot. Points on the dashed line = predictions match reality. This model is fine in the low-to-mid range but drifts below the line at high predicted risk — it over-predicts exactly where treatment thresholds live.

1. The calibration plot is the core tool. Bin (or smooth) patients by predicted risk, then plot observed event rate against mean predicted risk. Read the whole curve, not one number — many models calibrate well in the middle and fall apart at the high-risk extreme, which is precisely the range that drives treatment decisions.

2. Calibration-in-the-large / the O:E ratio is the overall level: total observed events ÷ total expected events. ~1.0 means the model is right on average. This is the single number recalibration (Section 8) is designed to fix.

3. The calibration slope is the spread: regress outcome on the linear predictor; slope = 1 is ideal. A slope < 1 means predictions are too extreme (high risks too high, low too low) — the classic fingerprint of overfitting.

Two traps to name in journal club. First, the O:E ratio and calibration-in-the-large are averages — a model can look perfectly calibrated overall yet be systematically off in a subgroup (a region, a sex, a disease bed). That is why SMART-REACH2 reporting calibration per region and per subgroup (Section 8) is the claim that matters, not a single pooled figure. Second, recalibration fixes the level (the O:E ratio), but it cannot rescue a bad slope — that needs refitting, not re-leveling. And a passing Hosmer–Lemeshow p-value is weak evidence: it is power-dependent, so it "passes" in small samples and "fails" in huge ones. Trust the plot over the test.

Now Apply It to SMART-REACH2

So how does this specific paper hold up against its own standard? Take the three tools to the actual model.

How they measured it. SMART-REACH2 assessed calibration exactly the way tool #1 prescribes — competing-risk-adjusted calibration plots of predicted vs observed risk — and, to its credit, reported them per region (the four ESC regions plus global regions), per sex, per CV-disease subtype, and across the age range. That breadth, not a single pooled figure, is the real calibration evidence. One honest caveat for journal club: the paper reports calibration only graphically, calling it "adequate." It gives precise C-statistics but no numeric calibration slope or calibration-in-the-large for the validation — so "adequate" is ultimately an eyeball verdict on the plots. Which means you should read them yourself.

Where to look on those plots: the high-risk tail. In several validation cohorts — the pooled very-high-risk region, the US Veterans Affairs cohort, BACS/BAMI, REACH–North America — the observed curve bends below the diagonal at the top end: the model over-predicts in the highest-risk patients, exactly the range that drives how hard you treat. (The VA gap was traced partly to CV events recorded outside the VA system; restricting to patients reliant on VA care alone restored agreement — a data-capture artifact, not pure model failure.) That is the calibration-plot reading skill, applied to this paper: the headline "adequate" is fair on average, but the tail tells you where to be cautious.

The O:E ratio isn't abstract here — it's the engine. SMART-REACH2's "recalibration" (Section 8) is tool #2 used as a fix: sex-specific expected/observed ratios computed per region and multiplied onto the baseline risk. That is precisely why the model can claim to travel. What re-leveling cannot do is repair a bad slope — and the paper deliberately keeps the original predictor coefficients, betting that the slope already transports across regions.

Section 5: Competing Risks — You Can Die of Something Else First

This is the methodological heart of the model, and it is where naive risk scores quietly go wrong — especially in exactly the older, sicker patients where the stakes are highest.

The Problem

A standard risk model implicitly assumes a patient stays alive until the event of interest occurs. But a 78-year-old with vascular disease, COPD, and a smoking history might die of lung cancer before they ever have their predicted recurrent MI. That non-cardiovascular death is a competing event — it removes the patient from the population that could have had a CV event.

What SMART-REACH2 Does About It

It fits two cause-specific models, not one: a model for recurrent CV events and a separate model for non-CV death. Their predicted hazards are combined through lifetables so that each year a patient can have a CV event, die of something else, or survive event-free. The lifetime CVD-free life expectancy falls out of running that competition forward to age 90.

Ignore Competing Risk → Overestimate

If you model CV events as if non-CV death did not exist, you give patients "credit" for years of CV risk they will never live to experience. The cumulative CV risk is inflated — most severely in elderly, comorbid patients. And because predicted benefit is proportional to predicted baseline risk (Section 2), you also overestimate treatment benefit in precisely the group where overtreatment is the real concern.

The Vocabulary, Demystified

"Cause-specific hazard" and "subdistribution hazard" (Fine–Gray) are two valid ways to handle competing events. The practical point for journal club: when a paper predicts absolute risk over a long horizon in an older population and does not mention competing risks, be suspicious that its risk — and any benefit derived from it — is inflated.

Why this is the right design for secondary prevention: the patients are older and multimorbid, the horizon is lifetime, and non-CV death is common. Competing-risk modeling is not statistical garnish here — it is the difference between a number you can act on and one that systematically pushes toward overtreatment of the frail.

Section 6: 10-Year Risk vs Lifetime Risk — the Horizon Changes Who You Treat

The same model can rank two patients in opposite order depending on the time horizon you ask about. This is not a bug — it is the most clinically important thing lifetime modeling does.

The Trap of 10-Year Risk: It Is Dominated by Age

Over a 10-year window, age swamps everything. A 75-year-old clears almost any "high risk" threshold on age alone, while a 50-year-old with terrible risk factors can look deceptively "low risk" — not because they are healthy, but because 10 years is too short a window for their risk-factor burden to express itself. Treat-by-10-year-risk systematically defers therapy in younger high-risk patients until they are older and have less to gain.

Two Patients, Opposite Conclusions (from the original SMART-REACH worked examples)

Patient A · age 55, smoker

26.7%

10-year risk (lower)

+2.0 years gained

from intensifying lipid therapy

Patient B · age 75

32.8%

10-year risk (higher)

+0.9 years gained

from the same therapy

By 10-year risk, you would prioritize Patient B. By lifetime benefit, Patient A — younger, with more years over which a high risk-factor burden does damage and over which treatment accrues benefit — gains more than twice as much. Same model, same therapy, reversed priority.

The meta-lesson: the horizon is a choice, and the choice has consequences. Lifetime estimation is the analogue of the QRISK-lifetime insight from primary prevention — it identifies patients who benefit most at a younger age, before a decade of avoidable exposure has already accrued. When a model reports both a 10-year and a lifetime number, look at both, and notice when they disagree.

Section 7: Treatment Benefit in a Currency Patients Understand

"Your 10-year risk drops from 32% to 27%" is hard to feel. "This adds about three years of life free of stroke or heart attack" lands. SMART-REACH2 outputs the second kind of number — and that is its real clinical contribution.

The Worked Example from the Paper

For a 50-year-old example patient, intensified preventive treatment — a 15 mmHg reduction in systolic blood pressure plus a 1.0 mmol/L reduction in LDL cholesterol — was estimated to add:

Low-risk region

+2.0 yrs

CVD-free life expectancy

Very-high-risk region

+4.4 yrs

CVD-free life expectancy

Same patient profile, same therapy — more than double the absolute benefit depending on the background event rate of where they live. That regional spread is the entire motivation for the recalibration work in Section 8.

Benefit and Harm in the Same Units

The lifetime framework's best trick is putting harm on the same axis as benefit. Applying the SMART-REACH engine to the COMPASS trial (de Vries et al., Eur Heart J 2019), adding low-dose rivaroxaban to aspirin gave a median 16 months of life free of stroke or MI gained — against a median 2 months of life free of major bleeding lost. Both ranged widely across individuals (gain 1–48 months; harm 0–20 months).

When benefit and harm are in the same currency — months of life in a given state — shared decision-making becomes an honest conversation instead of a clash of incommensurable numbers.

Why guidelines like this: the 2021 ESC prevention guideline calls for shared decision-making and for matching treatment intensity to absolute benefit. A tool that says "about three years, and here is the bleeding cost" operationalizes that recommendation far better than a risk percentage or an NNT a patient cannot picture.

Section 8: Recalibration and External Validation — Does the Model Travel?

A model derived in one Dutch cohort will mis-estimate risk in Spain, Poland, or Japan — not because the biology differs, but because the background event rates differ. SMART-REACH2 is, more than anything, the answer to "can we make this one model work everywhere?"

Recalibration, Precisely

Recalibration rescales the model's baseline risk to a target population's event rate — typically using the ratio of expected to observed events — while keeping the predictor effects (the coefficients) unchanged. It is not refitting the model. The relationships between risk factors and outcome are assumed transportable; only the overall level is adjusted, region by region (here, the four ESC/SCORE2 European risk regions plus other global regions).

The Validation, by the Numbers

Patients

2.09M

external validation

Countries

54

across global regions

Events

307,706

recurrent CV events

Discrimination held in the 0.66–0.72 band across regions, calibration was adequate after recalibration, and performance was consistent across sexes and across CV disease subtypes. That last point matters for a vascular audience: the model is not a coronary model wearing a vascular coat — it behaved consistently in cerebrovascular, peripheral, and aortic disease too.

What the Recalibration Actually Bought — a Concrete Test

The cleanest evidence that re-leveling mattered is the head-to-head against the original SMART-REACH model. In the European moderate-risk region, the old model underestimated risk — its calibration plot sat below the diagonal — while SMART-REACH2, re-levelled with that region's expected/observed ratio, stayed on the line. On decision-curve analysis, that translated into higher net benefit for the updated model. The recalibration was not cosmetic: a model that understates risk understates benefit, and would have left treatable moderate-risk patients under-treated.

The mechanism is worth stating precisely: the recalibration factors are sex-specific expected/observed ratios, derived per region from a representative cohort (e.g. CPRD for European low-risk, SWEDEHEART for moderate-risk, the Estonian Biobank for high-risk) and multiplied onto the baseline risk. The predictor coefficients never change — only the level does.

Why external validation is the whole ballgame. A model that reports only development-cohort performance is a hypothesis, not a tool — it has not been shown to survive contact with new patients, new geography, or a new era. Apparent performance is optimistic by construction (the model was fit to that data). Transportability — demonstrated here at scale — is what separates a usable instrument from a curve-fit.
What to still check at your bedside: even a well-validated, recalibrated model is conditional on the patient resembling the validation populations. Recalibration fixes the average level for a region; it does not rescue a patient whose phenotype (e.g., severe heart failure, end-stage renal disease) was sparse or excluded in development. The model informs the decision; it does not make it.

Section 9: How to Read a Prediction-Model Paper — Then Test Yourself

Trials have CONSORT; prediction models have TRIPOD. Here is the short checklist to carry into any modeling journal club, distilled from this paper.

The Pocket Checklist

Ask Why it matters
Derivation vs validation? Development performance is optimistic. Demand external validation.
Both discrimination AND calibration? A good C-statistic with bad calibration still misguides treatment.
Competing risks handled? Long horizon + older patients = mandatory, or risk is inflated.
Predictors routinely available? A model you cannot populate at the bedside is unused.
Applies to my patient? Conditional on the development/validation case-mix.
The honest limitations of SMART-REACH2 (and its family): risk factors are measured once at baseline and assumed constant for life (they are not); discrimination is modest by construction; and lifetime predictions extrapolate beyond the observed follow-up window, leaning on the age-as-timescale machinery. None of these sink the model — but a good journal club names them out loud.

Apply what you have learned. Five questions.

Key Takeaways

  • 1
    A model personalizes a trial's average effect

    SMART-REACH2 doesn't prove therapies work — it estimates each patient's baseline risk and converts a proven relative effect into an absolute, individual benefit. Guideline-referenced (ESC 2021), updated June 2026.

  • 2
    Absolute benefit scales with baseline risk

    A fixed RRR yields large absolute benefit in high-risk patients and trivial benefit in low-risk ones. Treat risk, not isolated risk factors.

  • 3
    Discrimination ≠ calibration

    C-statistic 0.68 is modest but near the ceiling for a homogeneous secondary-prevention population (and beats ~0.53 guideline labels). For treatment decisions, calibration is the property that counts.

  • 4
    Competing risks are non-negotiable here

    Two cause-specific models (CV events; non-CV death) combined via lifetables, age as timescale. Ignoring competing death inflates risk — and benefit — in the elderly.

  • 5
    Lifetime view can reverse the 10-year priority

    10-year risk is age-dominated; lifetime benefit favors younger high-risk patients with more years to gain. The 55-yr-old gained +2.0 years vs +0.9 for the 75-yr-old from the same therapy.

  • 6
    Benefit (and harm) in life-years

    50-yr-old example: +2.0 to +4.4 CVD-free years from BP+LDL lowering, depending on risk region. COMPASS: +16 months free of stroke/MI vs −2 months free of major bleeding — same currency, honest trade-off.

  • 7
    Recalibration + external validation make it travel

    Re-leveled to regional event rates without refitting; validated in 2.09M patients across 54 countries (307,706 events), C 0.66–0.72, adequate calibration, consistent across sexes and vascular beds.

The Bottom Line

SMART-REACH2 is a worked example of what a good prediction model is: a competing-risk, lifetime engine that turns proven relative effects into individualized absolute benefit, reports calibration as carefully as discrimination, and proves it transports across populations before asking you to trust it. For the patient with established vascular disease, it answers the question a trial cannot — "how much will this help you, over the years you have left?" — and it answers in a currency you can both understand.

Before You Go: A Quick Poll

A 52-year-old with PAD and a prior MI, still smoking, LDL 3.5 mmol/L, SBP 150 mmHg, normal renal function. Her 10-year recurrent-event risk looks only "moderate." What drives your decision to intensify preventive therapy?