A lifetime cardiovascular risk model for patients with established atherosclerotic disease — and a working tour of discrimination, calibration, competing risks, and individualized treatment benefit.
A randomized trial tells you what a therapy does on average. The patient in your clinic is not the average. A prediction model is the bridge — it turns a trial's average relative effect into this person's absolute benefit. SMART-REACH2, published in the European Heart Journal in June 2026, is the current guideline-referenced version of that bridge for established atherosclerotic disease.
Two observational cohorts produced short-horizon recurrence scores: the 10-year SMART score and the 20-month REACH score. Useful for ranking risk, but no competing-risk handling and no lifetime view.
Kaasenbrood et al. (JAHA) combined the cohorts into a competing-risk lifetime model estimating CVD-free life expectancy — and, crucially, life-years gained from a given therapy.
The ESC CVD-prevention guidelines named SMART-REACH as a tool to guide treatment decisions in patients with established ASCVD. Guideline endorsement raises the bar for transportability and calibration.
Re-derived and systematically recalibrated to four European and other global risk regions, then externally validated in over two million patients across 54 countries. The update is mostly about making the model travel — and proving it does.
For an individual with established coronary, cerebrovascular, peripheral, or aortic atherosclerotic disease: what is their lifetime risk of a recurrent cardiovascular event, and how many cardiovascular-disease-free life-years would a given preventive therapy add for them?
Note the shift in object. A trial asks "does the drug work?" A model asks "given that it works, how much does it do for this patient?"
The whole reason individualized models exist is one stubborn fact: relative risk reduction is roughly constant across patients, but absolute benefit is not. Absolute benefit scales with baseline risk.
Absolute Risk Reduction ≈ Baseline Risk × Relative Risk Reduction. A therapy with a fixed 25% RRR delivers a large absolute benefit to a high-risk patient and a trivial one to a low-risk patient — same drug, same RRR.
| Baseline 10-yr risk | RRR | Absolute benefit (ARR) | NNT |
|---|---|---|---|
| 10% | 25% | 2.5% | 40 |
| 30% | 25% | 7.5% | ~13 |
| 50% | 25% | 12.5% | 8 |
The drug never changed. Only the patient did. The model's entire job is to estimate the first column accurately for an individual.
Every prediction model has a development ("derivation") dataset, an outcome definition, a set of predictors, and a statistical engine. Get these four straight and you understand most of any modeling paper.
The predictors are things already in the chart: type and number of vascular beds involved, systolic blood pressure, lipids (non-HDL cholesterol), renal function (eGFR), smoking, diabetes, duration of disease, and an inflammation marker. No exotic assays, no genomics. A model that needs a test you cannot order at the bedside is a model nobody uses.
Most Cox models use time-since-enrollment as the clock. SMART-REACH-type lifetime models use attained age as the clock instead, with left truncation (patients enter the risk set at the age they joined). This is what lets a cohort with ~8 years of follow-up generate predictions across the whole 40–90 age span — you borrow information across people at different ages rather than waiting decades to observe one person age.
A prediction model has to do two different things, and they are not the same thing. Conflating them is the single most common error in reading these papers.
Can it rank? Does it give higher-risk patients higher scores than lower-risk patients?
Measured by the C-statistic (AUC). 0.5 = coin flip, 1.0 = perfect ranking.
Are the numbers right? When the model says 30%, do about 30% of those patients actually have events?
Assessed by predicted-vs-observed plots and expected/observed ratios.
The pooled C-statistic was 0.68 (95% CI 0.66–0.69), ranging from 0.66 in the European low-risk region to 0.72 in Latin America. By the usual rough labels that is "modest." A colleague will say "0.68, that's barely better than a coin flip's cousin — useless."
They are wrong, and here is why.
Discrimination depends on how spread out risk is in the population. In primary prevention you are separating healthy 40-year-olds from sick 70-year-olds — easy, C-statistics run high. In secondary prevention, everyone already has established disease. The population is homogeneously high-risk, so there is simply less spread to discriminate. C-statistics around 0.65–0.70 are typical and near the practical ceiling; adding more predictors has repeatedly failed to move them much.
For context: the categorical "very high risk" criteria in some guidelines discriminate at roughly 0.53–0.54. A calibrated 0.68 model is a real improvement over the yes/no labels clinicians use now.
"Calibration is adequate" is a claim, not a measurement. Here is what to look for so you can judge it yourself rather than take the authors' word.
A calibration plot. Points on the dashed line = predictions match reality. This model is fine in the low-to-mid range but drifts below the line at high predicted risk — it over-predicts exactly where treatment thresholds live.
1. The calibration plot is the core tool. Bin (or smooth) patients by predicted risk, then plot observed event rate against mean predicted risk. Read the whole curve, not one number — many models calibrate well in the middle and fall apart at the high-risk extreme, which is precisely the range that drives treatment decisions.
2. Calibration-in-the-large / the O:E ratio is the overall level: total observed events ÷ total expected events. ~1.0 means the model is right on average. This is the single number recalibration (Section 8) is designed to fix.
3. The calibration slope is the spread: regress outcome on the linear predictor; slope = 1 is ideal. A slope < 1 means predictions are too extreme (high risks too high, low too low) — the classic fingerprint of overfitting.
So how does this specific paper hold up against its own standard? Take the three tools to the actual model.
How they measured it. SMART-REACH2 assessed calibration exactly the way tool #1 prescribes — competing-risk-adjusted calibration plots of predicted vs observed risk — and, to its credit, reported them per region (the four ESC regions plus global regions), per sex, per CV-disease subtype, and across the age range. That breadth, not a single pooled figure, is the real calibration evidence. One honest caveat for journal club: the paper reports calibration only graphically, calling it "adequate." It gives precise C-statistics but no numeric calibration slope or calibration-in-the-large for the validation — so "adequate" is ultimately an eyeball verdict on the plots. Which means you should read them yourself.
Where to look on those plots: the high-risk tail. In several validation cohorts — the pooled very-high-risk region, the US Veterans Affairs cohort, BACS/BAMI, REACH–North America — the observed curve bends below the diagonal at the top end: the model over-predicts in the highest-risk patients, exactly the range that drives how hard you treat. (The VA gap was traced partly to CV events recorded outside the VA system; restricting to patients reliant on VA care alone restored agreement — a data-capture artifact, not pure model failure.) That is the calibration-plot reading skill, applied to this paper: the headline "adequate" is fair on average, but the tail tells you where to be cautious.
The O:E ratio isn't abstract here — it's the engine. SMART-REACH2's "recalibration" (Section 8) is tool #2 used as a fix: sex-specific expected/observed ratios computed per region and multiplied onto the baseline risk. That is precisely why the model can claim to travel. What re-leveling cannot do is repair a bad slope — and the paper deliberately keeps the original predictor coefficients, betting that the slope already transports across regions.
This is the methodological heart of the model, and it is where naive risk scores quietly go wrong — especially in exactly the older, sicker patients where the stakes are highest.
A standard risk model implicitly assumes a patient stays alive until the event of interest occurs. But a 78-year-old with vascular disease, COPD, and a smoking history might die of lung cancer before they ever have their predicted recurrent MI. That non-cardiovascular death is a competing event — it removes the patient from the population that could have had a CV event.
It fits two cause-specific models, not one: a model for recurrent CV events and a separate model for non-CV death. Their predicted hazards are combined through lifetables so that each year a patient can have a CV event, die of something else, or survive event-free. The lifetime CVD-free life expectancy falls out of running that competition forward to age 90.
If you model CV events as if non-CV death did not exist, you give patients "credit" for years of CV risk they will never live to experience. The cumulative CV risk is inflated — most severely in elderly, comorbid patients. And because predicted benefit is proportional to predicted baseline risk (Section 2), you also overestimate treatment benefit in precisely the group where overtreatment is the real concern.
"Cause-specific hazard" and "subdistribution hazard" (Fine–Gray) are two valid ways to handle competing events. The practical point for journal club: when a paper predicts absolute risk over a long horizon in an older population and does not mention competing risks, be suspicious that its risk — and any benefit derived from it — is inflated.
The same model can rank two patients in opposite order depending on the time horizon you ask about. This is not a bug — it is the most clinically important thing lifetime modeling does.
Over a 10-year window, age swamps everything. A 75-year-old clears almost any "high risk" threshold on age alone, while a 50-year-old with terrible risk factors can look deceptively "low risk" — not because they are healthy, but because 10 years is too short a window for their risk-factor burden to express itself. Treat-by-10-year-risk systematically defers therapy in younger high-risk patients until they are older and have less to gain.
10-year risk (lower)
+2.0 years gained
from intensifying lipid therapy
10-year risk (higher)
+0.9 years gained
from the same therapy
By 10-year risk, you would prioritize Patient B. By lifetime benefit, Patient A — younger, with more years over which a high risk-factor burden does damage and over which treatment accrues benefit — gains more than twice as much. Same model, same therapy, reversed priority.
"Your 10-year risk drops from 32% to 27%" is hard to feel. "This adds about three years of life free of stroke or heart attack" lands. SMART-REACH2 outputs the second kind of number — and that is its real clinical contribution.
For a 50-year-old example patient, intensified preventive treatment — a 15 mmHg reduction in systolic blood pressure plus a 1.0 mmol/L reduction in LDL cholesterol — was estimated to add:
CVD-free life expectancy
CVD-free life expectancy
Same patient profile, same therapy — more than double the absolute benefit depending on the background event rate of where they live. That regional spread is the entire motivation for the recalibration work in Section 8.
The lifetime framework's best trick is putting harm on the same axis as benefit. Applying the SMART-REACH engine to the COMPASS trial (de Vries et al., Eur Heart J 2019), adding low-dose rivaroxaban to aspirin gave a median 16 months of life free of stroke or MI gained — against a median 2 months of life free of major bleeding lost. Both ranged widely across individuals (gain 1–48 months; harm 0–20 months).
When benefit and harm are in the same currency — months of life in a given state — shared decision-making becomes an honest conversation instead of a clash of incommensurable numbers.
A model derived in one Dutch cohort will mis-estimate risk in Spain, Poland, or Japan — not because the biology differs, but because the background event rates differ. SMART-REACH2 is, more than anything, the answer to "can we make this one model work everywhere?"
Recalibration rescales the model's baseline risk to a target population's event rate — typically using the ratio of expected to observed events — while keeping the predictor effects (the coefficients) unchanged. It is not refitting the model. The relationships between risk factors and outcome are assumed transportable; only the overall level is adjusted, region by region (here, the four ESC/SCORE2 European risk regions plus other global regions).
external validation
across global regions
recurrent CV events
Discrimination held in the 0.66–0.72 band across regions, calibration was adequate after recalibration, and performance was consistent across sexes and across CV disease subtypes. That last point matters for a vascular audience: the model is not a coronary model wearing a vascular coat — it behaved consistently in cerebrovascular, peripheral, and aortic disease too.
The cleanest evidence that re-leveling mattered is the head-to-head against the original SMART-REACH model. In the European moderate-risk region, the old model underestimated risk — its calibration plot sat below the diagonal — while SMART-REACH2, re-levelled with that region's expected/observed ratio, stayed on the line. On decision-curve analysis, that translated into higher net benefit for the updated model. The recalibration was not cosmetic: a model that understates risk understates benefit, and would have left treatable moderate-risk patients under-treated.
The mechanism is worth stating precisely: the recalibration factors are sex-specific expected/observed ratios, derived per region from a representative cohort (e.g. CPRD for European low-risk, SWEDEHEART for moderate-risk, the Estonian Biobank for high-risk) and multiplied onto the baseline risk. The predictor coefficients never change — only the level does.
Trials have CONSORT; prediction models have TRIPOD. Here is the short checklist to carry into any modeling journal club, distilled from this paper.
| Ask | Why it matters |
|---|---|
| Derivation vs validation? | Development performance is optimistic. Demand external validation. |
| Both discrimination AND calibration? | A good C-statistic with bad calibration still misguides treatment. |
| Competing risks handled? | Long horizon + older patients = mandatory, or risk is inflated. |
| Predictors routinely available? | A model you cannot populate at the bedside is unused. |
| Applies to my patient? | Conditional on the development/validation case-mix. |
Apply what you have learned. Five questions.
SMART-REACH2 doesn't prove therapies work — it estimates each patient's baseline risk and converts a proven relative effect into an absolute, individual benefit. Guideline-referenced (ESC 2021), updated June 2026.
A fixed RRR yields large absolute benefit in high-risk patients and trivial benefit in low-risk ones. Treat risk, not isolated risk factors.
C-statistic 0.68 is modest but near the ceiling for a homogeneous secondary-prevention population (and beats ~0.53 guideline labels). For treatment decisions, calibration is the property that counts.
Two cause-specific models (CV events; non-CV death) combined via lifetables, age as timescale. Ignoring competing death inflates risk — and benefit — in the elderly.
10-year risk is age-dominated; lifetime benefit favors younger high-risk patients with more years to gain. The 55-yr-old gained +2.0 years vs +0.9 for the 75-yr-old from the same therapy.
50-yr-old example: +2.0 to +4.4 CVD-free years from BP+LDL lowering, depending on risk region. COMPASS: +16 months free of stroke/MI vs −2 months free of major bleeding — same currency, honest trade-off.
Re-leveled to regional event rates without refitting; validated in 2.09M patients across 54 countries (307,706 events), C 0.66–0.72, adequate calibration, consistent across sexes and vascular beds.
SMART-REACH2 is a worked example of what a good prediction model is: a competing-risk, lifetime engine that turns proven relative effects into individualized absolute benefit, reports calibration as carefully as discrimination, and proves it transports across populations before asking you to trust it. For the patient with established vascular disease, it answers the question a trial cannot — "how much will this help you, over the years you have left?" — and it answers in a currency you can both understand.
A 52-year-old with PAD and a prior MI, still smoking, LDL 3.5 mmol/L, SBP 150 mmHg, normal renal function. Her 10-year recurrent-event risk looks only "moderate." What drives your decision to intensify preventive therapy?