17 Comments
User's avatar
George's avatar

“because risk reduction keeps accruing in the years beyond the window during which it was measured”; how can we know this if not measured “ beyond the window “?

James H. Stein, MD's avatar

Because numerous large RCTs have measured long term follow up after the randomized phase ended, by tracking outcomes through national death and hospitalization registries. WOSCOPS (20-year linkage follow-up), LIPID (2-year extended follow-up with interviews), and ASCOT-LLA (11- and 16-year linkage follow-up) all did this. Also, 4S and the Coronary Drug Project's 15-year niacin follow-up.

Dr. Khadija Siddiqui's avatar

One of the hardest shifts is helping people think in probabilities instead of certainties. We tend to judge decisions by outcomes, but in prevention, the quality of a decision is determined by the information available at the time, not by whether the unfavorable event ultimately happened.

Frank Rockhold's avatar

The age old problem in estimating benefit-risk in preventative therapies (including vaccines). You always know who had an adverse event but never know who did not have en event due to the treatment so the patient benefit risk ratio is always skewed emotionally. Good graphics on this by David Spiegelhalter: "Patient reactions to a web-based cardiovascular risk calculator in type 2 diabetes: a qualitative study in primary care" https://pubmed.ncbi.nlm.nih.gov/25733436/

James H. Stein, MD's avatar

Well said! Thanks for the link!

Steve Cheung's avatar

I’m not sure I agree here.

The top-line quoted benefit is of an endpoint defined prospectively, but nonetheless tabulated/measured after the fact. For example, as a hypothetical, an event rate of 10% in the active treatment arm vs 20% in the control arm. From which we say there is a 10% ARR (and in this case, 50%RRR) on average. But of the 100 (for example) people in the active arm, none of them enjoyed a 10% absolute risk reduction, even counting from day 1 of the study; rather, nothing of note happened in 90 of them during the entirety of follow up, while something of note happened in 10.

I think you can quote the average effect on a population level but I don’t think you can do so on a patient specific level. (Same with the weather forecast: on 100 prior days with atmospheric conditions like today, past data shows it rained on 30 of them….that doesn’t mean it’s actually 30% chance of rain today….and afterwards if it DOESN’T rain, then today would have been just like any of the 70 days from historical samples when it also did not rain).

I agree with Frank Rockhold’s comment. If you take a treatment and nothing bad happens, you don’t know if it’s because of the treatment, or if nothing bad would have happened regardless (much like 80 of the 100 people in the control group of my hypothetical above). That is part of the known unknowns that we have to deal with every day. You only have certainty when an event occurs; there can be no certainty when it comes to an absence of an event.

James H. Stein, MD's avatar

Dear Steve, thanks for engaging with this. I am not entirely sure what part you disagree with, so let me restate my point and see whether it is clearer, or perhaps more agreeable.

We obviously agree that outcomes are binary, that ARR is a population-level measure, that we do not know anyone's personal risk, and that we cannot know whether any particular individual would have had an event without treatment or had one prevented by treatment. We only know group-level risks and treatment effects. And we know that no one has a fractional cardiac event.

But probability is not just a retrospective accounting tool. It is a real property of each person's situation before the outcome is known.

Take your own example. On day one of the trial, people like those enrolled had, on average, a 20% risk of an event without treatment and a 10% risk with treatment. Those estimates apply prospectively to each person, even though we cannot know any individual's eventual outcome or counterfactual outcome. The fact that 90 out of 100 had nothing happen does not mean the reduction was not real for them. It means we cannot identify which ten events were prevented.

Your weather analogy actually makes my point. On a day with a 30% chance of rain, you decide whether to carry an umbrella. If the day stays dry, the 30% was not wrong, and you were not foolish to carry one. Saying afterward that it did not rain, so the forecast was meaningless to you, confuses what happened with what was knowable when the decision was made. That is exactly what happens when someone says 90% of statin-treated patients got no benefit.

I hope that clarifies the distinction I am drawing. - Jim

Steve Cheung's avatar

Thanks for your response.

Let me first say that I had to look up weather forecasting, as I didn’t know how it’s done, and wanted to do so before I went too far towards removing all doubt that I’m an idiot. In my brief review, I was wrong earlier. The probability of precipitation is the product of the “confidence” of rain, and the fraction of geographic area affected. The confidence is based on reams of atmospheric data, then inputted into a computer model where multiple simulations are run, and the “confidence” is the proportion of simulations that suggest some rain of any measurable extent. I don’t know how the models are produced. Perhaps based on past data, but here again I’m starting down the “idiot” path once more so will remain silent on this instead.

Yes, there is much we agree on as to what trial data means (ie on a population/group level). However, as applied to patients, and specifically how to make a determination of the extent an intervention affects their individual risk, I am unsure we have that capacity, even when armed with the best data from the most methodologically rigorous clinical outcome studies (and I would also submit most studies do not belong in this rarefied category).

I think the crux here is when you say (based on our examples earlier) “which 10 were prevented”. How do you square an assertion that all subjects enjoy the average ARR of an intervention, with the observation that some of those subjects experienced an endpoint nonetheless? I don’t understand where the average benefit lies in that subject. They could not have started at >110% absolute risk of an outcome, such that even after a treatment benefit ARR of 10%, they were still at 100% and hence had an event.

This is why when I discuss with patients, I use “benefits on average”. And this comes up quite often with AF and DOAC. I tell them the data is clear they should take it (if chads score is appropriate), because of the size of the average benefit, but I can’t absolutely guarantee they themselves won’t still get a stroke on treatment.

I do agree with your concluding point. Even if an event occurs, it doesn’t in hindsight negate the worthiness of being on a therapy starting at some earlier point. But I think I disagree with how it would be framed on an individual patient level.

James H. Stein, MD's avatar

Hi Steve, your clinical framing to patients sounds exactly right. You asked how I square all 100 enjoying the RRR when 10 still had events. The answer is probability. A person whose risk drops from 20% to 10% still has a 10% chance of an event. Some of them will have events, not because the treatment failed to reach them, but because that is what a 10% probability means. In any given sample, fewer or more than 10 out of 100 might have events; that is the consequence of probability, not a failure of the treatment. Your “>110%” thought experiment assumes the outcome was preordained and the treatment either got there in time or did not. But outcomes are not predetermined. The treatment shifts the probability for everyone, and chance determines who lands in the residual risk. That is how I square it, and by the likeness of our clinical framing, it sounds like we ended up at the same place! Yay us! 😊

The Diagnostic Detective's avatar

I think you are both right. The key is how you communicate to a patient. I would say. For every 100 patients who don’t take this treatment for 5 years, 20 will have an event and 80 won’t. For every 100 that do take it 10 will have an event and 90 won’t. Would you like to take it? I think patients understand those kinds of figures and can make their own minds up. I certainly wouldn’t be talking to them about the weather forecast.

James H. Stein, MD's avatar

That is a reasonable way to explain it to patients. But at the risk of being pedantic, I want to point out that "20 will and 80 won't" is a deterministic framing of a probabilistic reality, and that distinction is what this whole discussion has been about. We do not actually know that 20 will have events and 80 won't. Those numbers describe the average outcomes across average groups, not a predetermined tally. There are confidence intervals around the future incidence in each arm, and accordingly, around the differences between them. In any given 100 patients, the actual number might be 17 or 23. That is the nature of probability, and it is the whole reason the decision is made under uncertainty. It's the reason I like to start these discussions with: "I can't predict the future; no one can. All we can do is make our best guess." As always, thanks for engaging. Moving on from this thread.

Kuma Folmsbee's avatar

Over my 15 years of practice there are two main issues that have developed from this line of thinking 1) The slow creep of low magnitude treatment effects being recommended 2) we teach everything as equivalent theraputics. We need to do everything all the time, all at once to satisfy all the check boxes.

A Majority of patients (not all) would probably want to take a treatment that has an ARR of 10, but what about a ARR of 1? (think of all cancer screening) and then think that the reduction is from an outcome that is less clinically relevant (disesae specific mortality).

I agree with everything said, I just am no longer sure how best to frame discussion (especially when we start to nitpick small numbers) as "shared decision making"

The Diagnostic Detective's avatar

On the flip side I would argue that sometimes the ARR is greater than 100 and the NNT less than 1. If you treat a single mother in rural Africa with anti retrovirals for HIV you're effectively keeping 4 or 5 people alive.

James H. Stein, MD's avatar

Your point about some interventions being powerful enough to help beyond the patient is well taken. Hiwever the numbers as you frame them are deterministic, which is the framing I am trying to move away from in this piece, though the broader point stands. I think we can wrap this thread up here. Thanks, everyone, for engaging.

James H. Stein, MD's avatar

Kuma, thanks for this. Two quick responses. First, the line of thinking I described has not led to slow creep of low-magnitude treatment effects, and to be honest, I am a little offended by that implication. The approach I described in this and my last post is what makes shared decision making possible, rather than a meaningless virtue signal in a guideline or the medical record. Without honest quantification of absolute benefits of any magnitude, bounded by the time frame, you cannot have an informed conversation with a patient about whether treatment is worth it to them. You cannot assume a patient would not accept a 1% ARR. That is their call, not ours. Some patients will want that reduction and some will not, and the whole point of preference-sensitive care is that we present the numbers and let the patient decide. Taking that option away on the assumption that the benefit is too small to matter is the opposite of shared decision making. And second, as I discussed in my last post, saying an X% ARR is misleading. It flattens the benefit into an average effect when what is really at stake is a chance of preventing a potentially large one (see Finegold et al. Open Heart 2016;3:e000343. doi: 10.1136/openhrt-2015-000343). It also ignores the time frame, because benefits measured over five years understate the value of treatment that continues beyond that window. EBM cuts both ways, and I fear the movement, which I consider myself part of, has slipped into therapeutic nihilism.

Kuma Folmsbee's avatar

Thanks. Your points are well taken. I’ll end by saying I do not think this line of thinking has led to low value treatments. My points are just me trying to reconcile the day to day growing list of metrics we primary care doctors are being tasked with, all of which should warrant careful shared decision making. Academic medicine speaks of shared decision making, yet in the real practice of medicine, my salary is partly tied to how many mammograms or statins I order. Anyway, really appreciate your comments and posts.