All Insights

Clinical Adoption

Your clinicians are right to ignore your risk scores

Dismissal is not an adoption problem. It is quality control. The industry has spent a decade calling it resistance because that was the more comfortable answer.

Oz Josefian, PresidentAugust 27, 20265 min read

There is a story health systems tell about clinical AI, and it goes roughly like this. We bought a good model. We deployed it. The clinicians would not adopt it. We need better change management.

Let me offer the less flattering version. The model was mediocre, it arrived with no reasoning attached, and the care team worked that out faster than the committee that bought it did.

I sell predictive software, so I am aware of how that reads coming from me. I am saying it anyway, because the accuracy arms race this industry has been running is the main reason clinical AI keeps getting deployed and then quietly abandoned, and I would rather argue against my own category's favorite sales metric than keep competing on a number that was never the bottleneck.

Dismissal is a measured behavior, not a cultural one

Clinical decision support has been studied for years and the pattern is consistent enough to be uncomfortable. A systematic review of alert override research found average override rates ranging from roughly 46% to 96% across studies. A separate meta-analysis of drug interaction alerts put the overall physician override rate at about 90%. One study found the likelihood of a clinician accepting a reminder dropped by about 30% for every additional reminder they received during an encounter. More alerting produces less action, not more.

The industry reflex is to treat this as a tuning problem. Raise the threshold, cut the volume, redesign the pop-up. Every one of those responses assumes the clinician is the malfunctioning component.

They are not. A nurse deciding whether an unexplained number is worth ten minutes of a twelve hour shift is running a cost benefit analysis, correctly, dozens of times a day. When most of those numbers turn out to be noise, the rational response is to stop reading them. That is not resistance to innovation. That is a well calibrated professional doing exactly what you would want them to do.

The case everyone watched and then mostly filed away

In 2021, researchers at Michigan Medicine published an external validation of a widely deployed proprietary sepsis prediction model. Across 38,455 hospitalizations, the model achieved an area under the curve of 0.63, with sensitivity of 33% and positive predictive value of 12%. That sat well below the 0.76 to 0.83 range cited in the vendor's own documentation. In practice it missed roughly two thirds of sepsis cases while producing enough false alarms that clinicians would review on the order of 109 alerts to find one true case.

The accuracy gap got the headlines. The finding I consider genuinely damning got much less attention. Because the model drew on whether antibiotics had been ordered, it was partly detecting that a clinician had already suspected sepsis. The tool was congratulating physicians on conclusions they had reached without it, and the industry called that a prediction.

None of that was visible from the bedside. A nurse saw a flag with no reasoning attached, chased it, found nothing, and repeated that experience often enough to learn what the flag was worth. The staff who learned to dismiss that alert were not obstructing the deployment. For a period of time they were the only functioning validation layer in the building.

The vendor eventually rebuilt the model and began advising hospitals to train it on local data. That was the right correction. What has not been corrected is the assumption underneath the original deployment, which is that a number is worth acting on because it is confident.

Change management is what you do when the tool is good

When a health system responds to low adoption by appointing champions and scheduling training, it is worth asking what the training is actually for.

If a tool earns trust, adoption is not a project. It spreads because the first three people who used it told the next ten. If a tool does not earn trust, a change management program becomes an expensive way to inform your staff that their professional judgment is the obstacle. I would call that gaslighting if the intent were not usually sincere.

The uncomfortable implication is that low adoption is often accurate feedback arriving through the only channel available. Health systems pay consultants for less useful signal than that, and then override it.

Three questions a score has to answer

A score that changes behavior is not a more accurate score. It is a score that arrives with enough context to be argued with.

Why this patient

Not the probability. The drivers. Which clinical and social factors pushed this person up the list, in language a discharge planner reads rather than feature importance values a data scientist reads. A clinician who can see the reasoning can also see when the reasoning is wrong, and that ability to disagree is the entire foundation of trusting the tool at all. A model you cannot argue with is not authoritative. It is unaccountable.

Why now

A signal that arrives after the decision point is not intelligence, it is documentation. The sepsis case makes the point precisely, since a warning that lands at or after the moment of clinical suspicion adds nothing to the workflow no matter how well calibrated it is.

What do I do

A risk score with no attached action transfers work to the person least able to absorb it. Follow-up call, home health referral, transitional care management, medication reconciliation. The recommendation does not need to be right every time. It needs to exist, so the clinician is editing a plan rather than building one at the end of a shift.

Explainability is not a compliance item

Most of the market files explainability somewhere near model documentation and bias auditing. Necessary, unglamorous, not a reason anyone buys. I think that is exactly backwards, and I think it survives because explainability is expensive to build and nearly free to claim.

For value-based care organizations and ACOs, shared savings are produced by care team behavior, not by model output. A model with excellent discrimination that nobody acts on generates exactly zero dollars. The reasoning attached to a score is not a trust feature bolted onto the prediction. It is the mechanism by which a prediction becomes money.

The same holds for health plans and payers evaluating therapy adherence intelligence. Identifying members at risk of discontinuing therapy is half a transaction. If the outreach team cannot see why a specific member was flagged, they cannot tailor the intervention, and a generic intervention on a correctly identified member still fails. You will have bought an accurate list and no outcome.

Use this to disqualify me

If you are evaluating predictive tools this year, do not open with the AUC question. Any vendor can produce a flattering number on a flattering cohort, and the sepsis validation showed exactly how far internal figures can drift from real world performance.

Ask instead to see a single score with its reasoning attached. Then ask what the nurse does in the sixty seconds after reading it. Apply that to my company as readily as to anyone else's. If a vendor cannot show you both, you are not buying intelligence. You are buying a number and the hope that somebody acts on it.

Clinicians are not resistant to prediction. They are resistant to being told to trust something that will not explain itself. That is not friction to be managed. It is the correct standard, and this industry will keep failing expensively until it stops treating the people at the bedside as the problem to be solved.

Sources

  • Wong A, et al. External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients. JAMA Internal Medicine, 2021. View study
  • Felisberto M, et al. Override rate of drug-drug interaction alerts in clinical decision support systems: a systematic review and meta-analysis. View study
  • Appropriateness of Overridden Alerts in Computerized Physician Order Entry: Systematic Review. JMIR Medical Informatics. View study

About the author

Oz Josefian is President of Preventra, a predictive care intelligence platform for hospitals, value-based care organizations, and health plans.

Then ask us the same question.

15 minutes is all it takes to see what Preventra can do for your organization.

Request a 15-Min Walkthrough